Open the source
The complete project includes the environment, training loop, tests, and dependency configuration.
Architecture
The boundary between the two systems is explicit:
The cookbook supplies the
FireworksAgent adapter, computes group-relative advantages, and converts
completed Runs into Fireworks training datums.
The example uses multiplication because it is small and easy to verify. HUD also represents
environments with browsers, code execution, games, APIs, and stateful tools; the adapter changes
needed for those environments are covered below.
Setup
Install Git, Python 3.11 or 3.12, anduv. Then clone the SDK and enter the
cookbook directory:
hud-python/cookbooks/fireworks-rl-training.
Set a Fireworks API key with Serverless Training access:
accounts/fireworks/models/qwen3p8-27b)
with tokenizer Qwen/Qwen3.8-27B. Fireworks deprecated the Qwen 3.5 9B and Qwen 3.6 27B
Serverless Training pools; see the
Fireworks changelog.
If you change the model, also provide the matching tokenizer and renderer:
qwen3_8_disable_thinking from tinker-cookbook>=0.5.7, disables thinking mode.
--max-tokens still limits the entire generated response; a response that reaches the limit may
not include a final answer.
Define the task and reward
The environment is a normal HUD environment. The task yields a prompt, receives the model response, and returns anEvaluationResult:
env.py
multiply task can be run in an evaluation, used with another model, or sent to
a different training backend without rewriting the reward.
Calibrate the reward
Group-relative training compares repeated attempts at the same task. If every rollout in a group receives the same reward, the normalized advantages are zero and the group cannot update the model. Calibration mode samples from the initial adapter and reports the reward spread without taking an optimizer step:
Positive
within_group_reward_std confirms reward variation in at least one group. It does not
prove that the grader is correct. --debug-samples prints responses with their rewards and token
counts; verify that better answers receive higher rewards. If all groups are correct, increase the
operand range with --min-a, --max-a, --min-b, and --max-b. If all groups are incorrect,
reduce the range or increase --max-tokens.
The default uses four-digit operands and a 2,048-token budget. Three-digit tasks were nearly
saturated on Qwen 3.8 27B, while smaller budgets cut off many answers. Inspect the full
--debug-samples responses and token counts so truncation is not mistaken for arithmetic difficulty.
Verify the complete path
After calibration shows useful reward spread, run one bounded training step:--require-update makes the command fail
instead of silently skipping the optimizer. A successful run verifies authentication, sampling,
HUD grading, one policy-gradient update, checkpoint creation, and held-out evaluation.
Sampling or grading errors fail the command in every phase, including calibration and evaluation.
Run the training loop
The default command requests 30 steps × 8 task groups × 8 attempts, or 1,920 training rollouts, followed by 16 evaluation rollouts:train.py; this is not standalone code.
FireworksAgent adapts the sampler to HUD’s Agent interface. It records the prompt tokens,
generated tokens, and sampling logprobs on each Run. make_training_batch then:
- Groups runs that attempted the same task.
- Standardizes rewards within each group.
- Drops groups with no reward variation.
- Assigns the group advantage to each generated token.
- Builds the datums expected by the Fireworks training client.
runs/fireworks-serverless/metrics.jsonl. Each row includes mean reward,
within-group reward spread, valid rollout count, retained groups, training datums, whether an update
was applied, loss, snapshot, and step duration.
Compare before and after training
Add--eval-before to evaluate the initial and final adapters on the same held-out tasks,
at temperature 0 with the same generation budget:
--seed controls the task split, not model sampling randomness.
The script saves config.json, eval-before.json, and eval-after.json. Each evaluation contains
the checkpoint, prompts, responses, rewards, grader info, output-token counts, and an
at_token_limit flag. Match results by prompt and inspect both improvements and regressions.
Report format failures and responses at the token limit alongside accuracy; they can explain a
change in reward.
Use a separate output directory per experiment. Without --eval-before, only the final
evaluation runs.
One run with this configuration applied all eight updates and produced:
Accuracy increased by 20.3 percentage points: 27 tasks improved and one regressed. Twenty
improvements came from responses previously at the token limit. This measures exact-answer success
within the fixed budget, with shorter completed responses contributing to the gain. It is one run
on the multiplication task distribution; repeat across training runs to assess reproducibility.
Save and resume
Sampling and training checkpoints serve different purposes:
The script prints the full path as soon as each training checkpoint is saved. Resume from a
printed
state-NNNN or final-state path:
--base-model; the startup banner and config.json record the resolved model, tokenizer, and
renderer. Defaults apply to Qwen 3.8 27B. Other models require both --tokenizer-model and
--renderer before sampling or training, because checkpoint metadata does not specify them.
Start a new run to change base models; resuming does not convert an older adapter to Qwen 3.8.
The resumed session creates a new Fireworks run and retains the optimizer state. The script prints
the final Sampler checkpoint path. Sampler checkpoints are session-scoped, so use that path to
identify and
promote the checkpoint
before the session is removed.
Use another HUD environment
For a local one-turn environment, pass its tasks and environment files:--tasks-per-step tasks. A training run needs
--tasks-per-step + --eval-tasks; the script shuffles them deterministically and creates disjoint
training and evaluation subsets.
For a hosted taskset, first
deploy the environment, sync the
tasks, and set HUD_API_KEY. Then pass the taskset name or id:
FireworksAgent
continues to call the Fireworks sampler from the training process.
The included FireworksAgent and batch builder support one generated assistant response per Run.
Use a different agent adapter and batch builder for tool-using or multi-turn environments so every
trainable assistant turn is recorded and user and tool-result tokens remain masked.
See also
Source code
Complete runnable project and unit tests.
Training agents
How HUD turns tasksets, grouped rollouts, and rewards into a training loop.
Designing tasks for training
Build rewards with enough signal to distinguish better trajectories.
Fireworks Serverless Training
Fireworks setup, lifecycle, pricing, and supported models.