A completed run of 80,000 steps and 327.68 million training tokens. We report learning curves and a separate capability evaluation, including the criteria that were not met.
Lower is better. 17 observations; training loss is a 250-step mean, development loss uses a held-out development sample. Lines connect measured points, not continuous observations.
We compared the base model and its fine-tuned version on the same tasks. Answer termination improved and severe degeneration decreased, but assistant acceptance criteria were not met.
Each model: 128 tasks × 2 modes = 256 responses. The axis starts at zero. Correctness, degeneration and termination are distinct measures, not parts of one total.
Failed model evaluations prompted us to revisit the data, tokenization and evaluation itself. We report measured changes and the preparation of a new corpus, still in progress as of 26 September 2026.
Lower is more compact. Equal-weight average across four text groups on the same sample, measured on 20 August 2026. This is not a measure of model response quality.
Experiment dates are checked against the working log. Publication dates record when papers appeared on this site. These studies are not a product release announcement or an independent model certification.