ORCNEIT Lab / № 01 / Model training
Training a language model from scratch: measurements and evaluation
- Experiment
- —
- Published
Abstract
A completed run of 80,000 steps and 327.68 million training tokens. We report learning curves and a separate capability evaluation, including the criteria that were not met.
Research question
How does prediction loss change when we train our own language model, and does that change translate into success on verifiable tasks?
Conditions and methods
The ORCNEIT model was trained from a fresh random initialization, without loading previously trained weights. The run began on 22 August and ended on 23 August 2026. All 80,000 steps completed, with no NaN/Inf numerical failures recorded.
327,680,000 counts training tokens presented to the model, including repeated passes. It is not the number of unique tokens in the source data.
The chart contains 17 telemetry observations: step 1,000, then every 5,000 steps. Training loss is the mean over the preceding 250 steps; development loss is measured on 8,192 tokens. Straight segments connect the observations; no additional smoothing is applied.
After training, a separate evaluation used tasks and scoring rules fixed in advance: 640 primary tasks, 64 context tasks, 640 generations and 256 held-out text windows. All 1,600 results were completed without modifying model weights during evaluation. Regressions were measured against a previously trained internal reference model on the same tasks: cases where the reference answered correctly and the evaluated model did not.
Measurements
Loss during training
LossTraining step
Data table
| Training step | Training | Development |
|---|---|---|
| 1,000 | 4.121886 | 4.778937 |
| 5,000 | 3.456732 | 4.340733 |
| 10,000 | 3.125353 | 3.67446 |
| 15,000 | 2.959091 | 3.385636 |
| 20,000 | 2.826652 | 3.283477 |
| 25,000 | 2.685255 | 3.178718 |
| 30,000 | 2.634553 | 3.151255 |
| 35,000 | 2.43181 | 3.006875 |
| 40,000 | 2.537828 | 3.055179 |
| 45,000 | 2.333696 | 3.045179 |
| 50,000 | 2.339327 | 2.961016 |
| 55,000 | 2.256835 | 2.975927 |
| 60,000 | 2.045726 | 2.912342 |
| 65,000 | 2.188968 | 2.966833 |
| 70,000 | 1.98551 | 2.98157 |
| 75,000 | 2.092879 | 2.896243 |
| 80,000 | 1.63502 | 2.912363 |
Result
Training completed, but the model did not pass the combined quality gate.
| Measure | Observation | Interpretation |
|---|---|---|
| Training steps | 80,000 / 80,000 | Completed |
| Training tokens, including repeated exposure | 327,680,000 | Not a count of unique tokens |
| Final development loss | 2.912363 | Separate from the held-out evaluation below |
| Held-out loss | 2.669892 | 262,144 tokens in a separate evaluation |
| Correct answers | 161 / 640 | Below the required 167 / 640 |
| Context tasks | 24 / 64 | Above the required 17 / 64 |
| Severely degenerate generations | 203 / 640 | Within the maximum of 271 |
| Paired task regressions | 178 | Above the maximum of 7 |
Discussion
Development loss fell from 4.778937 to 2.912363 between the reported checkpoints. This measures improved text prediction, not general understanding or dialogue readiness.
The separate primary evaluation recorded 161 correct answers out of 640. Its lower 95% bootstrap bound of 0.21875 did not exceed the 0.25 reference level. This criterion therefore does not establish performance above chance.
Some aggregate measurements improved, but paired evaluation revealed too many new errors. The gate remained failed: better loss and context performance cannot cancel those regressions.
Limitations
This is one internal experiment, not a comparison with commercial models or an independent assessment. Repeated runs to estimate variability are not reported here.
Development loss in the chart and held-out loss in the table use different evaluation samples. The two values must not be compared directly.
This report covers a completed August 2026 experiment, not current model availability. The published values alone do not allow independent reproduction of the complete run.
Data source: ORCNEIT training logs and final evaluations.