Skip to content
News
Research
← All research

ORCNEIT Lab / № 01 / Model training

Training a language model from scratch: measurements and evaluation

Experiment
—
Published

Abstract

A completed run of 80,000 steps and 327.68 million training tokens. We report learning curves and a separate capability evaluation, including the criteria that were not met.

Research question

How does prediction loss change when we train our own language model, and does that change translate into success on verifiable tasks?

Conditions and methods

The ORCNEIT model was trained from a fresh random initialization, without loading previously trained weights. The run began on 22 August and ended on 23 August 2026. All 80,000 steps completed, with no NaN/Inf numerical failures recorded.

327,680,000 counts training tokens presented to the model, including repeated passes. It is not the number of unique tokens in the source data.

The chart contains 17 telemetry observations: step 1,000, then every 5,000 steps. Training loss is the mean over the preceding 250 steps; development loss is measured on 8,192 tokens. Straight segments connect the observations; no additional smoothing is applied.

After training, a separate evaluation used tasks and scoring rules fixed in advance: 640 primary tasks, 64 context tasks, 640 generations and 256 held-out text windows. All 1,600 results were completed without modifying model weights during evaluation. Regressions were measured against a previously trained internal reference model on the same tasks: cases where the reference answered correctly and the evaluated model did not.

Measurements

Loss during training

Loss
040 00080 000

Training step

Lower is better. 17 observations; training loss is a 250-step mean, development loss uses a held-out development sample. Lines connect measured points, not continuous observations.
Data table
Loss during training
Training stepTrainingDevelopment
1,0004.1218864.778937
5,0003.4567324.340733
10,0003.1253533.67446
15,0002.9590913.385636
20,0002.8266523.283477
25,0002.6852553.178718
30,0002.6345533.151255
35,0002.431813.006875
40,0002.5378283.055179
45,0002.3336963.045179
50,0002.3393272.961016
55,0002.2568352.975927
60,0002.0457262.912342
65,0002.1889682.966833
70,0001.985512.98157
75,0002.0928792.896243
80,0001.635022.912363
Chart data · CSV ↓

Result

Training completed, but the model did not pass the combined quality gate.

Training a language model from scratch: measurements and evaluation — Result
MeasureObservationInterpretation
Training steps80,000 / 80,000Completed
Training tokens, including repeated exposure327,680,000Not a count of unique tokens
Final development loss2.912363Separate from the held-out evaluation below
Held-out loss2.669892262,144 tokens in a separate evaluation
Correct answers161 / 640Below the required 167 / 640
Context tasks24 / 64Above the required 17 / 64
Severely degenerate generations203 / 640Within the maximum of 271
Paired task regressions178Above the maximum of 7

Discussion

Development loss fell from 4.778937 to 2.912363 between the reported checkpoints. This measures improved text prediction, not general understanding or dialogue readiness.

The separate primary evaluation recorded 161 correct answers out of 640. Its lower 95% bootstrap bound of 0.21875 did not exceed the 0.25 reference level. This criterion therefore does not establish performance above chance.

Some aggregate measurements improved, but paired evaluation revealed too many new errors. The gate remained failed: better loss and context performance cannot cancel those regressions.

Limitations

This is one internal experiment, not a comparison with commercial models or an independent assessment. Repeated runs to estimate variability are not reported here.

Development loss in the chart and held-out loss in the table use different evaluation samples. The two values must not be compared directly.

This report covers a completed August 2026 experiment, not current model availability. The published values alone do not allow independent reproduction of the complete run.

Data source: ORCNEIT training logs and final evaluations.