Skip to content
News
Research

ORCNEIT Lab / Research & engineering papers

Research

How we train and evaluate our own model. Experiments, measurements and findings, with test conditions and explicit limitations.

Laboratory publications

№ 01Model training

Training a language model from scratch: measurements and evaluation

A completed run of 80,000 steps and 327.68 million training tokens. We report learning curves and a separate capability evaluation, including the criteria that were not met.

Experiment
—
Published
Read the study

Loss during training

Loss
040 00080 000

Training step

Lower is better. 17 observations; training loss is a 250-step mean, development loss uses a held-out development sample. Lines connect measured points, not continuous observations.

№ 02Fine-tuning experiment

Dialogue fine-tuning: fewer repetitions, but not enough correct answers

We compared the base model and its fine-tuned version on the same tasks. Answer termination improved and severe degeneration decreased, but assistant acceptance criteria were not met.

Experiment
Published
Read the study

Before and after fine-tuning

Responses out of 256

Correct responses ↑

Base model: 0 / 256
After fine-tuning: 43 / 256

Severe degeneration ↓

Base model: 238 / 256
After fine-tuning: 36 / 256

Normal termination ↑

Base model: 18 / 256
After fine-tuning: 239 / 256
Each model: 128 tasks × 2 modes = 256 responses. The axis starts at zero. Correctness, degeneration and termination are distinct measures, not parts of one total.

№ 03Engineering study · In progress

Why we are rebuilding the training corpus

Failed model evaluations prompted us to revisit the data, tokenization and evaluation itself. We report measured changes and the preparation of a new corpus, still in progress as of 26 September 2026.

Observation period
—
Published
Read the study

Vocabulary selection: text compactness

Tokens per UTF-8 byte

Small vocabulary

0.365322

Medium · selected

0.354379

Large vocabulary

0.348899
Lower is more compact. Equal-weight average across four text groups on the same sample, measured on 20 August 2026. This is not a measure of model response quality.

Experiment dates are checked against the working log. Publication dates record when papers appeared on this site. These studies are not a product release announcement or an independent model certification.