Skip to content
News
Research
← All research

ORCNEIT Lab / № 04 / Results overview · In progress

ORCNEITGPT: how the project began and where it stands

Observation period
—
Published

Abstract

From the first experiments with our own language model to a revised training approach and a new corpus. What we tested, which approaches fell short, and why the direction changed.

Research question

How did the approach to ORCNEITGPT change after working prototypes, broader evaluations and unsuccessful attempts to improve the model?

Conditions and methods

This account covers documented work from 13 June through 30 September 2026. Dates refer to experiments and decisions, not article publication. Early project work preceded company registration. This is a results overview, not a ledger of every development change.

We distinguish the language model, external response-control rules and the measurement instrument. Scores from different models, tasks and decoding modes cannot be combined into a single progress curve.

The values below were checked against preserved experimental results. Separate papers provide methods, tables and numerical supplements for individual questions. This overview connects those findings and decisions without replacing their limitations.

New-corpus status is based on the latest completed September records. Unfinished processing is not an accepted dataset. A target volume is not represented as ready data or completed training.

Measurements

Result

The project moved from a small prototype to its own training and stricter evaluation. A ready conversational model has not yet been obtained.

ORCNEITGPT: how the project began and where it stands — Result
MeasureObservationInterpretation
Project work began13 June 2026Own text processing, training and generation
Early expanded evaluation29 / 79 → 72 / 79Model → rule-controlled system; not improved weights
Final memory-learning evaluation0 / 246 tasksAfter repairing multiline answer extraction
Bounded rounding test7.76 × 10⁻⁸ → 8.58 × 10⁻¹⁰Absolute mean signed update-writeback error
Completed August base training80,000 steps; 327,680,000 token exposuresQuality gate failed
Dialogue fine-tuning0 → 43 correct responses out of 256Improvement, but acceptance criteria unmet
New large corpusPreparation continues4 billion tokens is a target, not accepted volume

Discussion

June: the project began with a small self-written core. Text representation, next-token training, generation and an interaction interface were implemented. Initial tests established that the pipeline worked, not broad knowledge or reliable performance on unfamiliar instructions.

Broader evaluation changed the picture: the early model passed 29 of 79 tasks; the rule-controlled system passed 72. In some scenarios rules replaced the model call with a prepared response. A good interface response therefore cannot automatically be attributed to the neural model. Later, external safeguards were measured separately while preserving the raw model's score.

Memory was the next question. A separate store could confirm, delete and supply records, but that did not establish model-level selection and mutation-protocol skills. Three fine-tuning attempts exposed forgetting and the partial benefit of replay. The final correctly measured test passed none of 246 memory tasks. This required reassessment, not a ready-memory claim.

In late June, a self-written billion-scale architecture was tested, followed by larger training experiments. Scale alone did not resolve the problems. July work examined small updates lost to rounding and compared tokenization through proxy training on identical source text, rather than compression alone.

August: continued training of an earlier model barely changed aggregate generation results: 497 versus 499 severely degenerate responses out of 1,600. Short controlled interventions did not produce an accepted repair. We chose to re-examine data and text processing and train the next model from fresh random initialization.

Base training completed on 22–23 August after 80,000 steps. Separate evaluation yielded 161 correct responses out of 640 against a required minimum of 167, while exceeding the allowed regression count. Dialogue fine-tuning improved termination and reduced repetition but solved only 43 of 256 responses. Both are published as completed experiments, not a ready-assistant release.

September: attention shifted to more varied data. Sources are checked for content and provenance, duplicates and overlap with evaluation sets. Final acceptance of the large corpus was not complete at September's end. Subsequent training must use accepted data and predefined quality criteria; this overview promises no model-release date.

Limitations

This is a retrospective of our own experiments, not independent certification or comparison with leading commercial models. Successful training execution and lower loss do not establish product readiness.

Early, later and dialogue evaluations use different samples and definitions. They are not the same test. Their corresponding studies provide details and caveats.

The overview does not claim that every method is a new scientific discovery. It reports measured project findings, including negative results, and reasons for decisions. Work after 30 September is outside this status snapshot.

Data source: ORCNEIT experimental measurements and evaluations.