Company · ORCNEIT
ORCNEIT is on GitHub: research, data and project development
We have opened the company's GitHub and published reports on ORCNEITGPT. Here is what is available, why numerical appendices matter and how we intend to develop public research reporting.
Published:
The company's public workspace
ORCNEIT now has an official GitHub organization, ORCNEIT-LLC. Its first public repository, research, presents ORCNEITGPT findings. It provides detailed materials readers can revisit, compare numerically and track as publications are revised.
LAB keeps research available as formatted articles. GitHub offers the same materials in Russian and English Markdown, alongside figures and CSV files. GitHub complements the website: readers choose a format while the substance and limitations remain the same.
What is available now
The research repository currently has eight publications: a project overview and seven research and engineering reports. The three existing papers on base training, dialogue fine-tuning and rebuilding the training foundation remain available. We added an overview and four studies.
The overview covers verified work from June through late September 2026: a small prototype, expanded evaluation, memory experiments, training from scratch and preparation of a new training foundation. It is an account of decisions and measured findings, not a ready-assistant announcement.
Four new research questions
Memory: how does learning a new task affect existing skills? General instruction success on a small test fell from 23 to 5 out of 25, then returned to 23 in the next attempt with replay. Yet the final memory-protocol evaluation passed none of 246 cases. The paper explains an early measurement defect and why its zero scores are not clean model evidence.
Numerical precision: how do we write updates below representable weight spacing? A bounded controlled comparison reduced absolute mean signed writeback error by about 90-fold through stochastic rounding. That is a local numerical effect, not 90-fold better responses. A separate high-precision-state run had no NaN/Inf failures but failed its behavioral gate.
Tokenization: is text compression enough for selection? With identical source documents, the preliminary compression leader lost on proxy training: 2.374638 versus 2.288238 bits per byte for the selected alternative. The paper reports text exposure, update counts and the single-seed limitation; equal text is not described as equal compute.
Response control: what do safeguards change without updating weights? An August comparison reduced severe failures from 499 to 433 out of 1,600 generations. Paired results included 141 improvements and 75 regressions. Eliminating invalid UTF-8 does not establish semantic correctness. The early system, whose rules sometimes bypassed the model entirely, is examined separately.
Reports with checkable numbers
Every paper includes its question, methods, result, discussion and limitations. Experiment and publication dates are distinct. Figures have tables and CSV downloads; appendices explain units, denominators and which measurements cannot be combined.
For example, 1,600 generations are multiple-mode outputs from 400 prompts, not 1,600 independent tasks. The 327.68 million training tokens count exposures with repeats, not unique text. A four-billion-token corpus target does not mean an accepted dataset already exists. These distinctions stay next to the findings.
Numerical appendices let readers check arithmetic and agreement between figures and published values. They do not replace independent training replication: aggregate reporting alone is not a complete experimental reproduction package.
Where we intend to go next
We intend to develop GitHub into a continuing library of completed studies: new comparisons, evaluation conditions, numerical appendices and explanations of decisions. When results need correction, we want an understandable revision history rather than silent changes to earlier conclusions.
Another direction is material on data preparation, quality checks and response evaluation. If standalone tools become ready for public use, their scope, terms and licenses will be announced separately. This reporting repository does not imply publication of the entire source code, weights or training data.
This is a direction, not a release calendar. Further materials will follow as findings become ready. ORCNEITGPT remains a research project: evaluation and conclusions come before publication, including approaches that did not work.
An overview, seven reports, figures and numerical appendices in Russian and English.
Explore the research on GitHub