Skip to content
News
Research
← All research

ORCNEIT Lab / № 05 / Fine-tuning study

Learning memory: retaining old skills does not establish new ones

Experiment
Published

Abstract

Three small-model experiments exposed forgetting, the benefit of replay, and the absence of a reliable memory protocol despite lower training loss.

Research question

Can a small language model learn memory operations while preserving ordinary instruction following?

Conditions and methods

We studied our own roughly 4-million-parameter model with 400 records: 320 training, 40 development and 40 held-out. Each of three attempts began from the same base model. This was a diagnostic sequence, not a randomized single-factor study.

The first attempt trained memory operations. The second added replay of prior training material and aligned training-response formatting with evaluation. The third strengthened memory training while retaining replay. Several conditions changed between attempts, so gains cannot be attributed to one setting alone.

General skills were evaluated with separate unchanged tests. The chart shows correct responses on 25 instructions, not memory tasks or knowledge on arbitrary requests.

An early measurement defect retained only the first response line although several protocols required multiple lines. Early zero scores on multiline operations therefore are not clean evidence of missing skills. Extraction was repaired and multiline preservation tested before the final attempt. Final results are reported separately.

Final memory evaluation contained 54 intent, 40 record-normalization, 48 safety, 56 change-proposal/confirmation and 48 abstention cases: 246 total. The 56 change cases contained 72 evaluated turns; turns are not added to case counts. Success required valid protocol output and task-specific correctness.

Measurements

Instruction retention

Correct responses out of 25

Before fine-tuning

23

Memory only

5

Memory with replay

23

Stronger memory with replay

20
The same general instruction test, not a memory score. Four model states; the axis starts at zero.
Data table
Instruction retention
MeasureCorrect responses out of 25
Before fine-tuning23
Memory only5
Memory with replay23
Stronger memory with replay20

Result

Replay helped retain prior skills, but the final model passed none of 246 memory cases.

Learning memory: retaining old skills does not establish new ones — Result
MeasureObservationInterpretation
General instructions: base model23 / 2592%
General instructions: memory only5 / 2520%; substantial forgetting
General instructions: memory with replay23 / 2592%; other conditions also changed
General instructions: stronger training20 / 2580%; some regression remained
Final development memory loss5.7350 → 2.5970Five epochs; not task success
Separate held-out memory loss2.6042Distinct from general instruction evaluation
Successful memory cases0 / 246Five categories; repaired response extraction
Intent recognition1 / 54 protocol-parsed; 0 successfulParseable format does not establish correctness

Discussion

Training only the new task reduced general instruction success from 92% to 20%. The second attempt, including replay, restored this score to 92%. This is an observed difference on a fixed small test, not a universal retention guarantee.

Final development memory loss decreased through 5.7350, 4.0609, 3.3282, 2.8774 and 2.5970. Yet the model still failed to emit the required protocol. Even the single parsed intent output did not pass its task. Lower loss and successful memory operations are different outcomes.

The failure does not establish that all small models cannot learn memory. One architecture, limited data and three regimes were tested. Stopping this direction under the existing conditions was a decision about these experiments.

Limitations

Early multiline scores are limited by extraction defects and are not pooled with the corrected final evaluation into a memory curve. Repairing a measurement instrument is not model improvement.

General tests are small, repeated runs for variability are absent, and several factors changed across attempts. The comparison does not isolate the independent causal effect of replay alone.

This evaluates model-generated memory protocols, not an external store, deployed-system security or finished-product quality. Independent experimental replication is not provided.

Data source: ORCNEIT experimental measurements and evaluations.