Language of Learning“Çünkü her şey dilde başlıyor.”“Because it all begins in language.”— İlkkan Yedinci

Continual learning from experience should be a capability of the model: learning signals should reach the model through language, the model should control learning through language, and what it learns should persist in its own memory, from context to weights.

Barış Deniz Sağlam · Duygu Ataman† · Altan Koçyiğit†

Graduate School of Informatics, Middle East Technical University · †Joint supervision

The problem

LLMs can solve hard problems.
Their ability to learn from experience has not kept pace.

Language models now achieve gold-medal scores at the International Mathematical Olympiad and help rewrite large production codebases. Yet even applying information from earlier interactions remains unreliable.

In DriftBench, models could accurately restate constraints they violated during multi-turn scientific ideation. The difficulty extends to memory systems: across 13 configurations, MemTrace found that, when systems failed, the evidence was retrievable about ten times as often as it was missing.

Making experience useful requires interpreting it and applying it when it matters. With frozen weights, past lessons affect later behavior through the context and memories supplied to the model. Storing a correction does not by itself turn it into a skill.

The broader challenge is continual learning: turning ordinary interaction into lasting improvements in what an agent knows and can do. A correction, a failed attempt, or a successful solution should help it handle future situations, including ones it has never encountered.

The historical pattern

AI history suggests a path: hand-designed components and procedures repeatedly become learned model capabilities.

Pretraining made representations reusable across tasks. Generated text replaced task-specific output heads, and prompting enabled adaptation without per-task fine-tuning. This recurring shift reflects Sutton’s Bitter Lesson: general methods that scale with computation can take over decisions once supplied by hand-designed procedures.

Reasoning provides the recent parallel. Strategies once orchestrated through external sampling and search can emerge within a model’s own generated trace. We argue that the decisions involved in learning from experience should undergo the same shift.

From hand-designed procedures to learned capabilities A timeline shows reusable representations from pretraining in 2018, generated text replacing task-specific heads in 2019, adaptation from prompts in 2020–21, and reasoning through generated tokens in 2024–25. A proposed fifth transition would turn learning decisions now made by training procedures into model-generated learning tokens. representations fromlanguage modeling outputs asgenerated text adaptation fromprompts reasoning bygenerating tokens learning by generating tokens 20182019 2020–212024–25 ? pretrainingtext-to-text pretraining / tuningreinforcement learning open ULMFiT, GPT-1GPT-2, T5 GPT-3, FLANo1, DeepSeek-R1 representationslearned per task outputs fromtask-specific heads adaptation byper-task fine-tuning reasoning throughexternal scaffolds learning bytraining procedures

The proposal

What if models had learning tokens, just as they have reasoning tokens?

Reasoning tokens let a model work through a problem. Learning tokens would let it work through what an experience should change: investigate a correction, gather evidence, test a revised understanding, and decide what should persist.

Language would carry both learning signals and the learner’s control over learning. A learning act is generated behavior directed at deciding or enacting how experience changes future behavior. Its tokens may interleave with reasoning, tool use, and world observations, as in the hypothetical episode.

world observation and agent response
world“Evaluate ∫ x ex dx.”
agent“Let u=ex. Then ∫ x ex dx=xex.”
world feedback
world“That answer is incorrect. Substitution is not the right method here.”
continuing agent–world interaction
agent“I see—the substitution leaves x unresolved. Which method should I use instead?”
world“Use integration by parts.”
agentLooks up the derivation from the product rule, then retries with u=x and dv=ex dx: ∫ x ex dx=xex−ex+C.
agentDifferentiates the result, then solves ∫ x cos(x) dx as a second test. Both checks succeed.
agent<learn> I chose a method from a familiar symbol instead of the expression’s structure. I should internalize two habits: consider integration by parts when one factor simplifies under differentiation; verify an antiderivative and include +C. ... </learn>

A plausible route

Plasticity supplies writable parameters. Meta-training would install the learning behavior.

In 1997, Schmidhuber, Zhao, and Wiering treated weight updates as ordinary policy actions, using integer instructions and a scalar reward signal; language played no role.

Much of the proposed behavior is already possible in context. External memory can preserve records for later retrieval. We argue that learning should also reach the weights, allowing experience to become an internalized skill.

Under this proposal, generated language would direct parameter updates during inference. Meta-training would reward the learner for transferring lessons to later, held-out situations after the original interaction has left the context. Installing this capability remains the central research challenge.

Learning from self-reflectionAgents can already reflect on failures in language and retain those reflections in external memory.Reflexion
Test-time plasticityModels can now update writable parametric state during inference, and some systems retain changes across contexts.SRWM · Titans · MesaNet · TTT-E2E · In-Place TTT · Nested Learning
Meta-trainingMeta-training can teach self-referential networks to learn new tasks while retaining earlier skills. Separately, models can be trained to learn from language feedback, with transfer to new domains and interaction settings.ACL · RL²F · SML

The prediction

Learning tokens could create a new scaling axis: just as reasoning tokens let a model spend more computation on a difficult problem, learning tokens would let it devote more computation to learning.

pretraining
Pretraining scaling Schematic: loss decreases as the number of tokens in the training corpus grows. loss tokensin the corpus
established
reasoning
Reasoning scaling Schematic: task performance improves with more tokens generated while thinking. task performance tokensgenerated while thinking
established
learning
Learning scaling prediction Predicted, schematic relationship: later transfer improves with more tokens generated while learning. later transfer tokensgenerated while learning
predicted

Requirements

Our position can be stated as five requirements:

  • Learning as capability — at test time, learning is performed by the agent itself rather than by a separate training procedure applied to it.
  • Parametric and non-parametric memory — the agent combines in-context learning with in-weight learning; external memory may extend its non-parametric memory.
  • Learning signals in language — information available for learning, including instructions, demonstrations, feedback, and binary or scalar outcomes, reaches the agent as content through its language interface.
  • Control in language — the agent uses generated language to determine what is retained and, when the substrate exposes these choices, its scope and placement in memory. The architecture may prescribe how that decision is executed.
  • Lifelong — the effects of learning persist across a continuing test-time stream without prescribed reset boundaries.

Opportunities

Learning throughout deployment could change not only what an agent knows, but how it develops over its lifetime.

Compounding learningAn agent may improve not only what it can do, but how efficiently it acquires new capabilities—requiring fewer interactions, less feedback, or less computation as its lifetime progresses.
Personalization, specialization, and expertiseAgents initialized from the same checkpoint could acquire distinct knowledge, skills, judgment, and strategies through work with people, organizations, or domains.
Open-ended discoveryAgents shaped by different experience streams could exchange discoveries and learning strategies, turning divergent expertise into stepping stones for further learning.

Cite

@misc{saglam2026languageoflearning,
  title  = {Language of Learning},
  author = {Sa{\u g}lam, Bar{\i}{\c s} Deniz and Ataman, Duygu
            and Ko{\c c}yi{\u g}it, Altan},
  year   = {2026},
  note   = {SSRN preprint},
  doi    = {10.2139/ssrn.7552418},
  url    = {https://doi.org/10.2139/ssrn.7552418}
}