Language of Learning“Çünkü her şey dilde başlıyor.”“Because it all begins in language.”— İlkkan Yedinci
Continual learning from experience should be a capability of the model: learning signals should reach the model through language, the model should control learning through language, and what it learns should persist in its own memory, from context to weights.
Graduate School of Informatics, Middle East Technical University · †Joint supervision
The problem
LLMs can solve hard problems.
Their ability to learn from experience has not kept pace.
Language models now achieve gold-medal scores at the International Mathematical Olympiad and help rewrite large production codebases. Yet even applying information from earlier interactions remains unreliable.
In DriftBench, models could accurately restate constraints they violated during multi-turn scientific ideation. The difficulty extends to memory systems: across 13 configurations, MemTrace found that, when systems failed, the evidence was retrievable about ten times as often as it was missing.
Making experience useful requires interpreting it and applying it when it matters. With frozen weights, past lessons affect later behavior through the context and memories supplied to the model. Storing a correction does not by itself turn it into a skill.
The broader challenge is continual learning: turning ordinary interaction into lasting improvements in what an agent knows and can do. A correction, a failed attempt, or a successful solution should help it handle future situations, including ones it has never encountered.
The historical pattern
AI history suggests a path: hand-designed components and procedures repeatedly become learned model capabilities.
Pretraining made representations reusable across tasks. Generated text replaced task-specific output heads, and prompting enabled adaptation without per-task fine-tuning. This recurring shift reflects Sutton’s Bitter Lesson: general methods that scale with computation can take over decisions once supplied by hand-designed procedures.
Reasoning provides the recent parallel. Strategies once orchestrated through external sampling and search can emerge within a model’s own generated trace. We argue that the decisions involved in learning from experience should undergo the same shift.
The proposal
What if models had learning tokens, just as they have reasoning tokens?
Reasoning tokens let a model work through a problem. Learning tokens would let it work through what an experience should change: investigate a correction, gather evidence, test a revised understanding, and decide what should persist.
Language would carry both learning signals and the learner’s control over learning. A learning act is generated behavior directed at deciding or enacting how experience changes future behavior. Its tokens may interleave with reasoning, tool use, and world observations, as in the hypothetical episode.
<learn> I chose a method from a familiar
symbol instead of the expression’s structure. I should internalize two habits: consider
integration by parts when one factor simplifies under differentiation; verify an
antiderivative and include +C. ... </learn>A plausible route
Plasticity supplies writable parameters. Meta-training would install the learning behavior.
In 1997, Schmidhuber, Zhao, and Wiering treated weight updates as ordinary policy actions, using integer instructions and a scalar reward signal; language played no role.
Much of the proposed behavior is already possible in context. External memory can preserve records for later retrieval. We argue that learning should also reach the weights, allowing experience to become an internalized skill.
Under this proposal, generated language would direct parameter updates during inference. Meta-training would reward the learner for transferring lessons to later, held-out situations after the original interaction has left the context. Installing this capability remains the central research challenge.
The prediction
Learning tokens could create a new scaling axis: just as reasoning tokens let a model spend more computation on a difficult problem, learning tokens would let it devote more computation to learning.
Requirements
Our position can be stated as five requirements:
- Learning as capability — at test time, learning is performed by the agent itself rather than by a separate training procedure applied to it.
- Parametric and non-parametric memory — the agent combines in-context learning with in-weight learning; external memory may extend its non-parametric memory.
- Learning signals in language — information available for learning, including instructions, demonstrations, feedback, and binary or scalar outcomes, reaches the agent as content through its language interface.
- Control in language — the agent uses generated language to determine what is retained and, when the substrate exposes these choices, its scope and placement in memory. The architecture may prescribe how that decision is executed.
- Lifelong — the effects of learning persist across a continuing test-time stream without prescribed reset boundaries.
Opportunities
Learning throughout deployment could change not only what an agent knows, but how it develops over its lifetime.
Cite
@misc{saglam2026languageoflearning,
title = {Language of Learning},
author = {Sa{\u g}lam, Bar{\i}{\c s} Deniz and Ataman, Duygu
and Ko{\c c}yi{\u g}it, Altan},
year = {2026},
note = {SSRN preprint},
doi = {10.2139/ssrn.7552418},
url = {https://doi.org/10.2139/ssrn.7552418}
}