Skip to content

How do neural networks acquire capabilities?

Toy Networks Lab runs controlled experiments on how neural networks
and language models acquire capabilities during training.

Capabilities have causes

Large language models develop capabilities whose origins remain poorly understood. This matters for AI reliability and safety. If we cannot explain where a capability comes from, we cannot reliably predict when it will appear, how it will behave out of distribution, or whether it might emerge from training data in unexpected ways.

  • Where capabilities come from

    We trace abilities such as in-context copying back to the properties of the training data that produce them.

  • When capabilities emerge

    We follow models across training checkpoints to see whether a capability forms gradually or all at once.

  • How to steer capabilities

    We change one property of the data at a time and measure what happens, so cause and effect are clear.

Open research

Our experiments are designed to be inspected, reproduced, and extended. We publish the code behind our research and make our code, models and datasets publicly available.

  • OPEN SOURCE

    Reproduce the experiments

    Our research code is public on GitHub, including the implementations and experimental tooling behind our work.

  • PUBLIC MODELS

    Explore our resources

    Models, datasets, and other research artifacts are available on Hugging Face for inspection.

Latest research

Experiments, results, and analysis.

View all research
  1. Part 1May 2026

    An introduction to our investigation into repetition capability in toy transformer models

    Why we want to study repetition in toy transformer models and what we aim to investigate

    toy-modelsinduction

  2. Part 2May 2026

    Repetition is surprisingly ubiquitous in tokenized natural language

    55% of tokens in the tokenized Pile dataset are part of repeated sequences, defined as either A or B in ...AB...AB, and we characterise the structure of those repetitions in detail.

    training-datainduction

  3. Part 3May 2026

    Is natural language special for learning repetition?

    We reverse all tokens in the Pile dataset and find that a transformer trained on completely unnatural data still learns to repeat sequences suggesting linguistic structure is not required for induction head formation.

    training-datainduction

  4. Part 4May 2026

    Token distribution drives repetition learning

    We surgically replace the tokens inside repeated sequences with random tokens, while keeping the sequence structure fixed to investigate the impact on repetition performance.

    training-datainduction

  5. Part 5May 2026

    How much data does a transformer need to learn repetition?

    We systematically degrade the repetition signal in the training data, token by token, and row by row, and find a critical threshold below which induction heads cease to form. Even 10% of tokens in repeated sequences is enough.

    training-datainduction

  6. Part 6May 2026

    Is induction a memorized or generalized capability?

    We probe whether the repetition capability of our toy transformer reflects genuine generalisation or memorisation of the training distribution. A single-token experiment reveals an apparent illusion of generalised induction, a cautionary finding for evaluations of larger LLMs.

    toy-modelsinduction