Journal7 entries

Writing on
machine
intelligence

Investigating the intersection of computational theory and real-world application. Technical notes on language, learning systems, agent architectures and the hardware they have to run on.

Most recentTokenization Decides What Your Model Can Count14 Jul 20268 min read

7 entries

  1. NLP8 min read

    Latest

    Tokenization Decides What Your Model Can Count

    The vocabulary you never think about sets a hard ceiling on arithmetic, multilingual cost, and how well Urdu is treated.

    Read entry

  2. Machine Learning7 min read

    Your Classifier Is Measuring Your Sampling, Not the World

    A ticket classifier at 94% offline and 71% live was not implemented wrong. The test set was drawn from the same snapshot as the training set.

    Read entry

  3. Agentic AI9 min read

    Agentic AI and the Compounding Cost of a Wrong Turn

    Per-step accuracy is a seductive metric. At 95% per step, a twenty-step task succeeds only about a third of the time.

    Read entry

  4. Agents8 min read

    Multi-Agent Systems Are a Communication Problem

    Adding agents does not add intelligence. It adds edges to a graph, and every edge is a place where context goes missing.

    Read entry

  5. LLM Developments7 min read

    Prompts Are Production Code and Should Be Versioned Like It

    A string that determines system behaviour, can be edited by anyone, and has no history is not configuration. It is an outage waiting for a calendar slot.

    Read entry

  6. Edge Computing8 min read

    Edge Inference Is a Memory Bandwidth Problem

    Teams optimise FLOPs and wonder why the model is still slow on device. For single-batch generation, arithmetic was never the bottleneck.

    Read entry

  7. Physical AI9 min read

    Physical AI Cannot Retry

    Software agents fail into a log file. Embodied ones fail into the world, where there is no exception handler and no undo.

    Read entry

Research

Notes from research
to production.

New entries cover NLP, machine learning, agentic systems, LLM developments, edge inference and physical AI — written from what actually held up in deployment.

Read the papers