Longitude / Labs

Artificial intelligence research · Byte-native language models · United Kingdom

Every frontier language model begins by throwing information away. We removed the part that does it.

Longitude Labs is a British artificial intelligence research company. We build byte-native language models that already outperform systems thousands of times their size, and that can run entirely inside the infrastructure of the organisation using them. Smaller and more precise is not a trade-off. It is the thesis.

The part in question is the tokeniser. Before most language models see text, a fixed vocabulary breaks it into fragments and decides what the model is allowed to see. The text can be reconstructed, but direct access to its characters, identifiers and scripts is gone. Kolmo removes that layer. It reads UTF-8 bytes directly. The difference is easier to see than to describe.

Where the information goes

Illustrative subword view
Byte-native

Model Kolmo v0.1
Scale 211M
Stage First 211M model
Status Scaling in progress

Purpose

The aim is not a smaller model that is good enough. It is a smaller model that is better.

The field's working assumption is that capability is bought with scale. We think a substantial part of what scale currently buys is compensation — for an architecture that commits to a fixed segmentation before model learning begins.

Our hypothesis is that removing that fixed segmentation moves the size required to be genuinely capable. Our first model, at 211 million parameters, already outperforms a 405 billion parameter system on four measured character-manipulation tasks. That is one axis on an early run. The ambition is to widen it until it is not one axis, and to keep the model small the whole way.

A model like that is not merely cheaper to run. It can go where the alternatives cannot — and almost nobody today has real control over the intelligence they depend on. Frontier systems run on someone else's hardware, in someone else's country, under terms someone else sets, and access can be withdrawn. For a bank, a hospital or a government department that is a dependency rather than infrastructure; hosted systems can be unsuitable or unavailable for the most sensitive workloads.

So the target is a model capable enough to be worth running, small enough to run inside the perimeter of the organisation that relies on it, honest enough to say when it does not know, and fully under the control of whoever runs it. We do not think those four things are in tension. We think the field has assumed they are.

The thesis

Compression is intelligence. We took it literally.

Solomonoff, Kolmogorov and Chaitin established the formal connection between prediction, description length and compression: the shortest description of the data is the best model of it. To predict is to compress. To compress is to understand.

Modern language-model training already minimises predictive loss, but model size and input representation are usually treated as separate engineering choices. Most frontier systems still begin with a separately trained, frozen subword vocabulary. The encoding is reversible, but it imposes boundaries and sequence-length trade-offs before a single parameter of the language model has learned anything.

Tokenisation is not ignored by the field, but it is normally fixed before the language model learns. Our question is what happens when the model retains byte-level access from the start.

We built a model without one, and measured what happened.

Each cell is one byte of UTF-8. Familiar Latin characters usually need one byte; accents, emoji and many other scripts need several. Brass cells are the bytes working together to encode a single character — structure Kolmo receives directly rather than through a fixed vocabulary. Try an emoji, an accent, or a non-Latin script.

Early results

First 211M model. A fraction of a competitive data budget.

Everything below is Kolmo v0.1 — an early, undertrained model of a new architecture at small scale. We are publishing it now because the direction is already clear, and because character-level manipulation is where the thesis makes its sharpest prediction: a model that reads bytes can address character structure directly, while a subword model must reconstruct it indirectly.

Full official CUTE test sets. Trained from scratch, no distillation, no foreign weights.

CUTE character manipulation · full test sets · accuracy
TaskKolmo v0.1 — 211MBest of 16 others, incl. Llama 3.1 405B
Substitute character91.5%63.8%
Delete character87.4%83.1%
Swap characters68.8%16.4%
Insert character33.6%19.1%

Kolmo is roughly two thousand times smaller than the largest model in that comparison. The gap is not a matter of scale — it is upstream of it.

Description length
−48%Against a parameter-matched tokenised twin on identical data: 1.90 vs 3.68 bits per byte. Held-out bytes received roughly half the predictive description length.
Training throughput
+32%Training bytes per second over the same tokenised twin, at matched parameters.
SciQ
80.7%Full official validation set.
Reproducibility
Byte-identicalSame input, same bytes out, every run.

Sovereignty

Nobody else holds the original.

Trained from zero, not adapted.

Kolmo inherits no weights from any other laboratory. No distillation, no teacher model, no fine-tuned foreign checkpoint underneath. A model adapted from someone else's open weights carries whatever was, and was not, built into them.

Foreign checkpoint→ continued pre-training→ “sovereign” model Inherits whatever a foreign laboratory built into those weights. Someone else holds the original.
Zero→ our corpus→ Kolmo Nobody else holds the original.

Documented provenance, source by source.

The corpus is assembled from openly licensed material with licence and origin recorded per source. Any source can be removed and the model retrained. That is what makes decontamination and causal ablation possible at all.

Small enough to run where the cloud cannot go.

For the regulated core of government, finance and health, a hosted foreign model may be unsuitable or unavailable. Models at this scale run inside the perimeter, on owned hardware, with nothing leaving the building.

Script-agnostic at the input layer.

A tokeniser trained predominantly on English can impose English-derived vocabulary boundaries on every language it is later asked to handle. Bytes carry no learned vocabulary, although the training data still determines learned capabilities. In tests, despite roughly one part in a thousand of non-English training text, Kolmo produced valid structured output containing Arabic, Cyrillic and CJK.

Models

Named for the people who worked out that description length is the thing that matters.

KolmoModel
The byte-native model family. Reads and writes raw bytes with no tokeniser. Currently 211M parameters; scaling in progress. After Andrey Kolmogorov.
SolomonSystem
The layer around the model: task routing, constrained decoding that guarantees output conforms to a given schema byte by byte, grounded retrieval, and abstention when confidence is low. After Ray Solomonoff.

Ambition

The UK does not need to win the race to the largest model. It needs to own one.

The most valuable position is not the biggest system. It is the one a country can run itself, on its own hardware, under its own control — and that is a position the UK can hold without matching anyone's capital expenditure, because it is won on precision rather than on scale.

An architecture with no fixed vocabulary carries no single language's segmentation assumptions at its input layer, although the training corpus still determines what the model learns. The same input approach can represent any script without rebuilding a vocabulary around it. Most countries that want AI on their own terms have no realistic route to one today, and what tends to be offered instead is adapted from weights another laboratory trained somewhere else. We are building the alternative — for the UK first, and then for whoever else needs one.

In 1714 Parliament offered a fortune for a solution to longitude. Every serious authority knew the answer lay in astronomy. John Harrison was a self-taught Yorkshire carpenter who saw a timekeeping problem instead, and won with a pocket watch. He got smaller, and more precise, and he was right.

Harrison's first attempt weighed thirty-four kilograms. The one that worked was thirteen centimetres across. He was opposed at every stage by an establishment that was not wrong about the sky — it simply could not see the other answer from where it was standing.

Contact

We are a small research company and we are early.

If you work on byte-level models, compression, or evaluation and want to compare notes — or you are responsible for a system where a model has to run inside your own infrastructure and prove what it did — we would like to hear from you.

hello@longitudelabs.co.uk