Weights in {−1, 0, +1}

Small models.Any CPU.Offline.

QLNI builds SHADOW: open language models trained in ternary from the first step, with a frozen binary vocabulary. A capable assistant in tens of megabytes, on the computer you already own.

Our aim

Useful AI that people can own

Most AI lives in data centres and is rented by the token. SHADOW goes the other way: models small enough to download in seconds, fast on an ordinary laptop processor, and open so anyone can adapt them to their own work.

Every word SHADOW knows is a fixed 512-bit code instead of a trained table of floats. The vocabulary that fills most of a small model becomes a few megabytes of bits.

Small

Tens of megabytes, not gigabytes.

CPU-first

Ternary weights, hand-written C kernel.

Private

Fully offline; your files stay on your disk.

Open

Weights, code and benchmarks released.

Released · built from scratch

Two models.
Zero cloud.

2026 · MIT licence

SHADOW‑250M Instruct

60 MBwhole deployment
~400tokens/s on CPU
100Mtoken offline archive

Trained on 30 billion tokens. A frozen 512-bit vocabulary of 131,072 tokens and a body stored below 2 bits per weight; reads an archive on disk far beyond its window.

2026 · proof of concept

SHADOW‑50M Instruct

19.8 MBternary model
1,400tokens/s on CPU
197 KBexecutable

44 million ternary parameters with exact circuits inside the model. Its frozen table learned 8,600 new words without retraining.

Traction
355k views

The SHADOW-250M launch, posted by the founder.

r/MachineLearning365 upvotes · 61 comments
136k
r/LocalLLaMA295 upvotes · 60 comments
107k
Third launch postviews
112k

139 GitHub stars · 360+ Hugging Face downloads · outside developers already benchmarking against SHADOW.

What’s next

Is the embedding table overpriced?

In a small model the vocabulary is the largest single part: 63% of Gemma 3 270M. Four models with the same body and the same 2 billion tokens differ only in their vocabulary.

A
Learned embedding table controlfloat body · the standard baseline
268Mtrainable
B
Frozen 512-bit semantic codesfloat body · what does freezing cost?
101Mtrainable
C
Frozen semantic codes, ternary bodyBitNet b1.58 · does it hold at 1.58 bits?
101Mtrainable
D
Frozen random codesfloat body · does meaning in the codes matter?
101Mtrainable
Data 2.05B tokens eachTable 16 MB for 262k tokensMeasured ~62 GPU-hours

Then: a SHADOW that sees and hears as well as reads, released as one offline package.

Contact

Build it with us

Grants, compute partners, research collaborators and early users are all welcome. QLNI is founded by Sai Kiran Bathula in New South Wales, Australia: independent and self-funded.

saikiranbathula@qlni.io