A small proof-of-concept language model (not an LLM) incorporating latent-space prediction, internal state using recurrent trace units, and byte-by-byte output, built with MLX.
Find a file
2026-08-12 16:11:30 +08:00
.gitignore test model thing, finally 2026-08-12 16:11:30 +08:00
benchmark.py test model thing, finally 2026-08-12 16:11:30 +08:00
LICENSE.md test model thing, finally 2026-08-12 16:11:30 +08:00
main.py test model thing, finally 2026-08-12 16:11:30 +08:00
README.md test model thing, finally 2026-08-12 16:11:30 +08:00

Hey! Thanks for being here.

Here's the video, if you came here from somewhere else -> Video

The semi-trained 4.5m model (it's not done, but also it's in millions of parameters, not billions) is available for you to use. So are the train and benchmark scripts (using MLX, but you can port to other platforms if you want).

However a larger 130m model is not provided, it is too large for GitHub to store. You can train it yourself and perhaps share it on a cloud storage provider instead.

Note that datasets are not included, and if you use a non-puretext dataset like wikipedia dump then feel free to write your own dataset extraction code or use an existing library.

The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore!