A small proof-of-concept language model (not an LLM) incorporating latent-space prediction, internal state using recurrent trace units, and byte-by-byte output, built with MLX.
Find a file
2026-08-12 16:13:45 +08:00
.gitignore test model thing, finally 2026-08-12 16:11:30 +08:00
benchmark.py test model thing, finally 2026-08-12 16:11:30 +08:00
LICENSE.md test model thing, finally 2026-08-12 16:11:30 +08:00
main.py test model thing, finally 2026-08-12 16:11:30 +08:00
README.md fix readme!! 2026-08-12 16:13:45 +08:00

Hey! Thanks for being here.

Here's the video, if you came here from somewhere else -> Video

The train and benchmark scripts are provided (using MLX, but you can port to other platforms if you want).

Note that datasets are not included, and if you use a non-puretext dataset like wikipedia dump then feel free to write your own dataset extraction code or use an existing library.

The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore!