test-model-thing/README.md
2026-08-12 16:11:30 +08:00

11 lines
No EOL
807 B
Markdown

Hey! Thanks for being here.
Here's the video, if you came here from somewhere else -> [Video](https://example.com)
The semi-trained 4.5m model (it's not done, but also it's in millions of parameters, not billions) is available for you to use. So are the train and benchmark scripts (using MLX, but you can port to other platforms if you want).
However a larger 130m model is not provided, it is too large for GitHub to store. You can train it yourself and perhaps share it on a cloud storage provider instead.
Note that datasets are not included, and if you use a non-puretext dataset like wikipedia dump then feel free to write your own dataset extraction code or use an existing library.
The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore!