From d4c0c88259ba7b0319e14206b271e2b9f169b422 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Wed, 12 Aug 2026 16:13:45 +0800 Subject: [PATCH 01/10] fix readme!! --- README.md | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/README.md b/README.md index 8314800..457e322 100644 --- a/README.md +++ b/README.md @@ -2,10 +2,8 @@ Hey! Thanks for being here. Here's the video, if you came here from somewhere else -> [Video](https://example.com) -The semi-trained 4.5m model (it's not done, but also it's in millions of parameters, not billions) is available for you to use. So are the train and benchmark scripts (using MLX, but you can port to other platforms if you want). - -However a larger 130m model is not provided, it is too large for GitHub to store. You can train it yourself and perhaps share it on a cloud storage provider instead. +The train and benchmark scripts are provided (using MLX, but you can port to other platforms if you want). Note that datasets are not included, and if you use a non-puretext dataset like wikipedia dump then feel free to write your own dataset extraction code or use an existing library. -The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore! \ No newline at end of file +The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore! From 6689a7d08fec80dbb98b08e19f77790900bb58e3 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Wed, 12 Aug 2026 16:28:54 +0800 Subject: [PATCH 02/10] update readme 2 --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 457e322..160d599 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ Hey! Thanks for being here. -Here's the video, if you came here from somewhere else -> [Video](https://example.com) +Here's the video, if you came here from somewhere else -> [Video](https://youtu.be/aXCaRem-vNA) The train and benchmark scripts are provided (using MLX, but you can port to other platforms if you want). From ee8f2ef98d041b680d3ca0a471f1766d7f483a02 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Sat, 22 Aug 2026 13:24:37 +0800 Subject: [PATCH 03/10] change link bc reupload --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 160d599..66c6f96 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ Hey! Thanks for being here. -Here's the video, if you came here from somewhere else -> [Video](https://youtu.be/aXCaRem-vNA) +Here's the video, if you came here from somewhere else -> [Video](https://youtu.be/mxYy9_4NBVg) The train and benchmark scripts are provided (using MLX, but you can port to other platforms if you want). From b598cd10b279eb98aa4c171fae7c4940bb13d43d Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Thu, 10 Sep 2026 14:37:21 +0800 Subject: [PATCH 04/10] redo the readme Now there's a lot more detail! and the video is down for now, because I need to rethink that. --- README.md | 38 +++++++++++++++++++++++++++++++++----- 1 file changed, 33 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index 66c6f96..404a306 100644 --- a/README.md +++ b/README.md @@ -1,9 +1,37 @@ -Hey! Thanks for being here. +# Test-Model-Thing (TMT) -Here's the video, if you came here from somewhere else -> [Video](https://youtu.be/mxYy9_4NBVg) +This is a small proof-of-concept language model (not an LLM) that incorporates the following (and some smaller features as well): +* Latent-space prediction +* Internal state + recurrent trace units (RTUs) +* Byte input/output +* Continuous data streaming +* Test-time training -The train and benchmark scripts are provided (using MLX, but you can port to other platforms if you want). +The model is built with MLX, so it should run fine on all Apple Silicon devices. It relies heavily on unified memory chips to achieve some things e.g. dynamic array resizing. -Note that datasets are not included, and if you use a non-puretext dataset like wikipedia dump then feel free to write your own dataset extraction code or use an existing library. +Being a proof of concept I have only trained a 4.5-million parameter model (keep in mind, GPT-1 was ~117m) for about 12 hours, but there are very promising results. The model tends to misspell characters (since it outputs byte-by-byte, rather than token-by-token) but it is able to close quotes/brackets and such. However given further training and scaling up the hyperparameters this could become much more powerful. My dataset is also tiny (only a few hundred MB), so there's a lot more world knowledge that can be fed into the model. -The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore! +Feel free to fork the training and benchmark code (everything is under MIT). + +_I used to have a video here, but I privated it for now._ + +3f7f1530-c0c7-43c4-9981-30e9023a19fb + +## How it works + +This model architecture was designed in about a month by me (a solo high school dev) and some Gemini. I write READMES myself though w/o AI. + +In detail, here are some of the main capabilities of the model that differ from LLMs: +* JEPA-style latent space prediction, as the decoder can be removed/disabled and the model still rolls out forward as is. The model is not trained explicitly on predicting the next byte, but rather on two separate goals (predicting the next 'thing' in latent space, and translating the current latent space vector to a byte). +* Theoretically infinite memory, as it does not have a context window and instead relies on RTUs to store internal state/memory. However it does decay old memories over time. Also I think this should be O(1) memory based on my implementation but I'm not 100% sure. +* Built-in multimodality, as the model outputs bytes (and thus should theoretically be capable of handling any binary data). +* Streaming data live, since the model only processes one byte at once at rapid pace. In fact it is completely 'blind' to everything that came before the current byte, only relying on the current processed byte and its internal memory to decide the next byte. This confirms the model is definitely learning to remember things. +* Continual training, as it keeps training on user input, training data, and its own output to improve its predictions automatically. Keep in mind the model can only output a byte (0-255) each pass anyways (in addition to updating its own state). + +The two important hyperparameters are the size of the latent vector (dim) and the amount of individual state layers the latent passes through before decoding (layers). For my 4.5m test these are ```dim = 512``` and ```layers = 16```. There are some other configurations you can change but I think they are less important. + +I think this probably will contribute significantly to solving continual learning and memory but I still need other people to review and verify my work! Please feel free to open GitHub issues to tell me what's wrong. If you have compute (e.g. you are a lab), feel free to fork my code and train larger models as well, with credit. + +Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. + +JEPA thing From 55779e90eae6bd98f57f7bcf154b2899a1071c67 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Thu, 10 Sep 2026 14:42:40 +0800 Subject: [PATCH 05/10] readme redo pt 2 --- README.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/README.md b/README.md index 404a306..b9e0903 100644 --- a/README.md +++ b/README.md @@ -17,6 +17,12 @@ _I used to have a video here, but I privated it for now._ 3f7f1530-c0c7-43c4-9981-30e9023a19fb +## Training your own model + +Model weights are not provided because GitHub doesn't like very large files. But, you can train your own model simply by initializing a ```venv``` and installing ```mlx```, no other libraries needed, then running ```main.py```. When you run it, you will be prompted with the mode, ```0``` being train on dataset and ```1``` being chat. You will have to configure your own dataset by modifying the code (to run dataset mode), but you should be able to run chat mode without modifying anything if you have weights already. + +Once it begins training, you can safely ^C the program and it will save weights. It should also periodically save weights if I'm not mistaken. The saved weights include the internal memory so the model will remember that the next time it runs. You can launch into chat mode and the memory should carry on from whatever it was learning in training. + ## How it works This model architecture was designed in about a month by me (a solo high school dev) and some Gemini. I write READMES myself though w/o AI. From bb2463a516494ce22cc4ea7fb03be277b255bccd Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Thu, 10 Sep 2026 14:48:21 +0800 Subject: [PATCH 06/10] fix readme pt 3 --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index b9e0903..38fd9b8 100644 --- a/README.md +++ b/README.md @@ -38,6 +38,6 @@ The two important hyperparameters are the size of the latent vector (dim) and th I think this probably will contribute significantly to solving continual learning and memory but I still need other people to review and verify my work! Please feel free to open GitHub issues to tell me what's wrong. If you have compute (e.g. you are a lab), feel free to fork my code and train larger models as well, with credit. -Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. +Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. Note that the arrow connecting the target latent to the CE loss should instead be the target byte to the CE loss. JEPA thing From aa3226708229f1d64f2a938bbf7a8ebddaa938f1 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Thu, 10 Sep 2026 14:49:02 +0800 Subject: [PATCH 07/10] readme 4 --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 38fd9b8..c6d27cf 100644 --- a/README.md +++ b/README.md @@ -19,7 +19,7 @@ _I used to have a video here, but I privated it for now._ ## Training your own model -Model weights are not provided because GitHub doesn't like very large files. But, you can train your own model simply by initializing a ```venv``` and installing ```mlx```, no other libraries needed, then running ```main.py```. When you run it, you will be prompted with the mode, ```0``` being train on dataset and ```1``` being chat. You will have to configure your own dataset by modifying the code (to run dataset mode), but you should be able to run chat mode without modifying anything if you have weights already. +Model weights (in ```.safetensors```) are not provided because GitHub doesn't like very large files. But, you can train your own model simply by initializing a ```venv``` and installing ```mlx```, no other libraries needed, then running ```main.py```. When you run it, you will be prompted with the mode, ```0``` being train on dataset and ```1``` being chat. You will have to configure your own dataset by modifying the code (to run dataset mode), but you should be able to run chat mode without modifying anything if you have weights already. Once it begins training, you can safely ^C the program and it will save weights. It should also periodically save weights if I'm not mistaken. The saved weights include the internal memory so the model will remember that the next time it runs. You can launch into chat mode and the memory should carry on from whatever it was learning in training. From b8ffe532942e15b0cc8dea127927befef349a968 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Thu, 10 Sep 2026 19:06:42 +0800 Subject: [PATCH 08/10] Readme redo part 5 there will be a part 6 stay tuned --- README.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index c6d27cf..a8b8b26 100644 --- a/README.md +++ b/README.md @@ -7,11 +7,13 @@ This is a small proof-of-concept language model (not an LLM) that incorporates t * Continuous data streaming * Test-time training -The model is built with MLX, so it should run fine on all Apple Silicon devices. It relies heavily on unified memory chips to achieve some things e.g. dynamic array resizing. +The model is built with MLX, so it should run fine on all Apple Silicon devices. MLX on Linux has not been tested, but feel free to try it. -Being a proof of concept I have only trained a 4.5-million parameter model (keep in mind, GPT-1 was ~117m) for about 12 hours, but there are very promising results. The model tends to misspell characters (since it outputs byte-by-byte, rather than token-by-token) but it is able to close quotes/brackets and such. However given further training and scaling up the hyperparameters this could become much more powerful. My dataset is also tiny (only a few hundred MB), so there's a lot more world knowledge that can be fed into the model. +Being a proof of concept I have only trained a 4.5-million parameter model (keep in mind, GPT-1 was ~117m) for about 12 hours, but there are very promising results. The model tends to misspell characters (since it outputs byte-by-byte, rather than token-by-token) but it is able to close quotes/brackets and such. Given further training and scaling up the hyperparameters this could become much more powerful. My dataset is also tiny (only a few hundred MB), so there's a lot more world knowledge that can be fed into the model. -Feel free to fork the training and benchmark code (everything is under MIT). +This model architecture was designed in about a month by me (a solo high school dev) and some Gemini (only pair programming, no agents). I wrote about a dozen prototypes before creating this architecture. I write READMEs myself without AI. + +Feel free to fork the training and benchmark code (everything is under MIT). I really encourage you to try things out, submit issues, and fork the repo. _I used to have a video here, but I privated it for now._ @@ -25,8 +27,6 @@ Once it begins training, you can safely ^C the program and it will save weights. ## How it works -This model architecture was designed in about a month by me (a solo high school dev) and some Gemini. I write READMES myself though w/o AI. - In detail, here are some of the main capabilities of the model that differ from LLMs: * JEPA-style latent space prediction, as the decoder can be removed/disabled and the model still rolls out forward as is. The model is not trained explicitly on predicting the next byte, but rather on two separate goals (predicting the next 'thing' in latent space, and translating the current latent space vector to a byte). * Theoretically infinite memory, as it does not have a context window and instead relies on RTUs to store internal state/memory. However it does decay old memories over time. Also I think this should be O(1) memory based on my implementation but I'm not 100% sure. @@ -36,7 +36,7 @@ In detail, here are some of the main capabilities of the model that differ from The two important hyperparameters are the size of the latent vector (dim) and the amount of individual state layers the latent passes through before decoding (layers). For my 4.5m test these are ```dim = 512``` and ```layers = 16```. There are some other configurations you can change but I think they are less important. -I think this probably will contribute significantly to solving continual learning and memory but I still need other people to review and verify my work! Please feel free to open GitHub issues to tell me what's wrong. If you have compute (e.g. you are a lab), feel free to fork my code and train larger models as well, with credit. +I think this probably will contribute significantly to solving continual learning and memory but I still need other people to review and verify my work! Please feel free to open GitHub issues to tell me what's wrong. If you have compute (e.g. you are a lab or just have GPUs lying around), feel free to fork my code and train larger models as well, with credit. I personally don't have enough compute and as such I can't really train very large models. Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. Note that the arrow connecting the target latent to the CE loss should instead be the target byte to the CE loss. From eaf8adc27355cbb54a989d3866c9e4da9db08237 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Sun, 13 Sep 2026 07:56:37 +0800 Subject: [PATCH 09/10] video -> readme --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index a8b8b26..ca30f69 100644 --- a/README.md +++ b/README.md @@ -15,7 +15,7 @@ This model architecture was designed in about a month by me (a solo high school Feel free to fork the training and benchmark code (everything is under MIT). I really encourage you to try things out, submit issues, and fork the repo. -_I used to have a video here, but I privated it for now._ +[YouTube Video](https://youtu.be/9UERVVwpNew) 3f7f1530-c0c7-43c4-9981-30e9023a19fb From aca7f3826e97373c572874844d292c7bb477fe43 Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Sun, 13 Sep 2026 07:57:02 +0800 Subject: [PATCH 10/10] video fix --- README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index ca30f69..bd9b00a 100644 --- a/README.md +++ b/README.md @@ -1,5 +1,7 @@ # Test-Model-Thing (TMT) +[YouTube Video](https://youtu.be/9UERVVwpNew) + This is a small proof-of-concept language model (not an LLM) that incorporates the following (and some smaller features as well): * Latent-space prediction * Internal state + recurrent trace units (RTUs) @@ -15,8 +17,6 @@ This model architecture was designed in about a month by me (a solo high school Feel free to fork the training and benchmark code (everything is under MIT). I really encourage you to try things out, submit issues, and fork the repo. -[YouTube Video](https://youtu.be/9UERVVwpNew) - 3f7f1530-c0c7-43c4-9981-30e9023a19fb ## Training your own model