diff --git a/README.md b/README.md
index bd9b00a..8314800 100644
--- a/README.md
+++ b/README.md
@@ -1,43 +1,11 @@
-# Test-Model-Thing (TMT)
+Hey! Thanks for being here.
-[YouTube Video](https://youtu.be/9UERVVwpNew)
+Here's the video, if you came here from somewhere else -> [Video](https://example.com)
-This is a small proof-of-concept language model (not an LLM) that incorporates the following (and some smaller features as well):
-* Latent-space prediction
-* Internal state + recurrent trace units (RTUs)
-* Byte input/output
-* Continuous data streaming
-* Test-time training
+The semi-trained 4.5m model (it's not done, but also it's in millions of parameters, not billions) is available for you to use. So are the train and benchmark scripts (using MLX, but you can port to other platforms if you want).
-The model is built with MLX, so it should run fine on all Apple Silicon devices. MLX on Linux has not been tested, but feel free to try it.
+However a larger 130m model is not provided, it is too large for GitHub to store. You can train it yourself and perhaps share it on a cloud storage provider instead.
-Being a proof of concept I have only trained a 4.5-million parameter model (keep in mind, GPT-1 was ~117m) for about 12 hours, but there are very promising results. The model tends to misspell characters (since it outputs byte-by-byte, rather than token-by-token) but it is able to close quotes/brackets and such. Given further training and scaling up the hyperparameters this could become much more powerful. My dataset is also tiny (only a few hundred MB), so there's a lot more world knowledge that can be fed into the model.
+Note that datasets are not included, and if you use a non-puretext dataset like wikipedia dump then feel free to write your own dataset extraction code or use an existing library.
-This model architecture was designed in about a month by me (a solo high school dev) and some Gemini (only pair programming, no agents). I wrote about a dozen prototypes before creating this architecture. I write READMEs myself without AI.
-
-Feel free to fork the training and benchmark code (everything is under MIT). I really encourage you to try things out, submit issues, and fork the repo.
-
-
-
-## Training your own model
-
-Model weights (in ```.safetensors```) are not provided because GitHub doesn't like very large files. But, you can train your own model simply by initializing a ```venv``` and installing ```mlx```, no other libraries needed, then running ```main.py```. When you run it, you will be prompted with the mode, ```0``` being train on dataset and ```1``` being chat. You will have to configure your own dataset by modifying the code (to run dataset mode), but you should be able to run chat mode without modifying anything if you have weights already.
-
-Once it begins training, you can safely ^C the program and it will save weights. It should also periodically save weights if I'm not mistaken. The saved weights include the internal memory so the model will remember that the next time it runs. You can launch into chat mode and the memory should carry on from whatever it was learning in training.
-
-## How it works
-
-In detail, here are some of the main capabilities of the model that differ from LLMs:
-* JEPA-style latent space prediction, as the decoder can be removed/disabled and the model still rolls out forward as is. The model is not trained explicitly on predicting the next byte, but rather on two separate goals (predicting the next 'thing' in latent space, and translating the current latent space vector to a byte).
-* Theoretically infinite memory, as it does not have a context window and instead relies on RTUs to store internal state/memory. However it does decay old memories over time. Also I think this should be O(1) memory based on my implementation but I'm not 100% sure.
-* Built-in multimodality, as the model outputs bytes (and thus should theoretically be capable of handling any binary data).
-* Streaming data live, since the model only processes one byte at once at rapid pace. In fact it is completely 'blind' to everything that came before the current byte, only relying on the current processed byte and its internal memory to decide the next byte. This confirms the model is definitely learning to remember things.
-* Continual training, as it keeps training on user input, training data, and its own output to improve its predictions automatically. Keep in mind the model can only output a byte (0-255) each pass anyways (in addition to updating its own state).
-
-The two important hyperparameters are the size of the latent vector (dim) and the amount of individual state layers the latent passes through before decoding (layers). For my 4.5m test these are ```dim = 512``` and ```layers = 16```. There are some other configurations you can change but I think they are less important.
-
-I think this probably will contribute significantly to solving continual learning and memory but I still need other people to review and verify my work! Please feel free to open GitHub issues to tell me what's wrong. If you have compute (e.g. you are a lab or just have GPUs lying around), feel free to fork my code and train larger models as well, with credit. I personally don't have enough compute and as such I can't really train very large models.
-
-Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. Note that the arrow connecting the target latent to the CE loss should instead be the target byte to the CE loss.
-
-
+The repo is MIT license, so feel free to fork the repo, I would be very happy to see that. Go ahead and explore!
\ No newline at end of file