From bb2463a516494ce22cc4ea7fb03be277b255bccd Mon Sep 17 00:00:00 2001 From: jrz97619761 Date: Thu, 10 Sep 2026 14:48:21 +0800 Subject: [PATCH] fix readme pt 3 --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index b9e0903..38fd9b8 100644 --- a/README.md +++ b/README.md @@ -38,6 +38,6 @@ The two important hyperparameters are the size of the latent vector (dim) and th I think this probably will contribute significantly to solving continual learning and memory but I still need other people to review and verify my work! Please feel free to open GitHub issues to tell me what's wrong. If you have compute (e.g. you are a lab), feel free to fork my code and train larger models as well, with credit. -Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. +Below is an approximate flow chart of the model architecture, made in Apple's Freeform app (excluding the wrapper for dataset cleaning and input/output handling) for reference. Note that the arrow connecting the target latent to the CE loss should instead be the target byte to the CE loss. JEPA thing