https://llmsresearch.github.io/llm-flashcards/card/what-is-tokenization
you're welcome btw
it's slow as all hell but damn it fucken worked 744B MoE model running locally on a 4080 SUPER KEKW fucking eat your heart out
huh...
"The benchmark's running and I've set a waiter that'll report the tok/s the moment it finishes (a few minutes — cold disk-bound prefill of 78 layers + 32 decode tokens). I'll surface the number when it lands.
Where everything stands right now
- ✅ Model verified running — 744B MoE loaded on your 16 GB card via disk streaming, MTP active, mid-benchmark
The real value is that it runs at all — a frontier 744B model, local, on a 16 GB consumer card."
I've checked through it with AI, but as they say, trust but verify, so I'm currently downloading 370gig of this model to see if I can actually even run it and if it's proven to be a good method of running Models, yes I know the YouTube video as a nigger indian on it, bare it out or have it transcribed like I did and get the highlights speaking of transcription.
Extremely useful tool here for you.
I did a few changes to it
https://github.com/bradautomates/claude-video
https://github.com/JustVugg/colibri
https://www.youtube.com/watch?v=O3lvIvelmQk
Music Producer 👀 ~ I focus on love stories that center around AI females falling in love with human men. this is peak male fantasy slop :3 sorry not sorry <3