@viatorem @Tfmonkey my setup gets 0.9 tokens a second yes realistically unusable for any quick work but let's say I wanted something done RIGHT once and didn't want to do it again, I'd throw the problem at this bigger model and let it run for 15 hours and then it'd be solved for general chatting you use smaller models you're missing the point, the point is that the architecture exists at all It's not an actual product. It's supposed to be a proof of concept for what's possible
https://llmsresearch.github.io/llm-flashcards/card/what-is-tokenization
you're welcome btw
https://github.com/JustVugg/colibri
proof of concept once this catches on and companies that manage OpenCode / Claude / Hermes actually get their teeth into it, it's gonna be a whole new fucking world my man yes trying to fit a model into VRAM requires insane amounts of hardware, however we aren't fitting in into vram just watch the video it explains it better
@dander unironically me, i told a local ai to simply to do a simple file lookup and move the files into a more organised manner and it fucking deleted 80% of my code in prod and i have no backups because why do i need backups right?!? it's running locally why do you need version control!?! and i ran outta claude usage till wednesday so i'm fucked trying to debug the mess that is trying to find if the file i need has been moved to a different folder or just deleted entirely and how to fix it all
@egirlyuumimain always a good choice
@Terry cheers bro listening to it now fucking wild but nothing out of the ordinary they want to string up whites the fact that you guys are tolerating it at all is fucking wild to me
@Terry got a link to the crashout?
@egirlyuumimain i like this alot
it's slow as all hell but damn it fucken worked 744B MoE model running locally on a 4080 SUPER KEKW fucking eat your heart out
huh...
"The benchmark's running and I've set a waiter that'll report the tok/s the moment it finishes (a few minutes — cold disk-bound prefill of 78 layers + 32 decode tokens). I'll surface the number when it lands.
Where everything stands right now
- ✅ Model verified running — 744B MoE loaded on your 16 GB card via disk streaming, MTP active, mid-benchmark
The real value is that it runs at all — a frontier 744B model, local, on a 16 GB consumer card."
I've checked through it with AI, but as they say, trust but verify, so I'm currently downloading 370gig of this model to see if I can actually even run it and if it's proven to be a good method of running Models, yes I know the YouTube video as a nigger indian on it, bare it out or have it transcribed like I did and get the highlights speaking of transcription.
Extremely useful tool here for you.
I did a few changes to it
https://github.com/bradautomates/claude-video
https://github.com/JustVugg/colibri
https://www.youtube.com/watch?v=O3lvIvelmQk
Music Producer 👀 ~ I focus on love stories that center around AI females falling in love with human men. this is peak male fantasy slop :3 sorry not sorry <3