@Tfmonkey @VooDooMedic Stop spreading false information. Kimi K3 needs 2.7 TB of VRAM to run. That is 2700GB Of ram, where a high end GPU will have 96 GB of VRAM.

K3 is a great model, that is on Claude Fable level. But you cannot run it locally. Maybe in 3 years on an Apple M7 Max with 3 TB of Unified Memory. That device will cost 70kUSD.

@viatorem @Tfmonkey

github.com/JustVugg/colibri

proof of concept once this catches on and companies that manage OpenCode / Claude / Hermes actually get their teeth into it, it's gonna be a whole new fucking world my man yes trying to fit a model into VRAM requires insane amounts of hardware, however we aren't fitting in into vram just watch the video it explains it better

youtube.com/watch?v=Uw6JBNR7gQ

@VooDooMedic @Tfmonkey

0.05 token per second btw (20 seconds to 1 token)

Good luck with that

@viatorem @Tfmonkey my setup gets 0.9 tokens a second yes realistically unusable for any quick work but let's say I wanted something done RIGHT once and didn't want to do it again, I'd throw the problem at this bigger model and let it run for 15 hours and then it'd be solved for general chatting you use smaller models you're missing the point, the point is that the architecture exists at all It's not an actual product. It's supposed to be a proof of concept for what's possible

Follow

@VooDooMedic @Tfmonkey

I Got it from newsletter from the framework laptop

“… In addition to bringing in the new Ryzen AI Max PRO 495+ processor, it uses slightly faster 8533 MT/s memory, has an open-end x4 PCIe slot to make it easier to plug in 50Gbps NICs for clustering, and optionally comes with Linux pre-loaded...”

It can run deepseek 4 flash locally, and it has rather fast memory controller up to 192 GB Ram,

Much slower than apple but it’s a start.

Apples memory is on the cpu die.

· · Web · 0 · 0 · 0
Sign in to participate in the conversation
Merovingian Club

A club for red-pilled exiles.