@Tfmonkey @VooDooMedic Stop spreading false information. Kimi K3 needs 2.7 TB of VRAM to run. That is 2700GB Of ram, where a high end GPU will have 96 GB of VRAM.
K3 is a great model, that is on Claude Fable level. But you cannot run it locally. Maybe in 3 years on an Apple M7 Max with 3 TB of Unified Memory. That device will cost 70kUSD.
https://github.com/JustVugg/colibri
proof of concept once this catches on and companies that manage OpenCode / Claude / Hermes actually get their teeth into it, it's gonna be a whole new fucking world my man yes trying to fit a model into VRAM requires insane amounts of hardware, however we aren't fitting in into vram just watch the video it explains it better
@viatorem @VooDooMedic @Tfmonkey Yeah, this is essentially what Revy was trying to explain a few months ago with major and minor AI models. This also explain the need for this insane amount of data centres. For a single person's use, we could likely set up a reasonably sized processing unit, but to serve millions of people at the same time, while also have a near instant response time... We need a lot more. And outsourcing kind of make economical sense for the user.