I'm going to idle in a Matrix room from now on for those who want to discuss local AI (LLMs as well as diffusion models and adjacents). I've noticed this is an area that can be quite difficult to get into and since I've spent a considerable amount of time on it myself I'll happily share ideas and practices with others.

matrix.to/#/#LocalLLaMa:argot.

Local models allow using them for many different workflows without having to think of token costs - as well as knowing that sensitive data don't leave your premises.

#LLM #AI #LocalLLaMa

Follow

@troed Which open model is currently the best allround one? And will I be able to fit anything meaningful onto a 32 gb ram laptop?

· · Web · 1 · 0 · 0

@h4890 It will likely be very slow, but you can fit a Qwen 3.6 35B A3B quant into that memory profile and it's quite useful. I use variants of it for both coding and sysadm-tasks.

You'd use llama.cpp to serve it, and try to use as large as possible to fit quant from the ones available here: huggingface.co/unsloth/Qwen3.6

Start with UD-Q4-K_XL and go down smaller Q4s if needed.

@troed Great! Thank you very much for the advice! =)

Sign in to participate in the conversation
Merovingian Club

A club for red-pilled exiles.