Rendered at 23:44:40 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
simonw 8 hours ago [-]
It looks to me like this won't speed up llama.cpp for everyone, just for users running it in this particular kind of Virtualization.framework VM.
The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.
frabonacci 8 hours ago [-]
> this won't speed up llama.cpp for everyone, just for users running it in this particular kind of Virtualization.framework VM.
correct. these figures apply to llama.cpp inside the macOS guest configuration we tested. Lume is the VM frontend we used, while Apple's Virtualization.framework provides the virtual GPU. bare-metal llama.cpp is unaffected.
> The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.
mostly, with one nuance: llama.cpp is selecting the correct kernels for the capability answers it receives. the stock guest reports an older Apple GPU family and a 32 KB threadgroup memory limit, so llama.cpp chooses slower kernels. Our process-scoped layer reports the tested Apple 9 and 64 KB values while allowing llama.cpp to select newer paths that the paravirtual GPU successfully execute
the layer itself though works at the Metal API boundary, independently of llama.cpp. other Metal compute and graphics apps now may select newer paths from the same capability answers, although this is still preliminary and each app needs separate testing. for example, MLX-LM stayed flat in our tests
Why do you expect an AI engineer to manually write prose?
b112 4 hours ago [-]
Don't post generated text or AI-edited text. HN is for conversation between humans.
Because it is not allowed here, that's why. See the guidelines.
throwaway173axp 5 hours ago [-]
Why do you think an AI engineer would go through the trouble lower casing everything except for proper nouns and abbreviations?
frabonacci 5 hours ago [-]
the better question is why a throwaway account is doing capitalization forensics
octocop 6 hours ago [-]
a win is still a win
engzaanin 8 hours ago [-]
That makes sense. The title initially sounded like a general llama.cpp speedup on Apple Silicon, but if the improvement comes from fixing kernel selection inside Virtualization.framework VMs, that distinction is pretty important.
> 11.08× faster and generated tokens 16.36× faster than the same workload in the same stock VM.
So this was the comparison, for me the title was a bit confusing
frabonacci 8 hours ago [-]
yeah fair point. it's always tricky to get the whole idea across within HN's title limit. tldr: we ran the same workload in the same Lume macOS VM on the same Apple Silicon host, first with stock Metal capability reporting and then with our process-scoped dynamic library. The 11.08x figure is prompt processing, while 16.36x is token generation. the mechanism technically extends to graphics workloads too but these figures are specifically from llama.cpp
aeriose 7 hours ago [-]
What I don't get, which this article doesn't talk about, why would Apple’s Virtualization.framework expose a lesser Metal profile instead of reporting all capabilities supported by the host GPU?
hugmynutus 6 hours ago [-]
Because nobody knows.
Apple doesn't let you "pass" the GPU through to a VM like most other ARM/x86_64 processors (forwarding interrupts and PCIe memory regions). There are symbols defined to do this within the kernel (if you dump the binary) but they aren't used in retail macos.
Instead you end up creating a paravirtual device that emulates the GPU acting like a 'normal PCI device' which you give to clients. This is usually reserved (by other hardware vendors) for when you're doing multi-tenat time sharing of higher end GPUs (like Nvidia enterprise cards can do).
These paravirtualized GPUs then just have 'less features' and Apple (being Apple) states no reason why.
frabonacci 2 hours ago [-]
my guess is apple chose a conservative profile for compatibility across different chips and guest releases
chorizo 6 hours ago [-]
All M-series chips support Metal 4. Wonder if we can fix this with a simple override somewhere.
adityazero 1 hours ago [-]
[flagged]
bestham 7 hours ago [-]
Because it cannot be safely virtualised?
7 hours ago [-]
frabonacci 6 hours ago [-]
[flagged]
b112 4 hours ago [-]
QEMU/kvm does this as a default, because keeping a more generic CPU / etc makes moving VMs between machines with different hardware possible. If you try to move a VM it won't work, of course, if the new machine doesn't support what the old did.
Not sure of this is why Apple does it. With KVM, you tend to pick a baseline that all your machines support.
w10-1 3 hours ago [-]
Related: does anyone have a basis for guessing whether the Neural Accelerators found in M5 Pro+ (accessed by Metal 4) will make their way into the M6 base processors?
wyzer 2 hours ago [-]
I see M1 Ultra host mentioned, are there any M1 pro or M3 pro results? Has anyne tried?
frabonacci 1 hours ago [-]
We've also seen similar improvements on a M5 max. no M1 pro or M3 pro results yet though - would love to see someone try those
azinman2 8 hours ago [-]
I don’t understand what Apple 1-9 are. At first I thought it was M series chips but there is no M9 (yet)
So those generation numbers aren't really anchored to Apple's hardware designs. It's just counting from when Apple introduced the Metal API, and the first several generations were when the GPU cores Apple was using were still nominally PowerVR designs.
frabonacci 6 hours ago [-]
yeah the naming is confusing. Apple family 9 isnt M9, it's a Metal GPU feature family. Apple maps family 7 to M1, family 8 to M2, family 9 to M3/M4, and family 10 to M5
shay_ker 8 hours ago [-]
I recall there was another YC startup that was working on Mac-specific ML optimizations for local inference (and perhaps fine-tuning).
I wonder if their work is related?
frabonacci 8 hours ago [-]
RunAnywhere or Conifer?
luciana1u 7 hours ago [-]
my whole setup is buy more RAM, run it on CPU, and tell myself the GPU is just a personality trait I'm working on.
cyanydeez 7 hours ago [-]
I'm hoping AMD wins when the RAM bubble bursts and their integrated AMD 395+ platform can keep getting faster and higher bandwidth.
bearjaws 5 hours ago [-]
Forget the 395+, I want the MI350P (or similar) long term.
PCIe card is the way forward IMO, AI keeps changing so you don't want static hardware.
petu 5 hours ago [-]
How PCIe card is better?
"Static hardware" is still fully featured computer with lots of RAM, could be easily reused for other purposes.
speed_spread 5 hours ago [-]
The unified memory architecture is interesting for toying with medium size models but will never offer as much bandwidth as a dedicated GDDR memory bank. Conversely, GDDR can't be used for general CPU purposes because access latency is just too high. Unless someones also comes up with dynamically programmable memory banks, something I'm not sure would even be possible.
gigatexal 7 hours ago [-]
All this work to get the Mac to be a platform useful for AI is being done despite Apple's efforts. They're famously pissed at Nvidia since the Nvidia + Intel Macs due to heat and other issues. But then the OS is a bit closed off and they move slow and are more focused on milking iOS and the App Store and services for money BUT the PA-Semi purchase and Apple Silicon and everything following it has made the hardware just so amazing and useful that despite all that people build for it.
I love the platform. I'm happy to see people building on it.
AND! if we ever get an M7 chip with the rumored 1.5TB of available ram all this work will not have been in vain. You think the ai acceleration is nice in the M5 wait till M7 and M8.
frabonacci 4 hours ago [-]
apple silicon is what made us start Lume in the first place last year. the hardware is so good (M1 is now 6 years old!) that people keep pushing through the gaps in the platform. just yesterday we ran a fully offline computer-use agent with Cua Driver and Muse Glimmer, all locally on Apple Silicon, an now the same kind of agent can run isolated inside a macOS VM and use Apple’s GPU path too
dxsecarch 8 hours ago [-]
[dead]
CurbStomper 5 hours ago [-]
[dead]
purplemoonx 8 hours ago [-]
[flagged]
kevin42 8 hours ago [-]
What hardware do you run? I have a first-gen mac studio, and I just run cmake and build with no special options. Same thing with llama-server, I just specify the model and use the built-in web UI.
For reference, I get ~26 tok/sec with the new Muse 30B model.
dofm 8 hours ago [-]
An M1 Max MBP manages roughly 10 tok/sec without the Dflash speculative draft support so that tracks; the M1 Max apparently has trouble actually saturating its memory bandwidth.
drittich 8 hours ago [-]
There are certainly challenges. When setting up a new model, I get AI to walk me through the commands using llama-benchmark that determine the best parameters for my particular configuration and needs. Once you've got that it's pretty easy to port those parameters to llama-server. It takes me about an hour to run through this process. It would be great if there was a registry of hardware, models, configuration parameters, and resulting tokens per second. Maybe one day we'll get there.
kevin42 8 hours ago [-]
What kind of parameters do you end up changing, and how much difference does it make. Perhaps I am missing something and get more tok/sec, but I usually just do a git pull, then rebuild the latest whenever I get a new model.
In the past, I had to play with chat templates for some models to work with agents for tool calling. But I've never had to do anything other than specify the model, and tweaking the context size in some cases.
purplemoonx 8 hours ago [-]
> It takes me about an hour to run through this process
Yeah not doing that
unglaublich 8 hours ago [-]
I think people generally throw Claude or Codex at the configuration challenge, so they don't know either.
purplemoonx 8 hours ago [-]
Maybe the llama.cpp dev loved webpack as a child, or just loves making the most simple thing complicated as hell for no reason lol
NamlchakKhandro 8 hours ago [-]
No
myshapeprotocol 8 hours ago [-]
[flagged]
woadwarrior01 8 hours ago [-]
The Claudish in the blogpost makes it really hard to ready. Also, TinyLlama 1.1B lol.
The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.
correct. these figures apply to llama.cpp inside the macOS guest configuration we tested. Lume is the VM frontend we used, while Apple's Virtualization.framework provides the virtual GPU. bare-metal llama.cpp is unaffected.
> The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.
mostly, with one nuance: llama.cpp is selecting the correct kernels for the capability answers it receives. the stock guest reports an older Apple GPU family and a 32 KB threadgroup memory limit, so llama.cpp chooses slower kernels. Our process-scoped layer reports the tested Apple 9 and 64 KB values while allowing llama.cpp to select newer paths that the paravirtual GPU successfully execute
the layer itself though works at the Metal API boundary, independently of llama.cpp. other Metal compute and graphics apps now may select newer paths from the same capability answers, although this is still preliminary and each app needs separate testing. for example, MLX-LM stayed flat in our tests
historically related limitations have been coming up across Apple Silicon VM frontends for a while e.g. Tart tracked MPS/GPU support back in 2023: - https://github.com/openai/tart/issues/501 - https://github.com/openai/tart/issues/1032
UTM also has related cases where apps detect the Apple paravirtual Metal device but falls back to software rendering: https://github.com/utmapp/UTM/issues/7671
Because it is not allowed here, that's why. See the guidelines.
So this was the comparison, for me the title was a bit confusing
Apple doesn't let you "pass" the GPU through to a VM like most other ARM/x86_64 processors (forwarding interrupts and PCIe memory regions). There are symbols defined to do this within the kernel (if you dump the binary) but they aren't used in retail macos.
Instead you end up creating a paravirtual device that emulates the GPU acting like a 'normal PCI device' which you give to clients. This is usually reserved (by other hardware vendors) for when you're doing multi-tenat time sharing of higher end GPUs (like Nvidia enterprise cards can do).
These paravirtualized GPUs then just have 'less features' and Apple (being Apple) states no reason why.
Not sure of this is why Apple does it. With KVM, you tend to pick a baseline that all your machines support.
I wonder if their work is related?
PCIe card is the way forward IMO, AI keeps changing so you don't want static hardware.
"Static hardware" is still fully featured computer with lots of RAM, could be easily reused for other purposes.
I love the platform. I'm happy to see people building on it.
AND! if we ever get an M7 chip with the rumored 1.5TB of available ram all this work will not have been in vain. You think the ai acceleration is nice in the M5 wait till M7 and M8.
For reference, I get ~26 tok/sec with the new Muse 30B model.
In the past, I had to play with chat templates for some models to work with agents for tool calling. But I've never had to do anything other than specify the model, and tweaking the context size in some cases.
Yeah not doing that