Apple Silicon

The crossing point between the two clusters: unified memory against dedicated VRAM.

Apple Silicon is not a third cluster. It is the second machine both clusters run on, which is the only reason a comparison here is worth reading: an M4 Pro and an RTX 4070 Ti are measured with the same model, the same quantization and the same context length, and the differences that remain are architectural rather than methodological.

Unified memory changes which failures happen. A 12GB CUDA card either fits a model or throws an out-of-memory error. An M4 Pro with enough unified memory tends to fit the same model and then run it slowly, so the limit shows up as duration instead. Both outcomes are recorded in the compatibility database, which currently holds 0 Apple Silicon runs against 3 CUDA runs.

Cross-platform articles arrive with the LLM cluster in week 7. Articles tagged 'apple-silicon' appear here automatically.