Rendered at 04:34:56 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
librasteve 9 hours ago [-]
As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org
feffe 6 hours ago [-]
I think tenstorrent architecture is more like this. A grid of RISC-V cores with local SRAM and vector units.
tolugenius 9 hours ago [-]
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
pjmlp 7 hours ago [-]
That was the whole point of Project Larrabee, whose ideas survive in AVX.
On the compute side there's the issue of scalability. CPUs are designed to perform a handful of operations at one time. They typically have a small number of dedicated integer, float, and other ALU configurations. Having dedicated matrix multiplication instructions would still lock that to how many matrix-capable ALUs there are in the CPU.
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
5 hours ago [-]
Dwedit 6 hours ago [-]
16 years ago, Intel CPUs finally got GPUs integrated inside of them. So they did get better at matrix multiplication 15 years ago.
alfiedotwtf 3 hours ago [-]
The tv ads for MMX made it feel like it was going to change the world
Dwedit 3 hours ago [-]
15 years ago was Sandy Bridge, which added the Intel HD Graphics GPU inside of the processor. MMX was almost 30 years ago.
actionfromafar 9 hours ago [-]
I wonder what "compute" would mean if CPUs were more efficient at matrix multiplication and vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.
pjmlp 7 hours ago [-]
It could be like the connection machine or something like that.
There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
aeve890 39 minutes ago [-]
>There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
Like in the existing SIMD technology?
vivzkestrel 1 hours ago [-]
- stupid question
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
aeve890 24 minutes ago [-]
Yes. there is real hardware architecture behind Apple’s “Unified Memory”, but the underlying idea is not uniquely Apple. What Apple did was design the CPU, GPU, memory controller, cache hierarchy, interconnect, package, OS, and graphics APIs together around the architecture. I mean, they can do shit like that because they own the entire product pipeline. They can fine tune the hardware in ways other OEMs can't.
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data.
You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_la...
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
Like in the existing SIMD technology?
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data. You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.