tomex
2 days ago
Crazy how little traction this kind of project gets on here. Posts about squeezing a 1+T model to seconds/token and you have this wave of optimism like "It's the effort that counts! We'll get there!". Sure, this particular project isn't really scalable in the same sense (PL fabric/use what ya got/cost/power) but IMO it's conceptually a brilliant thing to showcase comparatively. I have a strange feeling a decent chunk of people dismissing this project are the same who spent small fortunes on hobby llm inference setups/investments and see this as a useless exercise. Meanwhile dozens of $$$M startups in the CIM/analog compute/etc have been R&D'ing for years now that will make this same outcome a reality before we know it (crazy inference speeds on usable models within local reach). Anyways kudos to OP and really enjoyed the documentation and findings of this!
mikeayles
2 days ago
Appreciated. and yeah, agreed the interesting comparison isn't "is this scalable as-is" (it isn't, PL fabric, cost, power), it's that the CIM / analog-compute startups you mention are chasing exactly this endpoint with real money and years of R&D, the bar for making something useful is brutally high.
However, if no-one made anything that was useless on the same thesis, a lot of these concepts would have never got off the ground. I would hazard a guess that people like taalas would have started with a (much much bigger) fpga to validate whether the approach was possible before committing to designing a chip big enough to fit an 8B model in it.
I just nerd sniped myself...
VP1902 could fit around a 500m model in, whereas a cadence protium rack of them could squeeze in a ~6B at 8bit, or a ~13B at 4bit. So accounting for the headroom of distributed compute, Llama 3.1 8B at 4bit. I don't want to even estimate how long synthesis and place and route would take on that!