bArray
2 hours ago
That exchange is quite a good prediction all things considered.
> Note that "minor" implementation issues like die space, routing, and gate delays, especially of 128-bit adders & shifters are non-trivial, so people aren't going to rush out and build 128-bitters for fun, just as people matched timing dates of their 64-bitters to their expected markets.
I think we're stuck with 64 bit for quite a while. The circuit size jump from 64 bit to 128 bit is significant. There's no fixed scalar or apples to apples comparison (that I'm aware of). Just for adders though, 2 bits requires 2 full adders, 4 bits requires 4 adders, 8 requires 8, etc.
Another way to look at this is that 16 bit gives addressable memory up to 65k, and quite a few programs had to deal with paging in architectures like the 8086/8088. 32 bits gave us up to 4GB addressable memory, and not it's now not uncommon that a program such as a web browser exceeds this. 64 bits would give us up to 18 exabytes of addressable memory. I'm not aware of any common programs breaking into the terabyte category (even in most research), let alone petabyte and then exabyte.
Even iterating over that many numbers becomes a large computational task. Just a quick test program:
// gcc -O3 count.c -o count
#include <stdint.h>
int main(){
uint64_t z = 0;
for(uint64_t i = 0; i < UINT64_MAX; i++) z += i;
return (int)(z % 2);
}
Using uint32_t and UINT32_MAX, it returns almost instantly. For uint64_t and UINT64_MAX you will be waiting a long time. 128 bit values? Even longer. Maybe many many cores could break that memory up, but then it makes sense to have a 64 bit system with some kind of ability to occasionally change page.kalleboo
2 hours ago
> I'm not aware of any common programs breaking into the terabyte category (even in most research)
Wouldn't that be the commercial LLMs? ChatGPT, Claude etc are estimated to be in the 2-10 TB range, and a large portion of the population of developed countries are using those apps commonly. Not on their own systems, but if we're in the 1995 academic perspective, multiuser systems are assumed.
I imagine the strongest reason we won't need 128-bit memory addressing is because horizontal scaling is easier. If OpenAI had needed a single addressing plane to cover all their users, 128-bit might be required. Similar to how the internet hobbles along fine with 32-bit addressing by just adding a layer of indirection to the internet with NAT.
I guess the other example of 128-bit addressing is ZFS. What are the biggest ZFS file systems? And who has more data? S3? Once you get bigger than 64-bit you want to scale out horizontally anyway and not just put it all into one flat addressable plane.
bArray
38 minutes ago
> Wouldn't that be the commercial LLMs? ChatGPT, Claude etc are estimated to be in the 2-10 TB range, and a large portion of the population of developed countries are using those apps commonly. Not on their own systems, but if we're in the 1995 academic perspective, multiuser systems are assumed.
I think each part is generally treated like an individual large program - which makes sense because otherwise you have large parts of the model sitting there doing nothing. You want all of your silicon running hot, RAM sitting doing nothing for a period of time is a waste.
> If OpenAI had needed a single addressing plane to cover all their users, 128-bit might be required. Similar to how the internet hobbles along fine with 32-bit addressing by just adding a layer of indirection to the internet with NAT.
I think even then, 128 bit memory for OpenAI et al would be more of a headache than it is worth. Economically it's not worth building the 128 bit systems. The fact that servers and consumer CPUs have a lot of re-use within their designs reduces the costs for everybody. Sony for example stopped producing their 128 bit CPU [1], where it was mostly about pushing more data between CPUs and VPUs. Now we have specifications like PCIe where you can just adjust your pipe width to increase throughput.
> I guess the other example of 128-bit addressing is ZFS. What are the biggest ZFS file systems? And who has more data? S3? Once you get bigger than 64-bit you want to scale out horizontally anyway and not just put it all into one flat addressable plane.
Yeah exactly. The other thing to consider is that it doesn't matter how many bits wide you can go, the bottleneck is the smallest width. With ZFS, at some point you likely need to send it over a network, and I think fibre is up to 64 bits via parallel transceivers. If that points outwards, it'll probably be single bit by the time it gets to you.
torginus
an hour ago
Dunno, a 128bit integer add is essentially 2x as expensive as a 64 bit one, a multiply is 4x. We've already had SSE in the early 2000s that could do computations like this in a single cycle (tho on multiple 32 bit numbers, not 128-bit ones).
I think on x86, we're still limited to 48 bits of address space on 64 bit systems, not sure if this has changed, but even 64 bits is so vast (16exabytes) , that you'd need a supercomputer with a single address space to fill it, but at least the latter is practically realizable.