matt-p
25 minutes ago
This is super cool!
I agree the non-fungible nature of a H100 (or whatever) is a massive challenge. How do you actually verify usage/thermal history on these? Hours run and thermal violations aren't stored on the card IIRC, only ECC error counts and retired pages persist in the GPU's own InfoROM, host side monitoring like DCGM logs, are only as good as whatever the seller hands over. Is condition at inspection based on a real monitoring export from the deployment or just an InfoROM/diagnostic check at the time, and if it's the latter isn't hours run and thermal history basically unverifiable unless we assume good faith and the seller was logging the whole time ?
Essentially how do you avoid having the 'used car problem' without leaning heavily on seller reputation & warranties?