aroman
8 hours ago
Oh my god, and the author even supplied a "proof"[0] visual diff harness... that it replicates the original game pixel for pixel.
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
[0] https://github.com/terrapapagalli1516/quake-srp/tree/main/or...
malkia
20 minutes ago
There are several good algos used for diffing graphics for not so pixel-for-pixel rendering allowing thresholds (you are going to end up with differences between software and GPU rendering at least - it's inevitable):
https://github.com/NVlabs/flip
But also SSIM, perceptualdiff, and many others
For UI/web certain other methods are better, for video - I'm sure there are prefferences there too.
ChickeNES
8 hours ago
tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.
aroman
8 hours ago
If it is common in the world of video game porting, that just shows my ignorance. I'm familiar with visual diffs in CI for e.g. web development (comparing a static component), but to do that to compare frames over time in a video game/3D environment is new to me.
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
koito17
8 hours ago
Even then, visual diffs were pretty flakey for web development, because one's OS and browser choice would slightly alter the exact pixels blitted to the screen. At least this was the case for the tests that would simply match pixels instead of computing a sort of visual hash.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
tnolet
8 hours ago
For web UI tests this is mainly solved, at least when using Playwright. It allows setting thresholds, percentages and some other config items to allow some small differences in pixels. https://playwright.dev/docs/test-snapshots#options
ffsm8
18 minutes ago
I wouldn't call that solved solved, no.
That's an attempt at mitigation, but most definitely not solved.
It still causes both false positives and false negatives through that. The only true "solution" of to make sure generationalways happens on the same platform as your ci... And various mitigation strategies around that (eg fall back to structure tests on other platforms vs actual visual diffs in ci
Also not unique to playwright. Been available basically everywhere since the start
TeMPOraL
7 hours ago
I'm getting a 1997 PC game to run on modern hardware and fixing bugs and upgrading graphics as I go, and the amount of quality support tooling Claude is producing along the way is impressive. Fully headless in-memory execution (which, among other things, is used by it for per-pixel diffs too), logic VM devompiler and visualizer, asset explorer, CRT simulator... I just say what I'd like to see, and Claude does 120% job on it each time.
Keyframe
4 hours ago
I do that on my renderer. each commit when it's ready to be merged gets a class of visual diffs. Dolphin's way was a major inspiration how to structure it, but it was _the way_ in rendering way before it.
nitwit005
8 hours ago
Using pixel data might be unusual, but anyone porting a game with a replay feature is going to realize it's an easy way to compare implementations.
brobdingnagians
8 hours ago
The visual diff harness was probably how they got rid of a lot of visual bugs, just tell the LLM to keep going until the pixels match exactly as the verification criteria
maxyurk
8 hours ago
I suspect it's inspired by gbaeval. Both have the "oracle" and other similarities. https://gbaeval.com/
King-Aaron
8 hours ago
Yeah cool.
How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?
It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.
hn_submit
7 hours ago
Like I've said many times: LLMs are useful for this (i.e. porting software from one programming language to another).
Porting software is painstaking grunt work which still takes a moderate amount of intelligence. It's therefore extremely expensive to port say, COBOL banking software running on mainframes, to another language like Java. That's why a lot of COBOL software is still in use. I expect this to die out in the coming years as many of these systems will finally be ported to another language (could be Rust or any other language).
chii
7 hours ago
> spoofs a system so pivotal to modern human society that the plug needs to be pulled.
why would a recreated system be detrimental?
If currently there's a monopoly on a software, this AI recreation is a good outcome to poke holes in that monopoly. It's only bad if you are financially invested in said monopoly, and this would be a minority compared to the amount of benefits that society at large could obtain.
King-Aaron
7 hours ago
This is basically a digital era anarchist view - the problem you're overlooking is that a lot of critical infrastructure we rely on runs on systems that are considered security through obscurity. Software most people probably wouldn't even know or care that it exists. If you can break the trust of vendors by being able to spoof their proprietary platforms, a lot of the highly efficient networked systems becomes vulnerable to injection and abuse if you can't trust whos making calls to it.
In the past you'd need nation state actors with considerable budgets to do this kind of thing, and we're on a trajectory that could see any kid in his bedroom could do it.
chii
7 hours ago
revealing that security thru obscurity is broken can only lead to a better future, even if in the intermediate one there are lots of breakages. It's suffering that needs to happen, and better sooner rather than later imho.
And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded. Like using a photoshop replacement.
King-Aaron
6 hours ago
> And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded.
I am, actually, talking about MITM attacks that are much further in scope than just defrauding some people using their banking app. I work in resources and operate HMI systems that are networked, but not exactly the bleeding edge of modern software development. If you had an ability to decompile it and recompile your own version you could start sending instructions to infrastructure all over the country - the only thing stopping you is the keys, which if you're intent on hacking someone you'd have the means to obtain anyway.
I can see a lot broader attack vectors than just stealing peoples money. It's the erosion of trust in the api calls themselves.
brobdingnagians
8 hours ago
If someone recreates Amazon's logistics and distribution systems they could try to compete with Amazon? But they'd also need the connections, distributors, transportation, etc. same with banking software, you need capital to be a bank not just software, and if they have the capital then the technology is working we intended making it easier to make new things and innovate, or at least just compete?
King-Aaron
8 hours ago
No, I am not talking about "taking over" companies and trying to emulate them and do business yourself. You just need to be able to break trust in the api calls and no one knows if a purchase order or transaction is legitimate.
Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.
illwrks
7 hours ago
I think you’re saying ‘being able to do this means the opportunity for more fraud, by producing fake XYZ as proof’.
Photoshop has been around for around 35 years, fraud has always been an issue. There are plenty of reports of people selling things via Facebook marketplace and the ‘buyer’ showing them sending a payment on a fake baking app. Fraud will always exist and I don’t think tech will make it worse, everyone needs to be more cautious and tells friends and family to be the same.
King-Aaron
6 hours ago
Nah small potato stuff.
I'm talking about cloning hmi platforms to send fake instructions to offshore oil platform valve bodies or insert false market trades to collapse companies.
zrobotics
an hour ago
Or perhaps uploading faulty firmware to centrifuge controllers?
illwrks
5 hours ago
Ah. So in that instance are those platforms not validating that the things submitting information are correct and true. A bit like a utility company needing to do manual reads every now and then to ensure they are getting a correct signal.
dyauspitr
8 hours ago
Or they can just take your money and not send out anything.
Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.
pjmlp
7 hours ago
Example, parking places with QR codes for paying webapps.
Currently a plague in some European countries.
It looks like the real site, and you pay twice, in the fake app, and later the police.
jurgenburgen
7 hours ago
That’s an insecure design. The way we do it here is that you install an app and register your register number and payment card in it. Then when you drive in and out from the parking lot your license plate is scanned and you’re automatically charged. There’s only two providers so it’s not a huge hassle, if there was a single app per garage it would not really work from UX perspective.
No room for hostile social engineering.
pjmlp
5 hours ago
Are you sure to install the right app though?
https://www.bbc.com/news/articles/cwyjqg578e1o
And given your German nickname, here isn't safe either in a general way, when folks aren't regularly parking on the same place.
King-Aaron
8 hours ago
Yep. On a small scale, you could skim money off transactions. On a large scale, you could break global distribution and logistics chains.
konart
7 hours ago
>banking back ends
As someone who works in a bank: depending on the exacty subsystem of a bank the answer is from "already" to "in 3-5 years".
schleck8
7 hours ago
Making predictions for in 5 years with the current dynamic of the ecosystem seems rather speculative
cedws
7 hours ago
I’m already seeing videos of people who have used LMs to reverse engineer and clean room reimplement entire video games. I estimate this shit is minutes away from being shut down, because as we’ve all seen companies stealing is OK, but individuals stealing is heinous and a crime.
hypfer
8 hours ago
I mean some people (me included) have been begging society to pull that plug since over 10 years now.
The plug being "the cloud" and "hooking everything up to the same internet".
These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live. This was wrong even before LLMs.
amelius
4 hours ago
It won't be long until we can ask it to replicate iOS or MacOS.
Apple needs to up its game, fast.
varispeed
4 hours ago
It won't be long until corporations will start paying AI companies to block such requests or even lobby governments to make it illegal to do.
pjmlp
8 hours ago
Imagine the power of this magic superpower in the hands of business owners....
haukebri
8 hours ago
Adding to what you both said. I run pixel checks against live websites in a real browser, and I stopped doing full-page screenshots early on: scrollbars, lazy-loaded images and font timing shift pixels between runs like clockwork.
What works for me is diffing small stable regions, one component or one flow at a time, with a small per-pixel tolerance. I also keep a DOM-level assertion in front of the visual check, so when something fails I already know what changed structurally before I compare two images.
For the game port it's genuinely harder, the renderer decides the pixels, not the test. But the region idea transfers: lock the viewport and only diff what's stable.
7bit
5 hours ago
You must be fun at parties
magicalhippo
3 hours ago
Actually Claude just did something similar for me, as I'm working on something else with Quake.
On its own it decided to do demo playbacks and take periodic snapshots, and compare them pixel by pixel if the PNGs differ.
I was also working on a web-based port but it saw the original Quake code was not modified so it compiled a native version on its own to use for this.
It had to apply a small patch to make the game completely deterministic, but it figured that out on its own by reading the code.
It then asked me to record a demo with various elements, say an explosion or being under water, and visually verified that the screenshots had those elements present.
So now I have a solid set of tests to verify against.
Opus 5.5 Medium. Used at most 1% of the weekly limit of my $20 plan.