Surely the models, when well directed, can get most or all of the way there in short order.
"Hey ChatGPT, what would the sufficiently smart compiler argument look like in the LLM era?"
I did this a few months ago and it's not simple: different sensors have different pixel arrangements, files can have different data structures, and the default tone mapping is provided by the app (so opening the same raw file in darktable vs preview will have different results). There are different algorithms for how you assemble the raw data as well, and the issue is that you can get something _passable_ that isn't _optimal_, and it's really hard for the LLM to recognize that something isn't right. Is it the pixel processing, or is it the tone mapping, or is that just how the image looked?
And on top of that, you need a bunch of different files from different sensors to test with. Canon, sony, apple, fuji, nikon, olympus, etc. -- if you're building a truly general application, that's a concern.
Or if you're like me and you shoot fuji, you just build a fuji raw parser and call it good. The "TAM of 1" world of software development...