simonw
3 hours ago
The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.
Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":
https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f
Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
silasdavis
2 minutes ago
https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
> Aside from reading identically forwards and backwards down to the letter
No it doesn't.
simonw
2 hours ago
Here's quite an impressive follow-up. I have a tool which knows how to render Markdown documents with embedded SVG content - I use it for the pelican test.
Since this transcript has HTML in it, I decided to upgrade that tool to also render HTML.
I set Gemini 3.8 Flash the task, using my own VERY shonky coding agent tool (llm-coding-agent) - and it did a solid job.
So now you can see the "cool thing in html" rendered within the Markdown document using code that Gemini 3.8 Flash also wrote: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Transcript where it built that is here: https://gist.github.com/simonw/3e36b98292dfdc1b3baff158faa74...
badlucklottery
an hour ago
Definitely cool.
I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.
noir_lord
an hour ago
That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated".
Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.
senordevnyc
7 minutes ago
Easy, have another agent check it.
Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues.
The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.
kridsdale1
an hour ago
The new Bench-Maxxing!
trvz
an hour ago
Try turning the sound on, off, on again — not impressed by this bugginess.
w4zz
13 minutes ago
I suggest you fork it to improve
heliosAtwork
2 hours ago
Focus on speed and being OK with temporarily being #3/4 in intelligence might be the counterintuitive approach which makes Google win long term (whether accidentally or strategically). Can't wait to try Gemini Pro later this year!
bermudi
an hour ago
I honestly can't believe serious people are making this argument on a straight face.
Gemini 3.7 flash outputs so many tokens per answer it doesn't matter how fast its TPS is, sol will end up being both cheaper and faster than Gemini. So ppl are paying more for a given task, waiting longer and using a dumber intelligence because "TPS number shiny".
Gemini 3.8 outputs 11k more tokens PER TASK on average in AAII than 3.7 putting it dead last in output tokens per task in the leaderboard.
gundmc
an hour ago
There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them
https://artificialanalysis.ai/#cost-tabs
That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.
WarmWash
an hour ago
AA isn't the only benchmark
hglaser
3 hours ago
I saw your username, clicked the link without reading, and was very confused to see a cosmic vortex and not a pelican.
simonw
3 hours ago
Hah, the pelican is in this other comment: https://news.ycombinator.com/item?id=49537553#49538217
giancarlostoro
3 hours ago
> this thing is really good at HTML JavaScript is pretty exciting.
I would hope the people who make one of the most used JS engines in the world are capable of making a model good at JavaScript ;)
ericol
an hour ago
OK, but what about a pelican in a bycicle.
estetlinus
an hour ago
LLM: produces a toolbox of an id and a clock
User: use them both
Made me giggle.
jauntywundrkind
an hour ago
it's such a weird split how most AI companies are trying to be the best, but Google really has a different mission statement. they already have users. lots of users. they need to be working on building models they can deploy and use with the most number of people, as they already have the users.
i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-line with that split-of-focus.
pietz
3 hours ago
Mission accomplished. That's both cool and fast.
wayeq
2 hours ago
> That's both cool and fast.
and probably a barely modified knock-off of some github project that it trained on
sawjet
2 hours ago
You're so upset that you have to invent an imaginary hypothesis to make yourself feel better.
superze
2 hours ago
Yes, very imaginary to think that the code comes from pretrained data and copy pasting whole blocks. It's not like this is exactly how LLMs work.
snet0
2 hours ago
Correct, that's not how LLMs work [0].
simonw
an hour ago
Are you a frequent user of LLMs? That "copy pasting whole blocks" mental model doesn't hold up to regular usage, in my opinion.
ChickeNES
28 minutes ago
They probably saw that report years ago of copilot dumping out the fast inverse sqrt function, and assume that's all they can do. From experience most anti-LLM people have either never used them, or used them back in the 3.5-4 era and then never again, though you might have even more experience with those people than I do. :P
whateveracct
20 minutes ago
they're wrong but they are right that this isn't interesting
slopinthebag
an hour ago
The bar could not be any lower these days I guess