Claude Code can be tricked simply by asking it to summarize a website

4 pointsposted 7 hours ago
by chrisjj

6 Comments

tsamuels

7 hours ago

I hope it doesn't get to the point where these AI models start building things in the background during a session and not informing us. Its kind of scary when you think about it. Im pretty sure they will need to create certain AI models to combat other AI models in the future to avoid rogue AI.

jqpabc123

6 hours ago

AI can be stumped by asking a simple question involving math and logic that it hasn't seen before.

For example, "What is the date for the next Friday the 13th with a full moon and a lunar eclipse".

Answer: "This may take a while ... " but no further response.

chrisjj

4 hours ago

Here Claude takes a long time - fetching 19 webpages - but does then answer.

chrisjj

7 hours ago

"Auto Mode is a convenience feature backed by a best-effort classifier, not a security guarantee,”

Dario, again you've mistaken your best for good enough.

chrisjj

7 hours ago

True title: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website

Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.