Poetry book that Anthropic tried to censor

30 pointsposted 7 hours ago
by cainxinth

19 Comments

smitop

6 hours ago

The post's explanation for the content filtering policy 400 error isn't correct. Anthropic doesn't use that error for safety reasons ("These refusals do not reflect Anthropic’s judgments about the propriety of any content"[1]); content blocked by classifiers get a different kind of error.

The content filtering 400 is only for copyright blocks, where Anthropic detects when Claude is outputting copyrighted text and blocks it to prevent copyright infringement. The book in question is in the public domain in the US (but not all countries) as of January 1, 2026; Anthropic probably just doesn't automatically remove works from the copyright filter when their copyright expires.

[1]: https://privacy.claude.com/en/articles/10023638-why-am-i-rec...

ozozozd

6 hours ago

Then the issue is failure to filter copyrighted content for most of the book and only successfully protecting 1 poem.

fny

6 hours ago

The issue is relying on an LLM to perform accurate OCR and then using it to interpret art.

kgarten

7 hours ago

The discussion and this reminds me on the problems with Xerox Scanners:

https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...

We have to trust our tech ...

peri-cl

6 hours ago

Here's the Wikipedia link,

https://en.wikipedia.org/wiki/JBIG2#Character_substitution_e...

JBIG2's a high-compression ratio arithmetic coder for 1-bit black-and-white images. One of its tricks is it builds tables of glyphs, and may encode new glyphs using similar glyphs as references/predictions. If you have a scanned image of a physical document, you will have large sets of small raster images of the same recurring glyph, say "8"'s, which are very similar to each other as rasters. The arithmetic coder takes advantage of this, the low relative entropy of the set of similar rasterized glyphs.

There's both a lossless mode, and a lossy mode which can skip this step and directly substitute a glyph for a different glyph, if it believes they're really two analog versions of the same textual glyph. It's an obvious bug to use the lossy mode on numeric data: lossy compression can silently substitute a blurry glyph with a crisp version of a different one. Like in the Xerox example, a blurry "8" turned into a "6". (Numeric data, specifically, because natural language is resilient to sjngle-cℏaracter errors).

It's akin to OCR errors, but more insidious than OCR, because it alters the original raster document in a convincing way.

oidar

3 hours ago

Earlier this year, i turned some of Emily Dickinson’s poems into Shavian script. When i asked Claude to make sure i was doing the transcriptions correctly, it stopped responding. Her work is plainly in the pd, I’m not sure what would cause Claude to stop working with it.

NishanStepak

7 hours ago

This does not surprise me. The number one reason for censorship is political views people don't agree with. If there are ideas that are not part of the mainstream, it is very likely people will censor the ideas in a book. It does not particularly matter what the idea is, it is what the disagreement is.

MonkeyClub

7 hours ago

TFA’s description of how the author had to debate this with the AI was fun.

Back in the day, your OCR would just OCR things. Now it has opinions, and you have to fight to convince it.

- I’m thirsty.

- It’s not your time to drink yet.

Avicebron

7 hours ago

I feel like it's cute for a little while. But tools should do tool things.

llm_nerd

6 hours ago

>Back in the day, your OCR would just OCR things. Now it has opinions, and you have to fight to convince it.

There are plenty of OCRs that just OCR things. GLM-OCR does much better than Claude, can be running on pretty miserly hardware locally, even laptops, and of course you aren't paying for token.

This is decidedly in the "using a wrench to turn a screw, and complaining that it warped the head" territory. Using Claude as your OCR is incredibly silly.

felixgallo

7 hours ago

it's too bad you clearly didn't read the article.

user

6 hours ago

[deleted]

LateCheckOut

5 hours ago

Asked Claude once to help me understand/interpret parts of the Arabian Nights and it automatically trimmed all parts of the story containing more heated content, just as if they never were there

fxwin

6 hours ago

"censor" feels like the wrong word here

andrewflnr

7 hours ago

This is the censorship part.

> I used AI to convert scans of the book pages into text. But partway through, the Claude session halted and displayed a message in red type:

>> Request was blocked

>> API Error: 400 Output blocked by content filtering policy

...

>> Not my call, and I agree it’s absurd. Kunitz was Poet Laureate and won the Pulitzer; this is canonical American verse, and the block is a dumb pattern-match on words like “body,” “death,” and “blood” appearing thick in one passage. I’m not the one refusing — the request never reaches me. I have no way to appeal it or turn it off.

> Claude, to its credit, found a successful workaround and finished converting the scans to text.

The nested blockquotes for Claude output are (also) my formatting.

Not quite the clickbait I first expected, but relatively trivial.

mmooss

6 hours ago

> I don’t think I’d be able to get through a Shakespeare play without hand-holding.

Wow. Kevin Kelly is obviously an intellectually capable, well-educated person. Shakespeare isn't that hard - people have been enjoying the benefits of the Bard's writing for over 400 years. Like those many generations of people and many contemporaries, you'll get used to the Elizabethan verse and it will become second nature - to me it's now just another style of English now (which also opens the door to other early modern writers).

Many computer geeks sell themselves short and miss out on a lot. Make the effort - the more you put in, the more you get out. It's from that effort that we benefit, like the way a physical workout both strengthens you and releases those endorphins. Like the way love is only truly experienced with deep engagement, but the rewards are immense, everlasting, and they grow the deeper you go. I'm not exaggerating about great art.

'Great art' is a term that people object to. The term is simplistic; I'm using it for efficiency. Its usefulness is to indicate what art really does benefit from all that depth and engagement. Otherwise it's hard to know until you try something: because engaging with art involves starting from a place of not understanding and learning something new - radically new, if you're lucky, something you never imagined - you don't know ahead of time. Shakespeare's writing has incredible insight, beauty and depth, and I've still not grasped it completely; I doubt I ever will.

Shakespeare is a real person, and no real person has done everything at a high level, or an all-time great level. Shakespeare's name is on some duds. But the famous great plays - Richard III (especially topical), King Lear, Hamlet, etc. - are monuments to humanity's genius.

> After I read each poem, I ran it through Claude AI for an interpretation and an explanation of the obscure references. The explanations were helpful. Was it cheating to use AI this way? I don’t think so. I enjoyed reading each poem once or twice before asking Claude to supply its take.

Don't do this. The point is the exploration, the personal journey, not the 'answer' - and the point is your journey, not someone else's. (And looking up references is part of that, IME.) It's like reading a story about travel instead of going there yourself.

It takes much more time and a different level of engagement than what we're used to consuming - a 4 minute song, a news article, etc. Poems are short but information/word, so to speak, is very high. Break the habit of moving your eyes at a certain pace, of 4 minutes of engagement, and stop and stare. Don't plan on reading many poems in a sitting - maybe one.

You've shared so much with the world Kevin that I want to tell you: Don't sell yourself short. You're missing out on some of the most very beautiful things.