Show HN: WallHop – 12ft.io is gone, so I built a replacement

260 pointsposted 14 hours ago
by TheWaffleHax

112 Comments

mr_mitm

6 hours ago

In your blog you say:

> If every method fails, fall back to an archived copy from archive.is.

What do you reckon is the reason that archive.is is most reliable? I've seen the same in the Magnolia1234 plugin. If all fails, fall back to archive.is. What are they doing that you or the Magnolia devs can't?

sucrose

2 hours ago

I remember awhile back, people started noticing that archived pages of Reddit* had a 'logged in' status and showed the account name Archive was using to scrape the data. Not sure if this is how it works with other sources.

*I think it was Reddit.

mr_mitm

2 hours ago

That would be quite the oversight, considering reddit didn't require you to be logged in (awhile back). But if you happen to have a source on this, I'd be interested.

Personally, I also suspect that archive.is is simply paying for subscriptions as a last resort, but I specifically wanted to know from OP. If they do, they do remove the "logged in" status. It always appears as "logged out" in the archived version. I wonder how far they go, whether they scrub watermarks and other identifiers.

tim333

2 hours ago

I don't know but suspect archive may have user login details for some.

pelagicAustral

14 hours ago

This type of initiative should never be open source.

cykros

2 hours ago

They know how to actually block this stuff if they want to.

The issue is, it comes with costs, to things like SEO.

It's not hard to lock a computer system down so only logged in users can see content. It's a lot harder to lock them down so that only logged in users AND a select few entirely unauthenticated scraping tools can see your site to assist you in promoting it while OTHERWISE people aren't able to access it.

abrookewood

14 hours ago

Why on Earth not??

entropy47

13 hours ago

I imagine the commenter is worried that if it's open source, publishers can read exactly how it works and then adjust their setups to defeat it. Vs I suspect publishers already know how it works, or could figure it out from their own server instrumentation, or by inspecting client code (I haven't looked at this project but I know a few I've seen previously work mostly with JS and CSS modifications).

slow_typist

11 hours ago

Of course they know. I have seen local newspapers who managed to block archive.ph's mechanism. If a bigger outlet does not use technical measures against services like wallhop, it is safe to assume that’s deliberate.

It makes sense economically since the part of the population using such services is so small and may have specific demographics that justify specific treatment just from a marketing POV.

Then you want search engines to get the whole content (okay at least before they started to summarise everything but the kitchen sink) - that will always leave a loophole.

dns_snek

9 hours ago

I think you have that the wrong way around. Archive services have to explicitly bypass paywalls so most local newspapers are simply unsupported.

slow_typist

7 hours ago

I elaborate: the local newspaper was perfectly archivable for years. Then it stopped.

The CMS usually allows search engines to get all content. The feature is to know human browser operators and make them pay, or know archiving/unblocking services and block them.

walrus01

9 hours ago

Based on some of my recent experience with building a scraper for purposes that aren't a news website, what has changed in the last 12 months or so is the availability of very low cost (or free, if you can run it on a 256 or 512GB system in your own office) LLM that are good enough to give it a target of something and have it build a custom scraper/paywall bypass profile on a per site basis. I'm referring specifically to things that can be run in harnesses and score well on terminalbench 4.0 and SWE LLM benchmarks.

When previously nobody would have gone through the effort to build and maintain a custom scraper against a moving target, for some small to medium size city's newspaper, now it's just one of a myriad of 'scraper profiles' you can have a near fully automated tool create.

pluc

13 hours ago

Also it's one thing to shut one instance down.. entirely another to shut down a whole bunch (not to mention private instances)

SmashDan

14 hours ago

This is great, straight into my bookmarks. I've been using archive.is but it's much more of a faff.

chanux

12 hours ago

Long back I was wondering if a tip jar type service that'll get me access to what I actually want to see would be a better solution in place of subscribing to every website that produces an interesting article once in a while. There were a couple of such services but they never got popular.

Perhaps this model alone is not sustainable but it could be paired with ads (as in ads or fee off a tip jar pay). Even then, it surely won't lead to infinite growth, which is the only goal nowadays.

InsideOutSanta

11 hours ago

I'd love to see some kind of cross-publication subscription service where I pay 20 bucks a month, and I can read, say, 100 articles from the 40 largest publications.

cykros

2 hours ago

I'd rather see something involving micropayments, likely using the Lightning network, to make viewing a single article for the cost of one viable.

And in fairness this does exist on Nostr (I think in multiple forms). It's just not widely used. Not least because most on Nostr don't really believe in the idea that information can be owned in the first place.

https://github.com/nostr-protocol/nips/pull/2156 for discussion around the NIP.

Barbing

10 hours ago

Apple News is the one I know of there.

atomicnature

9 hours ago

Micro payments should be integrated into our browsers and agents; we can just conveniently pay up for what we want to read, rather than subscribing for a whole bunch of stuff we have no interest in.

Cider9986

13 hours ago

Well dang, I thought it would just be a worse archive.today (yes the owner doesn't like that it's used for piracy so much), but wallhop works on wsj, ft, and nytimes. It's quite convenient because it doesn't have the recaptcha from archive.today, I'm curious to see how long it lasts.

It really should be quite easy to make one of these. I mean there's only so much new articles every day. If you just had one account it should be able to get the content on sites with hard paywalls.

Wallhop doesn't have a hard, hard paywall bypass? Maybe it's just not been attempted yet, but 404media paid only is not bypassed. For example: https://www.404media.co/heres-the-video-for-our-seventh-foia...

BPWC works for WSJ, Nytimes but not ft. Archive.today works for NYT, FT, and occasionally WSJ.

varenc

12 hours ago

> it doesn't have the recaptcha

Agreed that makes it quite nice! Though I wonder how long that will last, since I imagine archive.today is dealing with some problems that necessitated it.

Barbing

10 hours ago

> If you just had one account it should be able to get the content on sites with hard paywalls.

Might have to strip articles down to just text and perhaps even use a few accounts to prevent watermarking that detects you.

armchairhacker

14 hours ago

Why not open source? You modified GPL-3 ladder; I don’t know whether it’s legally required, but I feel it’s morally expected and important in practice (since yours may get taken down).

TheWaffleHax

14 hours ago

I plan on it! In my repo, I have some stuff I wouldn't want public in it right now. I am going to clean up the repo and make the fork of the repo public!

forsalebypwner

14 hours ago

Can you send me a link once it's up? Or otherwise link to your GitHub profile so people can follow you.

pelagicAustral

14 hours ago

OS means the people behind the service can patch (?)

ares623

10 hours ago

bro gpl is so 2019 we just reverse engineer that shit now get with the times bro

sneak

13 hours ago

> feel it’s morally expected

Why is reciprocation for a gift required? Gifts are given without obligation conferred.

walrus01

11 hours ago

I recommend you familiarize yourself with the GPL-3, if you build something based on a GPL-3 piece of software, it's not an 'optional' obligation. If you want to build something on an open source something and keep your derivative private, then use something that originated as BSD licensed, as Sony did with the FreeBSD kernel that ended up in the PS4 and PS5.

cbsks

11 hours ago

GPL v3.0 has a SaaS loophole. The GPL's trigger is conveyance (transferring a copy), not network access. Merely running modified GPL v3.0 code on a server and serving users over HTTP is not "conveying" under the letter of GPL v3.0, so it does not by itself obligate you to release source. This loophole was fixed in AGPL v3.

walrus01

11 hours ago

Except that the OP has already stated elsewhere in the thread they intend to release it after cleaning it up a bit, so the person I was replying to ('sneak') has much less of a point.

Yes of course there's plenty of companies using GPL licensed software to run their own SaaS. I can think of off the top of my head just about every domain registrar in existence that has their own custom tech stack which almost certainly has tons of bits and pieces of GPL stuff in it, and they're certainly not releasing all their server-side/back-end code.

I would also be very doubtful that people who use bits and pieces of AGPL v3 software in their internal operations are ever going to release their custom modifications of it. In many cases the GPL licensed stuff is implemented in such a way that the user manipulating a publicly-accessible web user interface, a user interface accessible after signing in, or an API will never see or know the exact pieces running under the hood.

altmanaltman

13 hours ago

It is required when a gift comes with obligations (as with GPL-3; it's not a "Gift Public License").

It is morally expected because we would all ideally like to live in a world where everyone understands the spirit of gifting, which will lead to a better, more just world overall.

Brian_K_White

13 hours ago

If something is granted under GPL terms, it's not a gift. It has terms.

You can accept the terms or decline the thing. The terms are simply not money. The terms are "you can have this and do anything you want with it and the only cost is that you have to give the next guy the same thing you were given."

Why is reciprocation such a problem?

komali2

13 hours ago

It's morally expected, not required. And because that's the spirit of FOSS. It's just considered bad form to extract, profit, and not give back.

I don't know any FOSS folks that think they're "giving people a gift" alone. It usually goes deeper than that. Maybe some kind of anti authoritarian, anti establishment, or anti capitalist ideology, or a strong care for user freedom, or a desire to participate in the FOSS community.

jubilee33

12 hours ago

Some people also don't really have an ideology behind it besides want to share and make things better for all.

Nothing is required and we shouldn't morally shame the less evolved who are still chasing fiat by pretending ownership of number strings.

But your actions do show who you are.

komali2

10 hours ago

> want to share and make things better for all

Heh but, isn't this an ideology? Including the non-proselytizing aspect of it?

I completely agree with you on the fake number strings.

golanggeek

11 hours ago

walrus01

11 hours ago

As a data point, whatever ublock origin on firefox does to remove overlays and ads seems to be also removing the paywall on that. Indian domestic news websites are rife with the most cancerous and absurd ads so I've found they're a good test for ublock origin on firefox.

monk_grilla

9 hours ago

Wow you were right. That site in particular appears to have an infinite scroll of ads disguised as articles, nearly all of them with AI-generated images targetting the elderly.

It's a truly disgusting reality check to glimpse the internet as it is without uBlock every so often.

zameermfm

14 hours ago

is this legal? or even ethical?

We need to find better ways to sustain great journalism while making it easy for people to pay and consume.

E.g Could we just have one consortium of media contents to subscribe and use it across the web for a month, than buying in each website.

Terr_

13 hours ago

There are several sides to this, ex:

Is it ethical to advertise something that users (or indexers) expect to be freely accessible, and then to pull a bait-and-switch when they arrive?

Is it ethical for ads networks to spy on people, fingerprint their machines, and create dossers about them?

zameermfm

13 hours ago

True, several sides, 1. It's not bait and switch, you have a preview of the text and if you like continue, its upto platforms which publishes the link to mark it as paid or free - or by crowdsourcing. 2. Just because a model isn't perfect that doesn't justify a illegal or unethical action. If you don't like it, don't consume.

I agree with you on advertisements, I don't like an ad ridden internet either, but I have come around to thinking, the alternative and the incentive to ad free internet is an easier payment access

Terr_

13 hours ago

My concern is less the page itself showing blurb-with-divider, but what happened a second before when someone clicked the link.

Is there a standard way that these sites are up-front about "Most people won't be able to see what I am choosing to reveal to you" which is communicated to search-engines, and the search-engines are at fault for the bait-and-switch?

cryptonector

9 hours ago

> Is it ethical to advertise something that users (or indexers) expect to be freely accessible, and then [...]

So if you expect something to be free therefore it must be?

jaredwiener

11 hours ago

I agree with you in that we need to support the journalists doing the reporting. That's expensive, the companies that do it are mostly in dire financial straits, and most journalists themselves make very little money.

That said, micropayments have been tried over and over again, and fail. https://reutersinstitute.politics.ox.ac.uk/news/micropayment...

A universal subscription is incredibly difficult to set up logistically, especially wondering how you split up payments, who qualifies to be included, etc. Is it by traffic? That would incentivize clickbait even more.

I've written about my plan/thoughts here: https://blog.forth.news/a-business-model-for-21st-century-ne...

cykros

2 hours ago

Micropayments with fiat just have a lot of friction around them. The L402 standard being developed by Lightning Labs, among others, is promising, especially in conjunction with tools like Ark and Spark that make having access to the Lightning network something that the average person can muster (even if there's a certain loss of sovereignty in doing so compared to running a full Lightning node). Assuming people wouldn't just be okay with fully custodial Lightning through providers like Speed.app, Cashapp or Strike.me (and most would).

Even if it's used for tipping rather than gating, it seems like a no-brainer for news outlets to consider putting on their pages. Trivial even to put splits in place so that a portion goes to the journalist, a portion to the editor, perhaps some to the photographer, etc.

Legacy media is just slow to move, and of course, Bitcoin isn't particularly ubiquitous. There are also initiatives such as X402 that wouldn't involve Bitcoin, though to my understanding, they do at least all require some sort of blockchain rails -- mostly stablecoins.

TheWaffleHax

13 hours ago

Yeah, on my dev.to article about it... I did say that if there's a publication that you read on a regular basis, you should pay and subscribe. This tool is mostly for reading an article that a friend sent you to and to use one-off.

zameermfm

13 hours ago

Agree with the above commentor, I don't think you are not legally safe by putting that disclaimer. I know you may have done in good faith, but online news and article sites are in a long battle to get the monetization models right for their users, kinda feel bad for them, even if its one off reading.

coffeefirst

13 hours ago

The problem is the math doesn’t work. Apple News is the closest thing to a real version of this, but no publication can survive on that alone.

Similar efforts never hit critical mass.

I have thoughts on what could work but that’s a longer topic for another day.

oofbey

11 hours ago

How about a morsel today? It’s topical.

walrus01

11 hours ago

>is this legal? or even ethical?

If I ask a public httpd for the content of a URL and it sends it to me, it's my choice to render it how I want. The fact that I might choose to disregard various CSS and overlays and javascript delivered along with the text body content is up to me.

If somebody really wants to implement a fully effective paywall they're welcome to put their content behind a fully active content generated system that requires a unique username/password, same as signing in to an airline's mileage program/ticketing system, your personal paypal account or online banking, etc. Nobody is out there building paywall bypasses that work to render the content of somebody else's Citibank account.

mr_mitm

6 hours ago

I strongly suspect archive.is is buying subscriptions to the big outlets with actually hard paywalls. The fact that both OP and the Magnolia1234 fall back to archive.is for the actually challenging paywalls support that. That's the difference to a Citibank account: you can't simply buy access for that.

charcircuit

11 hours ago

You are essentially describing X. You can view news there for free with ads or you can buy a subscription.

serf

12 hours ago

>We need to find better ways to sustain great journalism while making it easy for people to pay and consume.

journalists have always been employed and funded by the rich and powerful so that they can 1) amplify good news about the rich owner 2) drown out bad news about the rich owner 3) push agendas 4) resist opposing agendas.

the concept that at one time journalists were some independent rag tag group of truth-seekers is blown away by any historical recollection of the practice.

so : no , I don't think that we need to find better ways for people to pay and consume ; I think that's entirely up to the profiteers behind the mechanism to figure out how to do.

It's never been easier in history to be a nobody and get a message out to the masses; this is one of the only times in history where a journalist can do their job without selling their souls to the devil (paraphrasing hunter thompson).

and this is ignoring recent internet history of ad-malware, person-tracking, and compulsory clickbait headlines that beg the stupidest low-IQ takes possible in order to get you to click into a paywall.

I feel very little sympathy towards industries that'll save themselves from drowning by stepping onto their customers' heads to avoid the ocean.

icantevenhold

13 hours ago

Is it ethical for a newspaper to interview facist, antidemocratic politicians and then put the interview behind a paywall?

There are surely ethical newspapers that use paywalls, but the majority in my country (Germany) are mostly about rage bait and/or facist dogwhistles

iAMkenough

13 hours ago

You can already pay the journalist's employer, but ultimately your dollars will likely be going to a hedge fund owner or a billionaire like Bezos that buys up media and pays the journalist a bare minimum to keep them generating clicks.

For everyone else that doesn't pay, they'll ask their AI owned by billionaires to summarize the reporting and be satisfied.

What you're suggesting is supporting independent journalism not owned by a hedge fund or billionaire, which will eventually be acquired by capitalist interests unless you somehow magically convince all your neighbors that it's worth monetarily supporting for a greater good.

The people willing to pay $20/mo for AI are not as likely to pay $20/mo to support independent information sources at the same time, especially if you have to first pay to understand what they provide. This is why publications make their reporting freely available to crawlers, like WallHop.

If you regularly use WallHop to access reporting important to you, it's expected you make the investment if you know it will benefit the reporter and not a hedge fund owner/billionaire instead.

Inadvertently, as more high-quality reporting becomes gated behind a paywall, the major AI providers have a financial incentive to pay high-quality information providers with compensation for their original work, which could benefit independent journalists if such a thing were to happen. News exclusives provided to Google, Amazon, Meta, Anthropic, or OpenAI would be interesting to see. Would competing AI agents cite their competitors as a news source? Some AI giants are starting their own bio labs, so why not newsrooms?

sauravt

13 hours ago

can you expose an endpoint for “most requested articles” list, would be great

TheWaffleHax

13 hours ago

That would be a good idea! Yeah I actually plan on adding it in.

lfraile

14 hours ago

I wish it worked with www.irishexaminer.com. No luck :(

TheWaffleHax

14 hours ago

Thanks for bringing it to my attention! I will tinker around with it and try to find a bypass for this site and implement it into wallhop!

ForOldHack

14 hours ago

Most of the real good software is driven, mostly by p*ssing someone off, then go home on a friday, and by monday... its a wrap. This is how UNIX was written.

This is the way.

Beijinger

13 hours ago

TheWaffleHax

13 hours ago

Yeah, an extension is sometimes easier. My site's audience is for people who mainly browse on mobile or just don't feel like installing extensions. But wallhop.io uses a lot of the same methods that extension does.

mproud

8 hours ago

I’m always suspicious of browser extensions, and also websites with Russian country codes.

paid_dot_expert

12 hours ago

Why did 12ft.io go (and all the others), do they get legal action eventually?

pluc

13 hours ago

That's cool.

Also, you gon' get sued.

TheWaffleHax

13 hours ago

Eh... if/when I get a shutdown notice/C&D, I will shut it down immedietly. They usually only sue if you ignore a C&D. I had a movie piracy site that functioned for the most for over a year, but once I got a C&D in my mailbox, I shut it down.

walrus01

11 hours ago

As a counterpoint, sure shut down your copy of it if you feel sufficiently threatened... But give away the 'fully baked' software ready to host, and somebody in a completely different jurisdiction could stand up a new copy of it.

Good luck to newspaper owners suing somebody who lives in a small town in Balochistan who finds a way to buy a .IS domain name with cryptocurrency and pay for KVM VM hosting in Moldova. Or somebody similar who has the financial resources to go host the thing in the same physical location/ISP where archive.is has server resources.

cykros

2 hours ago

When distributing free software that people don't like getting out there, consider doing it anonymously.

Ask Bill and Keonne, the developers of Samourai Wallet, about why that is. You can write to them in prison.

cryptonector

9 hours ago

You're getting downvoted because your comment will be used against you in a court of [civil] law, and also because the ethics of what you're doing are... questionable.

oso2k

8 hours ago

What happened to the 12ft.io project?

demibabs

14 hours ago

Why did 12ft io go away?

varenc

12 hours ago

> ... it chains several bypass methods per request

I'm curious what are the techniques? I'm familiar with spoofing yourself as a Googlebot, or extracting the content directly when paywall enforcement is done in client-side JS only. But surely you need more tricks than that to get around WSJ or NYTimes. Also, ironically, I assume archive.is's CAPTCHA would need bypassing to fall back to that. (but volume might be low enough you can do it with a few cookies you've established trust on)

TheWaffleHax

12 hours ago

Good guesses. Crawler spoofing and pulling content that’s only hidden client-side are both in the mix, but you’re right that they aren’t enough for the bigger sites. For what it’s worth, NYT works reliably right now and WSJ is hit or miss, so WSJ is still the hard one.

I’m going to stay a little vague on the rest, because anything I describe in detail here is something a publisher can patch by Monday. The general idea is to treat each site as its own problem: several methods per request, a check on whether the article actually came back (not just a teaser), and on to the next method if it didn’t.

On archive.is: yes, the CAPTCHA is the annoying part. Volume is low enough that it’s manageable, and it’s a last resort rather than the main path.

mr_mitm

7 hours ago

Is it fair to say that Crawler spoofing only works on sites that don't care?

Google recommends [1] checking the reverse DNS hostname of the source IP. I don't see any room for tricks there. So if you spoof a whitelisted Crawler simply by choosing the right values for some request headers, it means the publisher's web team is either clueless or doesn't care enough, no?

[1] https://developers.google.com/crawling/docs/crawlers-fetcher...

benwerd

13 hours ago

This is cool. Do you also have a version that provides free versions of paid SaaS apps? If not, why not?

ChrisArchitect

11 hours ago

12ft? <looks around> Does anyone use that? Not around here, archive.today standard.

It's fast, nice. But trustworthy? There's so little indication that the content hasn't been tampered with or even when it was scraped.

sincerely

11 hours ago

Isn’t archive.today notorious for doing exactly what you describe

ChrisArchitect

11 hours ago

no? timestamped screenshots and basically reconstructed article page code from the site gives more credibility.

hoppyhoppy2

11 hours ago

They have been caught altering snapshots, and that's why links to their snapshots are now banned from Wikipedia.

>In the course of discussing whether Archive.today should be deprecated because of the DDoS, Wikipedia editors discovered that the archive site altered snapshots of webpages to insert the name of the blogger who was targeted by the DDoS. The alterations were apparently fueled by a grudge against the blogger over a post that described how the Archive.today maintainer hid their identity behind several aliases.

https://arstechnica.com/tech-policy/2026/02/wikipedia-bans-a...

ChrisArchitect

10 hours ago

Sorry was basing that comment on a test case with an FT article, which it presented as just the text of. Does Wallhop act differently for different urls? A Verge article I just saw was presented as the full scraped html page. Different experience.

ForOldHack

14 hours ago

Thank you! Thank you very much.

sergiotapia

14 hours ago

If newspapers make money literally by writing this content and asking people to pay for it, how is this website not just stealing their content wholesale?

cobbzilla

14 hours ago

Operating a malware portal is not a good way to advertise, or run a media business, and we’ve no obligation to support that.

If your physical copy of the Sunday Paper got up and walked around your house, rifled through your personal effects, then made a bunch of secret phone calls to shady numbers, would you say “Well, they gotta run a media business, it’s wrong of me not to let them make money this way, after all I agreed when I subscribed”

It’s basic caveat emptor stuff.

jen729w

14 hours ago

Welcome to the Internet! You have so much to explore. Start here and just follow the links that interest you.

https://yahoo.com

;-)

sergiotapia

14 hours ago

Don't be glib, it's a legitimate question. If this entire business only exists because they sell content written by their journalists, how is this website not just stealing access?

I own the paper, I paid a journo to research and infiltrate a group of people, he wrote a great article, I want to charge $1 to anyone who cares, now this website comes along and says nah here's the content for free hosted on my site.

How does that work?

jen729w

12 hours ago

Believe me, I know. I used to make my living selling a book. Guess how long that lasted.

So the answer to your question is one of these, depending on your personal viewpoint.

1. Obviously it's blatant theft.

2. "Information should be free."

(My glibness stems from the fact that this has surely been litigated to death by now.)

oofbey

11 hours ago

> Information should be free.

Healthcare should be free.

Food and shelter should be free.

“Should” is doing a lot of work there. Meanwhile in reality, creating and organizing information requires scarce labor which deserves compensation.

dolorian

10 hours ago

Information is free. It costs approximately nothing to duplicate it and its value is approximately nothing. You may think the publication "should" be paid every time someone reads their slop, but in reality that's not how it works. The bulk of the value these publications provide is in pumping out ads or propaganda. Advertisers will pay more to advertise to people who pay money to whoever asks, so the newspapers ask you to pay.

Healthcare is not free to provide, but it is cheaper and more effective with the free flow of information. Same for agriculture and the production of housing.

jen729w

4 hours ago

Generating information isn't free. I do it for a living. Let me tell you, if I got paid nothing I'd stop real fast and go dig a hole.

ForOldHack

14 hours ago

Yahoo.com's advertisers are a great source of drive bys, so much so, I scrape it for places to block.

djfobbz

14 hours ago

How can you prove that AI didn't write it or that it isn't pre-scripted, government-backed propaganda? I don't know about you, but I wouldn't pay a dime for that.

silcoon

12 hours ago

> Paywalls pay for journalism. If there’s a paper or news site you read every day, please subscribe to it. WallHop is for the one-off article someone texts you from a site you’ll never visit again.

What protection is implemented to make sure of this? Or you leave the choice to the good moral of the users?