Settings

Theme

Who does Anubis actually stop?

fzakaria.com

52 points by type0 5 days ago · 83 comments

Reader

jdlshore 5 days ago

Hmm, I don’t think I agree. The author is claiming that Anubis is meant to stop individuals using LLMs, but I don’t think that’s its purpose. I believe its purpose is to reduce mass scraping of data that puts excessive load on systems. It does that by increasing the cost of scraping, not by preventing it entirely.

Specifically, bad actors were ignoring robots.txt and rotating IPs to make blocking difficult. Anubis serves to make new connections more expensive, so that people will reuse a connection/cookie, and then falls back to normal means of preventing bad actors.

(That mass scraping is the result of AI training companies, yes, but it’s the mass scraping that’s the problem, not the LLMs.)

  • TomatoCo 5 days ago

    There's plenty of arguments that mass scrapers have compute to spare but it seems to me that if Anubis makes it 100x more expensive to scrape then, for any given scraping budget, that means you get scraped 100x less. Which is the difference between your server buckling under the load or continuing to serve reliably.

    • VladVladikoff 5 days ago

      In my experience, not really. Each asset is made by a new proxy, eg some TV somewhere, they send every new request via a new IP address, which is a new machine, and it doesn’t matter how long the response takes, they have already rotated to the next ip immediately and sent another request. Most of these providers have millions of IPs, and it generally doesn’t take millions of requests to scrape a website (unless it’s really big!)

      • inigyou 5 days ago

        And each new IP triggers a new Anubis challenge. Isn't it great?

        • microtonal 4 days ago

          So, once the challenges are solved by the residential proxies (with optimized native implementations [1]). Some unsuspecting TV owner is paying with their electricity bill (not the companies doing the scraping) and we all suffer through challenges.

          Yes, great outcome! </s>

          [1] https://lock.cmpxchg8b.com/anubis.html

          • inigyou 4 days ago

            Proxy networks just proxy. They don't let their client run compute on the proxy. That isn't even a possibility in socks5 protocol.

    • dullcrisp 5 days ago

      If compute isn’t the bottleneck then it wouldn’t make it 100x more expensive to scrape. Of course if it does, then it’ll be effective.

  • inigyou 5 days ago

    It does it by blocking the scrapers because they are really stupid scrapers. go-away blocks them too, by seeing if they load images, a much quicker test.

  • jsnell 5 days ago

    Except it does not actually increase the cost of scraping meaningfully. Compute is really cheap. The compute for minting an Anubis cookie will cost less than a thousandth of a cent even assuming the attacker uses the same JS implementation of proof of work rather than an optimized native implementation. That cookie can then be used for hundreds of requests.

    How big of a deterrent is a millionth of a cent per page going to be? A non-existent one. Just the bandwidth from a residential proxy will cost the scraper orders of magnitude more.

    Proof of work just isn't a viable counter-abuse challenge, even for something as low-yield as scraping. (It might be economically viable as a counter against some types of DDoS attacks, since they're even lower yield. But in practice it just moves the attack surface to the proof of work validation service.)

    • xboxnolifes 5 days ago

      Your theoretical counterargument falls apart by the reality of just putting up anubis and comparing the before and after. I don't understand why this argument shows up in every thread about anubis. There are plenty of people and orgs who have empirical before and after results. We don't need theoretical arguments when there exists actual data.

      • jsnell 5 days ago

        The exact claim I was replying to "It does that by increasing the cost of scraping, not by preventing it entirely." In that context, quantifying the cost isn't theoretical nitpicking!

        Anubis works to the extent it does because in counter-abuse security by obscurity tends to be the best security of them all. Right now Anubis isn't popular enough for many scrapers to try working around, so they don't. But they will pour resources into working around Turnstile and Recapcha. In this regime, the exact challenge is irrelevant. Anubis would be just as effective if the challenge was running a javascript function to add two numbers instead of a proof of work.

        Proof of work is just a uniquely dismal basis for a counter-abuse challenge.

        • margalabargala 5 days ago

          Proof of work is fine if you can make it demonstrably take a few seconds. If 100 IPs are scraping you as fast as they can, making every 100ms request take 2.1 seconds means you're getting hit with 20x less traffic per unit time.

      • microtonal 4 days ago

        falls apart by the reality of just putting up anubis and comparing the before and after. [...] We don't need theoretical arguments when there exists actual data.

        And the data says it stopped working for some sites that actually have a lot of data.

        We apologize for a period of extreme slowness today. The army of AI crawlers just leveled up and hit us very badly. [...] It seems like the AI crawlers learned how to solve the Anubis challenges. [...] However, we can confirm that at least Huawei networks now send the challenge responses and they actually do seem to take a few seconds to actually compute the answers. It looks plausible, so we assume that AI crawlers leveled up their computing power to emulate more of real browser behaviour to bypass the diversity of challenges that platform enabled to avoid the bot army.

        https://social.anoxinon.de/@Codeberg/115033790447125787

      • inigyou 5 days ago

        It's not because of the PoW though, it's just because the scraper doesn't run JavaScript. The alternative package called go-away does these tests without the PoW.

      • TylerE 5 days ago

        Captchas worked for a few years too. Then computers got better at solving them than humans, and we're still dealing with the fucking useless things decades later.

    • ragall 5 days ago

      > How big of a deterrent is a millionth of a cent per page going to be? A non-existent one.

      In practice, this would seem to be false, according to the admins of said sites.

    • kstrauser 4 days ago

      Excellent! When I put Anubis in front of my home Forgejo server, it blocked about 600,000 IO-expensive queries a day. I’d love to think it’s costing some moron $2000 to scrape my site the most idiotic way possible instead of just running git clone and analyzing it to their heart’s content.

  • ddtaylor 5 days ago

    There is no distinction between the two. If one can do it the other can just as easily. I know multiple people right now that can scrape any a site very easily.

    I do this at scale and it took me very little time to set up and almost no resistance. So anyone who thinks that this is difficult or you're preventing people from doing this at scale, you're wrong.

novafunc 5 days ago

Anubis's primary goal is to prevent web scrapers from DDoSing a website. It's not meant to be an unbeatable challenge or only allow humans like Google's more privacy-invasive captchas.

You do the proof of work, you get the content. Not all web scrapers are willing to do the work, which reduces the strain put on web servers.

It's by no means a perfect system. It's goals in part prevent it from doing so. It tries to not be too annoying for humans, to not block real users, and not be privacy invasive.

  • what 5 days ago

    >tries not to be annoying

    The little anime girl is pretty off putting. I bounce when I see it.

    • arn3n 5 days ago

      It’s meant to be both funny AND highly unprofessional; Anubis makes money of licensing a version of the firewall where you can change the image.

      It’s a good strategy; personal websites and blogs can display the anime girl without fear, and companies that care about their image end up paying. Win-win.

      • ssl-3 5 days ago

        It appears in places where that are neither personal websites nor blogs, and that are places where professionals conduct work.

        For instance: The act of searching the Arch Linux wiki produces a picture of the anime girl.

        (Should I just not use Arch professionally?)

        • jdlshore 5 days ago

          If you care that much, you can donate money to Arch to buy a commercial license, or inform your sales rep that you find their conduct unprofessional.

          Or, yes, you can take the presence of the free version as a sign that the professional service is freeloading, and take your business elsewhere.

          • ssl-3 5 days ago

            I see.

            I take this to mean that we have some kind of (perhaps-fundamental) difference in the ways in which we understand how free software works.

            (That's OK, comrade. I'm not here to change you.)

            • jdlshore 4 days ago

              I suspect we do. In my view, people using free software take the gift they were given as-is, and if they don’t like it, they either contribute in a way the giver appreciates, or move on. They certainly don’t whine about their free gift online, and especially don’t complain that their free gift isn’t professional enough for them.

              • ssl-3 4 days ago

                So no bug reports. No wishing that things could be improved in any way.

                All users all must take what they are given. If they do not like what they are given, then they must only contribute in the ways that are prescribed from on-high. No freeloading. No complaining. No discussion, even in the form of a comment on a completely separate, independent forum. If they are unwilling or unable to follow these rules, then they must pound sand.

                Am I on the right track here with the intended restraint?

                • johneth 4 days ago

                  Yes, they have no obligations to you, just as you have no obligations to them (beyond whatever licenses apply to their content).

                  • ssl-3 4 days ago

                    Indeed. They owe me nothing, and I owe them nothing.

                    We are all free to have a slice of the infinitely-divisible cake, and also to talk about it when that behooves us. We can say positive things when it suits us, but we can also say whatever else we wish to as well.

        • solarkraft 4 days ago

          Arch is a largely non-commercial project. You’re free to not use it, if that’s where you draw your personal line. You must be a really principled person then, who no doubt also won’t tolerate much worse offenses.

        • pooploop64 4 days ago

          I don't see what's so offensive about it. If it had huge boobs or something that would be one thing, but it's just a regular mascot character. Out of all the "necessary evils" that get inserted to keep a service above water, that anime girl image is the least offensive one I can think of. It's still completely normalized to show ads for porn on sites like youtube. Captchas are an insult, with the "photo challenge" type being a full on slap in the face most of the time. I would look at 1000 anime girls to avoid one street sign captcha.

          • ssl-3 4 days ago

            I don't know what it is, either. But it creeps me out and makes me feel dirty every time it shows up.

            It does this in ways that a boring corpo captcha or Cloudflare prompt do not, so it's presence of this artwork more than it is the temporary time-wasting impediment that seems to do it.

            People have irrational reactions to things sometimes, and I am people.

      • what 5 days ago

        There’s nothing funny about it? Also seems like a lose for the personal websites and blogs when people bounce because of it.

        • kstrauser 4 days ago

          Anubis’s silly, harmless images do a good job of driving away the traffic I don’t want on my personal site. Anyone who couldn’t abide seeing a cartoon for half a second wouldn’t likely be fun to interact with.

    • himata4113 5 days ago

      People are mysterious indeed. Letting a widely accepted png cause that much discomfort sure is interesting.

yellow_lead 5 days ago

> The exact adversary Anubis targets defeats it trivially.

Wrong, anubis stops mass crawling of web pages, by requiring a proof of work, which makes accessing these websites more expensive.

Anubis was not built to stop individual users with llms.

  • drum55 5 days ago

    Claude Opus can make an optimized version of the proof of work solver that’s more than 10000x faster than the javascript one, in about 5 minutes time. Who is this stopping exactly? It’s the wrong tool in the wrong place.

    • Macha 5 days ago

      Practically, the people indiscriminately scraping don't bother to do the work to bypass it or implement the POW test, which results in reduced CPU load for all the properties that were having trouble with scrapers before. Until scrapers start implementing it en masse then, it still serves its purpose.

      • greyface- 5 days ago

        > Until scrapers start implementing it en masse

        If it becomes widespread (as it has been doing), they will. Anubis' strategy only works while it remains a niche approach only adopted by a small number of sites.

        • Macha 5 days ago

          And then people will move to a new solution.

          I think this is a pragmatist vs idealist debate. The pragmatic answer is that Anubis solves a problem now, and so it will be used until either it doesn't or a better solution presents itself. The idealist approach is that it's obvious that there's ways to get around Anubis, and so some people argue from there that it shouldn't be used. But the only other alternatives being offered are to either eat up the costs (in server resources or engineering time), or to go behind cloudflare with its own tradeoffs.

          • greyface- 5 days ago

            Anubis is the penicillin of web hosting.

            Every individual practitioner has a strong incentive to overuse it, because it's extremely effective for them individually. Every additional practitioner that uses it increases the selection pressure on their collective adversary to develop resistance. Eventually, it reaches a tipping point and becomes ineffective. But the ecosystem-level impact of its historical use remains, and makes things worse for everyone.

            Recently, the baseline expectations of the Web have shifted, and I need to enable JavaScript to read a non-CloudFlare-proxied simple HTML site like LKML. I do not expect this shift will revert once Anubis outlives its effectiveness.

            • Macha 5 days ago

              Viruses do not have motivations, so it’s hard to assign blame to them. The companies running badly behaving AI scrapers are run by people with more of a mind than viruses, so please direct your complaints about the second order effects of their actions that way, rather than on their direct victims.

            • Krutonium 5 days ago

              I do want to note that my website has Anubis, and does not require Javascript. You can configure Anubis to test in other ways, JS is just the default.

              • greyface- 5 days ago

                Thank you. I decrement my Anubis-ire counter by 1 whenever I encounter a site configured like this.

        • drum55 5 days ago

          The barrier is a two sentence prompt to Claude to add a patch to curl that does 500MH/s to the 50kh/s the javascript version does on the same hardware, it only works because it’s obscure and largely irrelevant.

        • inigyou 5 days ago

          Then they'll change what Anubis does.

      • PunchyHamster 5 days ago

        That was true but with AI that could be automated pretty easily. Sure, not worth for random blog but random blog won't get the traffic anyway

    • Kuinox 5 days ago

      The difficulty is increased when the server get heavy traffic.

      • drum55 5 days ago

        Then the only people who can use it are the ones with a GPU implementation.

        • Kuinox 5 days ago

          No, the value of scrapping the page is long gone and the scrapper would have moved on. The users would still access the page by waiting a bit.

nitwit005 5 days ago

This mistakes the goal. It's not to block the scrapers, but to discourage excessive (and costly) scraping.

The Anubis cookies are bound to particular IP address. The scrapers are often using a large set of IP addresses, so they'll be paying a far higher cost than this suggests.

  • ssl-3 5 days ago

    There are multiple goals at play. They're easy to find in HN comment sections whenever these topics arise.

    One goal is to reduce excessive scraping, usually for monetary or performance reasons. This is the goal you mention, and is motivated by a desire keeping the thing working at all.

    Another goal is to stop bots from ingesting the content, carte blanche. This goal is motivated by a desire to dictate how bots (and by extension, people) may use the information that is otherwise freely-available on the web.

    These are not the same goals.

  • cedws 4 days ago

    The cost is nothing. The browser implementation is too slow, mobile devices are too slow, and hash algorithms with hardware acceleration are too fast. You can’t balance these three constraints in a way that only keeps out the bad guys.

    The hashing is just elaborate obfuscation. Anubis uses SHA256 which isn’t ASIC-resistant, and thanks to Bitcoin you could probably buy one off the shelf.

    • eqvinox 4 days ago

      and yet, it works.

      • cedws 3 days ago

        The hashing has nothing to do with it. It's just an arbitrary hoop for clients to jump through that filters out the bots that haven't implemented that hoop. So why not just use the client's fingerprint and call it a day? Use their canvas or JA3 fingerprint and I'm willing to bet it would be just as effective.

        • eqvinox 3 days ago

          It's impossible to know without trying, but I suspect those checks would be too easy to bypass/fake in a crawler.

          You can submit a PR and then we can try…

          • cedws 3 days ago

            Most scrapers are unsophisticated, this is why Anubis "works." For this same reason, basic fingerprinting already employed by Cloudflare et al would have the same efficacy.

throw0101d 5 days ago

I was curious about the name:

> Anubis is a Web AI Firewall Utility that weighs the soul of your connection[1] using one or more challenges in order to protect upstream resources from scraper bots.

* https://anubis.techaro.lol/docs/

> The Weighing of the Heart would take place in Duat (the Underworld), in which the dead were judged by Anubis, using a feather, representing Ma'at, the goddess of truth and justice responsible for maintaining order in the universe. The heart was the seat of the life-spirit (ka). Hearts heavier than the feather of Ma'at were rejected and eaten by Ammit, the Devourer of Souls.

* https://en.wikipedia.org/wiki/Weighing_of_souls#Ancient_Egyp...

  • dlcarrier 4 days ago

    I'm not familiar with the firewall utility, and my initial thought, in response to the question, was "anyone who's heart weighs more than a feather".

  • aboardRat4 5 days ago

    >>I was curious about the name:

    That's knowledge usually learnt in primary school.

    • throw0101d 5 days ago

      > That's knowledge usually learnt in primary school.

      Anubis may be known generally as an Ancient Egyptian god, but what he was a god of specifically is a little more obscure. I.e., what 'trait' of Anubis led the developers to choose that name?

      I would have thought of something 'defensive', like a shield; to use a Greek example:

      * https://en.wikipedia.org/wiki/Aegis

      The association of weighing (human(!)) souls is not something that would have come to my mind.

    • theshackleford 5 days ago

      > That's knowledge usually learnt in primary school.

      Perhaps where you reside, I’m unsure why you would believe it to be universal.

    • Terr_ 5 days ago

      Grades 1-5 is a little hyperbolic, IIRC my "World Religions" class was probably middle/junior-high school.

      In any case, it's not the type of lifetime common-knowledge which warrants your scornful response.

    • Rendello 5 days ago
    • j-bos 5 days ago

      Or scholastic book fairs (showing my age)

    • ThrowawayTestr 5 days ago

      I mean I know about the heart weighing thing but I'm pretty sure I didn't learn it in school.

object-a 5 days ago

Maybe free market principles apply here: if Anubis fails to reduce scraper/bot load on servers, or blocks too many desired users, then the sites that adopt it would probably scrap it.

If they’re keeping it even after all these posts, it must be stopping _some_ sort of undesirable traffic without costing too much desirable traffic.

satvikpendem 5 days ago

Exactly, it's the same as Cloudflare captchas where only certain blessed devices and browsers can seem to actually pass it. Ironically, adding Anubis accelerates the death of the open web.

  • grim_io 5 days ago

    Open to whom?

    There can't be a truly open web if the guys with all the resources in the world have the incentive to absolutely crush you by draining all of your resources.

    They won't crush you out of malice, but by accident, like an ant.

  • nitwit005 5 days ago

    Cloudflare got it's early business from sites being taken down by DDOS attacks. The early users of Anubis were often sites taken down by scrapers.

    Not a fan of either solution, but the sites simply going offline does not promote an "open web" either.

  • pibaker 5 days ago

    A world without cloudflare does not necessarily mean a world where anyone can access any website. It could very well be a world where anyone can point a DDOS at any website and make it inaccessible to everyone.

  • dlcarrier 4 days ago

    It's getting to the point that the only way to get past the bot filters is to have bots do everything.

  • mike_hock 5 days ago

    No, this is not the same as Cloudflare fascism. You're free to use any JS runtime to solve the challenge.

cjd8 5 days ago

Funny, I'm working on a simple tool that pulls the atom feed of the latest patches from lore with Python, and I just slip a User-Agent header into the requests.get() and it works great.

inigyou 5 days ago

We don't know who is doing the really dumb global scraping attack, but that's who it's meant to stop.

patchtopic 5 days ago

perhaps the author, instead of the "I'm so clever" theoretical arguments from the client side, actually implemented Anubis on the server side and observe the results.

I have set it up on a few sites being relentlessly hammered by clearly idiotic bot traffic, and it drops bot the traffic levels from insane to manageable. I don't even care if the traffic is AI or bots, if the bot has gone to the same effort as the author has the bot may even just access the site in a responsible manner and that's fine.

PunchyHamster 5 days ago

Well, my call on this thing being useless waste of time of everyone involved was correct. Compute wasted on AI tokens alone is probably far more than some cpu for token solving, better invest time into caching it well (or blocking agentic traffic entirely if that's your jam)

m463 5 days ago

Anubis does not stop me from browsing a site (not scraping, browsing with a human at the wheel)

On the other hand, Cloudflare and the others stop me dead. "enable javascript and cookies." Ok. "your browser is too old". (I have an old OS with the newest firefox esr that supports it).

sigh.

  • nonamesleft 4 days ago

    Anubis also forces cookies and javascript, which deters me from many of the sites that use it that require neither.

lrem 5 days ago

Uh wait, a cookie that can be reused for a time to fetch the rest of the content? That's the exact opposite of what I'd intuitively want for this system...

  • recursivecaveat 5 days ago

    If you behave yourself and reuse the same cookie/IP instead of trying to hide your identity, other tools can be used to block you if your request volume is crazy. It also means regular people only have to solve once.

seba_dos1 4 days ago

The most interesting property of Anubis is how much content like this that's comically missing the point it inspires.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection