Settings

Theme

Half a Second – a book about the XZ backdoor

half-second.com

45 points by zvr 15 days ago · 53 comments

Reader

this_user 15 days ago

Well, the "About the Author" section should probably just be a link to claude.ai.

  • jllyhill 15 days ago

    For anyone wondering, here's the author confirming it

    https://news.ycombinator.com/item?id=48966159

    And here's their Claude skill for writing

    https://github.com/AdrianMastronardi/bookwright

  • sevg 15 days ago

    Yeah the (annoyingly) excessive use of colons feels like recent Claude models to me.

    Looks like author posted this themselves earlier, and even used Claude for the HN comment:

    https://news.ycombinator.com/item?id=48958457

    • infinite_spin 15 days ago

      I've always used a lot of semi-colons in my writing, especially in my technical writing. So I did a search on this pdf for semi colons, and there are 260 in a 264 page document. I then repeated this for the last book I read, Candide by Voltaire, and there were 507 semi colons in a 189 page book.

      This seems like a witch hunt.

      • rwmj 15 days ago

        I read parts of it, being first hand involved in all this, and it seems like it was written by AI to me. If you used an AI to write it or help with it, why not just admit to it?

        • infinite_spin 15 days ago

          the argument was that use of semi-colons indicates AI usage..

          what do you mean by "being first hand involved in all this"? what part did you take if you don't mind sharing, this was an incredibly interesting turn of events.

          • rwmj 15 days ago

            You'd know that if you'd done any research at all.

            Compare this to Jeff Guo (NPR) or Henry van Dyck (Veritasium) or my contact at the WSJ. They all investigated this story, interviewed key people (more people were interviewed for the Veritasium video than appeared), and fact-checked everything with subject matter experts. The NPR story took about 3 months of work and the Veritasium video took 5 months. I was interviewed in total for over 6 hours across all of them.

            These are real journalists who worked on these stories and they produced novel work and new findings. I have huge respect for them, especially after being involved with it and seeing what work it took to add something new to the story. At the same time they managed to tell that story to a lay audience which is another skill in itself.

            • infinite_spin 15 days ago

              > You'd know that if you'd done any research at all.

              I'm not in the habit of researching the people who respond to me before I ask them about what their role was in something they seem proud of. I thought I was being polite by asking about it.

              That's really cool that you were interviewed, which of the interviews do you think most represents your hand in this? I wouldn't mind checking it out. I also really like Veritasium's channel, that's pretty awesome.

              • bigDinosaur 15 days ago

                In this case all it required was to follow the link to their blog on their HN profile just FYI. Although your question was reasonable either way.

      • jstanley 15 days ago

        I pasted the introduction into Pangram and it said 100% AI.

      • sevg 15 days ago

        You know I said colons right? Not semi colons :)

        • infinite_spin 15 days ago

          ah my mistake, those are a bit uncommon in my writing, I feel like they seem pretentious. i just use the semi-colons because it's closer to the way I pause, and alternate, in my thoughts.

    • progbits 15 days ago

      In last month or two I noticed semicolons getting used where previously it would be em-dash, in both cases excessively and often incorrectly.

      I assumed some of my coworkers added "replace all em-dashes with semicolons" into their CLAUDE.md as a really crappy attempt at hiding their inability to write a single sentence without assistance.

      • tux3 15 days ago

        Come on now, you can't also take away my semicolons. What am I supposed to do now; I barely have any punctuation left.

        • dag100 15 days ago

          Embrace the inner Cormac McCarthy in you and abandon punctuation altogether

        • progbits 14 days ago

          Not taking it away.

          I hate those arguments "it has emdash therefore AI", of course humans also write that way.

          But poor and excessive usage of them is a pretty strong red flag, also together with other tells.

    • eth0up 15 days ago

      "free book" "no paywall and nothing to buy."

      Might this not actually be a reasonable purpose for AI? I can tolerate the quirky style and AI signatures for something honest and free.

      • rwmj 15 days ago

        But prompting an AI adds nothing. Journalism is about research, interviewing the people involved, finding new facts, and then presenting that to your audience in a way they can understand. Getting an AI to write about something is just diluting the signal for future researchers.

        • jersmtpd 15 days ago

          You can have an AI help write it while still having done research yourself, they're not mutually exclusive. I'm not claiming whether he actually did that or not.

          Like, for me personally, I'm terrible at writing, so I'll gladly have some AI throw together a draft, which I then verify and edit.

          • eth0up 15 days ago

            exactly. If I am writing a free book with solid information, well, I am not writing it, because I am not writing free books -- I don't have time. But if AI can arrange my data, thoughts, perspective, and articulate it effectively, then I will let AI write the free book. If the book contains new, or previously unarticulated data and insight, then why not? But I see the very thought is hotly contested here.

      • bandrami 15 days ago

        Why not just post the prompt you used? I have access to LLMs too.

        • rpdillon 15 days ago

          You think if he gave you the "prompt" you'd be able to produce that book?

          Heck, I should put money on this. Start a competition: I give folks a prompt, and tell them to write a book about a subject with AI. I'm the sole judge of quality. Winner gets $2000. Then let's examine the variance in the entries.

          In short: using AI is a skill. It's silly to pretend otherwise.

          • eth0up 14 days ago

            Are you the only reasonable person on HN today? Others are framing this author as some official journalist who must be held to rigorous standards and has violated every rule of the trade, with wanton contempt. The author has failed to produce new data. The author is bad. This seems deranged and extraordinarily unfair. The front page explicitly states the book is for general readers and is non-technical. I do not see anything claiming official original research or boasting academic credentials. No ads. Free. No download registration. Just there. The result of work, tokens, time -- didn't 'just happen'.

            The Gladwell-style of conceptual-synthesis is the foundation for half the popular-science books in the world, and the bedrock for the majority of general-reader material out there. And this book happens to have a truly fascinating and extremely relevant and rich theme at its core, sufficiently rich that any motivated idiot could churn out half-decent content with. Yet it gets flagged, insulted and framed as abuse of AI, a tool specifically designed as a force multiplier and tool. Damned if you do, damned if you don't, and how dare anyone generate something of value with it.

            I sincerely find this all highly suspicious. The only explanation is hypersensitive AI allergy in a highly unstable immune system, or the topic itself is making some nervous and, ironically, insecure...

            I could totally understand disagreement, and complaining. But a comment that simply states acceptance of the book as possibly of value, being flagged? Who not hiding something trounces on a friendly comment to flag it? It's gray already, but that's not dead enough? Kill it til it's dead!? I say this because on a completely different post, I mentioned Jia Tan and several of the very concepts highlighted in this book, and I was flagged for it. I have said some stupid shit around here and not been flagged. But this subject really seems to bring out seriously inspired opponents. One would think, gee wiz, security, potentially many lives or billions of dollars or FOSS at stake -- there's no such thing as too much security, by all means, obsess on it, stay focused and talk about it, there are real world dangers here. But no.

            Something seems fishy.

            Anyway, glad someone had the courage to question the outrage. And I will promptly warn my author friends of the coming witch-hunt and to be very quiet about the tools they're using.

            • rpdillon 14 days ago

              The HN crowd is generally skeptical, but "shallow dismissals", which are specifically called out in the guidelines, are still extremely common. These dismissals are not as clever as they tend to seem to those skimming the comments, but folks have developed a bunch of thought-terminating cliches to empower this: calling anything involving AI "slop" is the most common, but more catchy (but fundamentally incoherent) phrases have also emerged:

              > If you didn't take the time to write it, why should I read it?

              I mean, fundamentally you read things if they bring value. It doesn't actually matter if a human wrote it. It matters if it is correct and useful. One implementation is to have a human review, but you can imagine others. Fundamentally anyone who reads and leverages the output of AI knows this. The phrase seems reasonable because it is a good refutation of folks literally copying and pasting AI output from their ChatGPT session in the comments, which I think everyone agrees isn't good hygiene for online discussion. Because the guideline is "right" in that case, it seems easy to apply everywhere, but I think that's overly broad.

              > I have tokens, too. Just give me the prompt.

              As if everything generated with AI is created as a one-shot with ChatGPT. I feel like half the people writing this stuff have no clue how professionals are using AI. The model, reasoning level, skills, tools, harness all matter, and that's even if the whole thing is a one-shot and we ignore all the server-side variables (inference engine, temperature, top_k, quantization level, etc.). If not, it's the product of a (probably extensive) back-and-forth with an AI, possibly changing models for various subagent calls (Kimi K2.7 Code for the plan, MiniMax M2.7 for implementation, Deepseek v4 Flash as advisor, etc.), etc. Orchestration is as important as prompt and context, but the space has a lot of dimensions. This is why you get radically opposed takes on models, harnesses, and the even the utility of AI: everybody is doing different stuff, solving different problems, and reporting results all over the spectrum. There was a guy just today on HN claiming Opus 4.8 was terrible at agentic development, for example, with half a dozen responses asking how s/he could possibly say that.

              Anyway, my post was controversial, but overall negative on points. I just don't understand how anyone can actually use these tools and then post in such a black-and-white way about AI use. I spend a ton of time refining my use of AI and I'm getting better over time, but it's a lot of trial-and-error trying various combinations of tools and trying to match them to tasks. Appreciate your response and the discussion.

              • eth0up 14 days ago

                In posterity, your grayed comment will benchmark the integrity of the environment here. Grayed, dead, undisputed. Out of sight, out of mind, except for persisting anyway ;)

                Keep the courage. That's what it was once all about.

        • eth0up 15 days ago

          I have over 6gb of material resulting from interactions with AI. Even a micro-essay of mediocre quality requires dozens of prompts. Why post dozens and dozens of prompts if the final product is cleaner, and easier to process? I see utterly zero reason.

          Edit: One of the things I learned quickly while working with LLMs, is that the quality of the user, the input, determines the quality of the output. Not everyone's input is of equal quality.

OldMatey 15 days ago

Given the time and effort that went into this, and the luck that one diligent person noticed, investigated and discovered what was going on before it could get further... it seems very likely to me that this has happened already in other libraries without being discovered.

  • IshKebab 15 days ago

    The effort of gaining trust over an existing project isn't even really required. All you need to do is monitor when popular GitHub repos get archived. That's usually when the original authors don't want to work on it any more. Then just quickly make a fork to continue the project (think Phabricator -> Phorge), and if you're quick enough and authoritative sounding enough, boom control of the project!

    Maintain it for a bit so people switch to your version, and job done.

  • miramba 15 days ago

    Exactly, and I wonder since then: How closely did people in comparable situations look? Since nothing similar has been reported, I suspect not very close…

  • VladVladikoff 15 days ago

    I wonder if the nation state actors are doing people profiling on owners of important packages to find the most vulnerable for an attack.

  • akimbostrawman 15 days ago

    Anybody running opensnitch would notice

skippyfish 15 days ago

Here's the tool developed by the author that was almost certainly used to generate this book in its entirety - "A structured pipeline for writing long-form nonfiction, packaged as a Claude Code skill":

https://github.com/AdrianMastronardi/bookwright

There's nowhere near enough public information about the xz vuln to be worth turning into a book, so the merits of AI-generated text aside, this is just a very inefficient way to learn about the topic.

  • eth0up 15 days ago

    "There's nowhere near enough public information about the xz vuln to be worth turning into a book," As an actual person, who reads, and having glanced at this book, I do not agree. I can already see part of the purpose is to put perspective on the crazy reality of backdoors/exploits and the undermined implications thereof. Jia Tan alone could warrant a book or film. It also says "for the general reader" and non-technical, which is where I myself think the importance really is. Maybe the AI is catching blindspots.

    I downloaded the book. It's not fake. Perhaps no masterpiece, but far from worthless. Also, not "enough public information" really sells inference and imagination short. There is a rich amount of material to work with here. More than enough.

    I am going to go ahead and say that flagging this was more information suppression than crankiness about AI. A lot of folks get strange when Jia Tan or similar subjects come up. I guess it's wiser to just wait until the grid goes down, or the water supply gets a bit more chlorinated....

dhx 15 days ago

See [1] for April 2024 Clickhouse Github activity analysis of the xz backdoor.

I haven't seen anyone write up a proper analysis that includes consideration of:

- GitHub activity (e.g. all API actions on GitHub side including replying to comments) _and_ mailing list activity _and_ other public facing activity all considered together.

- Complexity of public actions e.g. was there a queue of code changes that would have taken 20 hours effort to put together that were all committed at once? Were there any long streaks of high activity where it might reveal how many people were involved?

- Latency of public actions e.g. if an issue was raised by some random person, how long did it take for the attacker to respond, and later resolve/commit a patch? Similar to the complexity of public actions, it might reveal how many people were involved by estimation of the time needed for an experienced developer to fix an issue vs. actual time taken, both in terms of level of effort and duration.

- International dispersement of a team in different timezones with some core hours for collaboration, review and public facing activity.

- Public holidays, country/region-specific work habits, etc--e.g. consideration of "summer holiday" periods or similar common holiday periods, consideration of unusual days of no/low activity versus snow days, power outages, etc which might have been experienced by the attacker.

Distribution of actions from Github indicates the attacker used a 6 day work week excluding Sunday, and almost all activity conducted between UTC 12:00-16:00. Within these 6 days, activity was uneven at 0.5, 1, 1, 1, 1, 0.5 effort per day. There are low activity periods too that line up with summer solstice (southern hemisphere) or winter solstice (northern hemisphere).

There are interesting patterns in the data not yet publicly analysed (I think?) that seemingly would reveal the true location of attackers, particularly because attacker actions are anchored to uncontrollable events such as a known-good contributor (such as Linux distro maintainer) raising a Github issue against a repository and the attacker replying an hour later. For such events with low latency of reply, it'd be well worth considering when a reply was made quickly, and when it wasn't, across a few years of data points.

[1] https://news.ycombinator.com/item?id=39905375

kreyenborgi 15 days ago

> The catch is where the book begins, not what it is about.

  • edblair 15 days ago

    Half expecting the next few paragraphs to contain, verbatim, "The smoking gun was a performance issue in a development build of Debian. It's was a sharp observation, and sharper than you may think."

ethanhawksley 14 days ago

I've also been working on a free book about the XZ backdoor called 500 milliseconds. However, unlike this author it is entirely human written. A first draft is ready and I'm aiming to publish in a couple weeks

lucasRW 15 days ago

It's only going to recycle old news. The missing piece is attribution. Had it been the usuals (DPRK, Russia, China), the attribution would have been made publicly. The fact that is has not points as a friendly - especially when Microsoft (who owns Github) had all that telemetry and very likely has the means to find out. Some serious OSINT (consistency of timezone across months if not years of commits) pointed to the Middle East. An obvious Unit name comes to mind.

  • diogocp 15 days ago

    Fun fact: the guy who reported it (Andres Freund) works for Microsoft.

    Another fun fact: Moscow is in the same time zone as the Middle East.

  • anonreplier 15 days ago

    "all that telemetry" doesn't count for much if it's behind a VPN

    • lucasRW 14 days ago

      Highly debatable.

      Threat-hunting at that level can easily use VPNs to make attributions, especially if those same VPN exit points happen to be correlated to other stuff that was attributed. And when you are Microsoft or Google (Jia Tan had gmail accounts), the telemetry they have goes way beyond "oh we can't see the real IP lolz".

      The group responsible for the xz attempted compromise is circulating in certain Chatham House rules conference. It's just that, as someone there said "no one has had the balls to say it publicly", which in itself gives a strong hint.

tryauuum 15 days ago

So Microsoft did something good? I thought they are too busy keeping my personal data in a prison and writing tight bash loops wasting 100 percent of a core

pflenker 15 days ago

It’s great to have the entire topic brought together into a cohesive book - but wow, I find it very annoying to read.

  • ptx 15 days ago

    Probably because LLM output only gives the appearance of cohesive writing, but the more you try to understand the author's intent and meaning, the more confusing and annoying it gets, as there is no author, no intent and no meaning there to understand.

    • pflenker 14 days ago

      That’s not it in this case, because the book is mostly a recounting of what happened. I narrowed it down to two things: One is that every aspect of the story is treated equally - unimportant facets get an equally profound sounding prose like the really important bits.

      And the other one is the repetition of phrases. Everything is worth slowing down on.everything is load bearing. That’s so annoying.

      • ptx 14 days ago

        Right, that's the sort of thing I meant.

        As you read it, you would naturally try to understand why the author used such profound-sounding prose in that particular part of the text. You would expect that the author's intent was to emphasize the important bits. When the author talks about slowing down or something being load-bearing, you would expect that they're going somewhere with this, that this particular bit is especially important in some way.

        But with an LLM-generated text, these attempts at understanding will get you nowhere and only produce frustration. The reader is left trying to interpret a signal that isn't really there.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection