This is the first in a series where New York Times CTO, Nick Rockwell, talks to leaders in the technology world about their work. Rockwell interviewed Charity Majors, the co-founder and CEO of Honeycomb, about observability software, debugging and “getting inside the software’s head.”
Nick Rockwell: Can you give us some background about the problems you’re trying to solve with Honeycomb?
Charity Majors: The market is saturated with monitoring startups, logging startups, analytics startups and APM (Application Performance Management). There are so many startups, it shows people aren’t terribly happy with the options available to them. And yet, the newer entrants aren’t terrifically different from the older products. Rearranging of deck chairs aside, monitoring hasn’t fundamentally changed in at least 20 years.
I don’t mean to say the existing tools are bad; some of them are actually great at answering questions — your known-unknowns — or at building dashboards for canned metrics and system-at-a-glance, or anything and everything you can predict you might want to look at.
The problem is, you can’t predict everything.
It used to be the case that with our monoliths, our one big database and our predictable user patterns, we were very rarely stumped. You could glance at your dashboard and see which component was at fault, or check for a recent deploy, and that got you through 90% of all problems. If it was a new problem, you would sift through dashboards and use your intuition about what was happening under the hood, maybe attach a debugger or look at the queries in your database.
Eventually you’d find the problem. Then you’d have a post mortem, write a monitoring check for whatever failed and maybe make a new dashboard to find that error condition instantly. Over time, you built up a solid repertoire and were very rarely stumped by what was going on in your systems.
But that model is falling to pieces. With the decentralization and ephemeral nature of modern systems, you can rarely tell in a glance where the problem originates. You rarely encounter the same problem twice. Once-in-a-million edge cases happen all the time, at scale. Most of your problems are unknown-unknowns, and they aren’t cropping up monthly or yearly but many times a day. Our tools just aren’t built for that uncertainty.
NR: Can you talk about what you like about troubleshooting? Where does that passion come from?
CM: I actually don’t like monitoring — it’s always felt like an after-the-fact cleanup job — but I’ve always loved debugging. There’s something so fascinating, open-ended and terrifying about encountering a problem or a bug for the first time — especially in the data world, where your mistakes can literally put a company out of business. There are no experts out there who can come to your rescue; if you don’t solve it, no one will solve it.
I love high stakes, and I love problem solving. I’ve grown to love doing both under pressure; but not everyone does, to say the least.
Anybody who does a lot of debugging knows about the dopamine hit you get from fixing things, especially if you can track it down before it impacts your customers. Or the surge of raw joy you get from figuring out some hard, intricate problem.
I always learn the most under extreme pressure. There’s no high quite like saving the world in the nick of time.
NR: We have bad tools for instrumenting, monitoring and debugging our apps. What have we been getting wrong, and why?
CM: They aren’t “bad,” but they were built to solve the last generation of problems. You need to know what you’re looking for in order to find it. Older tools prioritize the health of the system over the health of the event, they don’t handle high cardinality and they aggregate at write time.
In older architectures, the health of the system was reflected in your user’s experience because all components were shared. So, if your system’s availability was 99.5%, your user was experiencing a failure 0.1% of the time, and failures were pretty evenly distributed. The health of the system was a great proxy for the user experience, so you could just look at your dashboards and understand the user experience pretty well.
In newer systems, everything is sharded and horizontally distributed, so perhaps your availability is 99.5%, but 0.5% of users experience your site as completely down. That is far worse. And if your dashboards are aggregating over a time interval, your users’ bad experience will never show up.
The health of newer systems is all but irrelevant. Your system may be down 25% of the time, but you don’t actually care unless it’s impacting your users. What is important is the health of each individual service request, and every slice of those requests (such as UUID, request ID, shopping cart, shard and so forth).
High cardinality is something that none of our tools handle well. All the data you actually care about has high cardinality, whether that’s UUIDs, request IDs, first name/last name or ip host:port pairs. Debugging is like looking for needles in a haystack made of other needles, so you need to tag the needles with the highest cardinality data possible because it’s the most identifying. This is why I say that high cardinality will save your ass: you have to be able to identify all events with pinpoint accuracy, and you need all the high cardinality dimensions you possibly can.
Debugging is like looking for needles in a haystack made of other needles.
Write-time aggregation, while extremely high performing and efficient, robs you of access to raw events. I don’t think you can have true observability without being able to trace your way back to the source of truth: raw events. Once you’ve smooshed everything that happens over the course of an interval into a single number, you can never un-smoosh them again.
There is no such thing as right or wrong, there are only different sets of trade-offs. For decades we’ve made trade-offs based on the metric, which is the smallest point of data you can possibly have — the metric throws away literally all the context of the original event, and then tacks a few tags back on so it can be grouped with other metrics. The metric became king in an era of extreme resource scarcity and expensiveness: hard drives, CPU and RAM were extraordinarily pricey 10–20 years ago, so we just had to put up with all the terrible side effects, like the explosive write amplification of tags.
Hardware has gotten cheaper, user demands have gotten more outsized, architectures have gotten more complex; it’s Moore’s Law all the way down. We need a similar revolution in the observability space to take advantage of these changes. That’s why newer tools will, I believe, revolve around the (arbitrarily wide) event, rather than the metric. Events pack way more context, and say “all these values were true at once.” And events work the way our human brains do by helping us craft a narrative.
NR: What makes a great ops person?
Get Nick Rockwell’s stories in your inbox
Join Medium for free to get updates from this writer.
CM: Great operations engineers are made through curiosity and stubbornness. I know so many ops people who are dropouts, specifically liberal arts dropouts. Old school systems, particularly unix systems, are very friendly to people who love languages and have patience for poking around and exploring. A top-notch engineering education isn’t especially valuable when it comes to debugging hard problems. It’s much more about persistence and desire.
NR: How do you think about instrumentation? Should one be selective in what you instrument, or just try to grab “all the data” and sort it out later?
CM: Any time someone says, “just instrument everything” or “monitor everything,” I hear, “I am unwilling to make hard choices and have poor technical judgment.” You can’t watch everything. Instrumentation and monitoring have costs, just like anything else. You should gather as much as possible, and you should select tooling that encourages you to gather rich detailing and makes it cheap, but there are plenty of hazards attached to just dumping trash into logs and hoping someone else will sort out the signal from the noise later.
NR: How do you feel about mixing metrics, such as ops metrics, other engineering metrics or product/business metrics? Where do you draw the lines, if you do?
CM: Many of the lines we have drawn around tooling — metrics, monitoring, business intelligence, logging, APM — have been drawn because hardware was so incredibly expensive and we had to make trade-offs. As hardware gets cheaper, these lines will erode.
Tools build silos. Teams that don’t use the same tooling can’t speak the same language, and they cannot fluidly share insights. I think that over the next five years, you will see many of these categories dissolve and vanish and be replaced by the umbrella category of “observability.” If there’s one thing dev-ops has taught us, it’s that silos are inefficiencies made manifest.
Tearing down silos may be the work of generations, but data is exactly where it begins.
NR: I like your idea that given enough scale, black swans are the norm. Can you talk more about that? What are the implications? How do you prepare for the unexpected?
CM: Lots of things happen rarely, maybe once in a million. But at the scale we tend to run these days, that can be kind of a lot; certainly often enough to make your platform unusable for some people.
In a mature system, every time you get paged should really be about an unknown-unknown. You cannot predict all the possible ways a system may break, and you shouldn’t try. You should invest that energy into guard rails for your production systems, safety measures that make failures easy to detect and recover from. Invest in canarying, rolling deploys, feature flags and rich exploratory tooling.
These are the tools that Google and Facebook have had for years. Now that our problems are catching up to theirs, our solutions must too.
Every individual problem is a rarity, something that almost never happens. But in aggregate, failure is incredibly commonplace. We need to treat it like a fact of life — it’s when we fail, not if we fail — and practice until it’s boring.
Focus on making your critical path as narrow as possible, and make your systems resilient to all other failures. You should only have to wake up in the middle of the night for failures that will put you out of business, and those should be remarkably few. It’s amazing how much failure you can really tolerate without anybody even noticing.
NR: A lot of startups are trying to apply Machine Learning to ops in general and troubleshooting in particular. Are you? Do you think there is potential there?
CM: No. Not any time soon. I’ve seen their internal success metrics, and they’re pretty awful. The problem is,you first have to train the machine on your own corpus and the learnings aren’t really reusable across system boundaries. You also have to substantially retrain it every time it changes, which may be several times a day. And false positives are incredibly costly from a human perspective.
So you’ve got people with large systems generating thousands of alerts every minute, and you’re trying to get some signal from that noise. This can be done, but you know what can be done better, easier, more cleanly? Removing those flappy or non actionable alerts.
If you don’t have a system generating that many alerts, machine learning isn’t going to do much for you anyway. So in basically every case, I recommend more intelligent inputs rather than trying to make the robots cull the herd for you. It’s not at all hard, and the results are dramatically better.
NR: I jotted down in my notes the phrase, “getting inside the software’s head,” so maybe you said that. Did you? If you did, I love it. What have you learned about how systems think over the years?
CM: Yes. You have to look at the systems from the perspective of the application itself, because almost nothing else matters. This is a seismic shift in operations and observability. Nines don’t matter if users aren’t happy. User happiness has nothing to do with system health, but has everything to do with application health.
We’ve been learning to treat infrastructure as a disposable resources for 10 years; these shifts in observability signal that we’re very nearly there.
NR: So, what are you up to now? What is the thought behind Honycomb.io and what do you want to accomplish with the company?
CM: Well, I think the center of gravity is shifting rapidly towards the generalist software engineer — even for infrastructure products. A lot of people think of it as “ops engineers vs software engineers,” when actually it’s more like “infrastructure vs product” where both are made up of software and operations engineers. But while ops and database administrators aren’t going away any time soon, they increasingly live on the other side of an API from the software engineer’s perspective. And I don’t think anybody in the space is effectively building for the software engineering eye, everyone is building for operations. That’s short-sighted.
I also think the cloud vs on-premises wars are over (spoiler alert: cloud won), and that the lines are blurring between vendors and teams. A good vendor should feel like a services team inside your company, and a good services team inside your company should feel just like a vendor.
People are narrowing their focus. Engineering cycles are the scarcest resource any company ever has, so why are we wasting them on things that have no competitive advantage? It’s very costly to roll your own servers, email, metrics — that whole massive iceberg of support software that empower you to build your business and your core differentiators. People are increasingly willing to outsource to companies that specialize in those services, and for good reasons.
We are still in the early years of the distributed systems revolution. It’s only going to gather momentum over the next 5–10 years, which is why we are only trying to sell to the people who already know they have these problems: platforms, multi-tenant systems, microservices architecture. The way people’s eyes light up when you show them how to make impossible problems into easy problems — that’s what keeps us going.
I think that treating systems like they are complex and distributed is going to be the next wave of innovation akin to Continuous Integration/Continuous Deployment. The testing phase doesn’t stop when code hits production — it’s only begun to mature with real users, real data, real services. You’ve got unit and integration tests; then you deploy a canary to production and commence production testing. This simply hasn’t been possible without high cardinality, event-oriented observability.