DrDroid — AI SRE Agent for Incident Response & Root Cause Analysis

DrDroid

7 min read Original article ↗

Backed by Y Y Combinator

Build agents for what begins after you ship.

Get started and build agents in under 5 minutes to automate maintenance and toil.

A tool that automates ops shouldn't create more ops — nothing to host, nothing to maintain, and it costs less than an EC2 instance.

AGENTS OUR USERS BUILT

  1. BEFORE RESPONSE
    1. Alert Middleware

      The first point of response for every alert.

    2. Alert Tuning Agent

      Fine-tunes thresholds and files tech-debt tickets.

    3. Deployment Tracking Agent

      Helps with canary deployments where script-based rollouts aren’t enough.

  2. DURING INCIDENTS
    1. Investigation Agent

      Investigates any alert or incident end-to-end.

    2. Auto-remediation Agent

      Runbook-based remediation with human approvals.

    Your workflowBuild your ownStart in under 5 minutes
  3. AFTER INCIDENTS
    1. Observability Agent

      Analyzes metrics cardinality, high volume logs, and telemetry gaps in code.

    2. Reporting Agent

      Periodic insights on infra costs and monitoring gaps.

    3. Infrastructure Migration Agent

      Helps with long-horizon infrastructure migrations through tight planning, testing, and change management.

    4. Retrospective Agent

      Analyses past incidents to surface long-term strategic improvements.

CONTEXT-AWARE FROM THE START

The agent already understands your infrastructure.

DrDroid starts with an operating model of your services, owners, dependencies, deploys, dashboards, and past incidents—so every agent begins with context, not an empty canvas.

Live service context

● continuously updated

payments-apiSERVICE · PRODpayments-apiGITHUB REPOp95 latencyDATADOGpayments-prodKUBERNETESHigh latencyRUNBOOK

Owners mappedDependencies currentPast decisions retained

ALWAYS-ON CONTEXT ENGINE

Connect once. Context keeps building.

Background jobs continuously extract, process, and scrape metadata from every connected system, turning scattered tool data into durable organizational context and memory.

Connectors

  • Datadog
  • Grafana
  • AWS
  • Kubernetes
  • GitHub
  • PagerDuty
  • Prometheus
  • New Relic

events + metadata

Background jobs

  • Extract entitiesRUNNING
  • Process relationsRUNNING
  • Scrape metadataRUNNING

normalized context

Durable org memory

  • Services & ownership
  • Dashboard ↔ service links
  • Alert history & noise
  • Runbooks & past fixes
  • Deploy & change history
  • Topology & dependencies

available to every agent

AGENTS MEET YOU THERE

You don’t go to the tool; it comes to where you work.

Trigger an agent from an alert, a conversation, another agent, or your own automation. Results return to the system where the work already happens.

  • PagerDuty

    The page arrives with findings already attached — triage starts before you open your laptop.

  • Slack / MS Teams

    Ask, approve remediations, and close investigations without leaving the thread.

  • Webhook

    Fire an agent from anything that can POST — alertmanagers, cron jobs, internal tools.

  • MCP Server

    Expose your org's context and agents to Claude, Cursor, and any MCP-compatible client.

  • API Trigger

    Kick off investigations programmatically from CI/CD or your own platform.

  • A2A Protocol

    Your in-house agents delegate to ours, agent-to-agent.

ONE CONTEXT-AWARE AGENT · EVERY WORK SURFACE

Connect your first surface

Map your entire stack in one graph.

DrDroid maps every repo, dashboard, K8s pod and cloud resource into one live knowledge graph — then traces the blast radius across connected entities the moment an alert fires.

payments-api SERVICE Repo 3 payments-api Dashboard 7 p95-latency Alerts 124 p95 anomaly Issues 18 GH#4821 Resources 9 k8s pods Infra 14 aws rds

Infra dependency confirmed 3 pods healthy 14 correlations 1 anomaly 2 deploys today

01

Cross-tool correlation

A GitHub repo maps to a Datadog service, a Grafana dashboard, K8s pods, and AWS resources, automatically.

02

Decision engine

When an alert fires, the graph traces the blast radius across every connected entity in seconds.

03

Continuously learning

Every alert, deploy, and incident strengthens the graph, surfacing patterns no dashboard can.

How DrDroid learns your stack.

We connect to your existing tools, crawl all telemetry, and generate a knowledge graph of your stack.

01 / CONNECT

Read-only access to your entire stack.

OAuth into cloud, code, CI/CD, and observability. No agents. No code changes. Live in 30 minutes.

AWS · GCP · Azure GitHub · GitLab Datadog · Grafana · NR

02 / CRAWL

We crawl all telemetry and build your knowledge graph.

Metrics, logs, traces, cloud configs, repos, docs, runbooks. All crawled and mapped into a cross-tool knowledge graph. Which repo → which service → which dashboard → which pods. Always live, always learning.

knowledge graph context map always live

03 / ACT

Act with full context.

The knowledge graph powers proactive suggestions, root-cause diagnosis, and automated runbooks — all with full context.

proactive explainable guarded

Tighten retry budget · orders-svc SUGGEST

Cause of INC-4821 · sidecar OOM RCA · 9m

Auto-scale on memory pressure RUN

Drain node-12 · disk-full RUN

Self-learning agent

Every investigation makes
the next one faster.

Nothing gets thrown away. The second time the same failure appears, DrDroid replays the path that worked and skips the one that didn't. MTTR drops on its own, without anyone writing a runbook.

Runbooks that worked Dead ends skipped Your metric names

See how the agent learns

Same alert, twice

First encounter Searches the whole stack

Next time it fires Goes straight to the cause

Nothing it already ruled out gets checked twice.

One memory for your entire stack.

AI Memory holds your service graph, runbooks, docs, and every live signal, alerts, deploys, conversations, incidents. It builds patterns over time so every engineer starts with full context, not a blank slate.

drdroid.app / ai-memory

⌘K

Platform Knowledge 27,752 records · 23.5 MB

Infrastructure Components/ 689

Alerts & Activity 9,456 records · 60.2 MB

infra APITimeoutError on OpenAI API in podracer 2 alerts

Last: a few minutes ago Sentry

infra APITimeoutError on Azure cognitive services endpoint 2 alerts

Last: a few minutes ago sentry

code psycopg2 UndefinedColumn created_at protoproddb connector 2 alerts

Last: a few minutes ago sentry

code psycopg2 UndefinedColumn tool_calls protoproddb connector 1 alert

Last: a few minutes ago sentry

code PostgreSQL UndefinedColumn investigation_id protoproddb 1 alert

Last: a few minutes ago sentry

Last seen: 9 minutes ago sentry +3 more reports +2 more

Service Name Upstream Downstream Data Sources Created By Rule Source
azure_monitorinfra None None 3 sources DroidAgentV2 Rules managed
app_serviceservice None None 3 sources DroidAgentV2 Rules managed
addon-resizerinfra None None 9 sources DroidAgentV2 Rules managed
storageinfra None None 9 sources DroidAgentV2 Rules managed
network_watcherinfra None None 9 sources DroidAgentV2 Rules managed
metrics-serverinfra None None 14 sources DroidAgentV2 Rules managed

Connect your entire stack.

Cloud, code, observability, incident response and ticketing — wired in via integrations, read-only and reversible.

Cloud & Infra

AWS AWS

Google Cloud Google Cloud

Azure Azure

Kubernetes Kubernetes

Amazon EKS Amazon EKS

GKE GKE

Code & Delivery

GitHub GitHub

GitHub Actions GitHub Actions

Bitbucket Bitbucket

Jenkins Jenkins

Argo CD Argo CD

Observability

Datadog Datadog

Grafana Grafana

New Relic New Relic

Prometheus Prometheus

Elastic Elastic

SignOz SignOz

Incident & Response

PagerDuty PagerDuty

OpsGenie OpsGenie

Sentry Sentry

Rootly Rootly

Zenduty Zenduty

Rollbar Rollbar

Workflow & Ticketing

Slack Slack

MS Teams MS Teams

Linear Linear

Jira Jira

Notion Notion

Confluence Confluence

See DrDroid in action

Watch how engineering teams use DrDroid to cut MTTR and stay ahead of incidents.

Built for SRE engineers on call.

We measure ourselves on pages avoided and minutes saved during the incident, not dashboards rendered.

"Earlier, debugging meant hopping between logs, workflows, and infra dashboards trying to piece together what went wrong. DrDroid pulls the context together and points us in the right direction, even someone new to the system can figure things out."

Rahul Bhattacharya Rahul Bhattacharya · Co-founder & CTO, Adopt.ai

"One time I was woken up at 3am by a pager that escalated. I instantly asked DrDroid to investigate it and in a few minutes, I was able to close the issue directly from Slack."

Moiz Arsiwala Moiz Arsiwala · CTO, WorkIndia

"DrDroid understood our context too well. It gave recommendations which showed deep understanding of the infrastructure and helped reduce 20–30% cost."

Prateek Prateek · Head of Technology, Stanza Living

Enterprise-ready security and deployment.

DrDroid runs where your data lives, meets the bar your security team sets, and ties its pricing to outcomes you actually care about.

SOC2 SOC 2 Type II certified Read-only integrations SSO / SAML

01

Self-hosted deployment

Run entirely inside your VPC or on-prem. No data leaves your network. Deploy via Helm or Docker Compose with air-gapped support.

02

Outcome guarantees

We tie our success to yours — measurable reduction in MTTR and incident frequency, SLA-backed with quarterly reviews.

03

Security & compliance

SOC 2 Type II, encrypted at rest and in transit, read-only access to all integrations. Built to pass your vendor review on day one.

Generate your knowledge graph in minutes.

Connect your stacks and see your services mapped in minutes.