Settings

Theme

Show HN: FixBugs – Reproduce production bugs and verify fixes

fixbugs.ai

43 points by kirtivr · 45 comments · 2 min read

Reader

I built FixBugs, an agent that ingests the rich context surrounding production bugs to reproduce them in a sandbox and generate verified fixes. It's available in the form of a self-hosted VSCode extension and as a Github app:

VSCode Extension: https://fixbugs.ai/go/vscode-extension

  - full code and data privacy.
  - zero data retention models opted out of training.
GitHub App: https://fixbugs.ai/go/github-app

  - we do access your code temporarily.
  - pick a repo to install FixBugs on.
What motivated me to build FixBugs were my years being on-call at Google and VMware. How many hours did I spend gathering logs, traces, reviewing metrics, and reading code only to find that,

* Some context was missing.

* The bug wasn't reproducible.

* The alert was caused by a transient infrastructure issue.

Too many. Inefficiency in investigating staging/production bugs has a real cost, and it's paid both by developers and customers.

Current capabilities:

  - Reproduce the bug.

  - Identify the root cause.

  - Generate a fix.

  - Verify the fix.

  - Review the generated code using multiple AI models to help catch potential regressions.
Do try it and let me know what you think!

I'd especially love feedback from engineers who work with distributed systems or handle high-volume production bug triage.

21 threads
Rishi1238

Interesting project. How does it handle services with multiple dependencies (queues, caches, third-party APIs)? Can it recreate enough of the production environment to reproduce intermittent bugs?

  • kirtivrOP

    We either mock or fake out the interfaces.

    I much prefer faking to mocking, because that still preserves a lot of the real-world behavior relevant to prod bugs.

    A full prod reproduction would be a holy grail, but probably only attainable for complex distributed systems if we have access to a prebuilt staging-like environment.

    We're traversing the bridge between mocks, fakes and staging at this time.

kumari_b

Interesting approach. The investigaion phase is usually the most time-consuming part of debugging. Curious how well this works on large, distributed systems.

  • kirtivrOP

    Nginx, Caddy, Flink and Firefox are some applications where we we've managed to consistently fix and reproduce reported bugs.

    Of course, we did not send the PRs to the repos seeing how they're already overloaded with them.

    Nginx for example has 181 open PRs right now, but they only merge 2 or 3 in a day.

    • kumari_b

      Thanks for the clarification. Those are some interesting projects to benchmark against.

satyamtiwary

This is a really interesting direction. Bug fixing is always a painful part of software engineering. From reproducing issues, to identifying root causes, then generating fixes, and verifying them, it all adds to our project delivery timelines and this tool will definitely help in that aspect. Kudos to the development team for working on this product.

kirtivrOP

A different approach that we took to root causing bugs that you may find interesting is that we first try to reproduce the bug before coming up with a fix for it.

This is essentially a (RCA <-> Repro test case) loop until we're recreated the bug. If our attempts are not converging and we’re on the wrong track, we ask for human input.

sriramkalluri3

Great tool to triage the reported issues in bugs and fix them swiftly. This will definitely improves your productivity.

majestic8

This looks very useful! What type of sandbox are you using? How does the mocking work?

  • kirtivrOP

    On Linux, we rely on a chroot-ed workspace at this time - although we are working on a prototype using the new landlock kernel interface (https://github.com/Zouuup/landrun).

    OSX is the best, we use the in-built (seatbelt) sandbox via sandbox-exec.

    For Windows, we use WSL containers when available.

    By default, if a safe sandbox environment is not available, we inform the user that a repro is not possible in the current conditions.

  • kirtivrOP

    On how does the mocking work, that's a really interesting question.

    We do a lot of AST parsing - for both code and build configuration languages. Even then, we still have to rely on the LLM to figure out a lot of the details.

    Making this work reliably for non-frontier models and codebases that don't have existing test harnesses is where a lot of the design work goes in.

shashi_ranjan

Nice work. The architecture looks well thought out, and it's refreshing to see engineering focused on reliability and developer experience rather than just adding AI buzzwords.

  • kirtivrOP

    Thanks. As teams move faster with AI generated code, the focus on quality and reliability will have to increase as well.

    And we will have to build better tools if we want to accelerate shipping velocity.

Bheemaneni

This seems especially used for teams handling production incidents.

  • kirtivrOP

    It's really nice to hear that.

    FixBugs was built because while investigating prod incidents, I had an epiphany.

    SWEs build tools to solve all types of problems, but we ourselves use the flakiest tools.

    While working at Google for example I was surprised GDB support for any type of binary debugging was almost non-existent.

vasmiReddy

Nice idea! The reproduce → verify loop is what makes this stand out.

  • kirtivrOP

    Thank you!

    are you using any AI tools to debug productions bugs at this time?

nxtcoder17

seems really interesting @kirtivr . Looking forward to it.

Though, i wonder if it could also be an integration to alerting platforms directly (like NewRelic, Datadog etc.), so that for on-call alerts, it could cover the foundation work, and have some hypotheses ready for the on-call engineer to directly jump into

  • kirtivrOP

    This is on our roadmap!

    At this time we are focussing on evaluation benchmarks like SWE-bench (verified). This is a simpler benchmark and does not really map well to investigating alerts that have a huge amount of context. But its a start.

    I am wondering if we can improve upon foundation models with our reproduction <-> hypothesis loop approach.

    Foundation models tend to be precision first, and context limited, so they can get sidetracked by various things.

    This is an interesting problem to be working on right now!

Hemashrie

Really like the reproduce-first approach.

raghu16

Also, How is the VSCode extension reproducing the bug on my machine? That sounds dangerous.

  • kirtivrOP

    Oh the repro runs in an isolated sandbox, and all interactions outside the sandbox (with lets say other services or databases) are mocked. The repro harness doesn't have access outside of it.

    This also allows us to inject various types of faults, which is helpful with debugging more complex systems.

sriramkalluri3

Great tool to triage and fix the issues in bug. Improves your productivity massively.

mundanevoice

It seems like a cool idea.

  • kirtivrOP

    Thanks a lot!

    I think FixBugs is most useful during high volume bug triage. This is where having a low false-positive async debugging agent is most helpful.

ShaikMohammad

Very cool.Reproducing production bugs is usually the hardest part.

rKulaSekhar

Love the focus on verification instead of just generating fixes.

omanhar

was there a technical constraint that you did not expect to have to solve or work around?

divvsaxena

Awesome man cool!

parspunisher

But how is this different from copilot?

  • kirtivrOP

    A few different ways:

    - Copilot may be better at implementing features. We're better at investigating bugs and fixing them.

    - We handle huge context very easily. We specialize towards investigating large amounts of logs/metrics and traces.

    - We do a lot of work to generate a non-trivial reproduction test case in a sandbox, which allows us to verify bug fixes. Ship confidently, not just based on the best available hypothesis.

    - You can apply and remove code changes with the click of a button. We have an in-built VCS that allows you to add or remove sets of changes from a long coding session.

    - Our pricing model is different. While with the standard Copilot developer plan you get $10 of AI usage (nothing if you're using Claude or GPT), we allow you unlimited triages on a fixed number of bugs. It's not about AI credits, it's about investigating and fixing complex bugs.

yizluo

very nice tool! helpful and I love it

  • kirtivrOP

    Thank you so much.

    are you using any AI tools to debug productions issues at this time?

nxtcoder19

Interesting approach.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection