Settings

Theme

Show HN: Sigabrt.dev – cronjob monitor with an SSH TUI

sigabrt.dev

64 points by 4815162342 · 35 comments · 1 min read

Reader

Hello HN. I built this mostly to monitor the things I host myself. I know it's nothing too exciting.

Anyway, the TL;DR is: Create an endpoint, and if your script/cronjob fails to regularly ping it, you get notified (by email or ntfy). E.g.:

    0 * * * * ./script.sh && curl -fsS https://sigabrt.dev/pulse/<id>/beat
It also has an SSH TUI which is currently experimental and read-only, mostly because I'm not sure whether it is actually useful or just a gimmick :):

    ssh sigabrt.dev
To use it, simply add your SSH public key in your account settings.

Yes, there are services like this already, and this is minimalistic by comparison. Feedback is welcome.

10 threads
mosselman

Shameless plug for a server monitoring macOS app I built. I built that too for the servers I managed myself. Instead of working through a service or a self-hosted monitoring tool that I then over to monitor itself again, I thought running it on my laptop was pretty nice too. One of the next features I've been wanting to build is uptime monitoring.

https://kitaso.app/

Sure there is the downside of having to be at your laptop or having it on, but the upside is that it has very few moving parts and is very simple and it just has one set price.

p2004a

I'm myself very happy with https://healthchecks.io/ for this purpose.

  • altern8

    I'll never understand commenting to show HNs with a link to competing products.

    • RamblingCTO

      If it has context I appreciate them because it helps me make better informed decisions. Why not? HN isn't an editorials page exclusive for posters (except YC slop lol)

    • hmokiguess

      Why?

      • altern8

        Because OP is trying to promote his project and get feedback, and you post about a competing product, to me the message is: don't look at the project OP wants you to check out, I've been using this instead and and it's amazing!

        Who cares what you're using, I see show HN posts to be about OP's project, seems like a dick move to me to diverge attention to some other project. I just don't understand the motivation

        • circularfoyers

          The first thing I always want to know is, what does OP's project offer over pre-existing mature ones? Especially important when I'm seeing an increasing amount of projects that don't seem to even bother considering pre-existing projects.

        • hmokiguess

          I never read such comments in that way, the way I read them is actually like feedback to OP around benchmarking against competition.

          Personally I love knowing about my competition, and it helps me with how I can differentiate and win over certain niches.

          I don't know about OP, but I'm never offended by that, always grateful, helps me craft my message and delivery to the right audience and map out strategy.

          I try not to over read into the open internet motivation behind end users, they may become customers, so I rather welcome them and seek to understand rather than judge.

        • sgt

          It could help people. It's possible the sigabrt didn't even care to Google about Healthchecks.io before I decided to vibe code his own.

          Also, what do you really think the odds are that sigabrt.dev will still exist in 2 years from now? Or 5?

      • TheSkyHasEyes

        Just goes against the spirit of people writing their own utils and sharing to me. Good on OP. They're not looking for money or fame here. If this util benefits one reader on HN it's worth it.

    • bojanstef

      @altern8 you're dope for supporting folks putting themselves out there.

  • singpolyma3

    And it's open source!

raimue

You could just wrap cronic around the command and immediately receive the full output per mail when it fails. I don't see how a heartbeat alone without logs would help to identify temporary failures.

https://habilis.net/cronic/

Your service could accomplish something similar if it had such a wrapper to report both success and failures with logs to a remote server. That would take away the need to run a local MTA, while also detecting with the heartbeat whether the job ran at all.

  • 4815162342OP

    Thanks for the suggestion!

    There's currently no explicit "fail" ping, as that's not the main use case I care about. It's valid, though, and I've thought about adding it.

    Being able to pipe logs to curl would also be useful. I'm concerned that people might accidentally send me private/sensitive data, but I'll consider it.

  • mherrmann

    > I don't see how a heartbeat alone without logs would help to identify temporary failures.

    It does help. Sometimes jobs unexpectedly don't run at all.

    • raimue

      A temporary failure means that the job failed only once and then starts running again. If you want to investigate why it failed, you will need the output of that particular failed run.

mrweasel

If you need/want a dashboard it's kinda cool. There's a lot of other options that will do something similar, but not via SSH. Crontab can email you directly, no need for a service.

You could also just use systemd timers and do: systemctl --failed -t service

  • theoli

    You need some sort of external monitor for a missed pulse style alert. Personally I do something similar by collecting a "last success" metric with Prometheus and alerting on it with Grafana if the value is too far in the past. The local system cannot reliably alert if your job does not fail into the alert path or the system is just down.

    I can see something like this being a great intermediate option to a full monitoring and alerting stack.

  • inezk

    You are missing the biggest value offering here - service such as the one linked here (or healthchecks io that I personally use) will let you know that signal didn't arrive even when other things on your end failed (e.g. cron having a hard time sending failure email due to incorrect smtp credentials). I use those kind of healthchecks literally for everything, especially for backups - they saved my ass many times over.

    • mrweasel

      For things that actually matters to me, I use systemd timers, Prometheus and alert on failed services, but I do get that this might be a bit much for many setups.

linsomniac

I've been toying with a similar idea for ~6 months: A lightweight job status dashboard.

I wanted something that required no setup, but could just push success/failure messages to as part of various cron jobs, windows tasks, and shell scripts we run throughout our organization.

StatShed server: https://github.com/statshed/statshed-server StatShed go-cli: https://github.com/statshed/statshed-gocli

The idea is kind of like "ntfy.sh", but for jobs status. You can send a "started" message at the beginning, update a "status" message periodically throughout the job, then send a "failed" message (optionally with logs) or a "success". Then a web dashboard gives you an overview with ability to drill down.

We use Icinga for monitoring and paging, but this just gives an overview for a quick look at things we don't want heavy duty monitoring on. Like my laptop backups, information about ansible runs across our fleet, etc.

noja

I don’t like the wrapper idea. I would like it built into the cron daemon.

adityamishra241

The pulse idea is neat. How do you handle jobs where the expected runtime is longer than the heartbeat interval?

  • 4815162342OP

    Good question! There's no such thing as a "running" task/job at the moment (i.e., a start signal). The closest thing is lengthening the grace period, but that also means that the notification is delayed, so it's not really a solution. This is in the backlog.

    • msupuka

      The usual fix is a second ping at the start of the job: the schedule check watches the start ping, and a separate “max runtime” check watches the gap between start and finish. You get alerted promptly either way without inflating the grace period.

DylanMerigaud

Context helps, shows intent, fosters discussion.

ww520

People seem to miss the point of the project. It’s not emailing you on failure. It’s emailing on missing reports of scheduled runs. If you rely on the job to report failure, the machine could go down. In that case you won’t get any email on the failed run.

dorianmariecom

i have something similar at https://heartbeats.dorianmarie.com/

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection