Here is the tax nobody warns you about when you start building your own tools. Every system you stand up is one more thing you now have to babysit. Did that job go stale? Did a long session crash at 2am? Has a goal quietly gone cold because nothing touched it in two weeks? None of that pages you. It just rots, silently, until you happen to look. For a while I was the monitoring system, and I am not a reliable one.
So I built a watchman: a small Python daemon whose entire job is to notice, so I do not have to. It runs on APScheduler, wakes ten monitoring modules on their own schedules, keeps a heartbeat every sixty seconds, and when something crosses a line it worth caring about, it sends me a Telegram message. Everything else, it keeps to itself.
remembers → 12 SQLite tables (additive migrations, WAL mode)
exposes → 22 API endpoints for the rest of the fleet
alerts → Telegram push only when a threshold trips
size → 5,274 lines / 23 files / 605 tests, 0 regressions
That box looks calm. The build was not, and the interesting failures were never in the clever parts. They were in the parts that sound boring, which is, reliably, where the real engineering hides.
The hard part was teaching it to shut up
A monitor that fires every time it sees a problem is not a monitor. It is a car alarm. Notice the same stale job every sixty seconds and you will mute the whole thing inside a day, and then it might as well not exist. So before the watchman sends anything, it checks a log of what it already told me and when. Warned you about this in the last 24 hours? Stay quiet. Each watch carries its own cooldown: stale jobs at a day, a crashed session at an hour, a cold goal at a week.
last sent → 10:15, cooldown 24h
now → 14:30 (4h 15m in)
result → SUPPRESSED — it is not time to bother him yet
Then Windows told me a living process was dead
Every daemon needs to answer one question before it starts: am I already running? The standard way on any Unix box is a one-liner, poke the process ID and see if it flinches. Mine ran on the Microsoft Store build of Python, which lives inside a sandboxed app container, and that sandbox quietly blocks the poke. So the check threw a permission error, my code read "error" as "not running," and the launcher cheerfully started another copy. Every invocation. I ended up with eight orphaned daemons all writing heartbeats to the same database, each one certain it was the only one alive.
The fix was to stop trusting the sandboxed signal and ask the operating system directly, through tasklist, which runs outside the container and tells the truth. And I moved the launch itself to Windows Task Scheduler, because a background process spawned from a console does not reliably outlive the console window that started it. Unglamorous, specific, and exactly the kind of thing you only learn by getting burned.
An archive that never double-counts
One of the nightly watches sweeps up my conversation logs, parses them into clean records of prompts, replies, and tool calls, and files them away. The trick with any job that runs every night is that it must be safe to run twice. So the archive tracks which sessions it has already stored and skips them on the next pass, which means a rerun after a crash costs nothing and corrupts nothing. Same principle that keeps a field app from writing a scan twice: do the work once, remember that you did, never repeat it. On the first real run it pulled 110 conversation turns and 191 tool calls out of a single session without a hiccup.
Why this is on the portfolio
If you are reading this to size me up, here is the point. This is not a script or a cron one-liner. It is a real daemon: it manages its own lifecycle, cleans up after itself on a crash, persists across a dozen database tables, exposes a couple dozen endpoints for the rest of my systems to lean on, and ships with six hundred tests that stay green. It watches my own machine every day and only speaks when it has something worth saying. I built it, I run it, and the bug that spawned eight ghosts is left in the story on purpose, because that is the part that was actually engineering.