Here is the uncomfortable truth about a detection setup. It is a pipeline, and every stage has its own way to fail quietly. A rule that does not translate into your search engine's dialect is dead weight. A rule that translates but names the wrong field returns nothing, forever, and looks exactly like "all clear." An alert that fires with no plan behind it is just noise wearing a siren. So this was not really a "write three rules" job. It was a "make three rules survive the whole gauntlet" job, run against a home Active Directory lab I call lab.local. (Active Directory is the system that runs logins and permissions for a Windows network.)
The detections themselves are written in Sigma, a portable format for detection rules: you describe the suspicious pattern once, in plain YAML, and a converter translates it into whatever search language your log system actually speaks. They are up as a public repo, ddub227/sigma-detection-rules, so this is not a screenshot, it is code you can read.
First I had to fix a clock that was lying
Before a single detection could be trusted, I hit the kind of bug that makes you doubt your own eyes. In Splunk (the log-search system that stores events and raises alerts, the SIEM), an event that happened around 11:00 AM my time showed up as 3:50 PM in one column and 7:50 AM in the raw log body. An eight-hour spread on a single event.
The cause was three machines living in three different timezones and none of them saying so out loud.
DC01 (Server 2022) → Pacific, UTC-8
Splunk (Ubuntu/Docker) → UTC, UTC+0
the trap → Event Log writes local time with no timezone tag
result → Splunk reads Pacific stamps as UTC (exactly the 8h gap)
The Windows Event Log stamps each event in local time but leaves off the timezone label. Splunk, running in UTC, took those Pacific timestamps at face value and read them as UTC. The eight-hour gap is precisely the distance between Pacific and UTC. I switched DC01 to UTC under Settings, Time and Language, and assumed I was done. I was not. The Event Log service reads the timezone offset once at startup and then caches it, so Get-WinEvent kept reporting the old offset until a full reboot flushed the stale value. After the restart, both clocks agreed and every source in the lab was finally on UTC.
Rule one: the ticket request that screams lockpick
The first detection catches Kerberoasting, and it is my favorite because the tell is so clean. Modern Active Directory hands out Kerberos login tickets encrypted with AES, the strong modern cipher. Attacker tools like impacket-GetUserSPNs deliberately ask for the old, weak RC4 flavor instead, because an RC4 ticket can be carried off and cracked offline with a tool called hashcat. So a request for the obsolete crypto is the fingerprint. The rule watches for Event ID 4769 (a Kerberos ticket request) with encryption type 0x17 (that is RC4), and it filters out machine accounts, which legitimately generate that traffic all day. It maps to MITRE ATT&CK technique T1558.003. (MITRE ATT&CK is the industry's shared catalog of attacker techniques, so "T1558.003" is a name everyone in the field already recognizes.)
selection: EventID 4769, TicketEncryptionType 0x17 (RC4)
filter: drop ServiceName ending in $ (machine accts)
# after pySigma converts it to Splunk
index=main EventCode=4769 Ticket_Encryption_Type=0x17
| where NOT match(Service_Name, "\$")
| table _time, Account_Name, Service_Name, Client_Address
test → attack on svc_sqldb returned only the real hits
I did not take that on faith. I ran the actual attack, impacket-GetUserSPNs against a service account named svc_sqldb, confirmed the 4769 events came through with 0x17 encryption, and checked that the rule returned only the attack traffic once the machine accounts were filtered out. Notice the small but load-bearing detail in that converted query: Sigma calls the field TicketEncryptionType, but Splunk calls it Ticket_Encryption_Type. The splunk_windows conversion pipeline handles that renaming automatically, and this is exactly the "translates but names the wrong field" trap that returns a silent nothing if you skip it.
Two more tells: the hammering and the odd hour
Rule two is brute force, mapped to T1110. One failed login (Event ID 4625) is nothing, everyone fat-fingers a password. The signal is the pattern: five or more failures from the same source IP inside a five-minute window. That threshold-counting is not something Sigma converts cleanly on its own (it needs a separate correlation file), so I enforce it at the Splunk layer with a query that buckets events into five-minute bins and flags any IP that trips the count.
Rule three is after-hours privileged logons, mapped to T1078. It watches Event ID 4672 (special privileges handed to a login), throws away the SYSTEM account and machine accounts so it does not fire on routine Windows housekeeping, and then keeps only the events that land outside 08:00 to 18:00. Sigma has no built-in notion of "time of day," so again the temporal cut lives in Splunk (date_hour < 8 OR date_hour >= 18). An admin credential lighting up at 3am is a high-quality "someone is using stolen keys" signal, and now the lab notices it.
All three converted the same clean way, the standard toolchain (pip install pySigma pySigma-backend-splunk sigma-cli, then sigma convert -t splunk -p splunk_windows), and I diffed the machine-generated SPL against the hand-written queries already running to make sure they agreed.
Then I actually deployed them, and planned for the day it all dies
Rules that only exist in a repo do not defend anything. All three are now live alerts in Splunk, each on a fifteen-minute schedule: Kerberoasting (High), brute force (Medium), after-hours privileged logon (High). Between them they cover two attacker tactics, credential theft and the abuse of valid accounts after a break-in. The Kerberoasting alert I proved out end to end by re-running the attack and watching the trigger land in Splunk's Triggered Alerts view, and every alert only looks at the last fifteen minutes so it does not keep re-screaming about the same old event. I also grew the SOC Overview dashboard to five panels, adding a Kerberoasting view that surfaces RC4 requests in a table so an analyst sees them without hand-writing a query.
The unglamorous finale was backups. I had run myself against NIST CSF 2.0 (a widely used security checklist) and two boxes were still bare: no backup capability, and no written recovery plan. A single dead drive would have meant rebuilding the whole lab from memory. So DC01 got a VirtualBox snapshot and a portable export (an OVA, 10.6 GB compressed from a 60 GB disk), arranged as a 3-2-1 backup: three copies, two kinds of media, one of them offsite in Google Drive. Then I wrote an actual recovery plan with four failure scenarios and honest target times, from a five-minute snapshot rollback up to a one-to-three-hour full Splunk rebuild, each with step-by-step procedures and a checklist to confirm AD, DNS, log forwarding, and the alerts all came back. Both gaps moved off "Not Implemented," though I am calling them "Partial," not "done," until the weekly exports are automated and I have actually rehearsed each recovery instead of just documenting it.