status.tony-stark.example.com last refresh: 2 min ago

Tony Stark

Site Reliability Manager · Ljubljana, Slovenia
All systems operational 9 yrs in service · 99.94% career uptime
tony.stark@example.com +386 555 0148 example.com/tstark UTC+1 (CET)

Summary

I run reliability for teams that ship to a few million users, and I have spent 9 years turning noisy 3am pages into fewer, quieter ones. On my last two teams I cut mean time to recovery from 47 minutes to 12, and pushed the error budget policy from a slide deck into something engineers actually check before a deploy.

I like the boring parts: runbooks that stay current, dashboards nobody has to squint at, and on-call rotations that let people sleep. If a graph is red, I would rather find out why than argue about who owns it.

Uptime Overview

99.94%
Career availability
12 min
Median MTTR
6
Sev-1 led to close
40+
Engineers on-call w/ me

System Components / Core Skills

Incident Response & On-call Operational 99.98% · 90d
90 days agotoday
Observability (Prometheus, Grafana, OpenTelemetry) Operational 99.91% · 90d
90 days agotoday
Kubernetes & Terraform (AWS, GCP) Operational 99.96% · 90d
90 days agotoday
Capacity Planning & Load Testing Degraded (learning k6 internals) 99.72% · 90d
90 days agotoday

Incident History / Career

Site Reliability Manager Vodenka Cloud 2022-03 → present
resolved
  • Lead an SRE team of 7 covering 214 services; cut Sev-1 count from 9 in 2022 to 3 in 2024.
  • Rewrote the on-call rotation so no engineer takes more than 1 primary week per month, dropping pager fatigue reports by roughly 60%.
  • Set SLOs on the top 30 services and tied a real error budget policy to deploy freezes; teams now pause on their own before I have to ask.
Senior Reliability Engineer Bela Tok Systems 2019-06 → 2022-02
resolved
  • Owned the observability stack for a payments platform doing 4.1M requests/day; brought MTTR from 47 min down to 12 min over 18 months.
  • Built 38 runbooks and a trigger-linked alert catalog, cutting duplicate pages by 45%.
  • Ran the chaos testing program (monthly game days) that surfaced 11 latent failures before they hit production.
DevOps Engineer Morje Analytics 2016-09 → 2019-05
resolved
  • Migrated 60+ services from hand-rolled scripts to Terraform, cutting environment setup from 2 days to 40 minutes.
  • Stood up the first CI/CD pipeline for the team; deploy frequency went from weekly to about 15 times a day.

Subscriptions