status.tony-stark.example.com
last refresh: 2 min ago
Tony Stark
Site Reliability Manager · Ljubljana, Slovenia
All systems operational
9 yrs in service · 99.94% career uptime
Summary
I run reliability for teams that ship to a few million users, and I have spent 9 years turning noisy 3am pages into fewer, quieter ones. On my last two teams I cut mean time to recovery from 47 minutes to 12, and pushed the error budget policy from a slide deck into something engineers actually check before a deploy.
I like the boring parts: runbooks that stay current, dashboards nobody has to squint at, and on-call rotations that let people sleep. If a graph is red, I would rather find out why than argue about who owns it.
Uptime Overview
99.94%
Career availability
40+
Engineers on-call w/ me
System Components / Core Skills
Incident Response & On-call
Operational
99.98% · 90d
90 days agotoday
Observability (Prometheus, Grafana, OpenTelemetry)
Operational
99.91% · 90d
90 days agotoday
Kubernetes & Terraform (AWS, GCP)
Operational
99.96% · 90d
90 days agotoday
Capacity Planning & Load Testing
Degraded (learning k6 internals)
99.72% · 90d
90 days agotoday
Incident History / Career
Site Reliability Manager
Vodenka Cloud
2022-03 → present
resolved
- Lead an SRE team of 7 covering 214 services; cut Sev-1 count from 9 in 2022 to 3 in 2024.
- Rewrote the on-call rotation so no engineer takes more than 1 primary week per month, dropping pager fatigue reports by roughly 60%.
- Set SLOs on the top 30 services and tied a real error budget policy to deploy freezes; teams now pause on their own before I have to ask.
Senior Reliability Engineer
Bela Tok Systems
2019-06 → 2022-02
resolved
- Owned the observability stack for a payments platform doing 4.1M requests/day; brought MTTR from 47 min down to 12 min over 18 months.
- Built 38 runbooks and a trigger-linked alert catalog, cutting duplicate pages by 45%.
- Ran the chaos testing program (monthly game days) that surfaced 11 latent failures before they hit production.
DevOps Engineer
Morje Analytics
2016-09 → 2019-05
resolved
- Migrated 60+ services from hand-rolled scripts to Terraform, cutting environment setup from 2 days to 40 minutes.
- Stood up the first CI/CD pipeline for the team; deploy frequency went from weekly to about 15 times a day.
Subscriptions
Toolchain
Prometheus
Grafana
Kubernetes
Terraform
Go
Python
PagerDuty
OpenTelemetry
AWS
GCP
Bash
PostgreSQL
Certifications & Education
- CKA: Certified Kubernetes Administrator (2021)
- AWS Solutions Architect, Associate (2020)
- B.Sc. Computer Science, University of Ljubljana (2016)