The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering podcast cover
Fexingo Technology

The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your system is already in production? Lucas and Luna don't pretend to have final answers — they build the conversation so you can draw your own. If you've ever argued about whether a page was necessary or whether an SLO should be tightened, this is your show.

#SiteReliabilityEngineering#SRE#Uptime#ProductionEngineering#IncidentResponse#ErrorBudgets#SLOs#Postmortem#ToilAutomation#CapacityPlanning#Observability#DevOps#PlatformEngineering#Resilience#OnCall#FexingoBusiness#BusinessPodcast#Technology

Support Fexingo

Episodes

141 episodes

How SRE Teams Use Load Shedding to Protect Core Services

Aug 1, 2026 · 6:34
0:000:00

How Observability Pipelines Cut Costs and Noise in SRE

Jul 30, 2026 · 6:22
0:000:00

How Error Budgets Help SRE Teams Stop Bad Deployments

Jul 30, 2026 · 7:19
0:000:00

How SRE Teams Use Dependency Graphs to Prevent Cascading Failures

Jul 29, 2026 · 7:52
0:000:00

How SRE Teams Use Configuration Validation to Prevent Outages

Jul 29, 2026 · 9:02
0:000:00

How SRE Teams Use Incident Command Systems for Major Outages

Jul 28, 2026 · 12:03
0:000:00

How Toil Budgets Free SRE Time for Reliability Engineering

Jul 28, 2026 · 7:47
0:000:00

How SRE Teams Cut Mean Time to Repair With Automation

Jul 27, 2026 · 9:53
0:000:00

How Game Days Sharpen Incident Response

Jul 27, 2026 · 8:04
0:000:00

How Chaos Engineering Makes Systems More Resilient

Jul 26, 2026 · 8:14
0:000:00

How SRE Teams Use Incident Severity Classification to Prioritize Response

Jul 26, 2026 · 7:11
0:000:00

How SRE Teams Use AIOps to Accelerate Incident Detection

Jul 25, 2026 · 6:52
0:000:00

How SRE Teams Use Blameless Postmortems to Improve Reliability

Jul 24, 2026 · 9:28
0:000:00

How SRE Teams Use Runbooks to Reduce Incident Response Time

Jul 23, 2026 · 13:45
0:000:00

How SRE Teams Use Cost-to-Serve Analysis to Optimize Infrastructure

Jul 23, 2026 · 11:24
0:000:00

How SRE Teams Use Rolling Backups to Recover From Ransomware

Jul 22, 2026 · 8:32
0:000:00

How SRE Teams Use Synthetic Monitoring to Catch Outages Before Users Do

Jul 22, 2026 · 9:13
0:000:00

How SRE Teams Use Capacity Planning to Prevent Outages Before They Happen

Jul 21, 2026 · 8:23
0:000:00

How SRE Teams Use Feature Flags to Control Risk in Production

Jul 21, 2026 · 10:42
0:000:00

How SRE Teams Use Graceful Degradation to Keep Services Running

Jul 20, 2026 · 7:48
0:000:00

How SRE Teams Use Canary Deployments to Reduce Blast Radius

Jul 19, 2026 · 11:26
0:000:00

How SRE Teams Use Traffic Shadowing to Test in Production

Jul 19, 2026 · 10:29
0:000:00

How SRE Teams Use Observability Signals to Diagnose Production Issues

Jul 18, 2026 · 7:28
0:000:00

How SRE Teams Use Burn Rate Alerts to Stay Within Error Budgets

Jul 18, 2026 · 8:53
0:000:00

How Slack Cut Mean Time to Acknowledge by 60 Percent With On-Call Orchestration

Jul 17, 2026 · 10:14
0:000:00

How SRE Teams Use Service Level Objectives to Align Business and Engineering

Jul 17, 2026 · 8:46
0:000:00

How SRE Teams Use Load Shedding to Protect Critical Services

Jul 16, 2026 · 12:33
0:000:00

How SRE Teams Use Toil Budgets to Automate the Right Things

Jul 16, 2026 · 10:16
0:000:00

How SRE Teams Use Fault Trees to Root Out Latent Defects

Jul 15, 2026 · 8:50
0:000:00

How SRE Teams Use Game Days to Build Incident Muscle Memory

Jul 15, 2026 · 8:20
0:000:00

How SRE Teams Use Error Budgets to Balance Reliability and Velocity

Jul 14, 2026 · 10:51
0:000:00

How SRE Teams Use Incident Metrics to Improve Postmortem Quality

Jul 14, 2026 · 9:45
0:000:00

How SRE Teams Use Chaos Engineering to Find Hidden Failure Modes

Jul 13, 2026 · 9:05
0:000:00

How SRE Teams Use Latency SLOs to Improve User Experience

Jul 13, 2026 · 6:47
0:000:00

How SRE Teams Use Incident Postmortems for Systemic Improvement

Jul 12, 2026 · 10:34
0:000:00

How SRE Teams Use Capacity Planning to Prevent Outages Before They Happen

Jul 12, 2026 · 9:18
0:000:00

How SRE Teams Use Non-Abstract Large System Design to Prevent Outages

Jul 11, 2026 · 8:34
0:000:00

How SRE Teams Use Runbooks to Standardize Incident Response

Jul 11, 2026 · 12:01
0:000:00

How Airbnb Uses Traffic Shifting to Prevent Cascading Failures

Jul 10, 2026 · 8:17
0:000:00

How SRE Teams Manage Cognitive Load During Incidents

Jul 10, 2026 · 10:08
0:000:00

How SRE Teams Use Dependency Graphs to Prevent Cascading Failures

Jul 9, 2026 · 10:05
0:000:00

How SRE Teams Use Canary Deployments to Reduce Release Risk

Jul 9, 2026 · 7:36
0:000:00

How SRE Teams Use Incident Cost Metrics to Justify Reliability Investment

Jul 8, 2026 · 7:50
0:000:00

How SRE Teams Use Observability Pipelines to Reduce Data Costs

Jul 8, 2026 · 11:08
0:000:00

How SRE Teams Use Blameless Culture to Improve Incident Response

Jul 7, 2026 · 10:51
0:000:00

How SRE Teams Use Incident Severity Frameworks to Triage Faster

Jul 7, 2026 · 13:04
0:000:00

How SRE Teams Use Synthetic Monitoring to Catch Problems Before Users Do

Jul 6, 2026 · 10:23
0:000:00

How SRE Teams Use Incident Command Systems to Coordinate Response

Jul 6, 2026 · 7:05
0:000:00

How SRE Teams Use Saturation Metrics to Prevent Capacity Crises

Jul 5, 2026 · 8:20
0:000:00

How SRE Teams Use SLO Burn Rates to Detect Problems Early

Jul 5, 2026 · 8:55
0:000:00

How SRE Teams Use Game Days to Build Incident Muscle Memory

Jul 4, 2026 · 10:33
0:000:00

How SRE Teams Use Error Budgets to Balance Reliability and Velocity

Jul 4, 2026 · 11:49
0:000:00

How SRE Teams Use Incident Metrics to Improve Response

Jul 3, 2026 · 9:41
0:000:00

How SRE Teams Use Cost Optimization to Reduce Cloud Waste

Jul 3, 2026 · 8:36
0:000:00

How SRE Teams Use Toil Budgets to Protect Engineering Time

Jul 2, 2026 · 11:31
0:000:00

How SRE Teams Use Structured Fails to Learn Faster

Jul 2, 2026 · 10:56
0:000:00

How SRE Teams Use Post-Incident Reviews for System Improvements

Jul 1, 2026 · 8:47
0:000:00

How SRE Teams Use Capacity Planning to Prevent Outages

Jul 1, 2026 · 7:39
0:000:00

How SRE Teams Use Chaos Engineering to Build Resilient Systems

Jun 30, 2026 · 11:45
0:000:00

How SRE Teams Use Cost of Delay to Prioritize Reliability Work

Jun 30, 2026 · 12:48
0:000:00

How SRE Teams Use Latency Budgets to Meet Performance SLOs

Jun 29, 2026 · 9:06
0:000:00

How SRE Teams Use Runbooks to Streamline Incident Response

Jun 29, 2026 · 13:37
0:000:00

How SRE Teams Use Observability to Reduce Mean Time to Detect

Jun 28, 2026 · 8:56
0:000:00

How SRE Teams Use Service Level Agreements to Set Expectations

Jun 28, 2026 · 8:33
0:000:00

How SRE Teams Use Canary Deployments to Reduce Risk

Jun 27, 2026 · 10:50
0:000:00

How SRE Teams Use DORA Metrics to Measure DevOps Performance

Jun 27, 2026 · 10:23
0:000:00

How SRE Teams Use Service Level Objectives to Drive Reliability

Jun 26, 2026 · 10:53
0:000:00

How SRE Teams Use Blameless Culture to Improve Incident Response

Jun 26, 2026 · 8:26
0:000:00

How SRE Teams Use Blameless Postmortems to Build Trust

Jun 25, 2026 · 8:26
0:000:00

How SRE Teams Use Fault Tree Analysis to Prevent Root Causes

Jun 25, 2026 · 11:49
0:000:00

How SRE Teams Use AI for Incident Triage and Root Cause Analysis

Jun 24, 2026 · 11:02
0:000:00

How SRE Teams Use Game Days to Test Incident Response

Jun 24, 2026 · 6:55
0:000:00

How SRE Teams Use Error Budgets to Balance Reliability and Velocity

Jun 23, 2026 · 9:00
0:000:00

How SRE Teams Use Infrastructure as Code to Prevent Configuration Drift

Jun 23, 2026 · 11:03
0:000:00

How SRE Teams Use Incident Response Playbooks

Jun 22, 2026 · 7:54
0:000:00

How SRE Teams Use Readiness Checks to Prevent Bad Deployments

Jun 22, 2026 · 8:07
0:000:00

How SRE Teams Use Cost Attribution to Prioritize Reliability Work

Jun 21, 2026 · 8:54
0:000:00

How SRE Teams Use Toil Budgets to Automate Smarter

Jun 21, 2026 · 7:57
0:000:00

How SRE Teams Use Capacity Planning to Prevent Outages

Jun 20, 2026 · 9:21
0:000:00

How SRE Teams Use Capacity Planning to Prevent Outages

Jun 20, 2026 · 9:21
0:000:00

SRE Teams Are Using Chaos Engineering to Test Resilience

Jun 19, 2026 · 10:56
0:000:00

How SRE Teams Use Postmortem Action Items to Prevent Recurrence

Jun 19, 2026 · 8:15
0:000:00

How SRE Teams Use Incident Severity Classification to Prioritize Response

Jun 18, 2026 · 9:15
0:000:00

How SRE Teams Use Post-Incident Reviews as Learning Tools

Jun 18, 2026 · 9:25
0:000:00

How SRE Teams Use Cost of Delay to Prioritize Reliability Work

Jun 17, 2026 · 9:43
0:000:00

How SRE Teams Reduce Incident Noise with Intelligent Alert Routing

Jun 17, 2026 · 9:11
0:000:00

How SRE Teams Use Incident Cost Analysis to Prioritize Reliability Investments

Jun 16, 2026 · 9:07
0:000:00

How SRE Teams Use On-Call Compensation to Prevent Burnout

Jun 16, 2026 · 8:37
0:000:00

SRE Teams Use SLO Burn Rate Alerts to Detect Incidents Faster

Jun 15, 2026 · 9:09
0:000:00

How SRE Teams Use Software Bill of Materials for Supply Chain Security

Jun 15, 2026 · 9:43
0:000:00

How SRE Teams Use Feature Flags to Reduce Deployment Risk

Jun 14, 2026 · 9:37
0:000:00

How SRE Teams Use Stress Testing to Simulate Real Workloads

Jun 14, 2026 · 11:19
0:000:00

How SRE Teams Use Game Days to Build Incident Muscle Memory

Jun 13, 2026 · 8:46
0:000:00

How SRE Teams Use Error Budgets to Align Risk and Velocity

Jun 13, 2026 · 8:48
0:000:00

How SRE Teams Use SLIs to Define Reliability

Jun 12, 2026 · 7:12
0:000:00

How SRE Teams Use Cognitive Load Management to Prevent Burnout

Jun 12, 2026 · 9:47
0:000:00

How SRE Teams Use Observability to Find Unknown Unknowns

Jun 11, 2026 · 10:06
0:000:00

How SRE Teams Use Dependency Graphs to Predict Outages

Jun 11, 2026 · 7:50
0:000:00

How SRE Teams Use toil budgets to prioritize automation

Jun 10, 2026 · 9:03
0:000:00

How SRE Teams Use Service Level Objectives to Drive Daily Decisions

Jun 10, 2026 · 8:54
0:000:00

How SRE Teams Use Canary Deployments to Reduce Release Risk

Jun 9, 2026 · 8:32
0:000:00

How SRE Teams Use Chaos Engineering to Test Resilience

Jun 9, 2026 · 10:50
0:000:00

How SRE Teams Use Capacity Planning to Prevent Outages

Jun 8, 2026 · 10:19
0:000:00

How SRE Teams Use Immutable Infrastructure to Eliminate Configuration Drift

Jun 8, 2026 · 9:18
0:000:00

How SRE Teams Use Auto-Remediation to Resolve Incidents Without Humans

Jun 7, 2026 · 12:29
0:000:00

How SRE Teams Use Incident Command Systems to Coordinate Response

Jun 7, 2026 · 9:34
0:000:00

How SRE Teams Use Blameless Postmortems to Build Better Systems

Jun 6, 2026 · 8:58
0:000:00

How SRE Teams Use Postmortems That Actually Change Behavior

Jun 6, 2026 · 8:17
0:000:00

How SRE Teams Use Runbook Automation to Reduce Human Error

Jun 5, 2026 · 8:14
0:000:00

How SRE Teams Use Cost Optimization to Balance Performance and Budget

Jun 5, 2026 · 6:48
0:000:00

How SRE Teams Use Load Shedding to Survive Traffic Spikes

Jun 4, 2026 · 9:51
0:000:00

How SRE Teams Use Feature Flags to Reduce Incident Risk

Jun 4, 2026 · 11:00
0:000:00

How SRE Teams Use Incident Metrics to Reduce Mean Time to Resolve

Jun 3, 2026 · 6:38
0:000:00

How Cloud SREs Use Circuit Breakers to Prevent Cascading Failures

Jun 3, 2026 · 14:03
0:000:00

How SREs Use Error Budgets to Balance Reliability and Velocity

Jun 2, 2026 · 8:56
0:000:00

How SRE Teams Use Game Days to Build Muscle Memory for Incidents

Jun 2, 2026 · 8:13
0:000:00

How SRE Teams Use Error Budgets to Balance Reliability and Velocity

Jun 1, 2026 · 8:07
0:000:00

SRE Runbooks That Actually Get Followed

Jun 1, 2026 · 11:02
0:000:00

How SRE Teams Use Observability to Reduce Mean Time to Acknowledge

May 31, 2026 · 8:30
0:000:00

How SRE Teams Use Synthetic Monitoring to Catch Outages First

May 31, 2026 · 11:02
0:000:00

How SRE Teams Use Traffic Shadowing for Safe Testing

May 30, 2026 · 11:11
0:000:00

How SRE Teams Use Canary Deployments to Reduce Blast Radius

May 30, 2026 · 10:33
0:000:00

How SRE Teams Use Data to Predict Incidents Before They Happen

May 29, 2026 · 7:49
0:000:00

How SRE Teams Use Capacity Planning to Prevent Black Friday Outages

May 29, 2026 · 8:45
0:000:00

How SRE Teams Use Service Level Objectives to Drive Business Decisions

May 28, 2026 · 10:46
0:000:00

How SRE Teams Use Toil Budgets to Prioritise Automation

May 28, 2026 · 6:57
0:000:00

How SRE Teams Handle On-Call Burnout Without Burning Out

May 27, 2026 · 13:04
0:000:00

How SRE Teams Use Chaos Engineering for Non-Netflix Systems

May 27, 2026 · 8:33
0:000:00

How Microsoft SREs Automate Capacity Planning at Cloud Scale

May 26, 2026 · 10:59
0:000:00

How GitHub SREs Run Postmortems Without Blame

May 26, 2026 · 9:07
0:000:00

How Cloudflare Handles 46 Million Requests Per Second With SRE

May 25, 2026 · 7:25
0:000:00

How AWS Observability Detects Outages Before Customers Call

May 25, 2026 · 6:58
0:000:00

How Netflix Chaos Engineering Keeps the Stream Alive

May 24, 2026 · 7:41
0:000:00

How Stripe Cut Incident Remediation Time by 40 Percent

May 24, 2026 · 10:31
0:000:00

How Google Manages Incident Stress with SRE Culture

May 23, 2026 · 8:20
0:000:00

How PagerDuty Calculated the True Cost of Alert Fatigue

May 23, 2026 · 9:05
0:000:00

How Shopify Uses Game Theory to Tame Its Incident Roster

May 22, 2026 · 9:23
0:000:00

How Slack Cut Alert Noise by 90 Percent

May 22, 2026 · 9:08
0:000:00

Why Error Budgets Changed How SRE Teams Sleep at Night

May 21, 2026 · 9:19
0:000:00

Incident Response Playbooks That Actually Work

May 21, 2026 · 6:23
0:000:00

The Cost of a Second Downtime at a Major Bank

May 19, 2026 · 9:44
0:000:00