Logo
    Search

    Slight Reliability Episode 58 - Tackling Cloud Cost with Harinder Seera

    en-nzJune 27, 2023

    About this Episode

    In this episode Stephen Townshend and Harinder Seera explore how to monitor and manage the cost of cloud. They discuss FinOps as a cultural practice, anti-patterns for implementing in the cloud, keeping cost down through resources, pricing, and architecture... and much more.

    You can find Harinder on LinkedIn: https://www.linkedin.com/in/harinderseera/

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/

    Twitter: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/

    Recent Episodes from Slight Reliability

    Slight Reliability Episode 83 - An Unfulfilled Promise with Itiel Shwartz

    Slight Reliability Episode 83 - An Unfulfilled Promise with Itiel Shwartz

    This week I hear about all things Kubernetes from Komodor CTO and co-founder Itiel Shwartz. We chat about the promise that was made when Kubernetes first entered the industry, the challenge of getting developers engaged and capable of working in Kubernetes, my hate/hate relationship with Helm but its important contribution to the Kubernetes project, Kubernetes observability, and so much more.

    You can find the Kubernetes for Humans podcast here:
    https://komodor.com/blog/the-kubernetes-for-humans-podcast/
    Or find out more about Komodor here:
    https://komodor.com/
    Or find Itiel on LinkedIn: https://www.linkedin.com/in/itiel-shwartz-18542853/

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    Twitter: https://twitter.com/the_kiwi_sre
    YouTube: https://www.youtube.com/c/SlightReliability
    Instagram: https://www.instagram.com/slight_reliability/
    TikTok: https://www.tiktok.com/@the_kiwi_sre

    This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.

    Slight Reliability Episode 82 - CI/CD with Amin Astaneh

    Slight Reliability Episode 82 - CI/CD with Amin Astaneh

    This week I sit down and have a discussion with Amin Astaneh (from Certo Modo) about CI/CD. We cover the power of the standard change as a way to navigate ITIL while still implementing DevOps practices, what to monitor to make your CI/CD observable, single piece flow, testing in production, and so much more.

    You can find Amin on his company website https://certomodo.io, LinkedIn: https://www.linkedin.com/in/aminastaneh/ and Twitter: https://twitter.com/aastaneh

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    Twitter: https://twitter.com/the_kiwi_sre
    YouTube: https://www.youtube.com/c/SlightReliability
    Instagram: https://www.instagram.com/slight_reliability/
    TikTok: https://www.tiktok.com/@the_kiwi_sre

    This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.

    Slight Reliability Episode 81 - Incident Management in Non-Prod Environments

    Slight Reliability Episode 81 - Incident Management in Non-Prod Environments

    "Environment issues are just incidents that happened to occur in a non-production environment"... so why do we treat them so differently?

    In this first episode of the 2024 season I reflect on how we handle incidents in non-prod environments.

    (Note: Had a few issues with noise suppression in OBS Studio cutting off the start of some words, will sort it for the next episode)

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    Twitter: https://twitter.com/the_kiwi_sre
    YouTube: https://www.youtube.com/c/SlightReliability
    Instagram: https://www.instagram.com/slight_reliability/
    TikTok: https://www.tiktok.com/@the_kiwi_sre

    Slight Reliability Episode 80 - What's Been Bugging Niall Murphy

    Slight Reliability Episode 80 - What's Been Bugging Niall Murphy

    This week I speak with co-author of the original SRE book + the SRE workbook, and renowned speaker Niall Murphy.

    We chat about the state of SRE in the current macro-economic climate and how we're not yet doing a very good job at articulating the value of SRE to leaders, the relationship that velocity and reliability have, the value of new features versus reliability improvements, and *much* more.

    You can find Niall at:

    LinkedIn: https://www.linkedin.com/in/niallm/
    X: https://twitter.com/niallm
    Website: https://relyabilit.ie/

    (and his company Stanza: https://www.stanza.systems/)

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    X: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/

    Slight Reliability Episode 76 - Sampling Distributed Traces with Paige Cruz

    Slight Reliability Episode 76 - Sampling Distributed Traces with Paige Cruz

    Paige Cruz (from Chronosphere) is back. This week we discuss sampling. What is sampling? Why do it? What kinds of sampling are there?

    You can check out Chronosphere's cloud native observability platform here: https://chronosphere.io/

    You can find Paige on:

    LinkedIn: https://www.linkedin.com/in/paigerduty/
    X: https://twitter.com/paigerduty

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    X: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/

    Slight Reliability Episode 79 - Incident Story Time with Valeska Victoria

    Slight Reliability Episode 79 - Incident Story Time with Valeska Victoria

    This week Valeska Victoria returns to share some of her experiences working as an SRE at eBay.

    We look at the cascading effect of production issues in complex integrated environments (how there's often no single root cause), developer literacy of how infrastructure works, the importance of ownership and accountability of reliability, and much more.

    You can find Valeska on:

    LinkedIn: https://www.linkedin.com/in/valeska-victoria/

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    X: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/

    Slight Reliability Episode 78 - Developer Experience with Ankit Jain

    Slight Reliability Episode 78 - Developer Experience with Ankit Jain

    This week I chat with Ankit Jain from aviator.co about developer experience.

    We define developer experience and developer productivity, and how this applies to SRE. We discuss the growing expectation on developers and how this leads to frustration and burnout. We also explore how to measure developer experience and how to start working to make improvements.

    You can check out Aviator's developer experience platform here: https://www.aviator.co/

    You can find Ankit on:

    LinkedIn: https://www.linkedin.com/in/ankitjaindce/

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    X: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/

    Slight Reliability Episode 77 - SRE to DevRel with Liz Fong-Jones

    Slight Reliability Episode 77 - SRE to DevRel with Liz Fong-Jones

    This week I had the privilege of interviewing Liz Fong-Jones from honeycomb.io about DevRel, Developer Advocacy, and how that applies to SRE.

    We discuss the difference between Developer Relations (DevRel) and Developer Advocacy, how Liz got into advocacy, how DevRel helps companies and the community, and some tips on how to get traction with SRE practices in your organisation.

    You can check out Honeycomb's observability platform here: https://www.honeycomb.io/

    You can find Liz on:

    LinkedIn: https://www.linkedin.com/in/efong/
    Website: https://www.lizthegrey.com/ (all her social/links are here)

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    X: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/

    Slight Reliability Episode 75 - Enterprise SRE with Steve McGhee

    Slight Reliability Episode 75 - Enterprise SRE with Steve McGhee

    This week I had the honour of chatting with Steve McGhee (former Google SRE, current Google Reliability Advocate, and co-author of Enterprise Roadmap to SRE).

    We discuss the evolution of SRE from where it began at Google and how it is being adopted by enterprises around the world now (and why this is happening). We talk about getting leadership support and how we get reliability taken seriously, the lies we tell ourselves to justify incidents and issues, leveraging transformation projects to bring SRE to life, how SLOs can act as the fulcrum between dev and ops, the fallacy of the pyramid model of reliability... and so much more.

    You can find Steve at on:

    LinkedIn: https://www.linkedin.com/in/stevemcghee/
    X: https://twitter.com/stevemcghee

    You can find Steve's book "Enterprise Roadmap to SRE" here: https://sre.google/resources/practices-and-processes/enterprise-roadmap-to-sre/

    Steve also mentions the book "A Seat at the Table": https://itrevolution.com/product/a-seat-at-the-table/

    You can find the official Slight Reliability podcast website at: https://slightreliability.com/

    You can find Stephen at:

    LinkedIn: https://www.linkedin.com/in/stephentownshend/
    X: https://twitter.com/the_kiwi_sre
    Instagram: https://www.instagram.com/slight_reliability/