Senior Site Reliability Engineer
Intellias is looking for a Senior Site Reliability Engineer to join a large-scale digital healthcare initiative focused on building and operating a secure, highly available, and reliable cloud platform. This is a hands-on production engineering role for an experienced SRE/Dev Ops engineer who combines deep expertise in AWS and Kubernetes with infrastructure automation, observability, reliability engineering, and advanced troubleshooting skills. You will help operate and evolve a production healthcare platform running entirely on AWS, establish reliability standards, automate infrastructure and operational workflows, and ensure the platform remains scalable, observable, secure, and resilient.
Project Overview The project is a cloud-native healthcare data platform connecting data from EHRs, digital health applications, wearables, portals, and other healthcare sources to enable integrated and personalized healthcare experiences. The production platform runs on AWS and Kubernetes/EKS and supports distributed services handling sensitive healthcare data at scale. Reliability is treated as an engineering discipline built around SLIs, SLOs, error budgets, automation, observability, and well-defined operational practices. The infrastructure ecosystem includes AWS, EKS, Kubernetes, Terraform/Open Tofu/Pulumi, Argo CD, Helm, Git Hub Actions, Grafana, Datadog, Open Telemetry, and modern cloud-native technologies. The role includes significant production ownership and participation in an on-call rotation.
Key Responsibilities Operate and evolve a production healthcare platform on AWS and Kubernetes/EKS, ensuring reliability, scalability, security, and operational efficiency. Design, implement, and maintain infrastructure using Infrastructure-as-Code, minimizing manual configuration and operational dependencies. Automate infrastructure provisioning, deployments, monitoring, scaling, recovery, and recurring operational processes. Define and continuously improve SLIs, SLOs, error budgets, alerts, dashboards, and reliability metrics. Build and maintain observability solutions that proactively identify and resolve issues before they impact users. Operate and optimize Kubernetes/EKS environments, including cluster lifecycle, scaling, upgrades, networking, workload reliability, and troubleshooting. Build and evolve Git Ops-based infrastructure and deployment workflows using tools such as Argo CD and Helm. Design and maintain reliable, secure, and efficient CI/CD pipelines, preferably using Git Hub Actions. Improve platform resilience through fault-tolerant architecture, automated recovery, disaster recovery, and failure-testing practices. Investigate and resolve complex production issues across infrastructure, Kubernetes, networking, distributed services, and application layers. Participate in on-call rotations, incident response, root-cause analysis, and post-incident improvements. Reduce operational toil by replacing repetitive manual processes with reliable automation. Develop reusable infrastructure tooling and self-service capabilities for engineering teams. Maintain technical documentation, runbooks, operational procedures, and infrastructure standards. Apply security-by-design principles across infrastructure, access management, networking, secrets, and production operations. Collaborate with software engineering, security, data, and product teams to address infrastructure and reliability requirements early in the development lifecycle. Evaluate emerging cloud-native technologies and introduce them where they provide measurable improvements in reliability, security, scalability, or engineering efficiency.
Requirements5+ years of professional experience in Site Reliability Engineering, Dev Ops, Platform Engineering, Cloud Infrastructure, or a comparable production engineering role. Strong hands-on experience with AWS and production cloud-native environments. Deep practical knowledge of Kubernetes, preferably AWS EKS, including production operations, scaling, upgrades, networking, troubleshooting, and workload management. Strong Linux expertise, including system internals, networking, process management, resource troubleshooting, and production diagnostics. Strong Infrastructure-as-Code experience with Terraform, Open Tofu, Pulumi, or comparable technologies. Hands-on experience with Git Ops, using tools such as Argo CD, Helm, or equivalent. Strong automation and scripting skills with Python, Go, Bash, or comparable languages. Experience designing and maintaining production CI/CD pipelines, preferably with Git Hub Actions. Strong understanding of SRE principles, including SLIs, SLOs, error budgets, availability, resilience, capacity management, and incident response. Hands-on experience with observability and monitoring platforms such as Grafana, Datadog, Prometheus, or comparable technologies. Strong understanding of networking fundamentals, including TCP/IP, DNS, load balancing, firewalls, routing, and TLS. Experience operating and troubleshooting distributed microservices architectures in production. Strong understanding of cloud security principles, including least privilege, IAM, secrets management, network security, and secure infrastructure configuration. Experience designing for high availability, disaster recovery, fault tolerance, and operational resilience. Strong production troubleshooting skills with the ability to investigate complex issues across infrastructure, networking, Kubernetes, and application layers. Experience participating in production on-call and incident response rotations. Strong ownership mindset with a focus on automation and eliminating operational toil. Professional English sufficient for direct collaboration with U. S.-based engineering and product stakeholders.
Nice to Have Hands-on experience with Open Telemetry instrumentation and distributed tracing. Experience operating infrastructure in HIPAA, HITRUST, SOC 2, or similarly regulated environments. Previous experience in healthcare, Health Tech, or platforms handling PHI/PII and other sensitive data. Advanced experience with Prometheus, metrics, and alerting architectures. Experience designing and operating multi-region or highly available AWS architectures. Experience building internal infrastructure or platform tooling for engineering teams. Contributions to open-source cloud-native or infrastructure projects. Practical experience using AI-assisted engineering tools for infrastructure automation, incident analysis, troubleshooting, documentation, or operational workflows.
Why This Position? This role provides an opportunity to take significant ownership of the reliability and infrastructure of a production healthcare platform operating entirely in the cloud. You will work with modern cloud-native technologies including AWS, Kubernetes/EKS, Infrastructure-as-Code, Git Ops, CI/CD, and distributed observability, while solving real production challenges around availability, scalability, security, and operational efficiency. Reliability is treated as an engineering discipline rather than reactive operations, with a strong focus on automation, measurable SLOs, proactive observability, resilient architecture, and continuously reducing operational toil. You will have substantial technical autonomy and direct production impact. You will not simply maintain infrastructure — you will engineer the platform, automate it, measure its reliability, troubleshoot it at every layer, and continuously improve how it operates.