Senior Software Engineer, Cloud Infrastructure / SRE
Want this job? Aria tailors your résumé to this role and submits the application on the employer's own platform — on your behalf. Free to start.
About this role
Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team.
Oscar is the first health insurance company built around a full stack technology platform and a relentless focus on serving our members. We started Oscar in 2012 to create the kind of health insurance company we would want for ourselves—one that behaves like a doctor in the family.
About the role:
Our Core Technology teams build and maintain the foundational platform upon which all Oscar engineering is built. We are responsible for architecting a world-class, resilient ecosystem using a modern stack centered on AWS/GCP, Terraform, and Kubernetes with developer focused tooling and CI/CD. Our mission is to provide an automated, self-service infrastructure that empowers our engineering organization to move fast without sacrificing security or stability.
You will report into a Staff/Senior Staff Engineer.
Work Location: This is a remote position, open to candidates who reside in: Boston, MA. You will be fully remote; however, our approach to work may adapt over time. Future models could potentially involve a hybrid presence at the hub office associated with your metro area. #LI-Remote
Pay Transparency: The base pay for this role is: $180,504 - $236,911 per year. You are also eligible for employee benefits, participation in Oscar's unlimited vacation program, company equity grants, and annual performance bonuses.
Responsibilities:
• Become the expert on your team's business and technical domains such as DevOps, site reliability, and cloud best practices
• Lead the planning, execution and release of complex technical projects across multiple teams outside of Core Technology
• Work with partners, product managers, and designers to solve challenging problems
• Lead and mentor engineers on the team to improve technology and apply best practices
• Independently responsible for large or complex technology capabilities (set of components or services) within their team's domain or spanning multiple domains
• Facilitates, encourages, and enhances cross-team execution and collaboration; knows when cross-team projects are at risk and actively mitigates risk to deliver on time
• Prolific contributor to the objectives of their functional group, as well as organization-wide projects
• Drives prioritization of technical roadmap and influences prioritization of product roadmap and process enhancements within their team
• Actively identifies and reduces failure domains, designs and builds resilient systems, and strives to reduce adverse effects of an outage.
• Builds software to minimize effort and business impact during maintenance and failures
• Guides the development of Service-Level Objectives (SLOs) for systems they are responsible for
• Own medium to large features or infrastructure projects from technical design through completion
• Compliance with all applicable laws and regulations
• Other duties as assigned
Requirements:
• 6+ years of professional software engineering experience, working with a variety of technologies, and have increasingly impactful accomplishments
• Experience as a major contributor cross-pod or cross-company deliverables
• Experience leading technical contributions, improving the quality of what your teams create, and are excited to build fault-tolerant, and scalable software systems.
• Demonstrates expertise of the practical application of CS concepts within their team.
• Sets and enforces the standard for writing stable, correct, and maintainable code
• Experience mentoring and training more junior engineers
Bonus points:
• Cloud Proficiency: Deep expertise in managing production environments within AWS or GCP at scale.
• Infrastructure as Code: Advanced experience with Terraform or similar IaC tools to manage complex, multi-account structures.
• Orchestration & Delivery: Proven track record with Kubernetes and workflows using ArgoCD.
• SRE Discipline: Strong background in Site Reliability Engineering, including Service Level Objectives (SLOs), error budgets, and incident management.
• CI/CD & Automation: Experience building robust deployment pipelines via GitHub Actions.
• Observability: Proficiency with monitoring using tools like Prometheus, Grafana, or similar.
• Security & Networking: Knowledge of cloud-native security (IAM, VPC peering) and service mesh technologies like Istio.
• Programming: Understanding of at least o
Stop filling out applications one by one.