Site Reliability Engineer
Want this job? Aria tailors your résumé to this role and submits the application on the employer's own platform — on your behalf. Free to start.
About this role
(ID: 2025-1135)
Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH).
Benefits We Offer:
• 100% Medical, Dental & Vision Coverage for Employees
• Paid Time Off and Paid Holidays
• 401K match up to 5%
• Educational Benefits for Career Growth
• Employee Referral Bonus
• Flexible Spending Accounts:
• Healthcare (FSA)
• Parking Reimbursement Account (PRK)
• Dependent Care Assistant Program (DCAP)
• Transportation Reimbursement Account (TRN)
The Site Reliability Engineer role centers on modernizing and consolidating a complex multi-cloud environment across AWS, Azure, and GCP, building a scalable, secure, and observable platform from the ground up using Kubernetes, AI/ML infrastructure, and zero-trust principles. You'll combine DevOps and SRE practices to support mission-driven scientific and clinical programs, emphasizing automation, reliability, compliance, and proactive monitoring while enabling innovation through AI-driven tooling. The team culture is highly collaborative and growth-oriented, valuing experimentation, continuous learning, and cross-functional leadership, with opportunities to shape future multi-cloud and platform engineering solutions.
Responsibilities:
• Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and Open-telemetry tools
• Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements
• Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments
• Automate workload onboarding and offboarding processes, ensuring standardization and governance
• <span style="background-color: white; color: black; font-family: Tahoma;
Stop filling out applications one by one.