Roche Logo

Roche

Principal/Senior Site Reliability Engineer

Posted 3 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Senior level
Remote
Hiring Remotely in United Kingdom
Senior level
Design, build, and operate resilient cloud and on-prem infrastructure for MLOps and HPC at global scale. Implement Infrastructure as Code, disaster recovery, autoscaling, observability, and chaos engineering. Provide technical leadership, mentor engineers, define SLAs/SLOs/SLIs, and collaborate across teams to optimize ML and HPC workloads.
The summary above was generated by AI

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections,  where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

Join the Computational Sciences Center of Excellence as a Senior Site Reliability Engineer, where the platforms you build accelerate the discovery of transformative medicines. You will work alongside talented engineers in the Data & Digital Catalyst organisation to design resilient, cloud-based systems for MLOps and HPC workloads at global scale. This is a role for someone who wants their engineering craft to have real impact on science and patients.

The Opportunity:

  • You architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.

  • You design for resilience building disaster recovery and failover plans with auto-scaling and load balancing to keep critical systems available worldwide.

  • You strengthen reliability through chaos engineering running experiments that validate systems and surface weaknesses before they become incidents.

  • You build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.

  • You provide technical leadership to a team of engineers, fostering collaboration, innovation, and continuous improvement.

  • You partner across teams to align infrastructure with ML and HPC needs and to advance operational maturity through SLAs, SLOs, SLIs, and error budgets.

Who you are:

  • You bring deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or CloudFormation in AWS, Azure, or GCP for MLOps and HPC workloads.

  • You understand cloud-native and on-prem architectures including autoscaling, serverless, and multi-region deployments, and you are hands-on with Docker, Kubernetes, and Kubeflow.

  • You are an expert in automation scripting confidently in Python, Bash, or Go, with a strong grasp of GPU-accelerated computing and HPC workload scaling.

  • You lead through influence communicating and mentoring with clarity, and solving complex problems with a methodical approach.

  • You hold a degree in Computer Science or a related technical field or bring equivalent experience in software and site reliability engineering.

Preferred:

  • Experience with distributed ML frameworks such as Horovod or TensorFlow Distributed.

  • Familiarity with data engineering pipelines such as Apache Airflow or Apache Spark.

  • Knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.

Relocation benefits are available for this position.

If building the resilient platforms that power the next generation of medicines is your calling, apply now and help accelerate science for patients worldwide.

 

 

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.


Let’s build a healthier future, together.

The statements herein are intended to describe the general nature and level of work being performed by employees, and are not to be construed as an exhaustive list of responsibilities, duties, and skills required of personnel so classified. Furthermore, they do not establish a contract for employment and are subject to change at the discretion of Roche Products Ltd. At Roche Products we believe diversity drives innovation and we are committed to building a diverse and flexible working environment. All qualified applicants will receive consideration for employment without regard to race, religion or belief, sex, gender reassignment, sexual orientation, marriage and civil partnership, pregnancy and maternity, disability or age. We recognise the importance of flexible working and will review all applicants’ requests with care. At Roche difference is valued and we are proud to be an equal opportunity employer where you are encouraged to bring your whole self to work.

Similar Jobs

An Hour Ago
Easy Apply
Remote or Hybrid
UK
Easy Apply
Senior level
Senior level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Lead enterprise deployments of Samsara hardware and SaaS, own project plans and governance, run solution workshops and training, guide change management, coordinate cross-functionally, and manage multiple customer engagements to enable fast time-to-value.
An Hour Ago
Remote or Hybrid
Senior level
Senior level
Consumer Web • Information Technology • Mobile • Music • News + Entertainment • Software
Set product and brand design vision for TIDAL; lead complex end-to-end initiatives; define design systems and quality standards; partner with Product, Engineering, and ML; pioneer AI-driven design; mentor designers; identify new business opportunities.
Top Skills: AIMl
An Hour Ago
In-Office or Remote
Bute, Argyll and Bute, Scotland, GBR
Senior level
Senior level
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Monitor and operate single- and multi-channel playout systems to ensure error-free linear channel output and OTT delivery. Maintain schedules, verify media consistency against client/RBM requirements, log and escalate incidents, communicate changes to managers and clients, and execute emergency/manual processes to minimize outages.
Top Skills: Broadcast PlayoutOtt StreamingPlayout Systems (Single/Multi-Channel)RbmTv Applications

What you need to know about the Belfast Tech Scene

If asked to name the birthplace of the RMS Titanic, you might not say Belfast. Similarly, if asked to name Europe's leading destination for foreign direct investment in new software development, Belfast might not come to mind. Yet, both are true. The city has emerged as a tech powerhouse, recently ranked among the best in the U.K. for tech careers — especially for software developers. It also leads the U.K. with the highest percentage of software development jobs advertised.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account