Runware Logo

Runware

Senior Site Reliability Engineer

Posted 3 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Senior level
Remote
Hiring Remotely in United Kingdom
Senior level
Ensure reliability, performance, and resilience of Runware's distributed, GPU-backed production platform. Define SLIs/SLOs and observability, investigate incidents across services, databases, networking and queues, automate remediation and deployment safety, participate in on-call rotations, conduct RCA and implement long-term fixes, and collaborate on capacity planning and scaling.
The summary above was generated by AI

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.

What you’ll do
  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

Requirements
  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing
  • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
  • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
Bonus
  • Experience operating high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
  • Experience building automated scaling, capacity management or self-healing systems

Benefits

We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.

Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off – vacation, sick days, public holidays
  • Meaningful stock options – share in the upside you create
  • Remote-first setup – work from home anywhere we can employ you
  • Flexible hours – own your schedule outside core collaboration blocks
  • Family leave – paid maternity, paternity, and caregiver time
  • Company retreats – twice-yearly gatherings in inspiring locations

Similar Jobs

24 Days Ago
In-Office or Remote
Senior level
Senior level
Information Technology • Software
Lead operation and scaling of Linux-based infrastructure and production Kubernetes clusters across bare-metal and virtualized environments. Implement automation (Ansible, Bash/Python, GitOps), design complex networking (VLANs, L2/L3, VPNs), maintain observability stacks, run incident response and on-call rotations, define SLOs/SLIs and SOPs, manage virtualization (OpenStack/Proxmox/VMware) and bare-metal provisioning (MAAS), and collaborate with cross-functional teams to improve reliability and capacity planning.
Top Skills: AnsibleBashCloud-InitCloudflare ApiDebianDns AutomationElkGitGitopsGpu InfrastructureGrafanaGraylogIstioKubernetesL2 RoutingL3 RoutingLinkerdLinuxLokiMaasNetwork PoliciesOpenstackPreseedPrometheusProxmoxPxePythonRbacUbuntuVlansVMwareVpn
24 Days Ago
In-Office or Remote
Senior level
Senior level
Information Technology • Software
Operate and scale production Kubernetes clusters across bare-metal and virtualized environments; automate provisioning and ops with Ansible/Bash/Python and GitOps; design networking and multi-site connectivity; deploy observability stacks; lead incident response and on-call rotations; define SLOs/SLIs and SOPs; manage virtualization (OpenStack/Proxmox/VMware) and bare-metal provisioning (MAAS/PXE).
Top Skills: AnsibleBashCloud-InitDebianElkGitGitopsGrafanaGraylogKubernetesL2 RoutingL3 RoutingLokiMaasOpenstackPreseedPrometheusProxmoxPxePythonUbuntuVlansVMwareVpns
24 Days Ago
In-Office or Remote
Senior level
Senior level
Information Technology • Software
Operate and scale Linux (Debian/Ubuntu) infrastructure and Kubernetes clusters across bare-metal, virtualized and on-prem environments. Build automation (Ansible, Bash/Python, GitOps), networking (VLANs, L2/L3, VPN), observability stacks, SLOs/SLIs, incident response and on-call rotations. Manage virtualization (OpenStack, Proxmox, VMware), bare-metal provisioning (MAAS, PXE), create SOPs, and coordinate cross-functional teams to improve availability and performance.
Top Skills: AnsibleBashCloud-InitDebianElkGitGitopsGrafanaGraylogKubernetesL2/L3 RoutingLinuxLokiMaasOpenstackPreseedPrometheusProxmoxPxePythonUbuntuVlansVMwareVpn

What you need to know about the Belfast Tech Scene

If asked to name the birthplace of the RMS Titanic, you might not say Belfast. Similarly, if asked to name Europe's leading destination for foreign direct investment in new software development, Belfast might not come to mind. Yet, both are true. The city has emerged as a tech powerhouse, recently ranked among the best in the U.K. for tech careers — especially for software developers. It also leads the U.K. with the highest percentage of software development jobs advertised.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account