Carbon3.ai Logo

Carbon3.ai

Site Reliability Engineer

Reposted 18 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Senior level
Remote
Hiring Remotely in United Kingdom
Senior level
Build AI-driven SRE tooling and agentic automation to triage, diagnose, and remediate infrastructure issues. Integrate LLM-powered agents with observability, ITSM, and infrastructure APIs, develop self-service tooling and ChatOps, tune event/alert intelligence, and convert runbooks into auditable executable automations while contributing to operational standards and incident learning.
The summary above was generated by AI

Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations



Role Summary:

We’re hiring SRE/Platform engineers with an automation bias to help build Era4’s operations capability from the ground up. You’ll turn runbooks, alerts and operational workflows into safe, auditable automation and internal tooling that improves reliability across our AI infrastructure and datacentre platform.

 

This is a Platform / SRE role with software engineering, not an AI model-building role. You’ll work closely with operations, platform and engineering teams to reduce manual toil, improve alert quality, and speed up incident response.

 

Key Responsibilities:

  • Build Python-based automation for incident triage, runbook execution, and routine operational tasks.
  • Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate workflows.
  • Improve monitoring signal quality through correlation, enrichment, suppression and deduplication.
  • Build internal tools and self-service capabilities such as CLI utilities, ChatOps integrations and dashboards.
  • Maintain version-controlled runbook-as-code and automation libraries.
  • Translate post-incident learnings into better tooling, automation and operational standards.
  • Support safe, auditable automation for higher-risk actions with appropriate approval controls.

 

Essential Experience:

  • Experience in SRE, Platform Engineering, or production infrastructure operations.
  • Hands-on experience with observability/monitoring tooling (for example Prometheus, Grafana or similar).
  • Exposure to incident management / on-call and converting manual runbooks into automation.
  • Experience with Python for automation, APIs and integrations.

 

Nice To Have:

  • GPU, datacentre or colocation infrastructure experience.
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar).
  • ChatOps tooling (Slack or Microsoft Teams bots).
  • OpenTelemetry, logging or distributed tracing experience.
  • DCIM, IPAM or hypervisor-control-plane integrations.
  • Experience with LLM-assisted or agent-based operational automation.

 

Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

 

Diversity & Inclusion:

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

 

Similar Jobs

5 Days Ago
Remote or Hybrid
Mid level
Mid level
Cloud • Software
Design, operate, and scale large distributed systems for telemetry processing. Build automation, use AI tooling to reduce toil, ensure availability and disaster recovery, participate in on-call incident response, troubleshoot production AWS/Kubernetes environments, and collaborate with application teams to meet SLOs/SLAs.
Top Skills: Ai ToolingAWSGnu/LinuxGoKubernetesPythonTerraform
12 Days Ago
Easy Apply
Remote
United Kingdom
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Maintain and improve reliability, scalability, and automation for user-facing production systems. Build infrastructure tooling, operate Kubernetes-based services, write IaC, participate in on-call and incident response, and advance observability and runbooks to reduce toil and improve platform reliability.
Top Skills: AWSCi/CdGCPGitopsGoInfrastructure As Code (Iac)KubernetesKubernetes Operators/ControllersLoggingMetricsRubySlos/SlisTerraform
21 Days Ago
In-Office or Remote
Senior level
Senior level
Cloud • Software • Analytics
Run and improve production platform reliability by building automation, monitoring, and CI/CD tooling; troubleshoot distributed systems; provide primary operational support; participate in capacity planning, incident response, and postmortems to balance feature velocity with service-level objectives.
Top Skills: AnsibleAWSBashC#ChefCircleCICloudFormationCloudwatchDatadogDockerDynamoDBEc2EcsElkGitlab Ci/CdGoGrafanaJavaJenkinsKubernetesLambdaLokiMimirPagerdutyPowershellPrometheusPuppetPythonRundeckSplunkTempoTerraform

What you need to know about the Belfast Tech Scene

If asked to name the birthplace of the RMS Titanic, you might not say Belfast. Similarly, if asked to name Europe's leading destination for foreign direct investment in new software development, Belfast might not come to mind. Yet, both are true. The city has emerged as a tech powerhouse, recently ranked among the best in the U.K. for tech careers — especially for software developers. It also leads the U.K. with the highest percentage of software development jobs advertised.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account