Powiązane oferty

Brak wyników spełniających kryteria wyszukiwania.

company logo

Senior Site Reliability Engineer

Akamai TechnologiesIkona lokalizacjiGlobalnie

Żródlo publikacji: JustJoin.it
Rodzaj zatrudnienia
Rodzaj zatrudnieniaPełny etat
Doświadczenie
DoświadczenieSenior
Dodano
Dodano21 lipca 2026
Wykryte przez nas
Wykryte przez nas22 lipca 2026
Zarobki
ZarobkiDo uzgodnienia

Are you passionate about cutting edge technology?

Do solving some of the Internet's most difficult content delivery challenges interest you?

Join our highly skilled Site Reliability team

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications

Partner with the best

As an SRE, responsibilities include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform.

  • Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai's existing observability platform

  • Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability

  • Participating in on-call rotations, responding to production incidents, and contributing to post-mortem analysis

  • Building and improving runbooks for inference-specific operational procedures, integrating into Akamai's existing incident management processes

  • Contributing to SLO tracking and reporting, identifying trends and areas for improvement

  • Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures

  • Collaborating with product engineering teams to troubleshoot complex problems across the stack

Do what you love

To be successful in this role you will:

  • Have commercial experience in Site Reliability Engineering

  • Show proficiency in a programming language such as Python or Go, with experience creating automation solutions.

  • Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues

  • Show familiarity with Kubernetes and containerization concepts

  • Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar

  • Have exposure to CI/CD pipelines and infrastructure-as-code tools (Terraform, SaltStack, or equivalent)

  • Show a willingness to learn and grow, with genuine curiosity about AI infrastructure and distributed systems

Build your career at Akamai

Our ability to shape digital life today relies on developing exceptional people like you. The kind that can turn impossible into possible. We’re doing everything we can to make Akamai a great place to work. A place where you can learn, grow and have a meaningful impact.

With our company moving so fast, it’s important that you’re able to build new skills, explore new roles, and try out different opportunities. There are so many different ways to build your career at Akamai, and we want to support you as much as possible. We have all kinds of development opportunities available, from programs such as GROW and Mentoring, to internal events like the APEX Expo and tools such as Linkedin Learning, all to help you expand your knowledge and experience here.

Learn more

Not sure if this job is the right match for you or want to learn more about the job before you apply? Schedule a 15-minute exploratory call with the Recruiter and they would be happy to share more details.


 

Zainteresowany ofertą?

Aplikuj już teraz!