Brak wyników spełniających kryteria wyszukiwania.

Are you passionate about cutting edge technology?
Do solving some of the Internet's most difficult content delivery challenges interest you?
Join our highly skilled Site Reliability team
The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications
Partner with the best
As an SRE, responsibilities include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform.
Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai's existing observability platform
Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability
Participating in on-call rotations, responding to production incidents, and contributing to post-mortem analysis
Building and improving runbooks for inference-specific operational procedures, integrating into Akamai's existing incident management processes
Contributing to SLO tracking and reporting, identifying trends and areas for improvement
Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures
Collaborating with product engineering teams to troubleshoot complex problems across the stack
Do what you love
To be successful in this role you will:
Have commercial experience in Site Reliability Engineering
Show proficiency in a programming language such as Python or Go, with experience creating automation solutions.
Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues
Show familiarity with Kubernetes and containerization concepts
Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar
Have exposure to CI/CD pipelines and infrastructure-as-code tools (Terraform, SaltStack, or equivalent)
Show a willingness to learn and grow, with genuine curiosity about AI infrastructure and distributed systems
Build your career at Akamai
Our ability to shape digital life today relies on developing exceptional people like you. The kind that can turn impossible into possible. We’re doing everything we can to make Akamai a great place to work. A place where you can learn, grow and have a meaningful impact.
With our company moving so fast, it’s important that you’re able to build new skills, explore new roles, and try out different opportunities. There are so many different ways to build your career at Akamai, and we want to support you as much as possible. We have all kinds of development opportunities available, from programs such as GROW and Mentoring, to internal events like the APEX Expo and tools such as Linkedin Learning, all to help you expand your knowledge and experience here.
Learn more
Not sure if this job is the right match for you or want to learn more about the job before you apply? Schedule a 15-minute exploratory call with the Recruiter and they would be happy to share more details.
Zainteresowany ofertą?
Aplikuj już teraz!