Job Description
What You'll Do:
- Operate and improve our multi-account AWS environment running on Amazon EKS, deployed via ArgoCD and Helm
- Build and maintain infrastructure as code with Terraform / Terragrunt — everything goes through code review
- Support product teams' delivery pipelines: CI/CD design, build optimization, deployment reliability across GitLab and GitHub
- Drive infrastructure migration and improvement projects from design through rollout, with clear documentation and audit-ready evidence
- Bring hands-on AI expertise to the team, introducing agent-based automation and AI-driven cost/anomaly analysis practices, and mentoring teammates on effective AI-assisted workflows
- Monitor, troubleshoot, and tune production systems using Datadog
- Participate in the team's on-call rotation and contribute to incident response and post-incident improvements
- Identify toil and eliminate it through automation
Qualification
What We're Looking For:
Core requirements:
- Minimum 7 years of professional experience in DevOps, Site Reliability Engineering, or Platform Engineering roles
- Demonstrated, hands-on expertise in Kubernetes administration and operations; production experience with Amazon EKS strongly preferred
- Strong command of Infrastructure-as-Code practices using Terraform (experience with Terragrunt is a plus), with a track record of managing infrastructure exclusively through code and automated pipelines
- Practical, hands-on experience with AI-assisted developer tooling (e.g., Claude Code, GitHub Copilot, or equivalent), including demonstrated success driving adoption of such tools within a team
- Solid working knowledge of CI/CD practices and tooling, with experience in GitLab CI and/or GitHub Actions
- Proven ability to troubleshoot complex, multi-layered issues, with demonstrated depth in root-cause analysis across the application, container, node, and network layers
- Strong sense of ownership and accountability, with the ability to work independently, proactively communicate progress to stakeholders, and produce clear documentation of work delivered
- Hands-on experience with GitOps methodologies, particularly using ArgoCD
Strongly preferred:
- Karpenter or cluster autoscaling tuning experience
- FinOps / cloud cost optimization experience
- Observability tooling beyond dashboards (Datadog APM, log pipelines, custom monitors)
- Experience with API gateways and edge stacks (Nginx, APISIX, Cloudflare, or similar) as well as edge security components such as WAF
Certification required: Kubernetes and/or AWS certification, with AWS Certified Solutions Architect – Professional, DevOps Engineer – Professional, or Specialty level preferred
Recruiter
Talent Acquisition Team, Human Resources Management Department (Head Office)
Muang Thai Life Assurance Public Company Limited
250 Ratchadaphisek Rd., Huaykwang, Bangkok 10310
Website: www.muangthai.co.th
Line Official Account: @mtlcareer
LinkedIn: Muang Thai Life Assurance Public Company Limited
"หมายเหตุ: ตำแหน่งงานนี้จำเป็นต้องตรวจสอบประวัติอาชญากรรมของบุคคลเมื่อพิจารณารับเข้าทำงาน เพื่อความปลอดภัยและรักษามาตรฐานขององค์กร"
"Remark: This position requires a criminal record information check when consideration for employment to ensure safety and maintain standards of the organization."

