Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Site Reliability Engineering Lead - Taiwan

Home > Other programming jobs

Site Reliability Engineering Lead - Taiwan in Others

  • Obsidian Security
  • Full time
  • Email
  • Taipei, Taiwan

Responsibilities

  • Establish and lead the Taiwan SRE function, including its technical roadmap, operating model, hiring plan, and global-team relationships.
  • Improve the availability, performance, scalability, security, and cost efficiency of production services.
  • Define service-level indicators, service-level objectives, error budgets, and operational health metrics.
  • Build observability across applications, data pipelines, APIs, infrastructure, and customer-facing workflows.
  • Lead production incident response, incident coordination, customer-impact assessment, service recovery, and post-incident reviews.
  • Establish on-call practices, escalation paths, runbooks, and incident command processes.
  • Develop automation that reduces manual operations, deployment risk, recovery time, and repetitive work.
  • Partner with engineering teams on production readiness, resilience testing, failure-mode analysis, capacity planning, and safe service rollout.
  • Improve deployment and change-management practices through progressive delivery, automated validation, and reliable rollback mechanisms.
  • Coordinate infrastructure vulnerability and CVE management, including assessment, prioritization, remediation, validation, and reporting.
  • Mentor SREs and software engineers and lead cross-team improvements to operational practices.
  • Recruit and develop a high-performing Taiwan SRE team while collaborating across Taiwan, the US, the UK, and Australia.

Requirements

  • Significant experience in site reliability engineering, production engineering, cloud infrastructure, DevOps, or a closely related discipline.
  • Experience leading SRE, infrastructure, or production operations initiatives and mentoring engineers.
  • Strong hands-on experience operating cloud-native SaaS products or distributed systems in production.
  • Deep knowledge of several areas including public cloud infrastructure, containers and orchestration, infrastructure as code, deployment automation, observability, incident management, capacity planning, networking, storage, databases, and distributed systems.
  • Demonstrated experience defining and using service-level objectives and operational metrics.
  • Experience improving production reliability through engineering and automation.
  • Practical knowledge of infrastructure security, vulnerability management, CVE remediation, and patching processes.
  • Ability to troubleshoot complex failures across services, data systems, networks, and cloud infrastructure.
  • Strong judgment during high-severity incidents and clear communication under pressure.
  • Experience influencing application and platform teams to adopt stronger operational practices.
  • Strong written and verbal communication skills in English.
  • Experience operating security, identity, data analytics, or other high-volume enterprise SaaS platforms is preferred.
  • Experience with large-scale data ingestion, new SRE team creation, chaos engineering, resilience testing, automated remediation, policy-as-code, compliance programs, or follow-the-sun operations is preferred.
  • Experience working in Taiwan or with globally distributed APAC teams and Mandarin proficiency are preferred.

Obsidian Security

Every enterprise runs on software it doesn't own: third-party apps, AI copilots, and agents logging in with OAuth tokens nobody's reviewing. Obsidian Security secures all of it, discovering every third-party app and AI feature connected to your business, governing risky permissions, and detecting...

Similar positions

Senior DevOps Engineer, Infrastructure & Reliabili

  • Worth AI
  • Full time
  • USA
  • 09/13/2026
  • Remote

Software Engineer Leader

  • Entefy
  • Full time
  • USA
  • 09/13/2026
  • Remote

Artificial Intelligence Leader

  • Entefy
  • Full time
  • USA
  • 09/13/2026
  • Remote

Federated Authentication Engineer - OESIS Framewor

  • OPSWAT
  • Full time
  • Others
  • 09/13/2026
  • Ho Chi Minh City, Vietnam

Head of Data Engineering & Platform

  • Mercury
  • Full time
  • Remote
  • 09/13/2026
  • Salary: $290k-$362k
  • Canada, USA