OpenAI2 дня назад

Senior Support Engineer - Сингапур

Зарплата не указана
РЫНОК
12 592медиана по профессии
IT Support / Helpdesk · 117 вакансий с указанной зарплатой
5 250половина предложений: 9 583–16 667132 617
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьУдалёнка

Обязанности

  • 01Design and run operational processes to monitor top strategic customers and a 24x7 response team
  • 02Work closely with Infrastructure and Engineering teams to deliver the best possible experience to customers at scale
  • 03Be among the foremost technical and troubleshooting experts for the API platform at OpenAI, acting as the last line of defense before the core Engineering team
  • 04Proactively identify and implement opportunities to scale support operations by leveraging automation and advancements in AI technologies
  • 05Configure and use advanced monitoring and alerting workflows to proactively detect customer impacting issues in real time
  • 06In partnership with engineering, contribute to reliability reviews and preparedness for new features, launches, or strategic customer requirement updates, ensuring operational readiness (monitoring, alerting, fallback plans) is in place
  • 07Design and refine incident response processes and documentation across strategic customers, engineering and support teams
  • 08Analyze operational metrics and incident RCAs to identify areas for improvement, and proactively recommend and implement enhancements to monitoring dashboards, alert configurations, and support workflows
  • 09Provide support coverage during holidays and weekends based on business needs

Требования

  • 01Bachelor’s degree in Computer Science or a related field; strong software engineering foundation
  • 028+ years of experience in technical operations roles such as SRE/NOC, designing monitoring systems and resolving production issues in fast-paced, mission-critical environments
  • 03Strong track record of troubleshooting complex technical problems at the systems level
  • 04Deep familiarity with modern monitoring, alerting, and observability practices; hands‑on experience setting up or managing metrics, logging, and tracing for distributed systems (understanding of SLIs/SLOs, alert tuning, dashboard creation)
  • 05Proven experience leading incident response for high‑severity outages or service disruptions; able to perform real‑time incident coordination, root cause analysis, and drive follow‑ups (post‑mortems, action items) to prevent recurrence
  • 06Knowledge of industry best practices for incident management and fault diagnosis
  • 07Strong skills in scripting or software engineering (e.g., Python or similar) to automate repetitive tasks and integrate tools
  • 08Solid understanding of cloud infrastructure and distributed systems fundamentals; comfortable working with cloud services, load balancers, databases, and containerized applications
  • 09Effective at working cross‑functionally in a high‑trust environment; strong communication skills to explain technical issues and resolutions to both engineering and non‑technical stakeholders; able to coordinate efforts across teams and provide updates during ongoing incidents

Условия

  • 01Based in Singapore
  • 02Hybrid work model: 3 days in the office per week
  • 03Relocation assistance offered to new employees