OpenAI20.02.2026
Senior Support Engineer - San Francisco
Зарплата не указана
Полная занятостьSan Francisco
Обязанности
- 01Be among the foremost technical and troubleshooting experts for the OpenAI API platform, acting as the last line of defense before the core Engineering team.
- 02Proactively identify and implement opportunities to scale support operations using automation and AI advancements, contributing to the future of technical support in an AI-driven era.
- 03Configure and use advanced monitoring and alerting workflows to proactively detect customer-impacting issues in real time.
- 04Partner with engineering to contribute to reliability reviews and preparedness for new features, launches, or strategic customer requirement updates, ensuring operational readiness (monitoring, alerting, fallback plans).
- 05Design and refine incident response processes and documentation across strategic customers, engineering, and support teams.
- 06Analyze operational metrics and incident root cause analyses to identify improvement areas and recommend enhancements to monitoring dashboards, alert configurations, and support workflows.
- 07Provide support coverage during holidays and weekends based on business needs.
- 08Design and run operational processes to monitor top strategic customers and a 24x7 response team, working closely with Infrastructure and Engineering teams to deliver a scalable customer experience.
Требования
- 01Bachelor’s degree in Computer Science or a related field with a strong software engineering foundation.
- 028+ years of experience in technical operations roles such as SRE/NOC, designing monitoring systems and resolving production issues in fast-paced, mission-critical environments.
- 03Deep familiarity with modern monitoring, alerting, and observability practices, including hands‑on experience setting up or managing metrics, logging, and tracing for distributed systems (SLIs/SLOs, alert tuning, dashboard creation).
- 04Proven experience leading incident response for high‑severity outages, performing real‑time coordination, root cause analysis, and driving follow‑ups (post‑mortems, action items) to prevent recurrence.
- 05Strong skills in scripting or software engineering (e.g., Python) to automate repetitive tasks and integrate tools.
- 06Solid understanding of cloud infrastructure and distributed systems fundamentals, comfortable working with cloud services, load balancers, databases, and containerized applications.
- 07Effective cross‑functional collaboration in a high‑trust environment with strong communication skills to explain technical issues to both engineering and non‑technical stakeholders.
Условия
- 01Based in San Francisco, CA.
- 02Hybrid work model: 3 days in the office per week.
- 03Relocation assistance offered to new employees.