OpenAI02/20/2026
Senior Support Engineer - San Francisco
Salary not specified
MARKET
12,354 ₽median for this role
IT Support / Helpdesk · 64 jobs with disclosed pay
5,714half of the offers: 8,520–16,354145,833
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьSan Francisco
Responsibilities
- 01Be among the foremost technical and troubleshooting experts for the OpenAI API platform, acting as the last line of defense before the core Engineering team.
- 02Proactively identify and implement opportunities to scale support operations using automation and AI advancements, contributing to the future of technical support in an AI-driven era.
- 03Configure and use advanced monitoring and alerting workflows to proactively detect customer-impacting issues in real time.
- 04Partner with engineering to contribute to reliability reviews and preparedness for new features, launches, or strategic customer requirement updates, ensuring operational readiness (monitoring, alerting, fallback plans).
- 05Design and refine incident response processes and documentation across strategic customers, engineering, and support teams.
- 06Analyze operational metrics and incident root cause analyses to identify improvement areas and recommend enhancements to monitoring dashboards, alert configurations, and support workflows.
- 07Provide support coverage during holidays and weekends based on business needs.
- 08Design and run operational processes to monitor top strategic customers and a 24x7 response team, working closely with Infrastructure and Engineering teams to deliver a scalable customer experience.
Requirements
- 01Bachelor’s degree in Computer Science or a related field with a strong software engineering foundation.
- 028+ years of experience in technical operations roles such as SRE/NOC, designing monitoring systems and resolving production issues in fast-paced, mission-critical environments.
- 03Deep familiarity with modern monitoring, alerting, and observability practices, including hands‑on experience setting up or managing metrics, logging, and tracing for distributed systems (SLIs/SLOs, alert tuning, dashboard creation).
- 04Proven experience leading incident response for high‑severity outages, performing real‑time coordination, root cause analysis, and driving follow‑ups (post‑mortems, action items) to prevent recurrence.
- 05Strong skills in scripting or software engineering (e.g., Python) to automate repetitive tasks and integrate tools.
- 06Solid understanding of cloud infrastructure and distributed systems fundamentals, comfortable working with cloud services, load balancers, databases, and containerized applications.
- 07Effective cross‑functional collaboration in a high‑trust environment with strong communication skills to explain technical issues to both engineering and non‑technical stakeholders.
What we offer
- 01Based in San Francisco, CA.
- 02Hybrid work model: 3 days in the office per week.
- 03Relocation assistance offered to new employees.