OpenAI9 days ago
Senior Support Engineer - Toronto
Salary not specified
MARKET
12,375 ₽median for this role
IT Support / Helpdesk · 65 jobs with disclosed pay
5,714half of the offers: 8,583–16,667145,833
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьУдалёнка
Responsibilities
- 01Be among the foremost technical and troubleshooting experts for our API platform at OpenAI. You are the last line of defense before the core Engineering team.
- 02Proactively identify and implement opportunities to scale support operations by leveraging automation and advancements in AI technologies. Contribute to shaping the future of technical support in an AI-driven era.
- 03Configure and use advanced monitoring and alerting workflows to proactively detect customer impacting issues in real time.
- 04In partnership with engineering, contribute to reliability reviews and preparedness for new features, launches, or strategic customer requirement updates. Ensure that operational readiness (monitoring, alerting, and fallback plans) is in place for any such changes.
- 05Design and refine incident response processes and documentation across strategic customers, engineering and support teams.
- 06Analyze operational metrics and incident RCAs to identify areas for improvement. Proactively recommend and implement enhancements to monitoring dashboards, alert configurations, and support workflows.
- 07Provide support coverage during holidays and weekends based on business needs.
Requirements
- 01Have a Bachelor’s degree in Computer Science or a related field. A strong software engineering foundation is important for this role’s success.
- 02Have 8+ years of experience in technical operations roles such as SRE/NOC, designing monitoring systems and resolving production issues in fast-paced and mission-critical environments. A strong track record of troubleshooting complex technical problems at the systems level.
- 03Have deep familiarity with modern monitoring, alerting, and observability practices. Hands‑on experience setting up or managing metrics, logging, and tracing for distributed systems (e.g., understanding of SLIs/SLOs, alert tuning, dashboard creation).
- 04Have proven experience leading incident response for high‑severity outages or service disruptions. Able to perform real‑time incident coordination, root cause analysis, and drive follow‑ups (post‑mortems, action items) to prevent recurrence. Knowledge of industry best practices for incident management and fault diagnosis.
- 05Have strong skills in scripting or software engineering (e.g., Python or similar) to automate repetitive tasks and integrate tools.
- 06Have solid understanding of cloud infrastructure and distributed systems fundamentals. Comfortable working with cloud services, load balancers, databases, and containerized applications.
- 07Are effective at working cross‑functionally in a high‑trust environment. Strong communication skills to explain technical issues and resolutions to both engineering and non‑technical stakeholders. You can coordinate efforts across teams and are comfortable providing updates in the midst of an ongoing incident.
What we offer
- 01This role is based remotely in Toronto, Canada.
- 02The nature of this role will be low volume, high difficulty.