MongoDB08.01.2026

Cloud Operations Engineer

Зарплата не указана
Cork

Обязанности

  • 01Successfully coordinate and collaborate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base
  • 02Help scale the worldwide Cloud Operations Engineering team with the strategic implementation and refinement of new processes and tools
  • 03Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents
  • 04Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution
  • 05Automate routine monitoring and troubleshooting tasks
  • 06Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution
  • 07Assist in performing root cause analysis after incident recovered; identifying any breakdowns in processes or workflows that contributed to the event and what changes need to be made to prevent similar events
  • 08Contribute to documentation of corner case scenarios, troubleshooting workflows and SOPs.
  • 09Work alongside our product management, cloud engineering and support organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure
  • 10Inform executive leadership and escalation management personnel of major outages
  • 11Coordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (proactively from automated monitoring or through reactive alerts via our Technical Services team)
  • 12Creating and monitoring system’s alert dashboards
  • 13Reviewing critical events and system logs
  • 14Accessing customer instances that underpin their production databases
  • 15Performing server administration duties including performance troubleshooting

Требования

  • 01Experience with being an on call DevOps, SRE, or Cloud Operations engineer (at least 2 years)
  • 02Expertise with Linux system administration, configuration, troubleshooting
  • 03Experience in monitoring, system performance data collection and analysis, and reporting
  • 04Knowledge of database operations and concepts
  • 05Expertise with networking technologies like DNS, TCP/IP, etc.
  • 06Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)
  • 07Knowledgeable about a wide range of web and internet technologies
  • 08Capability to write small programs/scripts to solve both short-term systems problems
  • 09A CS/CE degree or equivalent experience
  • 10At least 1 of the following programming languages: Java, Go, Javascript
  • 11A keen interest in learning new things

Условия

  • 01Competitive salary, equity, pension and health insurance
  • 02Regular performance, compensation and development reviews
  • 0320 weeks Maternity & Paternity leave to spend time with new arrivals
  • 04Working out of our Dublin or Cork office under our in-office working model Monday to Friday
  • 05Due to the 24/7 nature of our support organization, certain events throughout the year will require volunteering for coverage outside one’s normal work days or work hours (i.e. regional offsites, regional holidays, etc)
Cloud Operations Engineer · Rekru