Skip to content
← Back to job listings

Senior Software Engineer I

Egug · Bengaluru, KA, India

Software DevelopmentExternal listingfull-timeabout 2 hours ago

About The Role

As a Senior Software Engineer I, you will help advance Mission Control’s Automation First strategy by engineering scalable, reliable automation capabilities that reduce operational toil, accelerate incident remediation, and improve service availability. This role will focus on building and evolving automation solutions across the Mission Control ecosystem, including using automation framework Incident Pilot, AIOps integrations, event-driven remediation workflows, self-healing patterns, observability-driven triggers, and automation enablement during onboarding. You will partner closely with Operations Engineering, Tools Engineering, SRE, App Support, and onboarding teams to convert repeatable production support activities into governed, measurable, and resilient automations.

  • Design, build, and support automation capabilities that enable Mission Control to detect, triage, repair, and validate incidents with minimal manual intervention.
  • Engineer event-driven workflows that connect observability signals, incident data, runbooks, and automation execution platforms to deliver consistent remediation outcomes.
  • Advance inhouse automation framework IncidentPilot and related automation platforms by improving configuration workflows, routing logic, automation triggering, failure handling, auditability, reporting, and self-service capabilities.
  • Partner with Mission Control onboarding teams to identify automatable SOPs and runbooks before or during onboarding, ensuring automation is embedded as a standard part of service transition.
  • Build reusable automation patterns for self-healing, job retriggers, pod scaling, component-based validations, health checks, reroutes, incident enrichment, and operational notifications.
  • Integrate automation solutions with platforms such as ServiceNow, event and incident pipelines, Kafka, Dynatrace, Splunk/ELF, Ansible, APIs, and enterprise workflow tools where applicable.
  • Use AI-assisted engineering and AIOps capabilities responsibly to accelerate troubleshooting, automate configuration support, attach relevant SOPs, and identify new automation opportunities.
  • Develop highly available, observable, and supportable services with clear telemetry, dashboards, alerts, operational runbooks, and failure recovery patterns.
  • Establish engineering practices that ensure automation is safe, governed, testable, auditable, and aligned with enterprise controls and production readiness expectations.
  • Monitor automation performance, adoption, success rates, exception paths, and operational savings to demonstrate measurable value and guide continuous improvement.
  • Provide technical leadership and mentorship to engineers by reinforcing coding standards, testing discipline, design reviews, and responsible AI-assisted development practices.
  • Collaborate across Mission Control, SRE, App Support, Observability, and enterprise platform teams to remove friction, standardize repeatable processes, and scale automation across domains.

Bachelor’s degree in Computer Science, Computer Engineering, or comparable experience. 6+ years of software engineering experience in a professional environment, with hands-on experience building production-grade services, integrations, scripts, APIs, or automation workflows using Java, Python, or similar technologies.

  • Experience designing, developing, testing, and supporting automation solutions for production operations, incident management, observability, or reliability engineering use cases.
  • Knowledge of distributed, multi-tiered systems, algorithms, NoSQL databases, relational databases, APIs, and event-driven integration patterns.
  • Experience implementing large-scale automations such as self-healing, pod scaling, job retriggers, component-based testing, health checks, incident enrichment, or automated rerouting.
  • Experience with CI/CD environments and frameworks, including GitHub, Jenkins, automated testing, build pipelines, and deployment pipelines.
  • Practical experience implementing integration solutions, including REST APIs, batch and real-time processing, Kafka, messaging platforms, or workflow orchestration tools.
  • Understanding of observability and cloud-native principles, including monitoring, alerting, service discovery, circuit breakers, distributed tracing, dashboards, and operational telemetry.
  • Experience translating SOPs, runbooks, incident patterns, and production support activities into reliable, repeatable automation workflows.
  • Working knowledge of AI-enabled development tools, AIOps concepts, and responsible use of AI-assisted engineering workflows.
  • Understanding of governance, access controls, auditability, error handling, and production readiness considerations for automation platforms.
  • Experience with automated, functional, integration, and performance testing to validate automation correctness and stability.
  • Ability to communicate technical designs, risks, tradeoffs, and operational impact clearly with engineering and operations partners.
  • Strong analytical thinking, ownership, curiosity, continuous learning, and continuous improvement mindset.

Preferred Qualifications

  • Experience building or supporting incident automation platforms, workflow engines, event-driven remediation tools, or self-healing operational capabilities.
  • Experience with Mission Control-adjacent tooling or patterns such as IncidentPilot, ServiceNow integrations, Event and Incident Pipelines, Kafka-based event ingestion, AIOps workflows, or observability-driven automation.
  • Experience working with automation platforms such as Ansible, EAP-style execution platforms, runbook automation, or enterprise orchestration tools.
  • Experience building production APIs, operational dashboards, reporting views, automation audit trails, or self-service configuration capabilities.
  • Experience using AI to identify automation opportunities, improve incident triage, support SOP discovery, assist with configuration creation, or accelerate engineering delivery.
  • Experience partnering with SRE, App Support, Operations Engineering, Observability, or onboarding teams to standardize support activities and reduce manual operational effort at scale.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing