Skip to content
← Back to job listings

Senior Site Reliability Engineer (Linux Systems & Application Observability)

tastytrade · Chicago, United States

RemoteImported listingfull-time12 days ago

About The Role

Join our team as a Senior Site Reliability Engineer, where you'll play a crucial role in hardening our systems for order execution and market data delivery. You'll work closely with our infrastructure and application engineering teams, contributing directly to our Ruby, Java, and Elixir services. This is a unique opportunity to shape a practice and culture from the ground up, with real client capital on the line.

  • Contribuer directement aux services Ruby, Java et Elixir en travaillant en étroite collaboration avec les équipes d'ingénierie des infrastructures et des applications.
  • Construire une infrastructure auto-réparatrice et tolérante aux pannes, ainsi que des outils internes qui automatisent le travail opérationnel répétitif.
  • Effectuer une analyse des lacunes dans notre pile d'observabilité pour identifier les points aveugles dans la télémétrie, la journalisation et la couverture d'alerte.
  • Strong programming skills in a language such as Python, Ruby, Java, or similar
  • Deep understanding of one or more: distributed systems, Linux systems, cloud-native architectures, containerization
  • Hands-on experience designing fault-tolerant, self-healing distributed systems — not just describing the patterns, but having shipped them
  • Experience running gap analyses on observability/telemetry systems: identifying what's not instrumented, not alerted on, or not visible until it's too late
  • A track record scaling systems under real production load, including capacity planning and architectural bottleneck identification
  • Hands-on experience with OpenTelemetry, Prometheus, and Grafana, with the ability to instrument services directly
  • Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis
  • On-call experience on production systems and comfort building a blameless post-incident review process
  • Working knowledge of SLOs and error budgets as a tool, not the job description; HashiCorp Nomad, Consul, or Vault experience is a strong plus
  • Don't meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they have every single qualification

This is an external listing. JobSpring does not represent or verify the employer. Report this listing