Skip to content
← Back to job listings

Site Reliability Engineer (SRE)

Rocket.net · Remote, Florida, United States

Software DevelopmentRemoteQuick applyfull-timeabout 1 hour ago

About The Role

At Rocket.net , reliability, performance, and customer experience are at the center of everything we build. We are looking for a Site Reliability Engineer to help maintain the health, stability, and performance of our hosting platform while providing advanced technical support to our customers.

The Platform Operations team acts as a critical escalation layer between WordPress Support and Engineering. This role combines infrastructure operations, platform monitoring, troubleshooting, and advanced customer support.

As a Site Reliability Engineer, you will help ensure Rocket.net 's servers, services, and customer environments are operating at the highest standards. You will assist WordPress Support Engineers with complex issues, support VIP customers with advanced technical requests, investigate platform-level problems, and work with internal teams to deliver fast and effective solutions.

Responsibilities

Platform Monitoring & Reliability

  • Monitor the health, availability, and performance of Rocket.net servers, services, and customer environments.
  • Proactively identify infrastructure issues, performance degradation, and potential service disruptions.
  • Investigate alerts and operational events to maintain platform stability.
  • Perform regular platform health checks and ensure critical systems are operating correctly.
  • Participate in incident response and coordinate troubleshooting during customer-impacting events.
  • Communicate platform issues, updates, and resolutions to relevant internal teams.

Advanced Technical Support & Escalations

  • Provide advanced technical support for VIP customers and customers with complex hosting-related issues.
  • Act as a senior escalation point for WordPress Support Engineers when issues require deeper technical investigation.
  • Troubleshoot complex issues involving servers, websites, networking, DNS, performance, caching, and hosting infrastructure.
  • Assist customers with advanced technical problems beyond standard WordPress troubleshooting.
  • Investigate and resolve issues involving server resources, application performance, connectivity, and platform behavior.
  • Work directly with customers when required to provide expert-level technical assistance.
  • Ensure escalated customer issues are handled with urgency, ownership, and clear communication.

Infrastructure Operations

  • Troubleshoot and maintain Linux-based production environments.
  • Investigate issues related to NGINX, Apache, PHP-FPM, MySQL/MariaDB, Redis, and other platform services.
  • Assist with server maintenance, configuration changes, and operational improvements.
  • Support security updates, system hardening, and infrastructure best practices.
  • Monitor resource usage and identify capacity or performance concerns.
  • Help improve monitoring, automation, and operational workflows.

Team Collaboration

  • Work closely with WordPress Support Engineers, Shift Leads, Site Reliability Engineers, and Engineering teams.
  • Provide technical guidance and knowledge sharing to Support teams.
  • Help create internal documentation, troubleshooting guides, and knowledge base articles.
  • Identify recurring issues and recommend improvements to reduce future incidents.
  • Participate in incident reviews and root cause analysis.
  • 3+ years of experience in SRE, DevOps, Platform Engineering, or similar roles.
  • Strong experience troubleshooting Linux production environments.
  • Experience supporting customer-facing technical environments.
  • Strong understanding of web hosting technologies including NGINX, Apache, PHP-FPM, MySQL/MariaDB, and Redis.
  • Advanced troubleshooting skills across WordPress, servers, DNS, networking, and performance issues.
  • Experience with Linux command line (SSH).
  • Strong understanding of DNS, HTTP/HTTPS, SSL/TLS, CDN, and caching technologies.
  • Experience with Cloudflare, WAF, and web performance optimization.
  • Experience with monitoring tools and incident response processes.
  • Ability to troubleshoot complex issues independently and communicate technical solutions clearly.
  • Excellent written and verbal communication skills (English).
  • Ability to work under pressure during customer-impacting incidents.

Nice to Have

  • Experience supporting managed WordPress hosting platforms.
  • Experience handling VIP customers or enterprise-level support.
  • Experience with high-traffic websites and performance optimization.
  • Knowledge of Bash scripting, yum/dnf , or automation tools.
  • Familiarity with observability tools such as Nedata or Datadog or similar.
  • Experience with incident management and postmortems.

So what are the advantages of working with a pretty amazing tech company?

  • Ability to work from anywhere in the world. You can travel and work in a new city every month. Work and travel without ever using your vacation time. This is a remote position.
  • Flexible Vacation . Never get denied a vacation request ever again.
  • Paid Education. We care about your career. We will help you gain new skills, and guide you on where you want to go.

Come work with a fun, driven, and amazing team at Rocket.net!

Rocket.net is an equal opportunity employer committed to diversity and inclusion. As a multicultural organization, we encourage individual achievement and recognize the strength of our diverse team.

Rocket.net is committed to providing accommodations for people with disabilities. If you require accommodation, we will work with you to meet your needs. Accommodation may be provided in all parts of the hiring process.

We would like to thank each applicant; however, only qualified candidates will be contacted for an interview.

This listing was posted by a verified recruiter at Rocket.net. Report this listing