Technical Lead - Platform
Rakuten Symphony India Private Limited · RSIN_Bangalore, India
About The Role
Rakuten Symphony is reimagining telecom, changing supply chain norms and disrupting outmoded thinking that threatens the industry’s pursuit of rapid innovation and growth. Based on proven modern infrastructure practices, its open interface platforms make it possible to launch and operate advanced mobile services in a fraction of the time and cost of conventional approaches, with no compromise to network quality or security.
Rakuten Symphony has operations in Japan, the United States, Singapore, India, South Korea, Europe, and the Middle East Africa region. For more information, visit: Rakuten Symphony | Intelligent Growth .
Building on the technology Rakuten used to launch Japan’s newest mobile network, we are taking our mobile offering global.
To support our ambitions to provide an innovative cloud-native telco platform for our customers, Rakuten Symphony is looking to recruit and develop top talent from around the globe. We are looking for individuals to join our team across all functional areas of our business – from sales to engineering, support functions to product development.
Let’s build the future of mobile telecommunications together!
About Rakuten Group, Inc. (TSE: 4755) is a global leader in internet services that empower individuals, communities, businesses and society. Founded in Tokyo in 1997 as an online marketplace, Rakuten has expanded to offer services in e-commerce, fintech, digital content and communications to 2 billion members around the world. The Rakuten Group has over 30,000 employees, and operations in 30 countries and regions. For more information visit https://global.rakuten.com/corp/ .
Role Overview
As an RCP Engineer, you will provide technical leadership in the design, implementation, and operation of an enterprise-scale Kubernetes platform across highly regulated on-premises environments. This role requires deep L3-level expertise in both Linux systems engineering and Kubernetes architecture. You will define the strategic direction for virtualization and OS lifecycle management, platform reliability, and advanced Linux management, working closely with security and development teams to deliver a secure, observable, and highly performant platform at extreme scale. You will also own the platform release lifecycle, ensuring that updates are stable, backward compatible, and rigorously validated.
Key Responsibilities
- Platform Development & Kernel Engineering (60-70%)
- Kernel-Level Fine-Tuning: Conduct deep-dive performance tuning using , , and to optimize RHEL for real-time (RT) workloads.
- Red Hat Expertise: Deep proficiency in RHEL 8/9, including RHEL for Real Time; manage the end-to-end OS lifecycle, including patching and upgrading the OS stack.
- RAN Workload Optimization: Configure and tune vCU and vDU components, ensuring efficient resource allocation for high-speed 5G network demands.
- Latency & Determinism: Implement CPU isolation (), IRQ affinity balancing, and NUMA alignment to achieve deterministic performance for vDU signal processing.
- Networking Acceleration: Optimize packet processing using DPDK (Data Plane Development Kit) and SR-IOV to bypass kernel bottlenecks for high-throughput RAN traffic.
- Advanced Troubleshooting: Lead Root Cause Analysis (RCA) for high-severity incidents, utilizing tracing tools like , , and to diagnose kernel-space latencies.
- Automation & Infrastructure: Develop and maintain complex Ansible playbooks to automate kernel hardening and performance profile deployments across large-scale RAN clusters.
- Kernel Subsystems: Leverage in-depth knowledge of process scheduling (, ), memory management, and I/O schedulers.
- Virtualization & Containers: Manage KVM/QEMU, Red Hat OpenShift Virtualization, and containerized vDU pods in Kubernetes environments.
- Hardware-Software Interaction: Utilize Intel RDT (Resource Director Technology), CAT (Cache Allocation Technology), and BIOS-level tuning (e.g., C-states, P-states).
- Monitoring & Hardware: Expert use of , , , , and hardware-specific utilities like ; manage hardware via Supermicro, Dell iDRAC, and HP interfaces.
- Platform Release, QA & Reliability (30-40%)
- Release Stability & Governance: Own the platform release lifecycle, ensuring new features are backward compatible. Define the "Definition of Done" for all platform releases.
- Automated Validation Pipelines: Design CI/CD gates that incorporate automated testing for infrastructure-as-code and K8s manifests to ensure zero-touch, stable deployments.
- Performance & Chaos Testing: Establish performance baselines using tools like or to prevent latency regressions; implement resilience testing (Chaos Mesh/LitmusChaos) to validate platform stability under failure scenarios.
- Compliance & Vulnerability Management: Integrate security scanning (OpenSCAP) and policy enforcement (OPA/Kyverno) into the release pipeline to maintain hardened security standards.
- SRE Practices: Apply SRE principles at scale: perform major incident leadership, develop capacity/scalability strategies, and execute advanced change management.
Mandatory Requirements
- Experience: 8 to 12 years of relevant telco industry experience in infrastructure engineering or SRE, with a proven track record of managing large-scale, complex environments.
- Kubernetes & Virtualization: Strong knowledge in bare metal Kubernetes deployments and Virtualization in development/production environments.
- Networking: Expertise in host-level networking, DNS/DHCP architecture, TLS management, load balancing strategies, and advanced container networking.
- SRE & Incident Management: Proven application of SRE practices at scale, including major incident leadership and capacity strategy development.
- Security Architecture: In-depth knowledge of security hardening principles, advanced RBAC models, pod security mechanisms, and secret management infrastructure.
- Technical Stack:
- Core: Kubernetes, Linux (RHEL/Rocky), Docker, Helm, Git.
- Automation: Ansible (mandatory), Python and Shell (good to have).
- Certifications: RHEL Certification, CKA (Certified Kubernetes Administrator).
Why This Role Matters
In the era of 5G, the performance of the radio access network (RAN) is directly tied to the determinism and stability of the underlying infrastructure. As a Platform Development Engineer (RCP), you are the architect of the "digital foundation" upon which mission-critical telco workloads run. Your work ensures that the platform is not only capable of handling extreme-scale traffic with sub-millisecond latency but is also resilient, secure, and maintainable. By bridging the gap between low-level kernel optimization and high-level platform release engineering, you ensure that our network remains a competitive advantage—delivering the reliability of legacy telecommunications with the agility of modern cloud-native development.
RAKUTEN SHUGI PRINCIPLES
Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
- Always improve, always advance. Only be satisfied with complete success - Kaizen.
- Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
- Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
- Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
- Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
