Software Engineer - Data & Scalability P,latform
123 Cisco Systems (India) Private Limited · Bangalore, India
About The Role
Meet the Team
The Splunk Agent Resilience team is defining the future of AI resilience. Our team provides scalable, cost-effective evaluation and guardrails that ensure AI agents behave as intended, improving reliability and reducing risks. This unified approach empowers our customers to confidently deploy and manage AI-powered applications with enhanced observability and control.
As a Senior Software Engineer on the Data & Scalability Platform team, you will own the data plane end to end: the stores that hold the data — relational, analytical/columnar, object storage, caches, and queues — the streaming and compute components that move and transform it, and the performance tooling that proves how all of it behaves under load. You will make that data plane correct, fast, cost-bounded, and tenant-isolated at production scale, own the capacity and performance model that says when and how it needs to grow, and partner closely with the product engineering teams that build on top of it.
Your Impact
- Own and evolve the platform's data stores — relational, analytical/columnar, object storage, and caching — including schema design, migration safety, and retention.
- Design and operate the streaming and queueing backbone that carries data from ingest to queryable.
- Build and scale the compute and pipelines that move and transform data, including stream processors, writers, schedulers, and distributed worker fleets.
- Own end-to-end data performance: query and write latency, indexing and sharding strategy, and hot-path optimization.
- Own the capacity model for the data platform, along with quotas, rate limiting, and multi-tenant isolation.
- Drive cost efficiency across the data platform, including cost per unit of telemetry and cost attribution.
- Build and own performance tooling — load testing, profiling, and benchmarking — used to validate scale and guide optimization.
- Build observability instrumentation and telemetry pipelines so platform signals are correct, complete, and affordable.
- Own data operations: backup and restore, retention and deletion, replication, migration and backfill, and data-quality signals.
- Design, develop, test, and maintain production services and internal tooling in Python and/or Go, using secure coding practices and automated tests.
- Debug and resolve complex production issues across data stores, queues, and services, and contribute to monitoring, on-call support, and root-cause analysis.
- Contribute to platform architecture direction and act as a technical resource and mentor through design reviews, code reviews, and documentation.
Minimum Qualifications
- Bachelor's degree with 7+ years of related experience, or Master's degree with 4+ years, or PhD with 1+ year of related experience, in Computer Science, Software Engineering, or a related field.
- Strong backend software engineering experience building, operating, and delivering highly scalable, reliable, production-grade data-intensive services and platforms.
- Proven experience with large-scale distributed systems, including designing and solving complex problems of scalability, availability, performance, reliability, and fault tolerance.
- Hands-on experience operating and tuning both an OLTP database (e.g., PostgreSQL, MySQL) and an analytical, columnar, or time-series store (e.g., ClickHouse, Druid, BigQuery, Snowflake) — schema design, indexing, query tuning, and migration safety.
- Production experience with streaming or queueing systems (e.g., Kafka, RabbitMQ, Pulsar, Kinesis) — partitioning, consumer-group semantics, delivery guarantees, and backlog and dead-letter handling.
- Experience scaling data-processing compute — queue consumers, async task workers, or streaming jobs — including concurrency tuning, batching, backpressure, and throughput behavior under sustained load.
- Strong proficiency in Python and solid experience with at least one additional backend programming language such as Go, Java, C++, or similar.
- Demonstrated performance engineering ability — profiling, benchmarking, and load testing real systems, and turning the measurements into capacity, design, and cost decisions.
- Experience with cloud-native technologies, containerized environments (Kubernetes), and public cloud platforms (AWS, GCP, or similar).
- Demonstrated ability to independently design, develop, debug, test, and maintain software with minimal guidance.
Preferred Qualifications
- Experience running ClickHouse or a comparable columnar store at scale — sharding and replication topology, materialized views, merge and mutation behavior, retention and TTL mechanics.
- Experience owning a capacity model or a FinOps practice: cost per unit of telemetry, cardinality governance, and quota or rate-limit design.
- Experience with multi-tenant isolation and noisy-neighbor mitigation in shared data systems.
- Experience with backup and restore, tested disaster-recovery drills against stated RPO/RTO, and data retention or deletion compliance (e.g., GDPR/DSR hard delete).
- Experience building observability instrumentation and pipelines (OpenTelemetry, collectors, metrics/tracing backends) rather than only consuming dashboards.
- Experience building or operating performance tooling as a product for other engineers — load-test harnesses, continuous profiling (e.g., Pyroscope, py-spy, flame-graph workflows), or benchmarking suites used to justify capacity and design decisions.
- Experience with distributed task frameworks and worker orchestration (e.g., Celery, Temporal, Flink, custom consumer fleets) at production scale.
- Familiarity with streaming databases, stream processing frameworks, or distributed query engines (e.g., RisingWave, Flink, Spark Streaming, Trino/Presto) and with model-serving infrastructure.
- Experience supporting stateful systems in both cloud/SaaS and air-gapped or on-prem deployments.
- Experience leading medium-sized features or projects end-to-end, mentoring engineers, and influencing technical decisions across teams.
- Strong communication skills, with the ability to turn complex performance, capacity, and cost findings into clear recommendations for engineering and product partners.
Why Cisco?
At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.
We are Cisco, and our power starts with you.
Disclaimer
To ensure that we hire the best talent in the right way, we follow a strict hiring process and recently, Cisco has been made aware of fraudulent recruiters claiming to be from the company. Please be advised that any communication from Cisco about careers will:
- be in direct response to an application you have submitted through the company career site
- begin with screening or an interview
- originate from a Cisco email address, and
- be conducted across email, phone, or WebEx
Cisco will never make a job offer without conducting an interview process or ask you for money in any way. If you have been requested to apply for a role or have received an offer from a site other than https://careers.cisco.com or cisco.wd5.myworkday.com , do not provide any personal identifying information, including your Aadhaar or other personal identifying number, birth certificate, banking information, driver's license, or passport.
If you are the target of a recruiting scam, consider filing a report with your local law enforcement authorities. Cisco bears no responsibility, and cannot be held liable, for any claims, damages, expenses, or other inconvenience resulting from or in any way connected to recruiting scams.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
