DevOps / Infrastructure Engineer - High-performance trading systems
Koton · Brno, South Moravian Region, Czechia
About The Role
We trade the world's most volatile markets and make them more efficient.
Koton is a proprietary trading firm and a top-tier market maker on the world's largest crypto exchange. Our core is high-frequency market making, where the speed and cost-efficiency of execution decide the outcome.
We build software that trades real capital in live crypto markets. Every microsecond, every allocation and every design decision has a measurable impact on PnL. Engineering isn't supporting the business here, it is the business. When the system performs well, we make more money. When it doesn't, we lose our own. That creates a very different engineering environment from most software companies.
Most of the codebase is .NET today. Parts of the performance-critical core are moving to C++, while the business continues to evolve with new trading features, infrastructure and exchange integrations. You'll work across both worlds — improving today's production systems while helping shape tomorrow's architecture. You'll own meaningful parts of the system from day one.
TL;DR
→ Own the platform our trading systems run on: Linux, AWS, our own on-prem infrastructure, networking, databases, deploys and observability.
→ You'd be the first person fully dedicated to infrastructure — but you're not starting from zero or alone. The engineer who built most of it is still here and will hand it over deliberately.
→ Markets run 24/7. There's no maintenance window, and production means live trading with real capital.
→ End-to-end ownership. You get problems and context, not a queue of infrastructure tickets.
→ You don't need to know trading. Most of us didn't. What you do need is curiosity. You can't write good code for a system whose commercial purpose you don't understand.
Note: This is a senior, hybrid, high-ownership individual-contributor role
What you'd actually be running
Cloud: AWS across two regions — Tokyo close to the exchange for latency, Dublin for internal tooling
IaC: Terragrunt + OpenTofu, in-house module library, remote state with locking
Config: Ansible — roles covering ~15 host types, from trading engines to database nodes
CI/CD: Self-hosted GitLab, tag-driven build → registry → deploy pipelines, autoscaled spot runners, custom builder images
Data: TimescaleDB with Patroni HA behind health-checked HAProxy, plus SQL Server
Observability: Prometheus, Loki and Grafana, plus some home-grown alerting that needs consolidation
Network: WireGuard, pfSense driven by API, split-horizon DNS between cloud and on-prem
On-prem: Our own infrastructure running Proxmox, Git and core databases
Workloads: Low-latency .NET services consuming real-time exchange data
No Kubernetes. We run containers directly on hosts with host networking. That's deliberate: extra layers aren't free when you're measuring latency in microseconds.
If your main ambition is to build a Kubernetes platform, this probably isn't the job for you.
What you'll do
Own the platform end to end. Linux, AWS, on-prem, networking, deployment, secrets and access. Infrastructure should be reproducible and understandable — not a collection of machines somebody configured once and hopes nobody reboots.
Make deployments safe in a live market. There's no maintenance window. Production systems are trading real capital while releases happen. One of the early opportunities is improving how releases are versioned, deployed, verified and rolled back so we always know exactly what's running and can confidently return to the previous state.
Build observability that settles arguments. When latency spikes, we want to know quickly whether it was us, the network or the exchange. Structured logging, meaningful metrics, latency histograms and disciplined clock synchronisation should make that answerable in minutes rather than hours.
Run the data layer. PostgreSQL and TimescaleDB handle continuous market-data workloads. You'll own replication, Patroni failover, retention and compression policies, performance and backups that have actually been tested by restoring them.
Own networking that matters commercially. Connectivity and routes to exchange endpoints, IP whitelisting, remote access and failover paths. Latency here isn't an abstract infrastructure metric — it can be the difference between a good fill and a missed opportunity.
Improve the internal platform. Access and identity across our systems is automated from Google Workspace by a Python service provisioning Git, monitoring, VPN and internal tooling. You'll extend it and help us simplify and de-risk identity management.
Handle incidents — and make the next one smaller. Occasionally something exceptional happens outside normal office hours. You'll help resolve it, understand why it happened and improve the system so the same incident doesn't happen twice.
Work directly with developers, researchers and traders. You get a problem and context, not a written-up ticket.
What ownership actually means here
We say ownership because it's the main thing we're offering.
- You get problems, not specifications. We expect engineering judgement rather than following predefined designs.
- You own the platform, not tasks. Larger architectural decisions are made together with our CTO; how the platform is built, operated and improved day to day is yours.
- There are no handoffs. Developers continue to own their services; you own the platform they run on. You're partners, not separate teams throwing tickets over a wall.
- Your results are visible immediately. Reliability, deployment speed and latency settle arguments much faster than presentations ever could.
Our CTO is hands-on and knows the current infrastructure because he helped build it. You'll have someone to challenge ideas with and learn the existing system from, without having someone managing you at arm's length.
Ownership also means freedom comes with responsibility. Nobody micromanages you, but nobody automatically picks things up if you drop them either. This role suits people who'd rather be trusted than managed.
Deep Linux knowledge. You can work out what a process or machine is actually doing when something goes wrong and know which tools will give you the answer.
Infrastructure as code and containers. Terraform/OpenTofu, Ansible, Docker or close equivalents. You value reproducibility over cleverness.
Real networking knowledge. TCP behaviour, routing, MTU, VPNs, firewalls and DNS. Enough to debug a problem that exists somewhere between two machines — including reaching for a packet capture when necessary.
Observability as a discipline. You know what to measure, how to make it useful and what noisy telemetry to delete.
PostgreSQL operations. Replication, failover, backup and restore, and performance under load.
CI/CD you've owned, not just consumed. You've built pipelines and been responsible for what happens when a release goes wrong.
Useful scripting. Python or Bash written well enough that someone else can safely depend on it, including integrations with APIs.
End-to-end ownership. You take a problem, decide how to solve it, build it, ship it and keep it running.
Diligence and discretion. You'll be working close to trading systems and real capital.
Willingness to be occasionally on call when an exceptional incident needs a fast response.
Bonus points
- TimescaleDB or experience operating time-series data at volume.
- Patroni or other PostgreSQL HA setups.
- Kernel and network latency tuning — NIC queues, interrupt affinity, sysctl, NTP/PTP/chrony.
- Operating self-hosted GitLab rather than simply using it.
- Bare metal and hardware — procurement, racking and capacity planning.
- Exchange connectivity or another environment where network latency carries a commercial cost.
- Security around sensitive systems — secrets management, least-privilege IAM, access control and auditing.
- Reading C#/.NET well enough to understand what the services you're operating are actually doing.
What we won't promise you
You won't inherit a perfectly tidy platform.
Some parts are well designed and mature. Others were built quickly by developers while their main focus was keeping trading systems moving. Documentation isn't complete, and your first months will involve reading, asking questions and deciding what should be improved — and what should simply be left alone.
You will get a real handover. The engineer who built most of the current infrastructure is still here, and transferring ownership deliberately is one of the main reasons this role exists. You're not being dropped into an undocumented system and wished good luck.
Occasionally you'll investigate production incidents outside office hours because crypto markets never sleep. It's not frequent, but it's part of operating systems that trade real capital.
If you're looking for a mature infrastructure organisation with established teams and narrowly defined responsibilities, this isn't it. If you'd rather take a working platform, understand it deeply and make it genuinely yours, it probably is.
- Paid time off to actually switch off
- Fully catered office: all-day snacks, daily lunch, coffee and drinks on us
- Work-from-home flexibility
- Training and development to keep sharpening your edge
- Performance bonuses tied to results
- Relocation support for candidates moving to Czechia
- Pick your own hardware, top-spec kit, no IT-catalogue limits
Terms:
- 100 000–180 000 CZK per month
- Contract
- Brno, hybrid (3 days in the office)
We believe in paying fairly, and we're always open to a conversation. We like to work with smart people and want them to feel appreciated and to enjoy the work they do here.
How to apply
No cover letter, please. We don't necessary need a CV either.
Instead, write us a few paragraphs about a system you ran or an incident you handled — what broke, how you found it and what you changed so it wouldn't happen again.
That tells us far more.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing