
Sovereign Engineering Platform SRE - T Cloud Public (REF5740Q)
Deutsche Telekom IT Solutions · Remote, Debrecen, Hungary
About The Role
Mission 
Design, build, and operate the secure infrastructure foundation used by Meridian engineering teams for AI-assisted software development, model experimentation, repository analysis, CI/CD execution, and controlled handover work in isolated or sovereignty-sensitive environments 
Role focus 
This infrastructure and operations role centers on Kubernetes-based engineering platforms, GitOps, private registries, internal model endpoints, observability, access control, and reliable operations for AI-enabled SDLC workloads. The candidate should enable engineering velocity while preserving security, auditability, and operational discipline 
Key responsibilities 
- Build and operate Kubernetes environments that host AI engineering tools, internal model gateways, retrieval components, workflow services, CI/CD runners, and documentation services 
- Implement GitOps and Infrastructure as Code patterns for reproducible provisioning, configuration, policy enforcement, platform upgrades, and disaster recovery readiness 
- Manage private registries, package mirrors, secrets, identity integration, network segmentation, storage classes, backup routines, and controlled connectivity models 
- Provide observability for engineering workloads, including metrics, logs, traces, GPU and CPU utilization, service health, cost signals, and operational runbooks 
- Work with software, security, and architecture teams to ensure the platform supports AI-assisted SDLC workflows without creating uncontrolled data exposure or audit gaps 
Examples of market tools, models, and platform components expected 
- Platform tooling such as Kubernetes, Helm, Terraform, Ansible, ArgoCD, Crossplane, GitLab runners, Jenkins agents, private registries, and internal package mirrors. 
- AI platform components such as vLLM, Ollama, OpenAI-compatible gateways, Qdrant or similar vector stores, Open WebUI, Continue-compatible endpoints, and workflow services. 
- Observability and operations stacks such as Prometheus, Grafana, Loki, OpenTelemetry, ELK/OpenSearch, Alertmanager, SRE runbooks, and incident management tooling. 
- Security and governance components such as Vault, Keycloak, network policies, RBAC, admission controls, image scanning, SBOM tooling, and audit logging. 
- Infrastructure awareness covering GPU-backed nodes, CPU-only fallback, storage performance, network isolation, proxy patterns, on-premise environments, and dedicated landing zones. 
Candidate profile 
- 5+ years in SRE, platform engineering, DevOps, cloud infrastructure, or operations roles with strong Kubernetes and Linux expertise. 
- Proven experience building and operating production-grade engineering platforms with GitOps, Infrastructure as Code, observability, and operational runbooks. 
- Hands-on skills in Terraform, Ansible, Helm, Python or shell scripting, CI/CD runners, private registries, and secure configuration management. 
- Good understanding of networking, storage, secrets, access control, monitoring, backup, disaster recovery, and operational hardening in high-security environments. 
- Comfortable supporting AI-enabled engineering workloads in sovereignty-driven contexts where isolation, controlled data handling, reliability, and auditability are mandatory. 
Please note: remote working is only possible from within Hungary due to European taxation regulations.
- Please be informed that our remote working possibility is only available within Hungary due to European taxation regulation.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring