Skip to content
← Back to job listings

Lead Engineer, AI Agent Systems

patsnap · Shanghai

LeadExternal listingfull-time15 days ago

About The Role

### Responsibilities
**Architecture Leadership and Evolution**

  • Lead the architecture and evolution of next-generation agent infrastructure designed for complex, knowledge-intensive work.
  • Define clear boundaries and collaboration mechanisms across three core layers: the execution engine, context and reasoning orchestration, and the agent capability foundation. Ensure high availability, reliability, and long-term extensibility in environments with a low tolerance for hallucinations and incorrect outputs.

**Agent Execution Engine**

  • Design and implement the Agent Loop runtime and its middleware pipelines.
  • Lead the execution and orchestration of planning and sub-agent workflows, including task decomposition, dependency management, concurrency control, and execution scheduling.
  • Build mechanisms for checkpointing, interruption and resumption, failure recovery, self-healing, authorization, and cost control to ensure the reliable execution of long-running and complex multi-step tasks.

**Context and Reasoning Orchestration**

  • Own the design and implementation of core context orchestration capabilities.
  • Develop strategies for input standardization, dynamic capability representation, and hierarchical context-budget management, including structured degradation when resource or context limits are reached.
  • Build structured task workspaces that support efficient organization of dynamic context. Address challenges including long-history compression, tool-output normalization, evidence traceability, and the management of information across different stages of a task.

**Agent Capability Foundation**

Lead the development of foundational agent capabilities, including

  • Secure sandboxed environments using technologies such as Docker, Kubernetes, and AST-based controls
  • Multi-layer memory stores
  • Retrieval and knowledge-access capabilities
  • An MCP (Model Context Protocol) Hub
  • Skill execution and management engines
  • File-processing and transfer pipelines
  • Multi-tenant isolation and security controls
  • End-to-end observability and diagnostics

**Technical Leadership and Team Enablement**

  • Remain hands-on and personally contribute code to critical platform modules.
  • Lead technical decomposition, architecture decisions, code reviews, and the development of automated evaluation systems and feedback loops.
  • Guide the engineering team in translating specific business use cases into reusable platform and infrastructure capabilities.

### Qualifications
**Engineering and Leadership Experience**

  • At least five years of professional software engineering experience.
  • Proven experience leading the design and delivery of complex software systems beyond standard CRUD applications or basic integrations with AI APIs.
  • Demonstrated experience operating as a Tech Lead, Staff Engineer, or equivalent technical leader.
  • Experience leading an engineering team of at least three people.

**Core Engineering Capabilities**

  • Strong Python software-engineering skills and the ability to independently own critical platform modules.
  • Deep experience with common engineering challenges such as streaming responses, asynchronous and concurrent execution, and multi-model routing and provider integration.
  • Strong judgement in balancing system reliability, security, cost, latency, and delivery speed.
  • Solid understanding of distributed systems, production architecture, debugging, and operational reliability.

**Depth in AI and Agent Systems**
Candidates must have substantial hands-on engineering experience with Agent and LLM systems, with deep expertise in at least two of the following three areas:

Execution Engine

  • Multi-step reasoning loops
  • Tool lifecycle management
  • Planning and sub-agent orchestration
  • Interruption and resumption
  • Failure recovery and self-healing

Context and Reasoning Orchestration

  • Input standardization
  • Context assembly
  • Context and token-budget governance
  • Provider-specific request shaping
  • Task-stage modelling
  • Long-context compression and evidence traceability

Agent Capability Foundation

  • Sandbox isolation
  • Memory and retrieval systems
  • MCP infrastructure
  • File-system and file-processing capabilities
  • Multi-tenant isolation
  • Security, monitoring, and observability

This is an external listing. JobSpring does not represent or verify the employer. Report this listing