AI Harness Engineer - Agentic & Autonomous Operations (m/f/x)
Quick Summary
Build the control layer around models: task decomposition, state management, model and tool selection, execution loops, checkpoints, retries, stopping conditions, and human handoffs.
not only notebooks, chat interfaces, prompt experiments, or a basic retrieval pipeline. You understand the engineering layer around a model: orchestration, state, context, tools, structured outputs,
Responsibilities
~1 min readWe’re looking for an AI Engineer to build the agentic capabilities at the heart of xHeron. You’ll help shape how our agents reason, act and learn, and build a company where AI doesn’t just support the work, but increasingly delivers it.
We are looking for an AI Harness Engineer to build the software layer that turns language models into dependable operational systems.
You will work on orchestration, context assembly, tool execution, state, permissions, evaluations, observability, failure recovery, and human handoffs. The goal is not merely to generate a good response.
- Agent harness architecture: Build the control layer around models: task decomposition, state management, model and tool selection, execution loops, checkpoints, retries, stopping conditions, and human handoffs.
- Decision and execution behavior: Design how agents interpret situations, choose actions, manage uncertainty, and complete bounded operational tasks—not only how they generate text.
- Context and memory: Determine what an agent needs to know at each step. Build context assembly, retrieval, memory, evidence selection, and knowledge-maintenance strategies while managing relevance, latency, and cost.
- Tools and real-world effects: Connect agents to APIs and operational systems through explicit tool contracts, permissions, validation, idempotency, and confirmation that the intended effect actually occurred.
- Evaluation and regression testing: Build representative evaluation suites and automated tests for decision quality, tool use, task completion, safety, escalation behavior, and recurring failure patterns.
- Observability and failure analysis: Make agent behavior inspectable through structured traces, state, outcomes, and error classification. Diagnose failures across models, prompts, context, tools, and surrounding software.
- Continuous improvement: Use evaluations, production traces, operator feedback, and controlled experiments to improve prompts, context, tools, model choices, and agent architecture.
- Production ownership: Ship and operate production Python software alongside our CTO and engineering team. Work closely with systems engineering on integrations, authorization, durable state, recovery, and safe execution.
Requirements
~1 min read- You are a strong software engineer who can design, build, test, and operate reliable production systems in Python.
- You have built a substantial LLM-powered, automation, workflow, or decision system: not only notebooks, chat interfaces, prompt experiments, or a basic retrieval pipeline.
- You understand the engineering layer around a model: orchestration, state, context, tools, structured outputs, permissions, retries, fallbacks, observability, and evaluations.
- You can connect intelligence to real work: selecting relevant evidence, making a bounded decision, invoking tools or APIs, and verifying that the intended result occurred.
- You design for partial failure and uncertainty. You think deliberately about when an agent should act, retry, wait, stop, ask for help, or escalate.
- You evaluate behavior systematically. You use representative cases, traces, outcome checks, comparisons, and regression tests rather than treating a plausible response as proof of success.
- You have shipped software that people or operational processes depend on and have worked through debugging, timeouts, inconsistent data, changing APIs, and production incidents.
- You can investigate ambiguous problems independently, define a useful system boundary, explain trade-offs clearly, and turn decisions into working software.
- You care about simplicity and control. You know when a deterministic workflow is better than an agent and when additional autonomy is justified by evidence.
- You bring useful expertise and an independent perspective. Strong adjacent experience in workflow engines, developer tooling, distributed systems, automation platforms, or integration-heavy backend software can be highly relevant when paired with credible LLM-system understanding.
Around 2+ years of professional software-engineering experience is a useful guide, not a hard requirement. What matters is the technical depth of what you built, the decisions and production outcomes you owned, and your ability to learn unfamiliar systems.
- Your LLM experience is mainly prompt engineering, chatbots, basic RAG, or connecting a model API to a user interface.
- Your background is primarily offline data analysis, model training, or experimentation, and you do not want to own production software and operational behavior.
- You treat retrieval as the complete agent architecture rather than one possible source of context within a larger execution system.
- You consider a coherent model response successful without verifying the decision, tool call, state change, or real-world outcome.
- You rely on an agent framework to provide the system design and are not comfortable reasoning about the control flow, state, permissions, and failure behavior underneath it.
- You are not interested in owning evaluations, trace analysis, regression testing, and production learning alongside feature development.
- You prefer fully scoped implementation tasks with stable requirements and limited responsibility for product or operational outcomes.
What We Offer
~1 min readReal ownership: you're not maintaining someone else's codebase, you're building a core, novel part of the product from the ground up.
- Intro call (15-30 min): a conversation with Friedrich (CTO) about your background, what you've built, and why this role.
- Technical take-home case study: a short, real engineering problem, close to what you'd actually be building.
- Case study interview (1 hour): you walk us through your thinking, not just the final code.
- Culture fit (30 min): meet both Founders and make sure it's a fit both ways.
Location & Eligibility
Listing Details
- Posted
- September 14, 2026
- First seen
- September 28, 2026
- Last seen
- October 7, 2026
Posting Health
- Days active
- 8
- Repost count
- 0
- Trust Level
- 20%
- Scored at
- October 7, 2026
Signal breakdown
4 other jobs at
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.