Über die Rolle
Why You’ll Love This Role
You will help build Cribl Engineering’s AI platform and productivity rails: the shared systems, runtimes, and workflows that let engineers use autonomous AI across the software delivery lifecycle. Instead of one-off prompts or single-feature AI experiments, you will focus on intent engineering, orchestration and harnesses, and production agent infrastructure that other teams can trust and extend.
You will work with a small, high-impact team alongside tech leads, principals, and partner engineering groups. You will ship systems that plan, implement, review, test, and operate work with strong guardrails, then drive adoption until those tools become part of how Cribl builds software. You will also bring a builder mentality: use the product yourself to solve real team and Engineering problems, then turn that lived experience into better systems. Cribl strives to be a great place to work for everyone.
As An Active Member Of Our Team, You Will…
- Design, build, and operate production agentic workflows and the platform harnesses they run on (orchestration, tool integrations, shared context, extension points for other teams)
- Practice intent engineering: turn goals into clear specifications, rules, constraints, and acceptance criteria that AI systems can execute reliably
- Stay current on AI tooling and practices, and evangelize what works across Engineering so partners adopt proven patterns
- Engage hard in design discussions and healthy debate while the direction is open, then pivot with equal energy into implementation once the team decides
- Build observability and guardrails: tracing, regression detection, human-in-the-loop controls, safe rollout, and operability for agentic systems others depend on
- Own critical pieces of agent runtime: job isolation, scheduling, execution environments, secrets and access, and production operations in the cloud
- Design and operate event-driven architectures where it fits (queues, streams, webhooks, async job fan-out) so agentic systems stay scalable and loosely coupled
- Ship safely and often: keep agentic systems on automated CI/CD paths with progressive delivery and clear rollback
- Bring a builder mentality: dogfood what we ship. Use the platform to unblock yourself and the team, surface sharp edges, and turn real usage into product and platform improvements
- Compress software delivery loops by improving how AI helps engineers write, test, review, debug, and validate changes in real repositories and pipelines
- Partner across Engineering to understand workflows, ship tools that fit how people work, and drive adoption through playbooks, examples, demos, and enablement
- Evaluate build vs. buy; stay current on models, agent frameworks, MCP-style tool protocols, and the broader AI tooling ecosystem
- Define and track success metrics for adoption, quality, reliability, and satisfaction; use feedback and data to decide where to invest
- Partner on security and data access so tools are useful while respecting permissions and company policy
- Take designed projects from zero to production: own the path from agreed design through implementation, rollout, and day-two operability
- This position may include stand-by, on-call, or off-hours duties for systems you own
If You’ve Got It - We Want It
- Staff-level (or equivalent) professional software engineering experience building and operating production distributed systems
- Strong TypeScript (and modern JavaScript) plus Node.js experience shipping production services; polyglot comfort is welcome, TypeScript is the primary stack for this team
- Strong software engineering fundamentals and the ability to ship quickly: design, testing, debugging, APIs/services, and code quality
- You have lived (or thrived in) a continuous deployment culture: automated pipelines, progressive delivery or safe rollout, monitoring and rollback, and a bias toward frequent, reversible production changes rather than infrequent manual releases
- Hands-on fluency with modern LLM and agentic coding workflows in production or serious internal platforms (not curiosity-only)
- Experience building backend services, integrations, automation, and internal tools; comfort across product surfaces when needed
- Professional experience with agent orchestration, tool-calling systems, evaluation or guardrail techniques, and connecting agents to reliable backend systems
- Familiarity with agent frameworks, orchestration layers, and integrating external tools and data sources into LLM-based systems (MCP or equivalent experience is a plus)
- Experience with event-driven systems:
queues, streams, pub/sub, webhooks, or similar async patterns in production
- Observability fluency: metrics, logs, traces, and using them to operate and improve production systems
- Clear communication, documentation, and teaching ability; comfort driving adoption, not only writing code
- Good judgment around security, permissions, data access, and safe tool rollout
- Ability to problem-solve from first principles, make sound trade-offs, and drive work independently through ambiguity
- Nice to have
- Hands-on Kubernetes in production
- Terraform or similar infrastructure-as-code for cloud provisioning
- Temporal or other workflow-platform experience
- Deeper AWS / cloud-native ops fluency (IAM, networking, running production workloads end to end)
Quelle: die eigene Karriereseite des Arbeitgebers.