Careers
Part-Time
Stockholm, Sweden
About the role
Construction doesn't buy "AI". It buys answers it can act on, and an answer can look completely plausible while being wrong in the way that costs money on a real project. The Construction Evaluation Engineer role exists to catch exactly that: to define what great looks like for our agents, find where they fail, and turn that judgment into something our systems can measure.
It starts hands-on and close to the data, and gets more technical as you grow, from evaluating outputs to owning eval pipelines, graders and datasets. You'll work alongside our forward deployed engineers, design partners and construction experts. Evaluation here happens close to the customer.
What you might do (depends on you):
Run our agents against real construction workflows and judge the outputs on correctness, completeness and practical usefulness
Turn real workflows, edge cases and production failures into repeatable eval cases: reference outputs, scoring rubrics, failure taxonomies and expert-reviewed datasets
Grow into the systems around the work - eval harnesses, automated graders, regression tests, model comparisons
Work out what genuinely needs human judgment, then engineer the rest away
Capture what we learn from customers, experts and agent failures so every deployment improves the underlying system
How to apply
Send a short note with:
What you’d build or improve at Brickanta in your first 30 days links to
Anything you’ve shipped or otherwise built (or a quick write-up if it’s private)
Are you keen on doing your life’s work - bringing one of the world’s largest and most impactful industries, the $15 trillion construction industry, into the AI era?