Back to careers

Data Engineering & Automation Intern
Job Description
Location | Hybrid, with a minimum of 2 days per week at our office in Clerkenwell, London, UK |
Hours | Full time or Part time (minimum 20 hours per week) |
Term | Up to 3 months, depending on availability |
Language | English |
Salary range | UK minimum wage |
About Kestrix
Kestrix uses drone-based thermal imaging and AI to map and quantify how heat escapes from buildings, assess heat pump and solar readiness, and generate energy retrofit plans for homes at scale. Our data and patent-pending algorithmic analysis helps landlords, local authorities, and energy suppliers understand the optimal upgrade pathway for each home – enabling the effective allocation of scarce time, labour, and financial resources into housing decarbonisation and regeneration.
To date, Kestrix have mapped more than 10,000 homes, serving 25+ enterprise customers across the UK, Germany, and the US. We have raised £3.5M in funding from leading PropTech and Climate VCs (Pi Labs, the Conduit EIS Fund, and Carbon13 among them) as well as government bodies such as InnovateUK and the Department for Energy Security and Net Zero (DESNZ). We are an agile team of 12, meaning you will have real ownership and impact from day 1.
The mission
Our machine learning is only as trustworthy as the data we validate it against. That validation data is called ground truth: verified, real-world information about how buildings are actually constructed and insulated, covering their walls, windows, and roofs.
A rich source of ground truth already exists in the documents produced during building surveys. Retrofit assessment reports give us construction and insulation detail across walls, windows, and roofs, while borescope reports tell us the cavity fill insulation status of a wall. These are dense, unstructured PDFs full of exactly the information we need. The problem is that this information is locked inside documents in formats that vary from one assessor to the next, and none of it is machine-readable.
Your mission is to unlock it. You will build the pipeline that turns these messy PDFs into clean, structured, queryable data inside our platform, and connect it to our existing image annotations so it can be used to validate our algorithms.
The end goal we are working towards: someone uploads any of these assessment or survey PDFs, and the relevant data is automatically extracted into structured, processable form, correctly, every time.
What you'll do
Around 90%+ of this role is data and software work. There may also be some operational support on our ground truth data collection campaign (described below), but only if and when it is needed, so this is not a confirmed or fixed part of the role.
Build the extraction pipeline (core focus)
Write Python to parse unstructured content out of multiple report types, including retrofit assessment reports and borescope reports (our source for cavity fill insulation status), handling the variation in layout and terminology across different formats. You do not need to build everything from scratch, and you are encouraged to use AI tooling integrated into Python (for example LLM APIs and document extraction libraries) where it helps you get accurate results faster.
Design a data schema that captures the building characteristics we care about (wall, window, and roof construction and insulation), in a form that maps cleanly onto our existing data model.
Implement storage of that structured data inside our platform so it is easy to query and extract downstream.
Link the extracted data to our image annotations, so each piece of ground truth is tied to the building and features it describes.
Treat correctness as the top priority. Ground truth that is wrong is worse than no ground truth at all, so you will build validation, checks, and review steps to get extraction toward 100% accuracy, and be honest about where confidence is lower.
Support the data collection campaign (if needed)
Help us actually acquire ground truth data, which may include preparing and organising incoming documents, chasing and coordinating data sources, and keeping our collection tracking up to date.
Stretch goal: the ground truth collection portal
If the core pipeline is in good shape, help design and build a simple website we can send to external partners, where they can upload their own ground truth data directly. The vision is a self-serve ground truth collection portal that feeds straight into the pipeline you have built.
Our tech stack
You will be working with, and learning from, the tools we use day to day:
Python as our primary language.
Google Cloud Platform (GCP) for infrastructure and BigQuery for data warehousing.
PostgreSQL as our relational database.
Poetry for dependency management and pytest for testing.
LLM APIs and document-extraction tooling, which you will use directly for the extraction pipeline.
You are not expected to know all of these already. We care that you can pick things up quickly.
What we're looking for
This is an early-career role, so we expect you to be learning fast rather than arriving fully formed. What matters is aptitude, curiosity, and a genuine ability to code.
Essential
A background in computer science or data science, or a genuine personal interest in coding. A degree is not essential. If you have built things in your own time and have personal projects or products to share, we would love to see them.
Comfortable writing Python. You should be able to read, write, and debug real code, not just tweak snippets.
Strong problem-solving instincts. You can take an open-ended, poorly-defined problem, break it down, and work your way to a solution.
A structured, precise way of thinking. You care about getting things right and you notice when something does not add up.
Comfortable working with messy, real-world data and turning it into something clean.
Able to communicate clearly and work independently, while knowing when to ask for help.
Nice to have (not required)
Any exposure to data modelling, databases, or working with structured data (SQL, graph databases, or similar).
Experience parsing documents or extracting text and tables from PDFs.
Familiarity with any of Python's data tooling.
Basic web development for the stretch goal (any modern framework is fine).
An interest in buildings, energy, or the climate problem we are working on.
What you'll get
Ownership of a meaningful, end-to-end problem that directly shapes how well our product works.
Hands-on experience across data engineering, software, and applied machine learning validation in a real production environment.
Close mentorship from experienced engineers in a small team where your work is visible and matters.
The chance to build something that helps decarbonise the built environment.
How we work
These are the values we hire and work by:
Bold. History’s greatest challenges were not solved by doing things the ‘old way.’ No idea or considered action is too bold or too crazy.
Accountable. We hold ourselves accountable to outcomes and trust that our teammates do the same.
Curious. We approach everything – from our customers to our technical challenges – with curiosity and openness. We are never done learning.
Humble. We champion transparency, make time for conversations up and down the organisation, and are never finished improving as people and teammates.