
2015–2019
Honda Research Institute: helping researchers find the moments that matter in hours of driving data
HRI’s researchers train and test driver-assistance and autonomous-driving algorithms on recordings from instrumented cars: camera, radar, laser scanners, GPS and the car’s own data, hours per drive. From 2015 to 2019 we designed and built the web application they use to find the moments they need, play them back and turn them into training data, on a graph data model built with Carbon LDP, our own linked data platform.
- 5+
- sensor data types in one view: camera, LIDAR, radar, GPS, vehicle CAN
- 2016
- in production; handed to Honda’s teams in Japan in 2019
Schematic — client UI under NDA
- Client
- Honda Research Institute
- Industry
- Automotive
- Scope
- Requirements workshops · Graph data model · Data visualization · Web application · Data processing
- Platform
- Carbon LDP · Hadoop · Angular
- Partnership
- 2015–2019
- Services
Challenge
A researcher teaching an algorithm to handle a car cutting in, a pedestrian stepping out or a highway in the rain needs those moments, and only those, from every sensor at once. Honda Research Institute (HRI) records them with test vehicles in Europe and the US, and a single drive produces hours of synchronized video and sensor streams, kept in big-data storage because of their volume and variety.
The annotated recordings are the benchmark HRI’s machine-learning work is trained and measured against. In 2015 HRI asked us for a web application where its engineers could search that repository by what happened, see the streams side by side and pull out exactly the data they needed for analysis.
Approach
Workshops with the people who record the data
We started in June 2015 with design workshops with HRI’s teams in Europe and the US who run the recording vehicles, and turned what they told us into requirements: search by what happened and by where each sensor sits on the car, as much as possible on a single screen, and visual rather than textual.
A graph that knows what is in every drive
We designed a graph data model that describes the big-data store: vehicles, sensors, drives, data streams and the tags researchers add, from weather to lane changes. Built on Carbon LDP, the linked data platform Base22 created, it lets the application find data by the parameters engineers give it and connect related data as new recordings arrive. Graph data was still an early approach at this scale; we had bet on it with our own platform.
Play a drive back the way it happened
The application, in production in 2016, lets researchers search drives, play them back with video, a map and timelines of each annotated stream, cut clips and select data sets to deliver for analysis. Release 2 added metadata tagging, an annotation taxonomy, an annotation editor that creates the training data sets machine-learning results are compared against, and collections of annotated images taken from video.
Processing handed to the machines
In 2017 we built a Job Manager on Hadoop and Apache Kafka that automates the processing, clipping and export of video and sensor data, so a clip requested in the interface is prepared without anyone running it by hand.
Handed over to Honda in Japan
In 2018–2019 we moved the application to Honda’s research infrastructure in Japan, upgraded its front end and graph platform, automated the import of new recordings and trained Honda’s teams on the Job Manager, so the teams in the US and Japan work from the same system.
Impact
With the application, HRI’s researchers search the repository for what happened, watch every sensor in sync and send exactly the clips and streams they need to analysis or training. The proof of concept and the first production release were delivered on schedule, between June and September 2016, and their architecture became the pattern for the Job Manager that followed.
In 2019 the application moved to Honda’s teams in Japan, with teams in the US and Japan working from the same data. We designed and built it, and its graph data model, from 2015 to 2019.

