Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

CYCLIST+IMU: A synchronized visual–inertial dataset for cyclist orientation and perception in urban environments [version 2; peer review: 2 approved with reservations]

Дата публикации: 23-07-2026 04:25:41

Background Cyclists are among the most vulnerable road users in urban traffic environments. For autonomous vehicles to interact safely and effectively with cyclists, perception systems must go beyond detection and segmentation to include an explicit understanding of cyclist orientation. However, most existing cyclist datasets lack synchronized inertial metadata describing body orientation, limiting their use in multimodal and orientation-aware perception studies. Methods This Data Note presents a multimodal visual–inertial dataset acquired using the CYCLIST+IMU framework, which synchronizes monocular RGB images captured from a vehicle-mounted camera with inertial measurements recorded by bicycle-mounted and vehicle-mounted inertial measurement units. Data were collected during multiple real-world urban acquisition sessions, resulting in 3,606 RGB images, each temporally aligned with inertial measurements, including cyclist orientation angles (yaw and roll). From these acquisitions, cyclist-centered image crops were generated and manually annotated, resulting in polygon-based semantic segmentation labels, region-of-interest detection files, and relative depth maps estimated from the RGB images. To improve angular coverage, a targeted data augmentation strategy based on horizontal image flipping was applied to underrepresented orientation ranges, resulting in the generation of 718 additional samples. The final dataset comprises 4,324 synchronized multimodal samples organized in a hierarchical directory structure that preserves one-to-one correspondence across all data modalities. Conclusions The CYCLIST+IMU dataset provides synchronized RGB image crops, inertial orientation metadata, semantic segmentation annotations, relative depth maps, and detection files for 4,324 cyclist instances captured under real urban traffic conditions. By explicitly integrating visual and inertial data with precise temporal alignment and detailed documentation, this dataset enables reproducible research on cyclist orientation estimation, semantic segmentation, and multimodal sensor fusion for intelligent transportation systems.

Основное содержимое страницы с новостью.

CROSSMARK_Color_horizontal.svg

Gómez-Meneses L, Arias-Correa M, Herrera-Ramírez J and Ballesteros JR. CYCLIST+IMU: A synchronized visual–inertial dataset for cyclist orientation and perception in urban environments [version 2; peer review: 2 approved with reservations]. F1000Research 2026, 15:527 (https://doi.org/10.12688/f1000research.177481.2)

Data Note

Revised

[version 2; peer review: 2 approved with reservations]

Luis Gómez-Meneses

https://orcid.org/0000-0002-0667-7472

1Mauricio Arias-Correa2Jorge Herrera-Ramírez3John R. Ballesteros4

Luis Gómez-Meneses

https://orcid.org/0000-0002-0667-7472

1Mauricio Arias-Correa2Jorge Herrera-Ramírez3John R. Ballesteros4

Author details Author details

1 Faculty of Engineering, Instituto Tecnologico Metropolitano, Medellín, Antioquia, Colombia
2 Design Engineering Research Group (GRID), Universidad EAFIT, Medellín, Antioquia, Colombia
3 Faculty of Exact and Applied Sciences, Instituto Tecnologico Metropolitano, Medellín, Antioquia, Colombia
4 Department of Computer Science and Decision Sciences, Universidad Nacional de Colombia Sede Medellin, Medellín, Antioquia, Colombia

Luis Gómez-Meneses
Roles: Conceptualization, Formal Analysis, Methodology, Project Administration, Software, Validation, Visualization, Writing – Original Draft Preparation

Mauricio Arias-Correa
Roles: Investigation, Resources, Writing – Original Draft Preparation

Jorge Herrera-Ramírez
Roles: Supervision, Validation, Writing – Review & Editing

John R. Ballesteros
Roles: Data Curation, Investigation, Validation

OPEN PEER REVIEW

REVIEWER STATUS

Abstract
Background

Cyclists are among the most vulnerable road users in urban traffic environments. For autonomous vehicles to interact safely and effectively with cyclists, perception systems must go beyond detection and segmentation to include an explicit understanding of cyclist orientation. However, most existing cyclist datasets lack synchronized inertial metadata describing body orientation, limiting their use in multimodal and orientation-aware perception studies.

Methods

This Data Note presents a multimodal visual–inertial dataset acquired using the CYCLIST+IMU framework, which synchronizes monocular RGB images captured from a vehicle-mounted camera with inertial measurements recorded by bicycle-mounted and vehicle-mounted inertial measurement units. Data were collected during multiple real-world urban acquisition sessions, resulting in 3,606 RGB images, each temporally aligned with inertial measurements, including cyclist orientation angles (yaw and roll). From these acquisitions, cyclist-centered image crops were generated and manually annotated, resulting in polygon-based semantic segmentation labels, region-of-interest detection files, and relative depth maps estimated from the RGB images. To improve angular coverage, a targeted data augmentation strategy based on horizontal image flipping was applied to underrepresented orientation ranges, resulting in the generation of 718 additional samples. The final dataset comprises 4,324 synchronized multimodal samples organized in a hierarchical directory structure that preserves one-to-one correspondence across all data modalities.

Conclusions

The CYCLIST+IMU dataset provides synchronized RGB image crops, inertial orientation metadata, semantic segmentation annotations, relative depth maps, and detection files for 4,324 cyclist instances captured under real urban traffic conditions. By explicitly integrating visual and inertial data with precise temporal alignment and detailed documentation, this dataset enables reproducible research on cyclist orientation estimation, semantic segmentation, and multimodal sensor fusion for intelligent transportation systems.

Keywords

Cyclist dataset, Visual–inertial dataset, Cyclist orientation, Yaw and roll angles, Vulnerable road users, Semantic segmentation, Depth maps, Urban traffic scene.

Corresponding author: Luis Gómez-Meneses Competing interests: No competing interests were disclosed.

Grant information: The author(s) declared that no grants were involved in supporting this work.

Copyright:  © 2026 Gómez-Meneses L et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. How to cite: Gómez-Meneses L, Arias-Correa M, Herrera-Ramírez J and Ballesteros JR. CYCLIST+IMU: A synchronized visual–inertial dataset for cyclist orientation and perception in urban environments [version 2; peer review: 2 approved with reservations]. F1000Research 2026, 15:527 (https://doi.org/10.12688/f1000research.177481.2) First published: 15 Apr 2026, 15:527 (https://doi.org/10.12688/f1000research.177481.1) Latest published: 23 Jul 2026, 15:527 (https://doi.org/10.12688/f1000research.177481.2)

Revised Amendments from Version 1

This revised version addresses all comments raised during the open peer review process. The Introduction was revised to improve accessibility by incorporating real-world autonomous vehicle–cyclist interaction scenarios and additional recent literature. The Data Acquisition Protocol was expanded to provide further details on participant characteristics and acquisition environments. The Dataset Validation section was strengthened by including additional information on annotation quality, synchronization verification, and practical usability. The manuscript also clarifies the current scope of the dataset regarding traffic-context annotations and identifies this as a potential direction for future dataset extensions. No changes were made to the underlying dataset or its associated repository.

See the authors' detailed response to the review by Giuseppina Pappalardo
See the authors' detailed response to the review by Ammar Al-Taie

1. Introduction

Time has proven that conventional cars are remarkably inefficient. They don’t just consume energy but also erode human health and productivity. The paradox of private ownership exacerbates this, as vehicles spend most of their existence idle, wasting precious space and materials. Driven by these shortcomings, a global interest in autonomous vehicles (AVs) has surged since the 1980s, mobilizing universities and industry leaders to rethink the very nature of mobility (Badue et al., 2019; Thrun, 2010; Narula & Tyagi, 2023).

The integration of AVs -whether for personal or public use- into heterogeneous traffic settings demands an interactive capability that transcends rudimentary obstacle recognition. These systems must interface harmoniously and safely with manual drivers, cyclists, and pedestrians. Within these shared environments, human participants navigate through an intricate web of implicit cues, such as nuanced adjustments in approach speed, alongside explicit signals like eye contact or hand gestures. These interactions establish a mutual consensus that enables the fluid synchronization of future maneuvers among road participants. However, contemporary AV architectures tend to prioritize a strict, rationalist framework of collision avoidance over social negotiation. Consequently, these vehicles frequently manifest non-human patterns, including abrupt halts, hesitant movements, or excessive delays at junctions. Such behaviors disrupt the temporal rhythm of traffic and can, paradoxically, undermine overall systemic safety (Brown & Laurier, 2017; Brown, Broth & Vinkhuyzen, 2023).

The operational scope of an AV necessitates the accurate identification of both road signage and traffic participants, specifically those lacking a protective mechanical framework—such as pedestrians and cyclists. Categorized as ‘Vulnerable Road Users’ (VRUs), these individuals are exposed to a disproportionate risk of sustaining severe injuries or fatalities in the event of traffic accidents (Flohr, 2018; Mannion, 2019).

Road traffic injuries have persisted as the twelfth leading cause of death across all age groups globally. Within this context, VRUs account for more than half of the 1.19 million annual fatalities reported by the World Health Organization (WHO, 2023). As seen in Figure 1, cyclists account for 5% of these global deaths, a percentage that has increased by nearly 20% over the last decade. This vulnerability is further intensified when cyclists must navigate mixed traffic environments, where safety is predicated on the mutual understanding of motion; a form of social coordination that contemporary autonomous systems still struggle to replicate (Ghoul & Sayed, 2025; Lu et al., 2025).

75df1662-a2a9-4733-885c-a06236d3c9a9_figure1.gif

Figure 1. Global percentage distribution of country-reported deaths by road users.

Source: (World Health Organization, 2023).

These mixed traffic environments frequently include intersections, roundabouts, lane-merging areas, and other shared-road situations in which autonomous vehicles must continuously interpret cyclist motion to support safe navigation (von Sawitzky et al., 2024; Al-Taie et al., 2025). In these scenarios, cyclist detection persists as a challenge for AV perception systems, primarily due to the inherent visual complexity associated with non-rigid articulations, highly variable aspect ratios, and a diverse range of spatial orientations. Beyond the technical impediments posed by occlusions and cluttered urban environments, contemporary research underscores that mere classification is insufficient. Systems must achieve a sophisticated understanding of behavioral intent through advanced frameworks (Corral-Soto et al., 2025b). Consequently, the estimation of orientation angles has transitioned from a secondary metric to a critical precursor for the ‘reflexive adjustment’ required for an AV to safely navigate and avoid collisions with cyclists (Brown et al., 2023).

Regardless of the previously addressed technical and social imperatives, a systematic examination of cyclist datasets published between 2023 and 2025 reveals a significant deficit in metadata fidelity. Contemporary repositories fail to provide three-dimensional orientation parameters (Roll, Pitch, Yaw) integrated with the cyclist’s posture within the image frame, as illustrated in Table 1. This lack of data limits the ability of AVs to interpret the cyclist’s body language and, consequently, delays the achievement of what Brown et al. (2023) term the ‘sociality of traffic’ for AVs.

Table 1. Comparative analysis of cyclist datasets (2023–2025).

Source: Authors.

Work (APA citation)Dataset characteristics Does it include orientation angles (Roll, Pitch, Yaw) and acquisition methodChiang, C. Y., et al. (2024). AllTheDocks road safety dataset: A cyclist’s perspective and experience. AllTheDocks: Collected in London through citizen science. Includes video (61.68 km), accelerometer, GPS, and gyroscope data.No (for cyclists in the image). The dataset includes gyroscope data (GyroX, GyroY, GyroZ), but these correspond exclusively to the ego-cyclist carrying the camera. Method: Telemetry extracted from helmet-mounted GoPro cameras.Yan, Z., Li, J., Hang, P., & Sun, J. (2025). OnSiteVRU: A high-resolution trajectory dataset for high-density vulnerable road users. OnSiteVRU: High-resolution trajectory data (0.04 s) collected in China. Covers intersections, road segments, and urban villages with 17,429 VRU trajectories.No (partial). Only includes the heading angle (direction of motion relative to the X-axis). Roll and pitch for the cyclist posture in the image are not provided. Method: Extraction using elevated vision cameras (YOLOv7/DeepSORT) and onboard vehicle sensors (LiDAR/IMU).Goren, D., & Caesar, H. (2025). BikeScenes: Online LiDAR semantic segmentation for bicycles. BikeScenes-lidarseg: LiDAR semantic segmentation dataset captured from a bicycle perspective. Contains 3,021 scans annotated into 29 semantic classes.No. Although the SenseBike platform includes an IMU for ego-motion compensation, cyclist metadata corresponds only to semantic segmentation labels, not orientation angles. Method: Offline LiDAR point cloud registration with GLIM and manual scan-level annotation.Li, M., et al. (2025). A benchmark for cycling close pass detection from video streams. Cyc-CP: Benchmark combining Victorian On-road Cycling (VOC) data and CARLA synthetic data. Focuses on close pass overtaking events.No (partial). Predicts allocentric orientation angle (θ) of the overtaking vehicle. Roll, pitch, and yaw for cyclist posture are not defined. Method: Monocular 3D detection using FCOS3D on single-view video.Desai, N. P., Etemad, A., & Greenspan, M. (2025). CycleCrash: A dataset of bicycle collision videos for collision prediction and analysis. CycleCrash: 3,000 dashcam videos with 436,347 frames depicting cyclist collisions and near-miss events.No. Direction annotations are limited to five discrete classes (forward, backward, left, right, stationary). Method: Web video curation and manual annotation based on traffic rules.Corral-Soto, E. R., et al. (2025a). 3DArticCyclists: Generating synthetic articulated 8D pose-controllable cyclist data for computer vision applications. 3DArticBikes/3DArticCyclists: Hybrid synthetic–real dataset addressing cyclist data scarcity for autonomous driving. Includes 11,086 cyclist–bicycle configurations.Yes. Provides full 3D orientation parameters (θx, θy, θz corresponding to roll, pitch, and yaw). Method: Synthetic generation using Blender and 3D Gaussian Splatting, with pose refinement via inverse kinematics based on real video data processed with CLIFF.

To address the identified lack of data, this paper presents a multimodal cyclist dataset that synchronizes real-world RGB imagery, depth maps, and semantic segmentation with precise, frame-by-frame inertial telemetry. Unlike contemporary synthetic frameworks—such as 3DArticCyclists (Corral-Soto et al., 2025a)—this dataset provides empirical ground truth for three-axis orientation (Roll, Pitch, Yaw) and triaxial acceleration (Ax, Ay, Az). By integrating these dynamic parameters, the proposed dataset enables the training of autonomous navigation models that move beyond rudimentary collision avoidance, facilitating the complex social coordination required to achieve the ‘sociality of traffic’ (Brown et al., 2023).

2. Materials and methods
2.1 Data acquisition system

Cyclist images and synchronized visual–inertial data were captured using the CYCLOPS system (cyclists’ orientation data acquisition system using RGB camera and inertial measurement units). This original development consists of a node located on a vehicle and another on a bicycle. The vehicle node includes a monocular RGB camera, an inertial measurement unit (IMU), an RF transceiver, and a microcontroller. The bicycle node includes an IMU, an RF transceiver, and a microcontroller. The system facilitates the acquisition of images of a moving cyclist and associates each image with both acceleration and orientation angles (Ax, Ay, Az, Roll, Pitch, Yaw), as illustrated in Figure 2. Similarly, camera acceleration and orientation angles are acquired at the vehicle to subsequently obtain relative values (cyclist relative to camera) and establish the cyclist’s real orientation in each image acquired from the vehicle while both are moving in an urban environment.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure2.gif

Figure 2. Frame assignment for both, the camera attached to a car’s windshield (over the vehicle) and the bicycle’s top tube (cyclist).

For the vehicle, the axes have been named Xv, Yv, and Zv, and rotations around the axis are ROLLv, PITCHv, and YAWv (respectively). Similarly, the frame for the cyclist has axes Xc, Yc, and Zc, and rotations around the axis are ROLLc, PITCHc, and YAWc (respectively). Source: Adapted from the original in Arias-Correa et al. (2024).

A significant achievement of the CYCLOPS acquisition system lies in the synchronization between images and 6-axis motion data for both actors. Full details regarding the open-source hardware architecture, the IMU-based sensor fusion, and the software suite (VideoCapture) are documented in Arias-Correa et al. (2024), ensuring complete experimental reproducibility. A general diagram of the hardware and software for each node (cyclist node and camera node) of the CYCLOPS system is presented in Figure 3.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure3.gif

Figure 3. Diagram of the hardware and software for each node of the CYCLOPS system.

The block Cyclist includes a printed circuit board (PCB), which comprises an IMU, an Arduino Nano board (Microcontroller-C), and a transceiver HC12 in transmission mode (Transceiver-C) with an antenna. Similarly, the block camera includes the RGB camera mounted on the vehicle’s windshield, an IMU (IMU-V), an Arduino Nano board (Microcontroller-V), and a transceiver HC12 (Transceiver-V) with an antenna. Both the camera and the PCB send data to the acquisition software running on a computer. Source: Adapted from the original in Arias-Correa et al. (2024).

2.2 Visual–inertial data synchronization

The CYCLOPS system implements a hardware-level synchronization protocol to ensure precise temporal registration between the visual–inertial data from the RGB camera and the high-frequency inertial data from the IMU sensors. This synchronization is executed at the point of acquisition via a deterministic software architecture that manages concurrent sensor triggering and logging across both bicycle-mounted and vehicle-mounted nodes. Specifically, each RGB frame is hardware-level synchronized to a discrete set of inertial measurements captured at the exact timestamp tk of the camera’s shutter release. This one-to-one temporal mapping constitutes the fundamental visual–inertial sample of the dataset. For every image captured by the vehicle node, the system captures the cyclist’s instantaneous kinematic state, including orientation variables such as yawb and rollb referenced to the coordinate frames established during system calibration. Crucially, this process is managed in real-time using a unified system clock, which precludes the need for subsequent corrections such as temporal interpolation, resampling, or offline alignment. By avoiding these post-processing steps, the CYCLOPS dataset maintains the integrity of the raw sensor data, providing a high-fidelity, temporally consistent snapshot of the cyclist’s pose and motion at the precise moment of visual capture as seen in Figure 4.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure4.gif

Figure 4. Result of the CYCLOPS acquisition process.

The figure shows the output generated by the CYCLOPS system for several consecutive RGB image frames acquired by the vehicle-mounted camera, together with their corresponding inertial measurements stored in format. Each image frame is temporally synchronized with a discrete set of inertial data captured at the exact acquisition instant, establishing a one-to-one correspondence between visual and inertial information. This synchronized visual–inertial data product represents the final output of the CYCLOPS framework.

All details regarding the CYCLOPS hardware components (including camera and IMU models), synchronization strategy, calibration procedures, and data verification experiments are fully described and validated in Arias-Correa et al. (2024) and are therefore not repeated here.

Each RGB image captured by the vehicle-mounted camera is associated with a corresponding set of inertial measurements describing the cyclist’s motion and orientation at the time of capture. Orientation variables, including yaw and roll, are derived from the inertial measurements and expressed in their respective reference frames as defined by the system configuration. Although the dataset also includes pitch measurements, this variable is not explicitly discussed here because the subsequent analysis and data augmentation procedures focus exclusively on yaw and roll.

2.3 Data Acquisition Protocol

Data were collected across multiple independent acquisition sessions conducted in urban environments. Each session is identified using the label Adq_x and corresponds to a continuous recording sequence acquired while the cyclist and the acquisition vehicle were in motion under real traffic conditions.

During each acquisition session, RGB images of the cyclist were recorded using the vehicle-mounted camera, while inertial measurements describing the cyclist’s motion and orientation were simultaneously captured by the bicycle-mounted and vehicle-mounted IMUs. A single sample is defined as an RGB image associated with its corresponding synchronized inertial measurements, including orientation angles such as yaw and roll, as described in Section 2.2.

The acquisition protocol was designed to capture natural cyclist behavior in unconstrained urban scenarios. Data were collected under daytime lighting conditions on public roads, without imposing specific trajectories or maneuvers on the cyclist beyond normal riding behavior.

Following data collection, a basic quality control process was applied to the acquired data. Samples exhibiting acquisition failures, severe occlusions of the cyclist, or loss of visual–inertial synchronization were excluded from the dataset. Only samples in which the cyclist was clearly visible, and the corresponding inertial data were successfully recorded and retained for further processing and inclusion in the dataset.

Data acquisition involved 10 adult volunteer cyclists participating in 84 independent acquisition sessions. Participants were anonymized using cyclist identifiers (c1–c10). No formal demographic information, such as age, sex, occupation, or cycling experience, was collected during acquisition.

To further characterize acquisition diversity, all acquisition sessions were manually reviewed and categorized according to their recording environment. Two main acquisition environments were identified: Campus Environment and Urban Street. Of the 84 acquisition sessions, 41 were collected in campus environments and 43 in urban street environments.

The distribution of acquisition sessions and synchronized images associated with each cyclist is summarized in Table 2.

Table 2. Distribution of acquisition sessions and images per cyclist.

Source: Authors.

Cyclist IDAcquisition sessions Imagesc110590c220911c33346c47547c510249c64249c79377c815354c9233c10462Total 84 3606

In this work, a total of 3,606 RGB images were acquired, each associated with its corresponding inertial measurements obtained from the IMU units of the CYCLOPS system. Images and inertial data are stored following a hierarchical file structure designed to preserve the temporal correspondence between each image and its associated orientation and acceleration records. After data acquisition, the RGB images were manually annotated using the DarkLabel tool ( https://github.com/darkpgmr/DarkLabel ) to identify and delineate the cyclist region of interest (ROI). The resulting annotations were exported in YOLO format (*.txt files), providing normalized bounding box coordinates for each cyclist and enabling their direct use in object detection and subsequent analysis pipelines.

2.4 Image processing

To construct a dataset suitable for computer vision tasks, additional processing was applied to the cyclist region-of-interest (ROI) images extracted from the original RGB images. This processing stage includes manual semantic segmentation of the cyclist and the generation of relative depth maps, as illustrated in Figure 5.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure5.gif

Figure 5. RGB image processing stages.

(a) Manual segmentation of the cyclist from RGB images using the LabelMe annotation tool, where the region corresponding to the cyclist is precisely delineated. (b) Depth map estimation from RGB images using an encoder–decoder model, generating a complementary geometric representation of the cyclist and the surrounding environment.

Polygon-based semantic segmentation of the cyclist was performed using the LabelMe annotation tool (Russell et al., 2008), which supports polygon-based annotation and is widely used in computer vision applications. As illustrated in Figure 5a, for each RGB image, the region corresponding to the cyclist was manually delineated using polygonal annotations. This procedure was applied to the entire dataset to ensure a consistent and precise definition of the region of interest across all samples. The resulting annotations were stored in JSON format, preserving the geometric information required for subsequent processing and reuse.

In parallel, the RGB images were processed to obtain depth maps using the Depth Anything model (Yang et al., 2024), as shown in Figure 5b. Depth Anything follows a base-model paradigm, and is trained at a large scale using unlabeled data, enabling robust generalization across diverse visual scenes. Although alternative approaches such as MiDaS have demonstrated strong performance through supervised and weakly supervised training on curated datasets (Ranftl et al., 2020; Birkl et al., 2023). The Depth Anything model was selected due to its ability to produce consistent relative depth estimates without relying on task-specific supervision. The depth maps included in the CYCLIST+IMU dataset represent relative depth and have been stored as 8-bit grayscale images, where depth values (non-metric) were normalized and linearly mapped to the range [0, 255] before saving in JPEG format. These depth maps are provided as complementary geometric context to the RGB images and semantic annotations.

2.5 Data augmentation

To enhance the coverage of cyclist orientations and increase the angular diversity of the dataset, a data augmentation strategy based on geometric transformations was applied. This strategy was designed to expand the range of represented orientations while preserving the physical coherence and visual consistency of the samples.

Data augmentation was performed using a horizontal flipping transformation applied to the RGB images. This operation generates additional samples by flipping the original images along the vertical axis, thereby increasing orientation diversity without introducing artificial visual artifacts or altering the geometric structure of the cyclist and the surrounding environment.

To maintain consistency between the visual content and the associated orientation labels, yaw and roll values were deterministically updated following the transformation. Under horizontal flipping, yaw angles were transformed according to: yaw′ = 360° − yaw. Meanwhile, roll values were inverted as: roll′ = −roll. This transformation was applied only to samples belonging to selected underrepresented yaw intervals, as identified during the dataset validation stage (Section 3.1).

An illustrative example of the data augmentation process is shown in Figure 6, where the original image, the augmented image, and the corresponding updates to yaw and roll angles are presented.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure6.gif

Figure 6. Example of data augmentation using horizontal image flipping.

The original RGB image and the corresponding augmented image obtained by horizontal flipping are shown. Yaw and roll angles are deterministically updated to preserve geometric consistency, with (yaw’ = 360°-yaw) and (roll’ = −roll ).

2.6 Dataset structure and organization.

The dataset is organized using a hierarchical directory structure designed to preserve traceability between original acquisitions and all derived data products. At the top level, the dataset is divided into two main directories: Original_Image, containing the raw RGB images, and Image_Crops, which stores all processed cyclist-centered data.

Within Image_Crops, data are grouped into acquisition-level subdirectories labeled Adq_x, where x denotes an independent synchronized recording session, integrating RGB imagery and inertial measurement unit (IMU) data. Each Adq_x directory contains four modality-specific subfolders: RGB_IMAGE (cropped RGB images), Detection (region-of-interest detection files), Polygons (polygon-based semantic segmentation annotations in JSON format), and Depth_Map (relative depth maps).

Each acquisition session also includes a metadata file (Adq_x.xlsx) consolidating inertial information, including cyclist orientation parameters (yaw and roll), and unique identifiers linking all data modalities. File naming follows the convention Adq_x (n), ensuring a strict one-to-one correspondence among RGB images, detection files, segmentation annotations, depth maps, and inertial measurements.

An overview of the dataset structure is provided in Figure 7, and a summary of dataset contents and file formats is presented in Table 3.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure7.gif

Figure 7. Hierarchical organization of the CYCLIST+IMU dataset.

The dataset is structured to preserve traceability between the original multimodal acquisitions and all derived data products. Raw RGB images are stored in the Original_Image directory, while processed cyclist-centered data are organized under Image_Crops by acquisition session (Adq_x). Each Adq_x represents a synchronized acquisition unit integrating RGB imagery and inertial measurement unit (IMU) data, including cyclist orientation parameters such as yaw and roll observed from the acquisition vehicle. For each session, cropped RGB images, region-of-interest detection files, relative depth maps, and polygon-based semantic segmentation annotations are stored in modality-specific subdirectories using a consistent naming convention that ensures a strict one-to-one correspondence across data types.

Table 3. Summary of the generated dataset structure.

Source: Authors.

Data typeFormat QuantityCyclist RGB images.JPG4324Segmentation polygons.JSON4324Depth maps.JPG4324Cyclist detection annotations.TXT4324IMU data.XLSX4324

Table 3 presents the complete description of the dataset generated in this work.

3. Dataset validation

A dataset validation procedure was conducted to assess the internal consistency, angular coverage, and coherence of the visual–inertial orientation metadata. Validation focuses on descriptive and structural properties of the dataset rather than on model performance, in accordance with the scope of a Data Note.

3.1 Annotation quality and angular consistency

To assess annotation quality, all cyclist detection annotations generated using DarkLabel and all polygon-based semantic segmentation masks generated using LabelMe were manually reviewed throughout the dataset construction process. Annotation inconsistencies, including inaccurate bounding boxes, incomplete cyclist contours, and imprecise polygon boundaries, were identified through visual inspection and corrected prior to dataset release. This iterative quality-control procedure required approximately three months of manual review and correction to ensure consistency across all annotated samples.

Following annotation verification, the visual–inertial orientation metadata were analysed to evaluate the angular consistency and coverage of the dataset. The CYCLOPS orientation angles include yaw and roll, which require different statistical treatments due to their measurement domains. Yaw is a periodic variable defined over [0°, 360°), while roll is restricted to a narrow, non-periodic range around zero (approximately −20° to 20°).

Yaw distributions were analysed using circular statistics to avoid discontinuities at the 0°/360° boundary, computing circular mean, median, and deviation with the pycircstat Python library. Roll values were characterised using standard linear descriptive statistics (mean, median, and standard deviation).

Figure 8 shows the yaw and roll distributions prior to data augmentation, and Table 4 reports the corresponding descriptive statistics, confirming adequate angular coverage of cyclist orientations under real-world riding conditions.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure8.gif

Figure 8. Distribution of angular variables acquired by the CYCLOPS system.

(a) Circular distribution of cyclist orientation angles (yaw), represented using circular statistics to account for the periodic nature of the variable over the [0°, 360°) range. (b) Linear distribution of cyclist inclination angles (roll), analysed using linear statistics due to their bounded range around 0°.

Table 4. Descriptive statistics of the angular variables (yaw and roll ) in the CYCLOPS dataset.

Source: Authors.

Yaw RollTotal samples: 3606Total samples: 3606Circular mean: 297.05°Mean: 0.44°Circular median: 283.81°median: −0.81°Circular deviation: 141.12°Standard deviation: 9.68°
3.2 Validation of the data augmentation strategy

The data augmentation procedure based on horizontal image flipping was validated to ensure geometric consistency between RGB images and orientation angles. After augmentation, yaw and roll values were deterministically updated following the transformation rules defined in the Methods section.

Consistency between augmented images, updated orientation angles, and the associated segmentation polygons, detection files, and relative depth maps was verified through visual inspection. Figure 6 illustrates the augmentation process, while Figure 9 shows the post-augmentation distributions of yaw and roll, highlighting the angular regions targeted to improve coverage.

75df1662-a2a9-4733-885c-a06236d3c9a9_figure9.gif

Figure 9. Post-augmentation distribution of cyclist orientation and inclination variables.

(a) Circular distribution of orientation angles (yaw) after data augmentation. Angular bins highlighted in red indicate the yaw intervals selected for horizontal flipping and sample augmentation. (b) Linear distribution of inclination angles (roll ) after data augmentation. Red-highlighted bins correspond to the roll range associated with the augmented samples, while blue bins represent the remaining original and augmented data.

Table 5 summarizes the descriptive statistics after augmentation, confirming enhanced angular coverage without altering the overall structure of the original dataset.

Table 5. Descriptive statistics of the angular variables (yaw and roll ) after data augmentation in the CYCLOPS dataset.

Source: Authors.

Yaw RollTotal samples: 4324Total samples: 4324Circular mean: 307.15°Mean: −0.34°Circular median: 295.56°median: −0.56°Circular deviation: 140.22°Standard deviation: 8.94°
3.3 Synchronization verification and practical usability

The visual–inertial synchronization quality of the dataset is supported by the CYCLOPS acquisition framework, which implements a hardware-level synchronization protocol between the RGB camera and the inertial measurement units. The synchronization strategy, calibration procedures, and validation experiments of the acquisition system were previously described and experimentally validated by Arias-Correa et al. (2024). During dataset preparation, samples exhibiting synchronization failures, incomplete inertial records, or acquisition anomalies were excluded as part of the quality-control process. Consequently, only synchronized image–IMU pairs with complete associated metadata were retained in the final dataset.

The practical usability of the CYCLIST+IMU dataset is supported by the strict one-to-one correspondence maintained across RGB images, inertial measurements, cyclist detection annotations, polygon-based segmentation masks, and relative depth maps. This multimodal structure enables direct use of the dataset for cyclist orientation estimation, cyclist detection, semantic segmentation, visual–inertial sensor fusion, depth-aware perception, and autonomous vehicle research involving vulnerable road users.

The current dataset release provides synchronized visual–inertial information together with cyclist orientation, detection annotations, semantic segmentation masks, and relative depth maps to support perception tasks in autonomous driving research. Although acquisition environments were characterized at the session level (Campus Environment and Urban Street), the dataset does not currently include fine grained semantic annotations describing specific traffic situations, such as intersections, roundabouts, or lane-merging events. Such contextual annotations could further support future research on cyclist orientation interpretation and autonomous vehicle perception across diverse traffic scenarios. Their incorporation would require an additional semantic annotation process beyond the scope of the current dataset release and is therefore considered a valuable direction for future dataset extensions.

Ethical considerations

The data presented in this Data Note were collected in public urban environments under natural traffic conditions. The acquisition protocol consisted of recording cyclist behavior in real-world settings without clinical intervention, behavioral manipulation, or collection of personal identifiers.

All cyclists recorded in this dataset were adult volunteers known to the research team and were fully informed about the purpose of the data acquisition and the intended public release of the dataset. Informed consent for participation and publication of anonymized data was obtained verbally prior to data collection. Written consent was not deemed necessary because no personal identifiable information was collected, and all visual data were anonymized prior to public release.

No personal data such as names, identification numbers, contact information, or biometric identifiers were collected or stored during acquisition. Prior to publication, all captured visual images included in both the publicly released dataset and this manuscript were automatically processed using a YOLO-based face detection model, and any detected facial regions were anonymized through Gaussian blurring (21 × 21 kernel) to prevent individual identification. This anonymization procedure was systematically applied to all applicable captured images before public release.

Derived data products such as segmentation masks, depth maps, and inertial metadata do not contain identifiable facial information.

Because the released dataset does not contain personally identifiable information and consists of non-interventional observational recordings conducted with informed adult volunteers in public environments, this study qualifies as research without risk according to Colombian national regulations governing health research involving human participants (Resolution 8430 of 1993, Ministry of Health of Colombia). Under these regulations and applicable institutional guidelines, formal approval from an Institutional Review Board (IRB) or ethics committee was not required.

The individuals shown in Figures 4, 5, and 6, as well as all individuals appearing in the publicly released dataset, correspond to the same adult volunteers who provided informed consent for participation and publication of anonymized images. No third-party individuals were intentionally included in the dataset.

The study was conducted in accordance with the ethical principles outlined in the Declaration of Helsinki, insofar as applicable to non-interventional observational data collection.

Data availability

Open Science Framework (OSF). CYCLIST+IMU: A synchronized visual–inertial dataset for cyclist orientation and perception in urban environments. https://doi.org/10.17605/OSF.IO/HVPKZ (Gómez-Meneses et al., 2026).

This project contains the following underlying data:

  • CYCLIST_IMU_Dataset.zip (Complete dataset including RGB images, cyclist-centered image crops, inertial measurement files (.xlsx), semantic segmentation polygons (.json), region-of-interest detection annotations (.txt), and relative depth maps (.jpg), organized by acquisition session.)

Data is available under the terms of the Creative Commons Attribution 4.0 International copyright (CC BY 4.0) license.

Acknowledgements

The authors have no acknowledgements to declare.

References
  •  Al-Taie A, Matviienko A, O’Hagan J, et al.: Around the world in 60 cyclists: Evaluating autonomous vehicle-cyclist interfaces across cultures. Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM; 2025. Publisher Full Text
  •  Arias-Correa M, Robledo S, Londoño M, et al.: CYCLOPS: A cyclists’ orientation data acquisition system using RGB camera and inertial measurement units (IMU). HardwareX. 2024; 18: e00534. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Badue C, Guidolini R, Carneiro RV, et al.: Self-driving cars: A survey. Expert Syst. Appl. 2019; 165: 113816. Publisher Full Text
  •  Birkl R, Wofk D, Müller M: MiDaS v3.1 – A model zoo for robust monocular relative depth estimation. arXiv preprint, arXiv:2307.14460. 2023. Reference Source
  •  Brown B, Laurier E: The trouble with autopilots: Assisted and autonomous driving on the social road. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. ACM; 2017; pp. 416–429. Publisher Full Text
  •  Brown B, Broth M, Vinkhuyzen E: The halting problem: Video analysis of self-driving cars in traffic. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. ACM; 2023; pp. 1–14. Publisher Full Text
  •  Chiang CY, Zhong R, Ding J, et al.: AllTheDocks road safety dataset: A cyclist's perspective and experience. 2024 IEEE 99th Vehicular Technology Conference (VTC2024-Spring). IEEE; 2024, June; pp. 1–5.
  •  Corral-Soto ER, Liu Y, Ren Y, et al.: 3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications. 2025 IEEE Intelligent Vehicles Symposium (IV). IEEE; 2025a, June; pp. 2114–2121.
  •  Corral-Soto ER, Liu Y, Ren Y, et al.: Monocular Visual 8D Pose Estimation for Articulated Bicycles and Cyclists. arXiv preprint arXiv:2510.20158. 2025b.
  •  Desai NP, Etemad A, Greenspan M: CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis. 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE; 2025, February; pp. 6688–6698.
  •  Flohr FB: Vulnerable road user detection and orientation estimation for context-aware automated driving. Universiteit van Amsterdam). UvA-DARE (Digital Academic Repository); 2018. (Doctoral dissertation).
  •  Goren D, Caesar H: BikeScenes: Online LiDAR Semantic Segmentation for Bicycles. arXiv preprint arXiv:2510.25901. 2025.
  •  Ghoul T, Sayed T: Cyclist safety assessment using autonomous vehicles. Accid. Anal. Prev. 2025; 212: 107923. PubMed Abstract | Publisher Full Text
  •  Gómez-Meneses L, Arias-Correa M, Herrera-Ramírez J, et al.: CYCLIST+IMU: A synchronized visual–inertial dataset for cyclist orientation and perception in urban environments [Data set]. Open Science Framework. 2026. Publisher Full Text
  •  Li M, Beck B, Rathnayake T, et al.: A benchmark for cycling close pass detection from video streams. Transportation Research Part C: Emerging Technologies. 2025; 174: 105112. Publisher Full Text
  •  Lu H, Zhu M, Lu C, et al.: Empowering safer socially sensitive autonomous vehicles using human-plausible cognitive encoding. Proc. Natl. Acad. Sci. 2025; 122(21): e2401626122. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Mannion P: Vulnerable road user detection: State-of-the-art and open challenges. arXiv preprint arXiv:1902.03601. 2019. Reference Source
  •  Narula M, Tyagi D: Autonomous cars: A comprehensive survey. 2023 Seventh International Conference on Image Information Processing (ICIIP). IEEE; 2023; pp. 586–590. Publisher Full Text
  •  Ranftl R, Li J, Lasinger K, Hafner D, et al.: Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2022; 44(3): 1623–1637. Publisher Full Text
  •  Russell BC, Torralba A, Murphy KP, et al.: LabelMe: A Database and Web-Based Tool for Image Annotation. Int. J. Comput. Vis. 2008; 77(1–3): 157–173. Publisher Full Text
  •  Thrun S: Toward robotic cars. Commun. ACM. 2010; 53(4): 99–106. Publisher Full Text
  •  von Sawitzky T, Löcken A, Grauschopf T: Enhancing cyclist safety in cyclist-vehicle interactions through early hazard notifications: A comparison of bi-modal cues at head level. Traffic Saf. Res. 2024; 7: e000070. Publisher Full Text
  •  World Health Organization: Global status report on road safety 2023.2023. Reference Source
  •  Yang L, Kang B, Huang Z, et al.: Depth anything: Unleashing the power of large-scale unlabeled data. arXiv. 2024. Reference Source
  •  Yan Z, Li J, Hang P, et al.: OnSiteVRU: A High-Resolution Trajectory Dataset for High-Density Vulnerable Road Users. arXiv preprint arXiv:2503.23365. 2025.

Comments on this article Comments (0)

Version 2

VERSION 2 PUBLISHED 15 Apr 2026

Comment

Grant information

The author(s) declared that no grants were involved in supporting this work.

Copyright

© 2026 Gómez-Meneses L et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

Open Peer Review

Current Reviewer Status: ?

Key to Reviewer Statuses VIEW HIDE

ApprovedThe paper is scientifically sound in its current form and only minor, if any, improvements are suggested

Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit.

Not approvedFundamental flaws in the paper seriously undermine the findings and conclusions

Version 1

VERSION 1

PUBLISHED 15 Apr 2026

Reviewer Report 01 Jun 2026

Ammar Al-Taie, KAIST, Daejeon, South Korea 

Approved with Reservations

VIEWS 0

  • Is the rationale for creating the dataset(s) clearly described?

    Yes

  • Are the protocols appropriate and is the work technically sound?

    Partly

  • Are sufficient details of methods and materials provided to allow replication by others?

    Partly

  • Are the datasets clearly presented in a useable and accessible format?

    Yes

Competing Interests: No competing interests were disclosed.

Reviewer Expertise: I am a Human-Computer Interaction researcher investigating how Automated Vehicles can successfuly and safely communicate their intentions to surrounding road users.

Close

Reviewer Report 13 May 2026

Giuseppina Pappalardo, University of Catania, Catania, Italy 

Approved with Reservations

VIEWS 0

  • Is the rationale for creating the dataset(s) clearly described?

    Yes

  • Are the protocols appropriate and is the work technically sound?

    Yes

  • Are sufficient details of methods and materials provided to allow replication by others?

    No

  • Are the datasets clearly presented in a useable and accessible format?

    Yes

Competing Interests: No competing interests were disclosed.

Reviewer Expertise: Cyclist safety

Close

Comments on this article Comments (0)

Version 2

VERSION 2 PUBLISHED 15 Apr 2026

Comment

Open Peer Review
Reviewer Status

Alongside their report, reviewers assign a status to the article:

Approved
The paper is scientifically sound in its current form and only minor, if any, improvements are suggested
Approved with reservations
A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit.
Not approved
Fundamental flaws in the paper seriously undermine the findings and conclusions

Reviewer Reports
Invited Reviewers
1 2
Version 2
(revision)
23 Jul 26
Version 1
15 Apr 26
read read

  1. Giuseppina Pappalardo, University of Catania, Catania, Italy

  2. Ammar Al-Taie, KAIST, Daejeon, South Korea


Comments on this article

Sign up for content alerts


Browse by related subjects

Alongside their report, reviewers assign a status to the article:

Approved - the paper is scientifically sound in its current form and only minor, if any, improvements are suggested

Approved with reservations - A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit.

Not approved - fundamental flaws in the paper seriously undermine the findings and conclusions

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Machine Learning Models for Predicting Long-Term Visual Acuity in Highly Myopic Eyes09.7901-12-2023
2MedGemma Evaluation for Fundamental Radiological Imaging Classification Tasks [version 1; peer review: awaiting peer review]09.1804-08-2026
3Canyon’s new ebike puts serious V2X safety smarts on 2 wheels3604-07-2026
4Review of Hybrid Localization Frameworks in Wireless Sensor Networks for Precision Agriculture Applications [version 1; peer review: awaiting peer review]09.5924-07-2026
5Why Static Biomarkers Often Fall Short: Circadian Immune Coherence as a Missing Dimension in Immunotherapy [version 2; peer review: 1 approved with reservations]013.0215-07-2026
6Equity impacts of active travel interventions: A systematic evidence map of the extent and nature of evaluation research [version 1; peer review: awaiting peer review]05.405-08-2026
7Green Space Morphology and School Myopia in China09.701-02-2024
8Гигантский российский геосервис научился видеть города и дороги там, где не работает GPS5829-06-2026
9Инновации для «умной» городской инфраструктуры01027-07-2026
10В Новосибирске создали ИИ для анализа дорожных сцен по видео0725-06-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 7.7. Источник: f1000research.com.