top of page

DATASET CURATION AND PIPELINE

In developing any perception model (in supervised or unsupervised learning), the quality of the training data plays a key role in the final performance. We take quality of the data very seriously and we have developed a rigorous pipeline to ensure the highest quality of data and we always have the training process in mind when collecting and curating the dataset. We call a dataset training ready if it is reviewed, cleaned, and annotated in a way that it can be used for training any deep learning model. All the data must follow the same format so any arbitrary architecture can be trained on it and any trained models can be evaluated on it. This is critical at our scale simply because you cannot run arbitrary processes, filters and transformations on millions of data frames in a reasonable amount of time for each individual model.
 

To ensure quality of the datasets, we have also built specialized sensor rigs with custom software to collect the data exactly the way we want it for training purposes. This alone saves us significant time and resources that would otherwise be spent on cleaning and annotating data. In this section, we describe some of the main aspects of our datasets which helps us guarantee the quality of the data.

crop boundaries

Hardware and Data Collection

All the data in FarmScapes is collected using a custom designed sensor suite retrofitted onto standard broadacre machinery (tractors, combines, sprayers). The suite in cludes :
 

  • Cameras : Multiple high resolution RGB cameras in multi-view surround configurations (10+ cameras) and dedicated stereo pairs with baselines ranging from 20cm to 70cm.

  • LiDAR : High resolution LiDAR sensors providing detailed 3D point clouds for distance and terrain mapping.

  • Localization : RTK GNSS and high precision IMU sensors providing 6-DOF pose and orientation.

  • Synchronization : All sensors are synchronized at sub millisecond level ensuring spatial-temporal alignment across modalities.

  • Calibration : All sensors are calibrated providing accurate intrinsic and extrinsic parameters.

Dataset Specifications and Diversity

The FarmScapes dataset encompasses :

  • Scale : hundreds of millions of frames, totaling over half
    a petabyte of datasets.

  • Temporal Diversity : Data collected over five consecutive farming seasons (2020–2025) across Alberta and Saskatchewan.

  • Operational Diversity : Coverage of all farming stages and operations, including yard driving, road driving, seeding, land rolling, harrowing, tillage, spraying, harvesting, and more.

  • Environmental Diversity : distribution of scenes across dawn, dusk, bright daylight, and night, as well as varying conditions such as rain, sunny, and dusty.

  • Data Authenticity : All data is collected from real farming operations and is not simulated with permission from the owners. We take privacy very seriously and our data complies with all the privacy regulations and we have obtained all the necessary permissions to collect and use this data.

Annotation and Quality Pipeline

To maintain consistency at this scale, we have developed a proprietary 100% in-house annotation pipeline as we didn’t want to rely on external services that often lack the quality and consistency that we require. Furthermore, 3rd party annotation companies often use untrained staff (especially for broadacre farming) which has an extremely high turnover rate that directly results in low quality data and inconsistent annotations. Every subset is tagged using our tagging system which contains over 150+ metadata labels and descriptors describing objects (e.g., specific crops, implements, power poles, lighting, weather, etc.). This allows for fine-grained retrieval of specific subsets for evaluation and testing purposes or generally for custom dataset creation. We utilize a dedicated internal team of trained annotators and reviewers. We have developed a standard guideline for all the human annotators and reviewers to follow. All annotations undergo a two-tier independent review process. This rigorous internal system helps us set a high bar for the quality of the data. With our query system, you can query any subset of the dataset using any of the metadata labels and descriptors. For example, you can retrieve 10,000 frames of data from canola fields in cloudy days from 2023 spring season. The datasets are annotated for different tasks such as object detection, semantic segmentation, depth estimation, and more. We have specialized datasets for specific tasks that also contain dense annotations for all frames.

mojow logo
  • Youtube
  • LinkedIn
  • X

EMAIL

LOCATIONS

Corporate Office:

White City, SK 

Canada

Development Office: 

Edmonton, AB 

Canada

For potential business or farm collaboration, contact: 

Doug Schmuland

VP of Business Development

+1 (403) 807-8349

doug@mojow.ai

© 2026 Mojow Autonomous Solutions Inc. Powered and secured by Wix

Subscribe to our mailing list!
bottom of page