openHSU logo
Log In(current)
  1. Home
  2. Helmut-Schmidt-University / University of the Federal Armed Forces Hamburg
  3. Publications
  4. 3 - Publication references (without full text)
  5. DiaData: a multi-modal, integrated time-series dataset for type 1 diabetes research (Version v3)

DiaData: a multi-modal, integrated time-series dataset for type 1 diabetes research (Version v3)

Publication date
2026-02-02
Document type
Research data
Author
Cinar, Beyza  
Maleshkova, Maria  
Organisational unit
Data Engineering  
DOI
10.5281/zenodo.17285631
URI
https://openhsu.ub.hsu-hh.de/handle/10.24405/23893
Publisher
Zenodo
Part of the university bibliography
✅
Additional Information
Language
English
Keyword
Diabetes
Type 1 Diabetes
Hypoglycemia
Continuous Glucose Monitoring
Glucose
Hyperglycemia
Demographic Features
Heart Rate
HbA1c
Abstract
Type 1 diabetes (T1D) is an autoimmune disorder that leads to the destruction of insulin-producing cells, as to why affected individuals depend on external insulin injections. However, insulin can cause low blood glucose levels of hypoglycemia (≤70 mg/dL), which is a severe event with dangerous side effects. Data analysis can significantly enhance diabetes care by identifying personal patterns and trends leading to adverse events. However, diabetes and hypoglycemia research is limited by the unavailability of large datasets. Thus, we present DiaData. DiaData integrates 13 different datasets and presents a large continuous glucose monitoring (CGM) dataset comprising data from individuals with T1D across various age groups. CGM data is reported every 5 minutes. The Maindatabase (MDB) contains CGM measurements of all 1720 subjects. From this, three subsets are extracted: Subdatabase I (SDBI) includes CGM data and demographics of age and sex for 1709 subjects, Subdatabase II (SDBII) includes CGM and heart rate data for a subset of 51 subjects, while Subdatabase III (SDBIII) includes CGM data and personal features of sex, age, race, height, weight, age of diagnosis (DiagAge), and HbA1c values of 1651 subjects. In addition, the datasets are provided with a 15 minute sampling frequency.
Description
DiaData is provided in .csv format, where each row represents a single CGM measurement. The Maindatabase includes CGM data for all subjects, with the following columns: timestamp (ts), patient identifier (PtID), glucose value (GlucoseCGM), and the source database name (Database). Subdatabase I adds demographic information, including AgeGroup, AgeGroupBroad, and Sex. Subdatabase II contains CGM data combined with heart rate (HR) measurements. Finally, Subdatabase III integrates CGM data and demographic information of AgeGroup, AgeGroupBroad, Sex, Hba1c, DiagAgeGroup, DiagAgeGroupBroad, Race, HeightCm, and WeightKg.

The age groups are stored categorically. The groups in the "AgeGroup" and "DiagAgeGroup" columns represent:
(0: 0-2, 1: 3-6, 2: 7-10, 3: 11-13, 4: 14-17, 5: 18-25, 6: 26-35, 7: 36-55, 8:56-100)
The groups in the "AgeGroupBroad", and "DiagAgeGroupBroad" represent (0: 0-13, 1: 14-24, 2: 25-44, 3:45-100)

This release presents the raw and preprocessed version of DiaData. The raw dataset has not undergone any cleaning or imputation procedures. For subjects using a CGM device with a sampling frequency of 10 or 15 minutes, the data were sampled to 5-minute intervals. Missing values introduced by this oversampling were not imputed. In contrast, the preprocessed dataset incorporates quality enhancement steps. Outliers in the CGM and HR signals were removed using the interquartile range (IQR) method. Missing values were imputed with linear interpolation for a gap length of less than 30 minutes, and with Stineman interpolation for a gap length of 30 to 120 minutes. Moreover, the raw, unsampled heart rate data is provided, as well as the demographics, HbA1c values, and physical screening in separate csv files.


The datasets used in this study were obtained from a variety of third-party sources. The code for data preprocessing and exploration can be found in https://github.com/Beyza-Cinar/DiaData.
The sources of the data are:
- the D1NAMO dataset (https://doi.org/10.5281/zenodo.5651217),
- the HUPA-UCM Diabetes Dataset (doi: 10.17632/3hbcscwz44.1),
- the Diabetes Adolescents Time Series with Heart Rate dataset (https://github.com/ictinnovaties-zorg/dataset-diabetes-adolescents-time-series-with-heart-rate/tree/main/data-csv),
- the ShanghaiT1DM dataset (https://doi.org/10.6084/m9.figshare.20444397.v3),
- the T1GDUJA dataset (https://doi.org/10.5281/zenodo.11284018),
- the CITY dataset (https://public.jaeb.org/dataset/565),
- the ReplaceBG dataset (https://public.jaeb.org/dataset/546),
- the RT-CGM dataset (https://public.jaeb.org/dataset/563),
- the DLCP3 dataset (https://public.jaeb.org/dataset/573),
- the SENCE dataset (https://public.jaeb.org/dataset/537),
- the Severe Hypoglycemia in Older Adults with Type 1 Diabetes dataset (https://public.jaeb.org/dataset/537),
- the WISDM dataset (https://public.jaeb.org/dataset/564),
- the PEDAP dataset (https://public.jaeb.org/dataset/599).

The sources of subsets of the data are the Barbara Davis Center, Jaeb Center for Health Research, Joslin Diabetes Center, T1D Exchange, University of Colorado, and University of Virginia. The analyses, content, and conclusions presented herein are solely the responsibility of the authors and have not been reviewed or approved by the before mentioned institutions.
License: Creative Commons Attribution-NonCommercial International License (https://creativecommons.org/licenses/by-nc/4.0/)
Version
Published version
Access right on openHSU
Metadata only access

  • Privacy policy
  • Send Feedback
  • Imprint