Skip to content
Menu

Infographic Dataset

  • The Doodle Desk
  • Infographics
  • SVG 1.1
  • v1.0.0
  • Research
  • Commercial AI
  • Enterprise
  • 1,800,000 assets
  • updated 2026-09-03

Single-page infographics with live text, chart geometry and a 33-field record per asset.

Assets
1,800,000
On disk
75 GB
Lines
4
Format
SVG 1.1
Live text
Yes
Vector only
Yes
Sample
122 files

Preview

Inspect a file from the dataset

Infographic page: The help desk at Ullswater Group:, structure layer

Every shape as its path geometry

A2_050001_support_Ullswater-Group_2022

A2_050001_support_Ullswater-Group_2022 · 909×1285 · 65 KB

<text>
88
<path>
61
<g>
18
Live text nodes
Text nodesizewt
ONE PRODUCT, ONE YEAR900700
The help desk at Ullswater Group:2600700
20222600700
67,8007000700
tickets1500700
The year's tickets, arranged by what they1080400
concerned and how the volume moved.1080400
011350700
+ 32 more

Palette declared in file

  • #494949
  • #a03024
  • #8e9500
  • #009296
  • #ffffff
  • #9399ff

Showing file 1 of 12: A2_050001_support_Ullswater-Group_2022

fig. 01A2_050001_support_Ullswater-Group_2022.svg · 909×1285 · 88 text · 61 path · 65 KB · 1 of 12 · rendered with resvg; raw SVG is not served from this origin
  • Infographic page: Ullswater Collective, 2022:
    fig. 02A2_050002_warehouse_Ullswater-Collective_2022.svg909×1285 · text 63 · path 84 · 53 KB
  • Infographic page: Readings in munnar, 2022
    fig. 03IB_IN_000544_tower4_raingauge_tabloid_p.svg1650×2550 · text 42 · path 0 · 122 KB
  • IGX344013 — IGK045 — IGK045-V04
    fig. 04IGX344013.svg1400×1400 · text 70 · path 2 · 107 KB
  • IGX588013 — IGH0385 — IGH0385-V07
    fig. 05IGX588013.svg1600×1200 · text 63 · path 4 · 92 KB
  • Infographic page: Every title counted at the bookshop in Murshidabad,
    fig. 06IB_IN_000756_trio_band_bookshop_p18x24.svg2700×3600 · text 39 · path 8 · 184 KB
  • Infographic page: Ullswater Trust counted
    fig. 07A2_050004_saas_Ullswater-Trust_2022.svg909×1285 · text 55 · path 50 · 48 KB
  • Infographic page: Reading Ullswater Partners
    fig. 08A2_050012_workshop_Ullswater-Partners_2022.svg909×1285 · text 59 · path 57 · 35 KB
  • Infographic page: lessons in dharmanagar, 2023
    fig. 09IB_IN_082672_band_quad_musicschool_a3_p.svg1754×2480 · text 47 · path 6 · 129 KB
  • IGX288013 — IGD089 — IGD089-V04
    fig. 10IGX288013.svg1000×2600 · text 78 · path 4 · 103 KB
  • Infographic page: Ullswater, 2023
    fig. 11A2_050017_fleet_Ullswater_2023.svg909×1285 · text 69 · path 83 · 39 KB
  • Infographic page: Every kilo counted at the apiary in Lucknow, sorted by what it was and when.
    fig. 12IB_IN_000810_stack_pair_beekeeping_a3_p.svg1754×2480 · text 34 · path 6 · 119 KB
  • Infographic page: tiffins in panipat, 2021
    fig. 13IB_IN_246615_lead_grid5_tiffin_a2_p.svg2480×3508 · text 64 · path 72 · 144 KB

24 of 122 preview files shown · the sample pack contains 122 originals

Overview

What this dataset is

Four production lines of editable SVG infographics. Each page reports one subject over one period, and every chart on it is a different view of the same figures: a share breakdown, a ranked comparison, a trend, a process. All copy is live SVG text. Figures, names and captions are written placeholders; the data teaches layout, chart construction and typographic structure, not real-world statistics.

  • Live text
  • Vector geometry
  • Chart families in slot order
  • Palette and font IDs
  • SHA-256 per asset

AI use cases

  • Text-to-SVG
  • Layout generation
  • Chart understanding
  • Document AI
  • Design AI
  • Evaluation

Specifications

Dataset specification

Format, DOM node types, interleaved text and graphics, and the content rules (no brands or trademarks, no personally identifiable information) are stated for every dataset on the specifications page.

Assets
1,800,000
Format
SVG, UTF-8
DOM nodes
<text> editable strings, <path> geometry, <g> layout groups; <image> only where stated under Raster content
Content rules
No brands or trademarks, no personally identifiable information, placeholder figures throughout
Text
Live SVG text throughout, no outlined type
Raster content
None. Files carrying an embedded bitmap are excluded from delivery, samples and previews; every delivered file is vector only
Page sizes
15 across lines, from 1000×2600 long-form to 1962×1104 widescreen
Metadata
33 fields (composite), 16 fields (poster-size, slide-deck), 9 listing fields plus manifests (Indian business)
Median file size
47 KB (poster-size), 39 KB (slide-deck)
Delivery
tar and tar.zst shards with index.csv and SHA-256 sums

Production lines

Production lines in this dataset
LineAssetsNotes
Composite700,0002,074 source layouts, 56 sectors, 854 subjects, 11 page sizes, 33 metadata fields
Indian business250,00042 subjects across 417 towns and cities, 2019 to 2026, 5 to 9 chart modules a page
Poster-size350,000A2 portrait, 97,609 distinct stories, 306 palettes, 31 grids, 120 typefaces
Slide-deck500,00016:9, 156,169 distinct stories, 131 chart families, 15 reading types

Taxonomy coverage

Sectors
56
Subjects
854
Chart families
131
Source layouts
2,074
Palettes
306
Typefaces
120 families

Schema

Per-asset record

Parquet, mirroring the field conventions of MMSVG-2M and Hugging Face datasets so existing loaders work unchanged. The texts[] and layout[] arrays are not offered by any public SVG dataset.

asset_id, dataset_id, dataset_version, line, file_path, sha256
svg_raw
svg_normalized       fixed viewBox, transforms baked, CSS inlined, M/L/C/Q/A/Z only
png_448, width, height, orientation, page_format
text_count, path_count, group_count, image_count, element_count, token_len, complexity_tier
texts[]              {content, role, font_family, font_weight, font_size, bbox, script}
layout[]             {element_id, type, bbox, z_order, parent_group}
palette_id, colours[], is_dark, font_ids[], font_licences[]
sector, subject, industry, geography, tags[], headline
caption_short, caption_medium, caption_detailed
provenance           {template_id, illustration_source_url, illustration_licence, generator, generated_at}
quality              {valid_svg, renders, has_live_text, has_raster, near_dup_group, phash}

Quality

Checks and results

Badges are published now. A composite SVGZO Quality Score follows once the formula is frozen and applied to every line. Method →

Byte-identical pages
0 of 850,000 audited
Repeated headline and subject pairs
0
Pages with a figure their story does not hold
0
Near-duplicate signatures
Perceptual hash across the Indian business line; SHA-256 per asset on composite
Audit date
3 September 2026

Read before licensing

  • About 4% of Indian business pages carry one embedded bitmap. A vector-only filter is provided.
  • The corpus teaches layout and chart construction. It is not a source of real-world statistics.

Provenance

Where the data comes from

  • Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators.
  • Every figure, name and caption is a written placeholder. No real organisation, person or measurement appears.

A copy-ready EU AI Act training-summary paragraph and a per-asset manifest ship with every licence. Trust and provenance →

Licence

Tiers available for this dataset

Research

For
Academic and non-commercial experimentation on subsets of 25,000 to 100,000 assets.
Rights
  • Train, fine-tune and evaluate models
  • Licensee owns models and outputs
  • Perpetual, worldwide
Limits
  • Non-commercial deployment only
  • No redistribution of raw data
  • Attribution required
  • No font-generation models

Commercial AI

For
Model training and commercial AI products on a sub-line or a full dataset.
Rights
  • Train, fine-tune and evaluate, including commercial models and products
  • Licensee owns models, weights, embeddings and outputs
  • Share with contractors under NDA
  • Safe harbour for incidental memorisation
  • Chain-of-title warranty, liability capped at fees
Limits
  • No redistribution or resale of raw data
  • No reconstructable copy of the dataset in a model
  • No font-generation models
  • No biometric or real-person inference

Enterprise

For
Custom volume, exclusivity, provenance audit, indemnity and delivery terms.
Rights
  • Everything in Commercial AI
  • Affiliates and named contractors
  • Optional exclusivity on custom or carved-out sets
  • IP indemnity, cap at 1 to 2x fees
  • Audit access and change notices
Limits
  • Negotiated

Full texts on the licensing page. Drafts pending counsel review.

Files

Loading the data

Machine-readable metadata: croissant.json. Version history: changelog.

from datasets import load_dataset

ds = load_dataset("parquet", data_files="metadata/*.parquet", split="train")
row = ds[0]
print(row["text_count"], row["path_count"], row["caption_short"])

# Full SVG files ship as tar shards; verify before extracting:
#   sha256sum -c checksums/SHA256SUMS

FAQ

Common questions

Is the content real?

No. Every name, figure and caption is a written placeholder. The dataset teaches layout, chart construction and typography, not real-world statistics.

Where does the data come from?

Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators. A per-asset manifest and a training-content summary paragraph ship with every licence.

Can I train a commercial model?

Yes, under the Commercial AI or Enterprise tier. The Research tier is limited to non-commercial deployment.

Can I redistribute the files?

No tier permits redistributing or reselling the raw data. Models trained on it are yours.

How is it delivered?

Sharded tar archives with SHA-256 sums, a parquet metadata index and a Croissant manifest, via signed object-storage URLs or a scoped bucket for rclone.

Licence

Request pricing

Priced per subset, sub-line or full dataset. Reply within one business day.

Assets
1,800,000
On disk
75 GB
Format
SVG 1.1, UTF-8
Version
1.0.0
Updated
2026-09-03
Live text
every file in the preview set
Vector only
every file in the preview set
Median file
76 KB
Median text nodes
64

Use this dataset

pip install datasets
load_dataset("parquet",
  data_files="metadata/*.parquet")

croissant.json · changelog

Need this at enterprise scale, or with custom taxonomy?

Bulk licensing, exclusivity, private delivery and transformation.

Talk to the data team