Infographic Dataset
- The Doodle Desk
- Infographics
- SVG 1.1
- v1.0.0
- Research
- Commercial AI
- Enterprise
- 1,800,000 assets
- updated 2026-09-03
Single-page infographics with live text, chart geometry and a 33-field record per asset.
- Assets
- 1,800,000
- On disk
- 75 GB
- Lines
- 4
- Format
- SVG 1.1
- Live text
- Yes
- Vector only
- Yes
- Sample
- 122 files
Preview
Inspect a file from the dataset

Every shape as its path geometry
A2_050001_support_Ullswater-Group_2022
A2_050001_support_Ullswater-Group_2022 · 909×1285 · 65 KB
- <text>
- 88
- <path>
- 61
- <g>
- 18
| Text node | size | wt |
|---|---|---|
| ONE PRODUCT, ONE YEAR | 900 | 700 |
| The help desk at Ullswater Group: | 2600 | 700 |
| 2022 | 2600 | 700 |
| 67,800 | 7000 | 700 |
| tickets | 1500 | 700 |
| The year's tickets, arranged by what they | 1080 | 400 |
| concerned and how the volume moved. | 1080 | 400 |
| 01 | 1350 | 700 |
| + 32 more | ||
Palette declared in file
- #494949
- #a03024
- #8e9500
- #009296
- #ffffff
- #9399ff
Showing file 1 of 12: A2_050001_support_Ullswater-Group_2022

fig. 02A2_050002_warehouse_Ullswater-Collective_2022.svg909×1285 · text 63 · path 84 · 53 KB 
fig. 03IB_IN_000544_tower4_raingauge_tabloid_p.svg1650×2550 · text 42 · path 0 · 122 KB 
fig. 04IGX344013.svg1400×1400 · text 70 · path 2 · 107 KB 
fig. 05IGX588013.svg1600×1200 · text 63 · path 4 · 92 KB 
fig. 06IB_IN_000756_trio_band_bookshop_p18x24.svg2700×3600 · text 39 · path 8 · 184 KB 
fig. 07A2_050004_saas_Ullswater-Trust_2022.svg909×1285 · text 55 · path 50 · 48 KB 
fig. 08A2_050012_workshop_Ullswater-Partners_2022.svg909×1285 · text 59 · path 57 · 35 KB 
fig. 09IB_IN_082672_band_quad_musicschool_a3_p.svg1754×2480 · text 47 · path 6 · 129 KB 
fig. 10IGX288013.svg1000×2600 · text 78 · path 4 · 103 KB 
fig. 11A2_050017_fleet_Ullswater_2023.svg909×1285 · text 69 · path 83 · 39 KB 
fig. 12IB_IN_000810_stack_pair_beekeeping_a3_p.svg1754×2480 · text 34 · path 6 · 119 KB 
fig. 13IB_IN_246615_lead_grid5_tiffin_a2_p.svg2480×3508 · text 64 · path 72 · 144 KB
24 of 122 preview files shown · hover a tile for its structure layer · the sample pack contains 122 originals
Overview
What this dataset is
Four production lines of editable SVG infographics. Each page reports one subject over one period, and every chart on it is a different view of the same figures: a share breakdown, a ranked comparison, a trend, a process. All copy is live SVG text. Figures, names and captions are written placeholders; the data teaches layout, chart construction and typographic structure, not real-world statistics.
- Live text
- Vector geometry
- Chart families in slot order
- Palette and font IDs
- SHA-256 per asset
AI use cases
- Text-to-SVG
- Layout generation
- Chart understanding
- Document AI
- Design AI
- Evaluation
Specifications
Dataset specification
Format, DOM node types, interleaved text and graphics, and the content rules (no brands or trademarks, no personally identifiable information) are stated for every dataset on the specifications page.
- Assets
- 1,800,000
- Format
- SVG, UTF-8
- DOM nodes
- <text> editable strings, <path> geometry, <g> layout groups; <image> only where stated under Raster content
- Content rules
- No brands or trademarks, no personally identifiable information, placeholder figures throughout
- Text
- Live SVG text throughout, no outlined type
- Raster content
- None. Files carrying an embedded bitmap are excluded from delivery, samples and previews; every delivered file is vector only
- Page sizes
- 15 across lines, from 1000×2600 long-form to 1962×1104 widescreen
- Metadata
- 33 fields (composite), 16 fields (poster-size, slide-deck), 9 listing fields plus manifests (Indian business)
- Median file size
- 47 KB (poster-size), 39 KB (slide-deck)
- Delivery
- tar and tar.zst shards with index.csv and SHA-256 sums
Production lines
| Line | Assets | Notes |
|---|---|---|
| Composite | 700,000 | 2,074 source layouts, 56 sectors, 854 subjects, 11 page sizes, 33 metadata fields |
| Indian business | 250,000 | 42 subjects across 417 towns and cities, 2019 to 2026, 5 to 9 chart modules a page |
| Poster-size | 350,000 | A2 portrait, 97,609 distinct stories, 306 palettes, 31 grids, 120 typefaces |
| Slide-deck | 500,000 | 16:9, 156,169 distinct stories, 131 chart families, 15 reading types |
Taxonomy coverage
- Sectors
- 56
- Subjects
- 854
- Chart families
- 131
- Source layouts
- 2,074
- Palettes
- 306
- Typefaces
- 120 families
Schema
Per-asset record
Parquet, mirroring the field conventions of MMSVG-2M and Hugging Face datasets so existing loaders work unchanged. The texts[] and layout[] arrays are not offered by any public SVG dataset.
asset_id, dataset_id, dataset_version, line, file_path, sha256
svg_raw
svg_normalized fixed viewBox, transforms baked, CSS inlined, M/L/C/Q/A/Z only
png_448, width, height, orientation, page_format
text_count, path_count, group_count, image_count, element_count, token_len, complexity_tier
texts[] {content, role, font_family, font_weight, font_size, bbox, script}
layout[] {element_id, type, bbox, z_order, parent_group}
palette_id, colours[], is_dark, font_ids[], font_licences[]
sector, subject, industry, geography, tags[], headline
caption_short, caption_medium, caption_detailed
provenance {template_id, illustration_source_url, illustration_licence, generator, generated_at}
quality {valid_svg, renders, has_live_text, has_raster, near_dup_group, phash}Quality
Checks and results
Badges are published now. A composite SVGZO Quality Score follows once the formula is frozen and applied to every line. Method →
- Byte-identical pages
- 0 of 850,000 audited
- Repeated headline and subject pairs
- 0
- Pages with a figure their story does not hold
- 0
- Near-duplicate signatures
- Perceptual hash across the Indian business line; SHA-256 per asset on composite
- Audit date
- 3 September 2026
Read before licensing
- About 4% of Indian business pages carry one embedded bitmap. A vector-only filter is provided.
- The corpus teaches layout and chart construction. It is not a source of real-world statistics.
Provenance
Where the data comes from
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators.
- Every figure, name and caption is a written placeholder. No real organisation, person or measurement appears.
A copy-ready EU AI Act training-summary paragraph and a per-asset manifest ship with every licence. Trust and provenance →
Licence
Tiers available for this dataset
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
- Negotiated
Research
- For
- Academic and non-commercial experimentation on subsets of 25,000 to 100,000 assets.
- Rights
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Limits
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
Commercial AI
- For
- Model training and commercial AI products on a sub-line or a full dataset.
- Rights
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Limits
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
Enterprise
- For
- Custom volume, exclusivity, provenance audit, indemnity and delivery terms.
- Rights
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Limits
- Negotiated
Full texts on the licensing page. Drafts pending counsel review.
Files
Loading the data
Machine-readable metadata: croissant.json. Version history: changelog.
from datasets import load_dataset
ds = load_dataset("parquet", data_files="metadata/*.parquet", split="train")
row = ds[0]
print(row["text_count"], row["path_count"], row["caption_short"])
# Full SVG files ship as tar shards; verify before extracting:
# sha256sum -c checksums/SHA256SUMSFAQ
Common questions
Is the content real?
- No. Every name, figure and caption is a written placeholder. The dataset teaches layout, chart construction and typography, not real-world statistics.
Where does the data come from?
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators. A per-asset manifest and a training-content summary paragraph ship with every licence.
Can I train a commercial model?
- Yes, under the Commercial AI or Enterprise tier. The Research tier is limited to non-commercial deployment.
Can I redistribute the files?
- No tier permits redistributing or reselling the raw data. Models trained on it are yours.
How is it delivered?
- Sharded tar archives with SHA-256 sums, a parquet metadata index and a Croissant manifest, via signed object-storage URLs or a scoped bucket for rclone.
Licence
Request pricing
Priced per subset, sub-line or full dataset. Reply within one business day.
- Assets
- 1,800,000
- On disk
- 75 GB
- Format
- SVG 1.1, UTF-8
- Version
- 1.0.0
- Updated
- 2026-09-03
- Live text
- every file in the preview set
- Vector only
- every file in the preview set
- Median file
- 76 KB
- Median text nodes
- 64
Use this dataset
pip install datasets
load_dataset("parquet",
data_files="metadata/*.parquet")Need this at enterprise scale, or with custom taxonomy?
Bulk licensing, exclusivity, private delivery and transformation.

























