Presentation Slide Dataset
- The Doodle Desk
- Presentations
- SVG 1.1
- v1.0.0
- Research
- Commercial AI
- Enterprise
- also in Infographics
- 500,000 assets
- updated 2026-09-03
Half a million 16:9 slides built from PowerPoint template families, every chart and headline live, each page leading with its own conclusion.
- Assets
- 500,000
- On disk
- 5.9 GB
- Lines
- 1
- Format
- SVG 1.1
- Live text
- Yes
- Vector only
- Yes
- Sample
- 120 files
Preview
Inspect a file from the dataset

Every shape as its path geometry
1032000_insurer_Ottershaw-Group_2017
1032000_insurer_Ottershaw-Group_2017 · 1962×1104 · 32 KB
- <text>
- 22
- <path>
- 37
- <g>
- 14
| Text node | size | wt |
|---|---|---|
| THE CASE IN SHORT | 768 | 700 |
| Claims notified fell short of target | 1579 | 700 |
| The case for Ottershaw Group: · 2017, in claims | 730 | 600 |
| A proposal read against the record: the same measures on both sides… | 787 | 400 |
| 17,110 | 3456 | 700 |
| claims | 864 | 700 |
| What made up the total | 1104 | 700 |
| 3 groups. Together they are the whole claims count. | 730 | 400 |
| + 14 more | ||
Palette declared in file
- #323434
- #ffffff
- #6f9d04
- #4e59a7
- #597e03
- #4a6a02
Showing file 1 of 12: 1032000_insurer_Ottershaw-Group_2017

fig. 021032001_airport_Ottershaw-Group_2017.svg1962×1104 · text 33 · path 20 · 28 KB 
fig. 031032023_support_Ottershaw-Group_2018.svg1962×1104 · text 50 · path 49 · 46 KB 
fig. 041032005_charity_Ottershaw-Group_2017.svg1962×1104 · text 38 · path 35 · 29 KB 
fig. 051032084_lending_Ottershaw-Group_2019.svg1962×1104 · text 46 · path 51 · 48 KB 
fig. 061032007_publishing_Ottershaw-Co-operative_2017.svg1962×1104 · text 37 · path 40 · 45 KB 
fig. 071032120_insurer_Ottershaw-Group_2020.svg1962×1104 · text 27 · path 45 · 27 KB 
fig. 081032003_farm_Ottershaw-Trust_2017.svg1962×1104 · text 58 · path 94 · 46 KB 
fig. 091032012_survey_Ottershaw_2017.svg1962×1104 · text 50 · path 50 · 35 KB 
fig. 101032060_social_Ottershaw_2018.svg1962×1104 · text 28 · path 95 · 62 KB 
fig. 111032066_saas_Ottershaw-Trust_2019.svg1962×1104 · text 55 · path 97 · 59 KB 
fig. 121032338_property_Ottershaw-Partners_2025.svg1962×1104 · text 32 · path 46 · 39 KB 
fig. 131032002_port_Ottershaw-Trust_2017.svg1962×1104 · text 33 · path 23 · 21 KB
24 of 120 preview files shown · hover a tile for its structure layer · the sample pack contains 120 originals
Overview
What this dataset is
The slide-deck production line, offered on its own for teams training slide-generation, chart-to-slide and presentation-layout models. Every page is a single 16:9 slide composed from a named deck template family (roadmap, timeline, iceberg diagram, funnel, comparison and others), a grid, a palette and up to several chart modules in slot order. Headlines state a conclusion drawn from the figures on the slide, so text and chart agree. All copy is live SVG text. Figures and organisation names are written placeholders. These 500,000 assets are also counted inside the Infographic Dataset; licence one or the other, not both.
- Live text
- Chart families in slot order
- Deck template family
- Grid and palette IDs
- Content digest per page
AI use cases
- Slide generation
- Layout generation
- Chart understanding
- Text-to-SVG
- Design AI
- Evaluation
Specifications
Dataset specification
Format, DOM node types, interleaved text and graphics, and the content rules (no brands or trademarks, no personally identifiable information) are stated for every dataset on the specifications page.
- Assets
- 500,000
- Format
- SVG, UTF-8, one slide per file
- DOM nodes
- <text> editable strings, <path> geometry, <g> layout groups; <image> only where stated under Raster content
- Content rules
- No brands or trademarks, no personally identifiable information, placeholder figures throughout
- Aspect
- 16:9. Declared page size 1962×1104 px; the viewBox uses a 67716×38090 internal coordinate space, so geometry scales without loss
- Text
- Live SVG text throughout, no outlined type
- Raster content
- None. Files carrying an embedded bitmap are excluded from delivery, samples and previews; every delivered file is vector only
- Distinct stories
- 156,169
- Chart families
- 131; busiest family on 8.1% of pages
- Readings
- share, rank, trend, funnel, build, process, history, comparison, tiers, board, geography, demography, target, forecast, cohort
- Metadata
- 16 fields: serial, file, story key, variant, deck template, grid, chart families in slot order, illustration, header treatment, scheme, typeface, hero, frame, size, content digest, fault
- Median file size
- 39 KB
- Delivery
- 10 tar.zst archives with one manifest each; every archive decompressed and page-counted at 50,000
Production lines
| Line | Assets | Notes |
|---|---|---|
| Slide-deck | 500,000 | 10 batches of 50,000; 16:9 at 1962×1104 px; 156,169 distinct stories |
Taxonomy coverage
- Deck template families
- 306
- Chart families
- 131
- Grids
- 33
- Palettes
- 306
- Illustrations
- 2,155
- Typefaces
- 120 families
Schema
Per-asset record
Parquet, mirroring the field conventions of MMSVG-2M and Hugging Face datasets so existing loaders work unchanged. The texts[] and layout[] arrays are not offered by any public SVG dataset.
asset_id, dataset_id, dataset_version, line, file_path, sha256
svg_raw
svg_normalized fixed viewBox, transforms baked, CSS inlined, M/L/C/Q/A/Z only
png_448, width, height, orientation, page_format
text_count, path_count, group_count, image_count, element_count, token_len, complexity_tier
texts[] {content, role, font_family, font_weight, font_size, bbox, script}
layout[] {element_id, type, bbox, z_order, parent_group}
palette_id, colours[], is_dark, font_ids[], font_licences[]
sector, subject, industry, geography, tags[], headline
caption_short, caption_medium, caption_detailed
provenance {template_id, illustration_source_url, illustration_licence, generator, generated_at}
quality {valid_svg, renders, has_live_text, has_raster, near_dup_group, phash}Quality
Checks and results
Badges are published now. A composite SVGZO Quality Score follows once the formula is frozen and applied to every line. Method →
- Byte-identical pages
- 0 in the 850,000-page engine audit
- Repeated headline and subject pairs
- 0
- Pages with a figure their story does not hold
- 0; every printed numeral is checked against its story
- Busiest illustration
- 0.16% of pages
- Regeneration
- Any page can be rebuilt from its manifest row
- Audit date
- 3 September 2026
Read before licensing
- Cross-listed: these 500,000 slides are the slide-deck line of the Infographic Dataset. A Commercial AI licence for the Infographic Dataset already covers them.
- One slide per file. Multi-slide deck sequences are not modelled in this release.
Provenance
Where the data comes from
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators.
- Every figure, name and caption is a written placeholder. No real organisation, person or measurement appears.
A copy-ready EU AI Act training-summary paragraph and a per-asset manifest ship with every licence. Trust and provenance →
Licence
Tiers available for this dataset
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
- Negotiated
Research
- For
- Academic and non-commercial experimentation on subsets of 25,000 to 100,000 assets.
- Rights
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Limits
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
Commercial AI
- For
- Model training and commercial AI products on a sub-line or a full dataset.
- Rights
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Limits
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
Enterprise
- For
- Custom volume, exclusivity, provenance audit, indemnity and delivery terms.
- Rights
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Limits
- Negotiated
Full texts on the licensing page. Drafts pending counsel review.
Files
Loading the data
Machine-readable metadata: croissant.json. Version history: changelog.
from datasets import load_dataset
ds = load_dataset("parquet", data_files="metadata/*.parquet", split="train")
row = ds[0]
print(row["text_count"], row["path_count"], row["caption_short"])
# Full SVG files ship as tar shards; verify before extracting:
# sha256sum -c checksums/SHA256SUMSFAQ
Common questions
Is the content real?
- No. Every name, figure and caption is a written placeholder. The dataset teaches layout, chart construction and typography, not real-world statistics.
Where does the data come from?
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators. A per-asset manifest and a training-content summary paragraph ship with every licence.
Can I train a commercial model?
- Yes, under the Commercial AI or Enterprise tier. The Research tier is limited to non-commercial deployment.
Can I redistribute the files?
- No tier permits redistributing or reselling the raw data. Models trained on it are yours.
How is it delivered?
- Sharded tar archives with SHA-256 sums, a parquet metadata index and a Croissant manifest, via signed object-storage URLs or a scoped bucket for rclone.
Licence
Request pricing
Priced per subset, sub-line or full dataset. Reply within one business day.
- Assets
- 500,000
- On disk
- 5.9 GB
- Format
- SVG 1.1, UTF-8
- Version
- 1.0.0
- Updated
- 2026-09-03
- Live text
- every file in the preview set
- Vector only
- every file in the preview set
- Median file
- 39 KB
- Median text nodes
- 37
Use this dataset
pip install datasets
load_dataset("parquet",
data_files="metadata/*.parquet")Need this at enterprise scale, or with custom taxonomy?
Bulk licensing, exclusivity, private delivery and transformation.

























