Skip to content
Menu

Presentation Slide Dataset

  • The Doodle Desk
  • Presentations
  • SVG 1.1
  • v1.0.0
  • Research
  • Commercial AI
  • Enterprise
  • also in Infographics
  • 500,000 assets
  • updated 2026-09-03

Half a million 16:9 slides built from PowerPoint template families, every chart and headline live, each page leading with its own conclusion.

Assets
500,000
On disk
5.9 GB
Lines
1
Format
SVG 1.1
Live text
Yes
Vector only
Yes
Sample
120 files

Preview

Inspect a file from the dataset

Slide: Claims notified fell short of target, structure layer

Every shape as its path geometry

1032000_insurer_Ottershaw-Group_2017

1032000_insurer_Ottershaw-Group_2017 · 1962×1104 · 32 KB

<text>
22
<path>
37
<g>
14
Live text nodes
Text nodesizewt
THE CASE IN SHORT768700
Claims notified fell short of target1579700
The case for Ottershaw Group: · 2017, in claims730600
A proposal read against the record: the same measures on both sides…787400
17,1103456700
claims864700
What made up the total1104700
3 groups. Together they are the whole claims count.730400
+ 14 more

Palette declared in file

  • #323434
  • #ffffff
  • #6f9d04
  • #4e59a7
  • #597e03
  • #4a6a02

Showing file 1 of 12: 1032000_insurer_Ottershaw-Group_2017

fig. 011032000_insurer_Ottershaw-Group_2017.svg · 1962×1104 · 22 text · 37 path · 32 KB · 1 of 12 · rendered with resvg; raw SVG is not served from this origin
  • Slide: Week 4 carried 38% of the year
    fig. 021032001_airport_Ottershaw-Group_2017.svg1962×1104 · text 33 · path 20 · 28 KB
  • Slide: The year concentrates in May
    fig. 031032023_support_Ottershaw-Group_2018.svg1962×1104 · text 50 · path 49 · 46 KB
  • Slide: H1 carried 35% of the year
    fig. 041032005_charity_Ottershaw-Group_2017.svg1962×1104 · text 38 · path 35 · 29 KB
  • Slide: A walkthrough, 2019:
    fig. 051032084_lending_Ottershaw-Group_2019.svg1962×1104 · text 46 · path 51 · 48 KB
  • Slide: Independents carries 31% of the copies
    fig. 061032007_publishing_Ottershaw-Co-operative_2017.svg1962×1104 · text 37 · path 40 · 45 KB
  • Slide: Only 24% reach settled
    fig. 071032120_insurer_Ottershaw-Group_2020.svg1962×1104 · text 27 · path 45 · 27 KB
  • Slide: The tonnes divide into 5 groups
    fig. 081032003_farm_Ottershaw-Trust_2017.svg1962×1104 · text 58 · path 94 · 46 KB
  • Slide: Usable keeps 20% of what entered
    fig. 091032012_survey_Ottershaw_2017.svg1962×1104 · text 50 · path 50 · 35 KB
  • Slide: Short video outweighs long form
    fig. 101032060_social_Ottershaw_2018.svg1962×1104 · text 28 · path 95 · 62 KB
  • Slide: Renewed keeps 20% of what entered
    fig. 111032066_saas_Ottershaw-Trust_2019.svg1962×1104 · text 55 · path 97 · 59 KB
  • Slide: Viewings and the decision:
    fig. 121032338_property_Ottershaw-Partners_2025.svg1962×1104 · text 32 · path 46 · 39 KB
  • Slide: Containers, as found:
    fig. 131032002_port_Ottershaw-Trust_2017.svg1962×1104 · text 33 · path 23 · 21 KB

24 of 120 preview files shown · the sample pack contains 120 originals

Overview

What this dataset is

The slide-deck production line, offered on its own for teams training slide-generation, chart-to-slide and presentation-layout models. Every page is a single 16:9 slide composed from a named deck template family (roadmap, timeline, iceberg diagram, funnel, comparison and others), a grid, a palette and up to several chart modules in slot order. Headlines state a conclusion drawn from the figures on the slide, so text and chart agree. All copy is live SVG text. Figures and organisation names are written placeholders. These 500,000 assets are also counted inside the Infographic Dataset; licence one or the other, not both.

  • Live text
  • Chart families in slot order
  • Deck template family
  • Grid and palette IDs
  • Content digest per page

AI use cases

  • Slide generation
  • Layout generation
  • Chart understanding
  • Text-to-SVG
  • Design AI
  • Evaluation

Specifications

Dataset specification

Format, DOM node types, interleaved text and graphics, and the content rules (no brands or trademarks, no personally identifiable information) are stated for every dataset on the specifications page.

Assets
500,000
Format
SVG, UTF-8, one slide per file
DOM nodes
<text> editable strings, <path> geometry, <g> layout groups; <image> only where stated under Raster content
Content rules
No brands or trademarks, no personally identifiable information, placeholder figures throughout
Aspect
16:9. Declared page size 1962×1104 px; the viewBox uses a 67716×38090 internal coordinate space, so geometry scales without loss
Text
Live SVG text throughout, no outlined type
Raster content
None. Files carrying an embedded bitmap are excluded from delivery, samples and previews; every delivered file is vector only
Distinct stories
156,169
Chart families
131; busiest family on 8.1% of pages
Readings
share, rank, trend, funnel, build, process, history, comparison, tiers, board, geography, demography, target, forecast, cohort
Metadata
16 fields: serial, file, story key, variant, deck template, grid, chart families in slot order, illustration, header treatment, scheme, typeface, hero, frame, size, content digest, fault
Median file size
39 KB
Delivery
10 tar.zst archives with one manifest each; every archive decompressed and page-counted at 50,000

Production lines

Production lines in this dataset
LineAssetsNotes
Slide-deck500,00010 batches of 50,000; 16:9 at 1962×1104 px; 156,169 distinct stories

Taxonomy coverage

Deck template families
306
Chart families
131
Grids
33
Palettes
306
Illustrations
2,155
Typefaces
120 families

Schema

Per-asset record

Parquet, mirroring the field conventions of MMSVG-2M and Hugging Face datasets so existing loaders work unchanged. The texts[] and layout[] arrays are not offered by any public SVG dataset.

asset_id, dataset_id, dataset_version, line, file_path, sha256
svg_raw
svg_normalized       fixed viewBox, transforms baked, CSS inlined, M/L/C/Q/A/Z only
png_448, width, height, orientation, page_format
text_count, path_count, group_count, image_count, element_count, token_len, complexity_tier
texts[]              {content, role, font_family, font_weight, font_size, bbox, script}
layout[]             {element_id, type, bbox, z_order, parent_group}
palette_id, colours[], is_dark, font_ids[], font_licences[]
sector, subject, industry, geography, tags[], headline
caption_short, caption_medium, caption_detailed
provenance           {template_id, illustration_source_url, illustration_licence, generator, generated_at}
quality              {valid_svg, renders, has_live_text, has_raster, near_dup_group, phash}

Quality

Checks and results

Badges are published now. A composite SVGZO Quality Score follows once the formula is frozen and applied to every line. Method →

Byte-identical pages
0 in the 850,000-page engine audit
Repeated headline and subject pairs
0
Pages with a figure their story does not hold
0; every printed numeral is checked against its story
Busiest illustration
0.16% of pages
Regeneration
Any page can be rebuilt from its manifest row
Audit date
3 September 2026

Read before licensing

  • Cross-listed: these 500,000 slides are the slide-deck line of the Infographic Dataset. A Commercial AI licence for the Infographic Dataset already covers them.
  • One slide per file. Multi-slide deck sequences are not modelled in this release.

Provenance

Where the data comes from

  • Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators.
  • Every figure, name and caption is a written placeholder. No real organisation, person or measurement appears.

A copy-ready EU AI Act training-summary paragraph and a per-asset manifest ship with every licence. Trust and provenance →

Licence

Tiers available for this dataset

Research

For
Academic and non-commercial experimentation on subsets of 25,000 to 100,000 assets.
Rights
  • Train, fine-tune and evaluate models
  • Licensee owns models and outputs
  • Perpetual, worldwide
Limits
  • Non-commercial deployment only
  • No redistribution of raw data
  • Attribution required
  • No font-generation models

Commercial AI

For
Model training and commercial AI products on a sub-line or a full dataset.
Rights
  • Train, fine-tune and evaluate, including commercial models and products
  • Licensee owns models, weights, embeddings and outputs
  • Share with contractors under NDA
  • Safe harbour for incidental memorisation
  • Chain-of-title warranty, liability capped at fees
Limits
  • No redistribution or resale of raw data
  • No reconstructable copy of the dataset in a model
  • No font-generation models
  • No biometric or real-person inference

Enterprise

For
Custom volume, exclusivity, provenance audit, indemnity and delivery terms.
Rights
  • Everything in Commercial AI
  • Affiliates and named contractors
  • Optional exclusivity on custom or carved-out sets
  • IP indemnity, cap at 1 to 2x fees
  • Audit access and change notices
Limits
  • Negotiated

Full texts on the licensing page. Drafts pending counsel review.

Files

Loading the data

Machine-readable metadata: croissant.json. Version history: changelog.

from datasets import load_dataset

ds = load_dataset("parquet", data_files="metadata/*.parquet", split="train")
row = ds[0]
print(row["text_count"], row["path_count"], row["caption_short"])

# Full SVG files ship as tar shards; verify before extracting:
#   sha256sum -c checksums/SHA256SUMS

FAQ

Common questions

Is the content real?

No. Every name, figure and caption is a written placeholder. The dataset teaches layout, chart construction and typography, not real-world statistics.

Where does the data come from?

Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators. A per-asset manifest and a training-content summary paragraph ship with every licence.

Can I train a commercial model?

Yes, under the Commercial AI or Enterprise tier. The Research tier is limited to non-commercial deployment.

Can I redistribute the files?

No tier permits redistributing or reselling the raw data. Models trained on it are yours.

How is it delivered?

Sharded tar archives with SHA-256 sums, a parquet metadata index and a Croissant manifest, via signed object-storage URLs or a scoped bucket for rclone.

Licence

Request pricing

Priced per subset, sub-line or full dataset. Reply within one business day.

Assets
500,000
On disk
5.9 GB
Format
SVG 1.1, UTF-8
Version
1.0.0
Updated
2026-09-03
Live text
every file in the preview set
Vector only
every file in the preview set
Median file
39 KB
Median text nodes
37

Use this dataset

pip install datasets
load_dataset("parquet",
  data_files="metadata/*.parquet")

croissant.json · changelog

Need this at enterprise scale, or with custom taxonomy?

Bulk licensing, exclusivity, private delivery and transformation.

Talk to the data team