Poster and Print Dataset
- The Doodle Desk
- Posters and print
- SVG 1.1
- v1.1.0
- Research
- Commercial AI
- Enterprise
- 991,585 assets
- updated 2026-09-07
Print-ready posters, flyers, brochures, ads, menus and rack cards at known trim sizes, every line of copy live.
- Assets
- 991,585
- On disk
- 51 GB
- Lines
- 3
- Format
- SVG 1.1
- Live text
- Yes
- Vector only
- Yes
- Sample
- 79 files
Preview
Inspect a file from the dataset

Every shape as its path geometry
Vector image of manss hand raised up
123471_vector-image-of-manss-hand-raised-up_a5 · 559×794 · 44 KB
- <text>
- 4
- <path>
- 2
- <g>
- 3
| Text node | size | wt |
|---|---|---|
| MANSS HAND RAISED | 47.5 | 800 |
| UP | 47.5 | 800 |
| A raised rendered as a solid silhouette, in a single tone. | 20.1 | 400 |
| A portrait composition with room around the subject. | 20.1 | 400 |
Palette declared in file
- #ebeff0
- #084a5b
- #9dbcc5
- #f9d95a
Showing file 1 of 12: Vector image of manss hand raised up

fig. 02OV100391_ad_coupon.svg1275×1650 · text 14 · path 26 · 104 KB 
fig. 03OV103090_ad_landscape_split.svg2480×1754 · text 11 · path 43 · 82 KB 
fig. 04OV100679_flyer_checker.svg874×1240 · text 10 · path 65 · 86 KB 
fig. 05OV348381_flyer_duotone.svg825×1275 · text 11 · path 5 · 76 KB 
fig. 06OV114146_poster_square_dusk.svg1800×1800 · text 4 · path 2 · 53 KB 
fig. 07OV116376_poster_big_type.svg1754×2480 · text 9 · path 11 · 47 KB 
fig. 08OV114388_ad_coupon.svg1275×1650 · text 14 · path 17 · 99 KB 
fig. 09OV149374_dl_ticket.svg583×1240 · text 10 · path 6 · 114 KB 
fig. 10OV559804_poster_wide.svg2480×1754 · text 13 · path 16 · 119 KB 
fig. 11OV125015_ad_gradient_hero.svg1275×1650 · text 11 · path 56 · 75 KB 
fig. 12OV125707_ad_gradient_hero.svg1275×1650 · text 12 · path 1 · 87 KB 
fig. 13OV452968_dl_ticket.svg583×1240 · text 11 · path 414 · 151 KB
24 of 79 preview files shown · hover a tile for its structure layer · the sample pack contains 79 originals
Overview
What this dataset is
Four lines of single-page print products. Each page is one finished design composed from an illustration, a palette and a layout family, with the headline, body copy, call to action and contact line left as live SVG text at declared sizes. Colour comes from a palette held on the page, not baked into the artwork. Nothing is outlined, and no file references anything outside itself.
- Live text
- Vector illustration
- Palette on page
- Dublin Core metadata in file
- SHA-256 verified archives
AI use cases
- Text-to-SVG
- Design generation
- Layout generation
- Typography understanding
- Design AI
- Evaluation
Specifications
Dataset specification
Format, DOM node types, interleaved text and graphics, and the content rules (no brands or trademarks, no personally identifiable information) are stated for every dataset on the specifications page.
- Assets
- 991,585
- Format
- SVG with embedded Dublin Core RDF: title, creator, tags, category, niche
- DOM nodes
- <text> editable strings, <path> geometry, <g> layout groups; <image> only where stated under Raster content
- Content rules
- No brands or trademarks, no personally identifiable information, placeholder figures throughout
- Text
- Live SVG text at declared sizes; 40 of 40 sampled files carried live text
- Raster content
- None. Files carrying an embedded bitmap are excluded from delivery, samples and previews; every delivered file is vector only
- Products
- Posters 175,528 · Flyers 167,472 · Brochures 58,162 · Ads 57,264 · Social 23,150 · Menus 10,186 · Rack cards 6,897 (illustration line)
- Print formats
- 23, from DL rack card to A2 poster
- Typefaces
- Embedded or referenced per file; notices ship with every delivery
- Delivery
- 42 tar.gz archives, all SHA-256 verified and reconciled
Production lines
| Line | Assets | Notes |
|---|---|---|
| Illustration posters and print | 498,659 | 205 business niches, 23 print formats, 56 typographic themes |
| Brand posters | 286,613 | 305 layout families over 111 palettes and 10 formats |
| Poster and full-page ad | 206,313 | 282 layout families, 408 palettes, 92 verticals, 15 print sizes |
Taxonomy coverage
- Business verticals
- 205
- Layout families
- 587
- Colour palettes
- 519
- Print formats
- 23
- Typographic themes
- 56
- Posters, strict
- 433,054
Schema
Per-asset record
Parquet, mirroring the field conventions of MMSVG-2M and Hugging Face datasets so existing loaders work unchanged. The texts[] and layout[] arrays are not offered by any public SVG dataset.
asset_id, dataset_id, dataset_version, line, file_path, sha256
svg_raw
svg_normalized fixed viewBox, transforms baked, CSS inlined, M/L/C/Q/A/Z only
png_448, width, height, orientation, page_format
text_count, path_count, group_count, image_count, element_count, token_len, complexity_tier
texts[] {content, role, font_family, font_weight, font_size, bbox, script}
layout[] {element_id, type, bbox, z_order, parent_group}
palette_id, colours[], is_dark, font_ids[], font_licences[]
sector, subject, industry, geography, tags[], headline
caption_short, caption_medium, caption_detailed
provenance {template_id, illustration_source_url, illustration_licence, generator, generated_at}
quality {valid_svg, renders, has_live_text, has_raster, near_dup_group, phash}Quality
Checks and results
Badges are published now. A composite SVGZO Quality Score follows once the formula is frozen and applied to every line. Method →
- Checksum coverage
- 100% of archives, SHA-256
- Live text sample
- 40 of 40 files, averaging 25 vector path elements each
- External references
- None; every file is self-contained
Read before licensing
- Trim sizes are declared in the viewBox; bleed is not included.
Provenance
Where the data comes from
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators.
- Every figure, name and caption is a written placeholder. No real organisation, person or measurement appears.
A copy-ready EU AI Act training-summary paragraph and a per-asset manifest ship with every licence. Trust and provenance →
Licence
Tiers available for this dataset
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
- Negotiated
Research
- For
- Academic and non-commercial experimentation on subsets of 25,000 to 100,000 assets.
- Rights
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Limits
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
Commercial AI
- For
- Model training and commercial AI products on a sub-line or a full dataset.
- Rights
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Limits
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
Enterprise
- For
- Custom volume, exclusivity, provenance audit, indemnity and delivery terms.
- Rights
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Limits
- Negotiated
Full texts on the licensing page. Drafts pending counsel review.
Files
Loading the data
Machine-readable metadata: croissant.json. Version history: changelog.
from datasets import load_dataset
ds = load_dataset("parquet", data_files="metadata/*.parquet", split="train")
row = ds[0]
print(row["text_count"], row["path_count"], row["caption_short"])
# Full SVG files ship as tar shards; verify before extracting:
# sha256sum -c checksums/SHA256SUMSFAQ
Common questions
Is the content real?
- No. Every name, figure and caption is a written placeholder. The dataset teaches layout, chart construction and typography, not real-world statistics.
Where does the data come from?
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators. A per-asset manifest and a training-content summary paragraph ship with every licence.
Can I train a commercial model?
- Yes, under the Commercial AI or Enterprise tier. The Research tier is limited to non-commercial deployment.
Can I redistribute the files?
- No tier permits redistributing or reselling the raw data. Models trained on it are yours.
How is it delivered?
- Sharded tar archives with SHA-256 sums, a parquet metadata index and a Croissant manifest, via signed object-storage URLs or a scoped bucket for rclone.
Licence
Request pricing
Priced per subset, sub-line or full dataset. Reply within one business day.
- Assets
- 991,585
- On disk
- 51 GB
- Format
- SVG 1.1, embedded Dublin Core RDF
- Version
- 1.1.0
- Updated
- 2026-09-07
- Live text
- every file in the preview set
- Vector only
- every file in the preview set
- Median file
- 81 KB
- Median text nodes
- 10
Use this dataset
pip install datasets
load_dataset("parquet",
data_files="metadata/*.parquet")Need this at enterprise scale, or with custom taxonomy?
Bulk licensing, exclusivity, private delivery and transformation.

























