Indian Residential Floor Plan Dataset
- The Doodle Desk
- Architecture
- SVG 1.1 with layer groups; plans.jsonl geometry
- v1.0.0
- Research
- Commercial AI
- Enterprise
- 500,000 assets
- updated 2026-09-04
Dimensioned, furnished apartment plans from studio to 5 BHK (bedroom, hall and kitchen, the Indian unit convention) as layered vector drawings with schedules, title blocks and per-plan geometry JSON.
- Assets
- 500,000
- On disk
- 13.5 GB sheets, 69 GB with geometry and metadata
- Lines
- 2
- Format
- SVG 1.1 with layer groups; plans.jsonl geometry
- Live text
- Yes
- Vector only
- Yes
- Sample
- 120 files
Preview
Inspect a file from the dataset

Every shape as its path geometry
STUDIO floor plan, 398 sq ft carpet area
00008182_furnished · 18872.3×14097.1 · 21 KB
- <text>
- 88
- <path>
- 131
- <g>
- 28
| Text node | size | wt |
|---|---|---|
| LIVING | 329 | 700 |
| 18'9" X 14'9" | 269.8 | 400 |
| 5714 X 4507 mm | 243.5 | 400 |
| 277 SQ.FT | 243.5 | 400 |
| TOILET | 183.7 | 700 |
| 5'1" X 6'8" | 150.6 | 400 |
| 1541 X 2024 mm | 135.9 | 400 |
| 34 SQ.FT | 135.9 | 400 |
| + 32 more | ||
Palette declared in file
- #2b2926
- #3d3b37
- #78736c
- #f6f5f3
- #33312e
- #eceae7
Showing file 1 of 12: STUDIO floor plan, 398 sq ft carpet area

fig. 0200326182_furnished.svg20252.6×22765.8 · text 158 · path 273 · 42 KB 
fig. 0301765626_furnished.svg19429.4×13190.6 · text 99 · path 162 · 25 KB 
fig. 0403057550_furnished.svg33208.5×19610.8 · text 186 · path 348 · 48 KB 
fig. 0531112348_furnished.svg19384.1×15459.7 · text 111 · path 166 · 27 KB 
fig. 0631297597_furnished.svg27040.5×17359.8 · text 152 · path 302 · 41 KB 
fig. 07LUX1002255785_furnished.svg43291.2×29368.2 · text 220 · path 616 · 80 KB 
fig. 0800520504_furnished.svg19626.4×23473.1 · text 166 · path 341 · 45 KB 
fig. 09100337954_furnished.svg33316.6×20249 · text 189 · path 455 · 56 KB 
fig. 10100846645_furnished.svg30686.6×16674 · text 169 · path 397 · 49 KB 
fig. 11101408735_furnished.svg23538.6×24324.5 · text 177 · path 325 · 50 KB 
fig. 1250021236_furnished.svg12478.1×18972 · text 102 · path 160 · 25 KB 
fig. 13LUX1002432522_furnished.svg42463.2×29596.3 · text 219 · path 582 · 74 KB
24 of 120 preview files shown · hover a tile for its structure layer · the sample pack contains 120 originals
Overview
What this dataset is
Two production lines of residential floor-plan sheets drawn to Indian practice. Every sheet is a complete drawing: title block, carpet, built-up and super built-up areas, scale and scale bar, north point, door and window schedules, an area statement per room, and every room labelled with its dimensions in two systems. Walls are hatched, doors carry swings, windows carry marks, rooms are furnished. Geometry is organised on architectural layers (A-WALL, A-DOOR, A-GLAZ, A-FURN, A-ANNO) and every plan ships with a JSON row of rooms, clear rectangles, doors, windows, toilets, balconies and wall set. No real property or person is depicted.
- Vector walls and openings
- Live room labels and dimensions
- Door and window schedules
- Layer groups
- Geometry JSON per plan
AI use cases
- Floor-plan generation
- Spatial reasoning
- CAD understanding
- Layout optimisation
- Architecture AI
- Evaluation
Specifications
Dataset specification
Format, DOM node types, interleaved text and graphics, and the content rules (no brands or trademarks, no personally identifiable information) are stated for every dataset on the specifications page.
- Assets
- 500,000 (300,000 standard + 200,000 luxury)
- Format
- SVG 1.1, real editable text, no rasters; layers A-WALL, A-DOOR, A-GLAZ, A-FURN, A-ANNO
- DOM nodes
- <text> editable strings, <path> geometry, <g> layout groups; <image> only where stated under Raster content
- Content rules
- No brands or trademarks, no personally identifiable information, placeholder figures throughout
- Sheet
- Title block, areas (carpet, built-up, super built-up), scale bar, north point, door schedule, window schedule, area statement
- Dimensions
- Two systems per sheet: feet-inches and metres; dim_format recorded per plan
- Geometry
- plans.jsonl per shard: rooms with type, label, zone, centre-line rect and clear rect (mm), area; doors with from/to room, centre, width, kind; windows; toilets with role and ventilation; balconies with host and form; wall set
- Room vocabulary
- 17 types per line: living, dining, kitchen, bedroom, master bedroom, toilet, balcony, dry balcony, passage, foyer, utility, puja, study, store, wash, servant, dressing (standard); powder and family lounge (luxury)
- Wall sets
- External/internal mm: 200/100, 230/115, 230/150, 250/125
- File size
- 16 KB smallest, 44 KB median, 69 KB largest (standard line)
- Rooms per plan
- 3 to 22, median 12
- Metadata
- index.csv with 40 columns per plan (unit, areas, rooms, toilets, footprint, facing, vastu, entry, scores, layout signature, structural and canonical keys, seed) plus a listing row
- Delivery
- 1,000 numeric shards per line, each with plans.jsonl and *_furnished.svg; root index.csv, manifest.jsonl, metadata_generic.csv
Production lines
| Line | Assets | Notes |
|---|---|---|
| Standard, studio to 4 BHK | 300,000 | 3 BHK 153,454 · 2 BHK 72,372 · 4 BHK 54,955 · 1 BHK 15,315 · studio 3,904; carpet 237 to 2,186 sq ft, median 1,122 |
| Luxury, 3 to 5 BHK | 200,000 | 3 BHK 66,666 · 4 BHK 66,667 · 5 BHK 66,667; foyer entry, attached baths and powder room on every plan; luxury score 83.3 to 92.5 |
Taxonomy coverage
- Unit types
- 6 (studio to 5 BHK)
- Footprints (standard line)
- 6 (L 209,210 · T 65,526 · rect 18,373 · stepped 4,001 · U 1,521 · notched 1,369)
- Circulation types (standard line)
- 6 (corridor spine 271,893 · open plan 17,112 · central lobby 6,975 · foyer hub 2,012 · side lobby 1,853 · short corridor 155)
- Orientation (standard line)
- W 105,384 · S 97,034 · N 78,549 · E 19,033
- Vastu aligned (standard line)
- 66,900 plans, 22.3% of that line
- Toilets per plan (standard line)
- 1: 19,219 · 2: 201,770 · 3: 69,500 · 4: 9,511
Schema
Per-asset record
Parquet, mirroring the field conventions of MMSVG-2M and Hugging Face datasets so existing loaders work unchanged. The texts[] and layout[] arrays are not offered by any public SVG dataset.
asset_id, dataset_id, dataset_version, line, file_path, sha256
svg_raw
svg_normalized fixed viewBox, transforms baked, CSS inlined, M/L/C/Q/A/Z only
png_448, width, height, orientation, page_format
text_count, path_count, group_count, image_count, element_count, token_len, complexity_tier
texts[] {content, role, font_family, font_weight, font_size, bbox, script}
layout[] {element_id, type, bbox, z_order, parent_group}
palette_id, colours[], is_dark, font_ids[], font_licences[]
sector, subject, industry, geography, tags[], headline
caption_short, caption_medium, caption_detailed
provenance {template_id, illustration_source_url, illustration_licence, generator, generated_at}
quality {valid_svg, renders, has_live_text, has_raster, near_dup_group, phash}Quality
Checks and results
Badges are published now. A composite SVGZO Quality Score follows once the formula is frozen and applied to every line. Method →
- Duplicate ids
- 0 in 300,000; 0 in 200,000
- Distinct drawings
- No two plans in either line are the same drawing
- Quality floor
- Standard: every plan scores 90.0 or above out of 100 on the generator's gate. Luxury: gate passed at 100% on every mandatory requirement.
- Regeneration
- Every plan is a pure function of its seed and can be rebuilt from its manifest row
- Index date
- 3 September 2026
Read before licensing
- Plans are drawn to Indian residential practice and do not describe any real property.
Provenance
Where the data comes from
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators.
- Every figure, name and caption is a written placeholder. No real organisation, person or measurement appears.
A copy-ready EU AI Act training-summary paragraph and a per-asset manifest ship with every licence. Trust and provenance →
Licence
Tiers available for this dataset
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
- Negotiated
Research
- For
- Academic and non-commercial experimentation on subsets of 25,000 to 100,000 assets.
- Rights
- Train, fine-tune and evaluate models
- Licensee owns models and outputs
- Perpetual, worldwide
- Limits
- Non-commercial deployment only
- No redistribution of raw data
- Attribution required
- No font-generation models
Commercial AI
- For
- Model training and commercial AI products on a sub-line or a full dataset.
- Rights
- Train, fine-tune and evaluate, including commercial models and products
- Licensee owns models, weights, embeddings and outputs
- Share with contractors under NDA
- Safe harbour for incidental memorisation
- Chain-of-title warranty, liability capped at fees
- Limits
- No redistribution or resale of raw data
- No reconstructable copy of the dataset in a model
- No font-generation models
- No biometric or real-person inference
Enterprise
- For
- Custom volume, exclusivity, provenance audit, indemnity and delivery terms.
- Rights
- Everything in Commercial AI
- Affiliates and named contractors
- Optional exclusivity on custom or carved-out sets
- IP indemnity, cap at 1 to 2x fees
- Audit access and change notices
- Limits
- Negotiated
Full texts on the licensing page. Drafts pending counsel review.
Files
Loading the data
Machine-readable metadata: croissant.json. Version history: changelog.
from datasets import load_dataset
ds = load_dataset("parquet", data_files="metadata/*.parquet", split="train")
row = ds[0]
print(row["text_count"], row["path_count"], row["caption_short"])
# Full SVG files ship as tar shards; verify before extracting:
# sha256sum -c checksums/SHA256SUMSFAQ
Common questions
Is the content real?
- No. Every name, figure and caption is a written placeholder. The dataset teaches layout, chart construction and typography, not real-world statistics.
Where does the data come from?
- Assets are composed, created and processed into SVG from source files by The Doodle Desk's team and network of creators. A per-asset manifest and a training-content summary paragraph ship with every licence.
Can I train a commercial model?
- Yes, under the Commercial AI or Enterprise tier. The Research tier is limited to non-commercial deployment.
Can I redistribute the files?
- No tier permits redistributing or reselling the raw data. Models trained on it are yours.
How is it delivered?
- Sharded tar archives with SHA-256 sums, a parquet metadata index and a Croissant manifest, via signed object-storage URLs or a scoped bucket for rclone.
Licence
Request pricing
Priced per subset, sub-line or full dataset. Reply within one business day.
- Assets
- 500,000
- On disk
- 13.5 GB sheets, 69 GB with geometry and metadata
- Format
- SVG 1.1 with layer groups; plans.jsonl geometry
- Version
- 1.0.0
- Updated
- 2026-09-04
- Live text
- every file in the preview set
- Vector only
- every file in the preview set
- Median file
- 45 KB
- Median text nodes
- 164
Use this dataset
pip install datasets
load_dataset("parquet",
data_files="metadata/*.parquet")Need this at enterprise scale, or with custom taxonomy?
Bulk licensing, exclusivity, private delivery and transformation.

























