A 35-person Kent groundworks contractor showed us their building regulations process in detail. Each project required assembling a submission pack from 12–18 separate documents: floor plans exported from AutoCAD, structural calculations delivered as the engineer's PDF, drainage details, fire safety schedules, SAP energy calculations, and party wall agreements. A project manager was spending 2.5–3 person-days per project pulling these together, checking version consistency, formatting everything to BCO requirements, and uploading to the relevant building control portal. An extraction and assembly pipeline cut that to four hours — most of which is still a human reviewing the output before submission.
What building regulations submissions actually contain: the document set UK contractors assemble
Building regulations in England and Wales are governed by the Building Regulations 2010 (SI 2010/2214) and, for higher-risk buildings, the Building Safety Act 2022. The exact documents required depend on the notifiable work type, but for a typical residential new build or commercial groundworks project, a full submission pack includes:
- Location and site plan — block plan at 1:1250 minimum, site plan at 1:500
- Architectural floor plans and elevations — from the architect's AutoCAD export or Revit model
- Structural calculations and engineer's drawings — usually a separate PDF from the structural engineer
- Drainage and foul water design — sometimes embedded in architectural drawings, sometimes a separate document
- SAP or SBEM energy calculations — Part L compliance
- Fire safety strategy — required for buildings over two storeys or with sleeping accommodation
- Ventilation strategy — Part F
- Electrical installation certificate — Part P, for notifiable electrical work
For the Kent contractor, the average project touched eight of these document categories, spread across four or five separate suppliers sending files in different formats on different timelines. That fragmentation is where the time goes.
Where the manual time goes: extraction, reformatting, and version-control failures
We timed the manual process across five projects. The breakdown was consistent:
| Task | Time per project |
|---|---|
| Collecting documents from email, shared drives, and supplier portals | 3–4 hours |
| Checking drawing revision marks and cross-referencing specification sheets | 2–3 hours |
| Reformatting to a consistent file-naming convention | 1–2 hours |
| Filling in BCO cover sheets and form A/B where required | 2 hours |
| Version-control checks (confirming latest revision vs. what was submitted previously) | 2–3 hours |
| Final assembly and upload | 1 hour |
Total: 11–15 hours per project, or 2.5–3 person-days at typical project management throughput.
The most expensive failure mode was not the time — it was the version-control problem. In the 18 months before the pipeline, the contractor submitted a superseded structural drawing on three separate projects. Two of those led to BCO requests for clarification, adding 1–2 weeks to approval. One required a full resubmission. The version-check step should have caught these but was being done manually by comparing file timestamps and title-block revision codes.
OCR and vision-model extraction from structural drawings and specification PDFs
The extraction layer handles two distinct document types differently.
Clean PDFs — text-based structural calculations, SAP reports, specification documents — go through a standard text extraction path. We use pdfplumber for layout-preserving extraction, which handles multi-column layouts and embedded tables better than PyPDF2 for this type of document. A 30-page structural calculation PDF extracts and identifies fields in under four seconds.
Scanned drawings and image-based PDFs — AutoCAD exports, annotated drawings, sketches — need a vision model. We pass these to the vision API with a structured extraction prompt targeting title-block fields: drawing number, revision, date, project reference, and engineer stamp. For complex drawings with handwritten revision notes, the vision model outperforms Tesseract OCR by a meaningful margin on field accuracy.
The model-selection logic in the pipeline config looks like this:
{
"extraction_routes": {
"text_pdf": {
"method": "pdfplumber",
"fallback": "vision_model",
"confidence_threshold": 0.85
},
"scanned_drawing": {
"method": "vision_model",
"model": "claude-opus-4-5",
"prompt_template": "drawing_titleblock_v3",
"extract_fields": [
"drawing_no", "revision", "date",
"project_ref", "scale", "engineer_name"
]
},
"energy_calculation": {
"method": "pdfplumber",
"parser": "sap_worksheet_parser",
"target_fields": [
"dwelling_emission_rate",
"target_emission_rate",
"sap_rating"
]
}
}
}
The confidence_threshold on the text path matters. If pdfplumber returns low-confidence field values — common on scanned-and-reprinted PDFs — the document reroutes automatically to the vision model. This is the same human-in-the-loop principle described in our invoice OCR pipeline case study, where field-level confidence scoring determines whether a document goes straight through or queues for review.
For document classification before the extraction step — determining whether a file is a structural drawing, an SAP report, or a fire safety schedule — see our post on document classification with vision models, which covers the model selection trade-offs in more depth.
Document assembly pipeline: combining extracted data into compliant submission packs
Once extraction is complete, the assembly step builds the submission pack from a manifest and templates.
{
"project_ref": "KGW-2026-0047",
"submission_type": "full_plans",
"authority": "Dover District Council",
"documents": [
{
"type": "site_plan",
"source_file": "s3://proj-docs/KGW-2026-0047/site-plan-rev-C.pdf",
"revision": "C",
"hash": "a3f2c9d1e8b4...",
"included": true
},
{
"type": "structural_calculations",
"source_file": "s3://proj-docs/KGW-2026-0047/struct-calcs-v2.pdf",
"revision": "2",
"hash": "7bd1e4f22a91...",
"included": true
}
],
"cover_sheet_data": {
"applicant_name": "Thornfield Developments Ltd",
"site_address": "14 Ashford Road, Folkestone, CT20 1BX",
"work_description": "Erection of single-storey rear extension with structural alterations",
"estimated_cost": 85000
}
}
The assembly step merges extracted data into BCO cover sheet templates (one per local authority — they vary more than you would expect), applies the correct file-naming convention, and generates a compiled PDF. For professional services firms, we use a similar manifest-driven approach in client onboarding document pack automation, where the same principle of loosely coupling source documents from output format makes version management manageable.
Handling drawing revisions: version control and automatic diff detection
Drawing revisions are the specific failure point that breaks manual processes. A structural engineer issues revision B two days after revision A. The project manager updates one reference but misses another. The BCO receives an inconsistent set.
The pipeline handles this with hash-based change detection. Every source document has a stored SHA-256 hash in the manifest. When a new file lands in the watched S3 folder — or via webhook from the project's file-sharing tool (the Kent contractor uses Fieldwire) — the pipeline:
- Computes the new file's hash
- Compares against the stored hash for that document type and project reference
- If different, runs a page-by-page comparison to identify changed pages
- Passes changed pages to the vision model to summarise the changes in plain English
- Updates the manifest with the new revision and hash
- Flags only the submission pack sections that reference the changed document for regeneration
The project manager gets a Slack notification: "Structural drawings updated: Revision B → C. Changed: foundation pad dimensions on Sheet 4. Submission pack sections affected: 3 (structural overview), 7 (drainage summary). Regenerate? [Yes / Review first]". Regeneration takes 8 minutes. The previous manual equivalent took 3–4 hours.
Building control portal submission: API options and manual upload workflows for England and Wales
England and Wales still lack a unified building control submission portal. The Planning Portal handles planning applications with a reasonably mature developer API, but building regulations submissions work differently: some local authorities use Planning Portal's building regs submission service, others use their own portals, and a minority still accept email or post.
For the 12 local authorities the Kent contractor works across regularly:
- 7 accept submissions via Planning Portal's building regs service
- 3 use their own web portals (Medway, Swale, Folkestone and Hythe)
- 2 still accept email submission with PDF attachments
For Planning Portal submissions, the pipeline uses Playwright to automate the upload workflow — not a formal API, but a headless browser session that fills the submission form and uploads documents. This is brittle compared to a proper API integration and breaks when the portal updates its layout. One Medway portal change in early 2026 broke the upload script for six days. Always run portal automation with failure alerting and a documented manual fallback.
What changed in 2025–2026: AI-native building regulation checking and Planning Portal API updates
Two developments matter for contractors building or operating document automation in 2026.
First, the Building Safety Regulator has expanded its higher-risk buildings (HRB) regime under the Building Safety Act 2022. For buildings over 18 metres, the gateway 2 process now requires a substantially more detailed documentation set — 30–50 documents rather than the typical 12–18 for a standard full plans submission. This makes extraction and assembly automation more valuable, but the compliance stakes are proportionally higher. We have extended the extraction pipeline to handle the HRB document set, but we always recommend a specialist building safety consultant reviews the assembled output before gateway submission. Automation reduces assembly time; it does not replace technical judgement.
Second, Planning Portal published updated developer documentation in late 2025 covering building regulations submission endpoints. The API is still in limited rollout and not all local authorities have opted in, but if your contractor works primarily with authorities that have, a proper API integration replaces Playwright scraping and is significantly more reliable. Check current coverage against your local authority list before committing to either approach.
On the checking side: tools including Xaela AI and Corelink (beta at time of writing) aim to run preliminary compliance checks against Part L, F, and B requirements before submission. These are worth monitoring but are not yet reliable enough to remove human review from the process. BCOs retain full discretion — no software changes that. This is a fair counterpoint to the general AI-in-compliance narrative: automated checking is useful as a pre-flight scan, not as a substitute for the judgement call.
Good / Bad / Ugly: three document automation approaches for UK construction compliance
Good: Manifest-driven assembly with hash-based version control. Each document is tracked individually. Revisions trigger targeted regeneration. The project manager reviews a diff, not a full rebuild from scratch. This is what we built for the Kent contractor and it holds up at 15–20 submissions per month. The key is that the manifest acts as a single source of truth for what version of each document is in the pack — it makes the review step fast and the audit trail complete.
Bad: RPA screen-scraping for portal submission without monitoring. Playwright automation against local authority portals works until the portal changes. Run it without failure alerting and you will not know a submission failed until a BCO deadline is missed. RPA is a reasonable interim approach — we use it ourselves — but it requires a documented manual fallback and 24-hour alert coverage on upload steps.
Ugly: Bulk PDF merge with no manifest. Several off-the-shelf construction document tools offer a "combine PDFs" feature. Project managers use these to improve formatting consistency, which helps. But you still spend two hours checking drawing revisions by eye because there is no version tracking, no hash comparison, and no field extraction. This approach is worse than full manual in one specific respect: it creates a false sense of consistency. The pack looks clean. The revision cross-references may not be.
For construction companies also looking at automating bid preparation alongside building regulations work, the AI tender bid response automation post covers the adjacent problem of assembling compliant tender documents from previous submissions and specification PDFs. The extraction patterns overlap significantly with what is described here.