A 25-person UK agency had automated timesheet capture, approval, and export to Xero. But every month, their operations manager still spent four hours opening each employee record, verifying tax codes, cross-checking NI categories, and running payroll manually. The automation had stopped at the door to the payslip. That four hours per month is typical for UK SMEs that have automated the upstream steps but left the statutory document layer untouched — because payslips carry legal weight that no generic workflow tool acknowledges out of the box.
The gap is specific. Upstream tools handle hours and approvals well. Downstream, HMRC requires eight specific fields on every payslip, correct NI category assignment, accurate statutory payment labelling, and RTI submissions matching the paper record. None of that is hard to automate — it needs a pipeline that knows UK payroll rules, not one that treats a payslip as a formatted CSV export.
UK payslip legal requirements: the eight mandatory fields and what HMRC considers non-compliant
The Employment Rights (Itemised Pay Statement) Amendment Order 2019 extended payslip obligations to workers (not just employees) and added a mandatory hours field where pay varies with hours worked. The current eight required fields are: gross pay, net pay, fixed deductions as a single or itemised figure, variable deductions each itemised separately, method and amount of payment where an employee is paid in different ways, and — since April 2019 — either the total number of hours worked or a per-rate breakdown where pay varies by hours.
HMRC enforcement targets specific failure modes: missing NI contribution breakdown (showing both employee and employer NI separately), absent tax code references, and payslips that show a total deduction figure without itemising statutory deductions. A payslip that bundles PAYE and NI into a single "deductions" line is technically non-compliant under current rules.
The spec your pipeline validates against should check for all eight fields explicitly. We add a JSON Schema validation step before PDF render — if the check fails, the record is flagged for human review rather than discarded, because the failure is typically a data quality issue upstream (a missing tax code in BambooHR) rather than a calculation error.
Connecting payroll data sources: approved timesheets, Xero payroll, and BambooHR as pipeline triggers
The pipeline has three data sources. BambooHR holds the employee master record: tax code, NI category, employment type, start date, benefit entitlements. Xero Payroll holds the approved pay run data and bank payment instructions. Timesheet data — approved by a line manager — is the trigger event that starts the chain.
We trigger the payroll pipeline via a Xero webhook on pay run approval. Here is the n8n webhook receiver configuration that starts the process:
{
"trigger": "xero_payrun_approved",
"source_payrun_id": "{{$json.body.payRunID}}",
"tax_year": "{{$json.body.taxYear}}",
"pay_period_end": "{{$json.body.payPeriodEndDate}}",
"bamboohr_sync": {
"endpoint": "https://api.bamboohr.com/api/gateway.php/{{company_id}}/v1/employees/all",
"fields": ["taxCode", "niCategory", "employmentStatus", "startDate", "salaryBasis"],
"cache_ttl_minutes": 0
},
"pipeline_steps": [
"validate_ni_categories",
"calculate_statutory_pay",
"generate_payslip_narrative",
"render_pdf",
"deliver_secure"
]
}
Setting cache_ttl_minutes to zero forces a fresh BambooHR pull every run. This matters: HMRC issues P6/P9 coding notices mid-year, and using a cached tax code from six weeks ago will produce an incorrect payslip that creates an RTI discrepancy. The approved timesheet data from Harvest or Clockify (depending on client setup) feeds the gross pay calculation; draft or disputed timesheets are excluded at the trigger step, not mid-pipeline.
See the timesheet and payroll data extraction pipeline for the upstream approval flow this webhook depends on.
NI category and tax code assignment automation: handling W1/M1 codes, starter declarations, and mid-year changes
NI category assignment is deterministic but has several edge cases that break naive automation. The ones that appear most often in UK SME payrolls:
| Scenario | Correct handling |
|---|---|
| Employee reaches State Pension age mid-year | Switch to Category C from the week after birthday; recalculate from that date forward |
| New joiner without P45 (starter declaration) | Apply Category A; use 1257L W1/M1 emergency tax until HMRC issues a coding notice |
| Married woman holding CA4139 certificate | Category D only if certificate is on file and pre-dates employment start |
| Director with annual earnings period election | Apply Category M or A based on salary/dividend structure; flag for accountant review |
| Employee aged under 21 | Category M (zero employer NI up to UEL) unless already Category D or C applies |
| Mid-year P6 coding notice received | Apply from the effective date on the notice, not from the next full pay period start |
| W1/M1 emergency codes active | Do not carry forward deductions from prior periods; each period is calculated in isolation |
The W1/M1 handling trips pipelines most often. An emergency tax code means each pay period is calculated standalone — you cannot cumulate tax. Any pipeline that applies cumulative PAYE logic to a W1/M1 employee will over- or under-deduct, producing an FPS discrepancy and an unhappy employee on the same day.
LLM-assisted payslip narrative generation: deductions, benefits-in-kind, and statutory payment labels in plain English
The LLM's role in this pipeline is narrow: it converts structured payroll calculation output into the plain-English field labels that appear on the employee-facing payslip. It does not make calculations. It does not decide NI categories. It takes a validated JSON object from the calculation engine and produces human-readable descriptions for each line.
This matters most for benefits-in-kind and statutory payments. "Statutory Maternity Pay (weeks 1–6 at 90% of AWE)" is more useful to an employee than "SMP." "Private medical — BUPA — P11D value £1,200 pro-rata" is more useful than "Benefit in Kind." We use a Claude claude-sonnet-4-5 call with a strict system prompt that prohibits hallucinating figures (it must use only the values from the input JSON), requires field-for-field label output matching the schema, and flags any deduction it cannot categorise for human review.
The LLM-generated narrative goes through a second validation pass before PDF render. If any generated label does not map to a known field type in the schema, it is rejected and queued for human correction. This keeps the compliance-relevant calculation output entirely in the deterministic layer, while using the LLM for what it does well: producing readable, plain-English descriptions from structured data.
PDF generation and secure delivery: payslip formatting requirements and GDPR-compliant employee delivery
Payslip delivery is a GDPR consideration. Employee payslips contain NI numbers, tax codes, gross pay figures, and benefit details. Email delivery without encryption does not meet the ICO's guidance on personal data in transit.
The approach we use: generate the PDF with WeasyPrint from an HTML template, protect the PDF with a password derived from the employee's date of birth and NI suffix — not something the ops manager types manually, it is set programmatically during generation. The PDF is stored in S3 with a 24-hour pre-signed URL. The employee receives an email containing the pre-signed link and a reminder of their password format. The email body contains no payroll data. The link expires. The PDF is independently password-protected.
For the template, HMRC's employer guidance specifies that the document must be legible and must clearly distinguish gross pay, deductions, and net pay — it does not mandate a specific layout. Our HTML template enforces named sections; WeasyPrint renders the PDF consistently across pay periods and HR tool updates.
Statutory pay calculation: SMP, SSP, SPP — the rules that break generic payroll automation
Statutory payments are where generic automation fails most visibly. SMP, SSP, and SPP each have different calculation bases, different employer recovery rates, and different qualifying conditions.
SMP (Statutory Maternity Pay): 90% of Average Weekly Earnings for weeks 1–6, then £184.03/week (2025/26 rate) or 90% of AWE if lower, for weeks 7–39. AWE uses an 8-week reference period ending on the Saturday before the 15th week before the expected week of childbirth. Getting the reference period wrong is the most common SMP error in practice.
SSP (Statutory Sick Pay): £116.75/week (2025/26), payable from the fourth qualifying day of sickness. The Linked Periods of Incapacity rule means two absence periods within 8 weeks are treated as one — the pipeline must check absence history before applying the three waiting days.
SPP (Statutory Paternity Pay): Same rate as SSP at £116.75/week, payable for 1 or 2 consecutive weeks only.
We implement these as a Python calculation library with unit tests for each statutory type. The LLM does not touch these calculations. HMRC's payroll technical guidance provides the definitive rules your implementation must be validated against.
The Chartered Institute of Payroll Professionals takes a measured view: their guidance on payroll technology recommends that statutory pay calculations be verified by a qualified payroll professional before deployment, particularly for SMP. That is sound advice. The calculation library should be reviewed by a payroll professional before it runs against a live payroll.
Year-end P60 generation: automating the annual summary from monthly payslip records
P60s must be issued to all employees in post on 5 April, by 31 May of the following year. The figures on a P60 must match the cumulative RTI submissions made to HMRC across the full tax year: total gross pay, total PAYE deducted, total employee and employer NI contributions, student loan deductions, and any statutory payments received.
The pipeline generates P60s by aggregating the monthly payslip records stored in Postgres. Each monthly record already contains all required fields; the P60 is a summation with a different HTML template and a different PDF output. The critical validation: the Postgres aggregate must match the cumulative FPS submissions to HMRC to the penny. If they diverge — because of a mid-year correction, a late starter declaration, or an EPS adjustment — the discrepancy is flagged for human review before the P60 is issued.
P60 generation runs automatically on 6 April each year via a scheduled n8n workflow. The ops manager receives a list of flagged discrepancies — typically one or two per 25 employees, usually a starter who joined partway through a period.
The AI employee onboarding document automation post covers how starter declarations and P45 data are captured and stored at onboarding, which is where the P60 audit trail begins.
What changed in 2025–2026: HMRC Making Tax Digital for PAYE and RTI submission automation
HMRC's Making Tax Digital for PAYE programme moved into its next operational phase in April 2026. The development most relevant to SME payroll automation: HMRC now provides a programmatic API endpoint for employers to query current employee tax code data directly, replacing the manual process of waiting for P6/P9 notices to arrive in the employer's HMRC online account. A pipeline built to poll this endpoint at the start of each pay run will catch coding notice changes before payroll is processed — rather than discovering a stale tax code after the fact.
The RTI submission process itself is unchanged: Full Payment Submission on or before the payment date, Employer Payment Summary for months with no payments or to recover statutory pay costs. What changed is acknowledgement speed — HMRC now returns RTI acknowledgements in under 60 seconds in the majority of cases, making it practical to wait for the confirmation reference before triggering payslip PDF generation and delivery.
The 2025/26 statutory pay uprating came into effect on 6 April 2025: SMP weeks 7–39 at £184.03/week, SSP and SPP at £116.75/week. Any pipeline that hardcodes statutory rates rather than reading them from a configuration table must be manually updated each April. The Making Tax Digital VAT return automation post covers the broader MTD compliance landscape that sits alongside the PAYE changes described here.
Good / Bad / Ugly: three payslip automation approaches and their HMRC audit readiness
Good: Calculation-first, LLM-last pipeline
The deterministic engine handles all statutory maths. The LLM produces only plain-English labels for the employee-facing document. Validation gates run before PDF render and before RTI submission. A human review step covers flagged records — typically W1/M1 new starters and any statutory pay calculations. The full audit trail lives in Postgres and S3: timesheet ID to gross pay calculation to deduction breakdown to RTI submission reference. Build time: 15–20 days with an existing timesheet and HR data source. Our Invoice OCR automation case study shows the same validation-before-output pattern in document processing more broadly.
Bad: Spreadsheet-to-PDF formatter wrapped in an API
Several payroll tools take a CSV of pay figures and produce a formatted PDF. They do not handle NI category logic, W1/M1 tax codes, or statutory pay calculations — the ops manager still has to produce the correct figures manually. The document layer is automated; the compliance layer is not. HMRC audit exposure is unchanged.
Ugly: Routing statutory calculations through an LLM
We have seen attempts to calculate SMP and PAYE directly through an LLM. Do not do this. LLMs are non-deterministic, and statutory calculation rules change each April. Edge cases — AWE reference periods, linked incapacity periods, Category D NI — require exact arithmetic against specific legislative thresholds. An LLM-generated SMP figure that is off by £2.50 creates an FPS discrepancy that triggers an HMRC query. Use a tested calculation library; use the LLM only for text formatting where label variation has no compliance consequence.
For GDPR obligations in automated document handling, including secure delivery requirements for payroll data, the GDPR data subject access request automation post covers the regulatory framework that governs this delivery pipeline.