Overview / Method

Method, parameters, and how to read these numbers

This page is the study's specification. It states what was collected, how it was measured, what the measurements can and cannot support, and where the reader should be sceptical. Nothing in the database should be interpreted without it.

1 · Research question

Do Pakistan's four provincial textbook boards and its Federal Board teach materially different things — and if so, is the difference provincial policy, or an artefact of which curriculum generation each board happens to be printing? The second half of that question turned out to matter more than the first, and it determined the study's design.

2 · Corpus construction

The sampling frame is a fixed matrix: five boards × classes 6–10 × the subjects that carry social content (English, Urdu, Islamiyat, Pakistan Studies, General Science, Biology, Chemistry, Physics, and Social Studies or its History/Geography split). Where a board publishes an English-medium edition it was preferred, because English is the tier that supports direct quotation; Urdu was taken only where no English edition exists.

Provenance differs by board and this bears on reliability. Sindh operates a functioning official e-book portal and all Sindh titles came from it. Neither Punjab nor KP serves textbook PDFs from an official domain — KP's public downloads interface carries only tenders and a price list — so those titles came from free public mirrors, cross-checked between independent mirrors where possible. Every download was verified to be a genuine PDF by magic-byte check, not by file extension.

Two acquisition failures were caught and corrected, and they are instructive. One widely-mirrored KP Islamiyat file proved truncated at roughly half the book. Separately, one mirror's copy of a title was larger in bytes than the complete edition while containing half the pages, because each PDF page held two printed pages. Completeness was thereafter verified by parsing the PDF page tree, never by file size.

3 · Evidence tiers — the constraint that governs every claim

Only 15 of 167 current books carry a machine-readable text layer. The rest — including current official editions from every board — are scanned page images, rendered at 300 dpi and processed with OCR. This produces three tiers, and each claim in the study is graded by the tier it rests on:

Where a claim about an Urdu-medium book was load-bearing, the printed page was rendered as an image and read directly — a method that bypasses OCR entirely. This was used for the study's two most consequential Urdu claims.

Absence is the weakest claim the study makes. Where a term could not be located in an Urdu file it is recorded as "not established", not "absent" — except where the term is a proper noun likely to survive degraded script. The reason is concrete: one 104-page KP file containing an entire chapter of Sufi hagiography returned a single OCR hit for "حضرت" and none for "محمد". Term-frequency counts on Urdu files are unusable, and were not used.

4 · The rubric — fixed before scoring

Seven axes with observable indicators were written down before any book was scored, so that scoring could not drift toward a preferred conclusion. Each runs 1–5. 1 is always the secular / pluralist / open pole; 5 the religious / exclusivist / closed pole. The poles describe properties of text. They are not verdicts, and a high score is not an accusation.

AxisBand 1Band 3Band 5
A1 Religious saturation
outside Islamiyat
Nothing beyond a bismillah on the imprint pageAt least one dedicated religious lesson, plus recurrent referenceSubject matter subordinated to religious instruction
A1b Doctrinal register
Islamiyat only
Ethical, comparative or historical framingDevotional and ritual instruction, non-polemicalPolemical, boundary-policing or supremacist
A2 Civic thickness
reverse-scored
Constitution, rights instruments, elections, civil society and dissent all taughtCivic content present but thin or purely institutionalNo civic content; ideology occupies the slot
A3 Received truthContested questions presented as contested; evidence weighedInquiry verbs present but conclusions pre-loadedConclusions handed down to be defended or memorised
A4 Narrative closureMultiple actors; internal causes and wrongdoing admittedStandard nationalist arc with some internal criticismExternalised blame, no internal fault, providential inevitability
A5 MilitarismArmed forces absent or purely institutionalWars narrated; forces praised in historical contextMartyrdom taught as aspiration; forces in non-history subjects
A6 Othering
scored per out-group
Out-group treated warmly; shared heritage acknowledgedOut-group as historical adversary onlyEssentialised hostility; framed as inherently threatening
A7 Gender conservatismWomen as agents across roles, including intellectualWomen present, mostly commemorative or domesticWomen largely absent, or roles prescriptively bounded

A6 is scored per book × out-group, not once per book — Hindus, the Indian state, Pakistani religious minorities, the West, and internal others are tracked separately, because the boards rank differently depending on which is asked about. The table view shows the Hindus figure; the full set is in the board view.

5 · Scoring procedure

The unit of analysis is the book, never the board. Board figures are aggregates of book scores, and each axis has its own denominator because each lens scored only the books relevant to it. Blank means "not scored", never "zero". Civics scored only civics-bearing books. Narrative closure was scored only on books that deliver a national narrative at all — Punjab's History 6 and Sindh's Social Studies 6 and 7 deliver none, and that absence is reported as a finding rather than recorded as openness, which would have flattered them.

Where two boards run the same book, the score is counted once and identified as shared. This was established for Islamiyat at classes 6–7 (Sindh and KP), four science subjects (Sindh and KP), and the class-9 Urdu anthology (Punjab and KP). Reporting those as board differences would have manufactured variation that does not exist.

Every verbatim quotation in the study was re-verified against its source file by exact string match before inclusion. A verification pass on intermediate working notes found that 5 of 14 spot-checked quotes had drifted through line-wrap or paraphrase; those were corrected or dropped, and the discipline was then applied to all six lens reports.

6 · Vintage control — the study's central parameter

Four curriculum generations are in simultaneous use across these boards. An uncontrolled comparison would measure publication era and mislabel it provincial ideology. Every board-level difference therefore states whether the compared books share a vintage, vintage-matched comparisons are given wherever the corpus permits one, and anything that cannot be matched is marked confounded and is not used to support a conclusion.

An edition audit confirmed the gap is structural rather than an artefact of what was collected: the corpus already holds each board's newest edition of every book in the matrix. Sindh has not adopted the Single National Curriculum at all, and KP's adoption covers six subjects at one grade. So Punjab frequently cannot be matched against a contemporaneous Sindh or KP book, because none exists.

Consequence for interpretation: where no matched pair exists, the defensible statement is not "Punjab is more X than Sindh" but "Punjab's current books are more X than the books Sindh and KP are currently printing" — which is in any case the comparison that matters to a child in a classroom this year.

7 · Traps identified and controlled

Communal vocabulary is not devotional vocabulary. One book carries 167 instances of "Muslim" against zero of "Quran", "Prophet" or "Hazrat" — there, religion functions as a national-identity marker, not a devotional subject. A saturation score driven by the first word would misread it badly, so chapter fraction and register are scored on separate axes.

Personification inflates gender counts. Referring to the nation as "her" inflates female pronoun ratios in Pakistan Studies; 24 of 43 computed ratios were flagged uninterpretable for this and related reasons rather than reported.

A recent printing year is not a recent edition. KP's Biology 9 is a 2018 book reprinted for 2025–26; its Pakistan Studies 9 is a 2010-approved book printed for 2024–25. Where a record's approval note says "print run, not a new edition", the curriculum field is the one that carries meaning.

8 · How to interpret a score

A score is a measured property of a text against a published band definition, assigned by reading the book and citing the page. It is not a judgement of the board, the authors, the teachers, or the children who use it. Two further cautions:

Do not collapse the axes into one number without saying so. The study's clearest result is that the secularity index and the conservatism index rank the boards differently — KP is the most religiously saturated curriculum and simultaneously carries the most open historical narrative. "How religious is this curriculum" and "how conservative is this curriculum" are separate empirical questions here, and any single ranking that merges them will mislead.

Variance matters more than means. Sindh's othering scores run 1, 1, 2, 2, 4, 5 across its books — a mean conceals that spread entirely, and the spread is the finding. Open the individual records before quoting a board average.

8b · How much of this text was produced by machine, and what that costs

Most textbook criticism quietly presents its quotations as if they were read off a page. In this corpus they usually were not. Stating the size of that intervention is not a disclaimer — it is a figure a reader is entitled to before deciding how much weight a quotation carries.

Books in the current corpus167 across 5 boards, plus 43 historical editions
Carrying a publisher's own text layer15 (9%)
Reconstructed by OCR from page images152 (91%)
Characters of text extracted in total24,581,557
Of which Urdu or Sindhi script — never quoted6,501,238 (26%)
Books scored against the rubric137 of 167

What OCR does to a page, specifically. Punjab's PDFs carry a diagonal Web Version / Not for sale watermark that cuts through words mid-line and splits them into fragments. KP mirror copies carry a perfect24u.com watermark plus mirror-site advertising pages that are not part of the book. Multi-column layouts interleave: a two-column page can emerge with the columns woven line by line. Ligatures drop. Digits in tables are the least reliable characters on the page, so no statistic — reserves, tonnages, populations — is quoted from OCR without checking it against the PDF.

The rule that follows. Every quotation published on this site was re-verified by exact string match against the extracted text before publication, and no sentence is ever quoted from Urdu or Sindhi material — that 26% of the corpus supports statements about structure, chapter coverage and proper nouns only. Where a claim rests on something not appearing in Urdu material, the study says not established rather than absent, because degraded Nastaliq OCR cannot distinguish the two.

One silent failure was found and fixed, and it is worth recording. The OCR script passed absolute file paths to xargs -I, which on BSD caps a constructed argument at 255 bytes. Books with long filenames exceeded it; xargs failed, and the script then wrote a zero-character file that every later completeness check read as a finished extraction. Twenty files were affected. The pipeline now runs from inside the working directory so path length cannot matter, and refuses to write any extraction under 200 characters — an empty result is recorded as a failure rather than as an extraction.

8c · How this method fails, and how each failure was caught

Four failures during construction produced output that looked entirely normal and was wrong. They are recorded because they are properties of the method rather than accidents, anyone reproducing this work will meet them, and in every case the tell was subtle enough that a reasonable person would have missed it.

FailureWhat it looked likeHow it was caught, and the guard now in place
A language model applied to the wrong script. Sindhi-medium books read with the Urdu model. Fluent-looking Urdu output at normal length. Sindhi uses roughly sixteen consonants Urdu lacks — ٻ ڀ ٺ ٽ ٿ ڏ ڊ ڍ ڌ ڪ ڳ ڱ ڻ ڇ ڄ ڃ — and the model, having no glyphs for them, substituted plausible Urdu letters throughout. A Sindhi-specific character count returned zero across 213,785 characters, which is impossible for real Sindhi. Re-run with the Sindhi model the same book yields 19,219 — ڪ alone 6,684 times. The pipeline now selects the model from the medium, and the same error had earlier been made in reverse, running the English model over Urdu-medium books.
A shell limit silently producing empty files. Extraction reported success. The files were zero bytes, and every later "is this book done?" check answered yes. BSD xargs caps a constructed argument at 255 bytes; long book names exceeded it. The pipeline now runs from inside its working directory so path length cannot matter, and refuses to write any extraction under 200 characters — an empty result is a failure, not an extraction.
A mirror mislabelling medium. Six books listed as English medium. Plausible catalogue entries. OCR'd with the English model they produced 18,000 words of confident noise per book. They were byte-identical to the Urdu editions already held — same SHA-256. The downloader now hashes every file against the corpus and discards exact duplicates before they land. Medium is confirmed from the book's own text, never from a listing.
Words run together in authored prose. carrythe, the1965, Arabicscript. Ordinary sentences with a missing space, invisible at a glance and invisible to a regex when there is no case change. Caught by a reader. A case-change regex finds only half of them; detection needs a dictionary, asking whether a token is a real word and whether it splits into two. scripts/find_glued_words.py now runs over all authored prose.

The common shape is worth naming. None of these announced itself. Each produced output that was well-formed, plausible and wrong, and three were found only because a number was implausible rather than because anything failed. A pipeline that reports success is not evidence of success, which is why completeness here is verified by parsing a PDF's page tree and rendering its final page rather than by trusting an exit code.

9 · Limitations

A full, computed inventory of what is missing — books that could not be found, text that cannot be read at quotation grade, axes that were not scored, and material deliberately excluded — is at what is missing. It is generated from the catalogue at build time rather than written by hand.

76 of 167 current books are Urdu-medium and scored at theme grade; KP carries the heaviest Urdu load, so KP's scores are the least secure of the five boards. Absence-driven scores may reflect OCR failure rather than editorial choice. The single largest evidentiary gap is KP's class-8 History, which carries that board's entire colonial and Pakistan-Movement narrative at a register the present scan cannot resolve — a Nastaliq-tuned OCR pass on that one file would move more of this analysis than any other acquisition. Boards reprint continuously and mirrors lag, so all findings are tied to the specific editions recorded in each book's record.

10 · Reproducibility

Every source URL, the acquisition and verification scripts, the extraction pipeline, the rubric as written before scoring, six per-lens reports with page-cited per-book tables, and the assembled dataset behind this database are all retained with the project. This page's numbers are generated from that dataset, not typed in.