Psychological and social implications of Pakistan’s four provincial curricula and its federal one, stated as a testable research programme
0. What this document is, and what it is not
It is not a finding about children. Nothing in this corpus — 120 current textbooks from four provincial boards, classes 6 to 10 — shows what any pupil believes, has believed, or will believe. No child was surveyed. No classroom was observed. No examination script was read. Every sentence here that concerns a mind is a prediction, and is labelled as one.
It is also not a refusal to engage. The question that brings most readers to a study of school textbooks is the forward-looking one: what does this do to the people formed by it? Declining to address it would leave the field to people who will answer it badly. The defensible position is the middle one, and it has three parts.
The document has two halves, and they carry different weight. Sections 4 to 6 predict attitudes — what a leaver thinks about 1971, minorities, the army — and rest on the observation that an examinable proposition is compelled rather than merely offered. Section 4A predicts personality structure: identity foreclosure, need for cognitive closure, ingroup glorification, moral foundations, political efficacy, integrative complexity, and a small set of psychoanalytic propositions given behavioural operationalisations. That second half requires an additional inferential step — that compelled cognitive operations, repeated for years, generalise into dispositions — and it is the part most likely to be wrong. It is written so that the pattern rather than the level carries every claim, because a flat profile across life domains would exonerate the curriculum entirely.
- The corpus observes a treatment. What each board teaches, how often, in what order, in which subjects, and — the strongest variable available — what it requires a pupil to produce in an assessable form. That is measured, page-cited and reproducible by anyone with the same files.
- Established psychology supplies mechanisms. Not results: mechanisms. These are described in §3 in ordinary language and without attribution, because the point is to invoke well-known ideas rather than to borrow the authority of particular studies, and because a fabricated citation would destroy the value of the whole exercise. No study, statistic, effect size or survey finding external to this project is cited anywhere in this document.
- Together they generate region-differentiated, falsifiable predictions. Four boards administer measurably different treatments to demographically overlapping populations. If the mechanisms operate, the four regions should differ in specified ways and specified directions.
The document is written as a research programme, not a verdict.
A note on the hypothesis actually being tested. An earlier and cruder version of this analysis treated poverty, media, politics and social media as confounds — rival explanations to be conceded in a closing section. That framing is too weak. Those variables interact with curriculum, and several of them sharpen the argument rather than blunting it: exposure to the curriculum is itself unequal and unequally distributed across the four provinces; the textbook’s relative authority rises where competing sources are thin; media plausibly activates frames that schooling installed rather than competing with them; and the pattern of curricular divergence is probably a political artefact before it is a pedagogical one. Section 6 develops these as analysis, not as apology. The consequence is that the hypothesis under test is not textbooks matter but textbooks matter conditionally — on dosage, on the density of competing sources, and on whether later media reinforces or contests the frame. That is a better hypothesis because it says where effects should be largest, and is therefore easier to falsify.
Standing limits, stated here rather than at the end. Quotation is restricted to English-medium sources, and two of the four boards teach Islamiyat, and two teach History, only in Urdu — so a large fraction of the treatment is invisible to the strongest instrument. Four curriculum generations are simultaneously in print, so most board comparisons are vintage-confounded. Balochistan is now scored on all seven axes, but eight of its books could not be found at all, twelve of twenty-eight are Urdu-medium and unquotable, and its classes 6–9 volumes are pilot editions printed for one academic year — so Balochistan statements below are grounded but fragile. The corpus covers only public-board textbooks, while a large and unquantified share of Pakistani children are educated in madrassa, private-elite or low-cost private streams.
1. The design: four boards as a quasi-natural experiment
The analytical backbone is not the content of any one book. It is the comparison.
Since curriculum was devolved to the provinces, the four boards have diverged — not rhetorically but measurably, on scored axes, book by book. The populations they teach overlap substantially in religion, national media environment, political system and economic structure. They are not identical, and §6 and §7 give that its full weight; but they are far closer to one another than to any external comparison group, and the thing that most sharply distinguishes them for a child in a public secondary school is which board’s books arrive in the classroom.
The test. If curriculum shapes attitudes through the mechanisms in §3, the four provinces should differ from one another in specified, predictable ways on the dimensions where their curricula diverge — and should not differ on the dimensions where their curricula are identical. If the observed pattern of differences and non-differences does not match, the mechanism is weaker than claimed, or is conditioned by something the design has not modelled.
The design has a built-in control condition, which is unusual and valuable. The corpus establishes several corpus-wide constants — propositions and absences shared by every board at every vintage. Across roughly 500,000 words of English-medium science there is effectively no woman presented as an agent of discovery; no woman is profiled as a scholar in any Islamiyat volume; no woman is credited as author of any Urdu selection across twelve readers. The “Hindu teachers” explanation for the loss of East Pakistan survives in all three scored boards across 2004, 2023 and 2024–25 editions, rising in KP to a section heading — “7. NEGATIVE ROLE OF HINDU TEACHERS” (KP, Pakistan Studies 9, p. 95) — and appearing in Punjab as “Education sector in East Pakistan was totally under the control of the Hindus. They poisoned the Bengalis against Pakistan” (Punjab, Pakistan Studies 9, p. 56). Partition violence is absent from every board, as is any economic critique of colonialism. Inquiry apparatus collapses at matriculation in every board regardless of vintage.
(That internal disagreement is now resolved, against the synthesis. A targeted recount, separating Marie Curie the person from the curie unit of activity and the Curie temperature, finds four mentions of women scientists in three books from two boards: Marie Curie in KP’s Physics 10 and Punjab’s Physics 10, Madam Curie in KP’s Chemistry 9, and Lise Meitner in KP’s Physics 10 — credited with the theoretical explanation of nuclear fission and given a portrait. Sindh and Balochistan name none. The control dimension holds, but the premise is now the corrected one: women as agents of knowledge are vanishingly rare rather than absent, and the board that supplies all but one of the mentions is the board that scores worst on gender overall.)
Those constants function as a placebo dimension. On them the mechanism predicts no provincial difference. Testing the non-differences is as important as testing the differences, and it is the part such studies usually omit.
2. What is actually observed: the treatment, region by region
Axes run 1–5, where 1 is the secular / pluralist / open pole and 5 the religious / exclusivist / closed pole. A2 is reverse-scored, so 1 is the thickest civics. Figures are means of book-level scores and are measurements of textual properties, not verdicts.
n is given because coverage is uneven by design. A4 rests on so few books that it must not be read as a board ranking, and every board’s A1 is pulled down by the matriculation sciences, which are close to secular everywhere.
| Axis | Punjab (PCTB) | Sindh (STBB) | KP (KPTBB) | Balochistan (BTBB) |
|---|---|---|---|---|
| A1 Religious saturation (EM, non-Islamiyat) | 2.32 (n22) | 2.00 (n17) | 2.62 (n16) | 1.94 (n16) |
| A2 Civic thickness (1 = thickest) | 4.00 (n13) | 3.38 (n8) | 3.71 (n7) | 3.50 (n8) |
| A3 Inquiry vs received truth | 3.56 (n9) | 3.00 (n7) | 3.50 (n6) | 2.92 (n12) ‡ |
| A4 National-narrative closure ⚠ | 4.50 (n4) | 5.00 (n1) | 3.00 (n2) | 3.00 (n3) |
| A5 Militarism | 2.46 (n13) | 1.38 (n13) | 1.43 (n7) | 1.78 (n9) |
| A6 Othering — all targets | 2.46 (n13) | 1.80 (n10) | 1.71 (n7) | 2.67 (n9) |
| A6 Othering — Hindus | 3.17 | 2.50 | 2.67 | 3.00 |
| A6 Othering — Indian state | 3.40 | 2.25 | 3.00 | not disaggregated |
| A6 Othering — Pakistani minorities | 1.50 | 1.83 | 1.33 | not disaggregated |
| A6 Othering — the West | 2.50 | 2.20 | 1.33 | not disaggregated |
| A7 Gender conservatism | 4.18 (n22) | 3.85 (n20) | 4.31 (n16) | 3.75 (n16) |
⚠ A4 rests on one Sindh book and two KP books and cannot carry a provincial comparison. Every prediction below that depends on narrative closure is correspondingly weak, and P1 in particular should be read as resting on Punjab’s four books against a very thin comparison set.
‡ Balochistan’s A3 of 2.92 is the most misleading number in this table. It averages genuinely open geography and science tasks against justify-that prompts on contested history, and describes neither. See §4.4.
Punjab runs the most closed national narrative and the most militarised, in its newest books. Religion is placed by design in fixed slots — every English reader opens with a Seerah or Sahaba unit, while the geographies and physical sciences are as secular as anything in the corpus. Providential causation appears in the book’s own voice. Martyrdom is taught as aspiration, not history, and occupies non-history subjects: Rashid Minhas Shaheed is the model essay for teaching essay structure. The othering register reaches its corpus maximum in an assessment item: an MCQ stem that asserts “the enmity of Hindus against Muslims” before the pupil reaches the options, with the keyed answer “Hindus are not our friends” (Punjab, Pakistan Studies 9, p. 18). Punjab is also the only board printing an examination blueprint and a full specimen board paper, and the only board running the same ideological content through assessment in two media in the same year.
Sindh is the most secular lower secondary and the most religious upper secondary. English 6, 7 and 8 contain zero religious selections between them; the religious load is concentrated in Islamiyat and arrives at matriculation, where A1 reaches 4.00 — the steepest climb in the corpus. There is no military martyr anywhere in Sindh’s books. It is the only board that teaches “In a democracy, political control of the armed forces is with the civilian institutions” (Sindh, Social Studies 6, p. 71); the only board that names protest as legitimate — “Demonstrating peacefully through marches, boycotts, sit-ins or other forms of protest” (p. 84); and the only board with a sustained regime of assessment asking a pupil to reach a conclusion the book does not supply, including a requirement that groups of eleven-year-olds “use a scale to assess how well the current government of Pakistan is fulfilling its purpose” (p. 64). It is simultaneously the board whose Islamiyat carries the corpus’s most doctrinally assertive assessment item, a set debate on “Suggestions for the dominance of religion in the present age” (Sindh, Islamiyat 9–10, p. 34), and whose class-8 Social Studies carries the corpus’s hardest bloc characterisation of a living population.
Khyber Pakhtunkhwa has the highest religious saturation and, at class 9, roughly twice either other board’s religious instructional volume — two compulsory religious subjects totalling 371 pages against Punjab’s 170 and Sindh’s 176. Religion is distributed rather than placed. And yet KP has the most open national narrative in the corpus: it names its own military rulers’ overreach, prints the 1971 prisoner-of-war figure, treats Kargil as a loss of face, and records that after the 1971 military action East Pakistanis “became enemies of the Army”. Its English readers, across four years and roughly seventy-five selections, contain zero national-martial content, and carry the corpus’s only named non-Muslim Pakistani child protagonist — “I am a Sikh and a proud Pakhtun Pakistani” (KP, English 6, p. 173) — with a follow-on exercise requiring the class to write about him as the new boy in their class. Its assessment apparatus is the thinnest in the corpus, and across both Pakistan Studies volumes evaluate, assess, in your opinion and do you think return zero. KP narrates frankly and asks nothing. It also delegitimises protest where Sindh legitimises it — “When people are angry about something they organize a protest, they burn and destroy public buildings, buses and trains” (KP, English 7, p. 47) — and is the worst board in the corpus on gender.
Balochistan — the corpus’s clearest case of a template running on autopilot.
The treatment Balochistan administers is the justify-that template as standing chapter architecture. Its History 8 runs it 32 times, with zero open-opinion prompts anywhere in the volume, and it is applied mechanically enough to generate prompts for propositions the same book refutes. Two documented instances, and they are the evidential core of P10:
- “Justify that the British’ decision of settling Jews in Palestine was based on fair justice” (p. 82) — against the same volume at p. 89: “The British Policy of settling Jews in Palestine was unjust for the following reasons:”.
- “Justify that the Treaty of Sevres is based on fair justice” (p. 32) — against a narrative that elsewhere treats Sèvres as the grievance driving the Khilafat Movement.
Also present: “Conclude that the Partition of Bengal was the turning point of the Hindu Muslim Unity” (p. 19) — the conclusion supplied in the instruction to conclude.
A pupil working these prompts is not being taught what to believe about Palestine or Sèvres. The book has already told them what it believes, and then required them to argue the opposite. What is being trained is the production of supporting argument on demand, decoupled from the arguer’s own position — which is why Balochistan supplies the test in §4.4 that could most cleanly establish the demand mechanism, and could most cleanly produce a null on belief.
The assessment-demand variable. English-medium markers per 100,000 words, deduplicated:
| Board | teacher-instruction blocks | fill-in-the-blank | evaluation verbs | open-opinion prompts |
|---|---|---|---|---|
| Punjab | 6.5 | 3.4 | 9.1 | 3.3 |
| Sindh | 42.0 | 23.3 | 3.6 | 6.6 |
| KP | 0.1 | 7.3 | 2.6 | 1.6 |
| Balochistan | 11.5 | 8.3 | 13.1 | 1.3 |
The two right-hand columns point in opposite directions, and only the second tracks the presence of items that let a pupil judge. Balochistan has the corpus’s highest density of the word “evaluate” and its lowest density of items permitting an unsupplied conclusion.
3. The mechanisms
Described plainly, in the author’s own words, with no study cited and no magnitude claimed.
Social identity and the in-group boundary. People derive part of their self-concept from group membership, and the mere act of sorting the social world into “us” and “them” produces differential evaluation. A curriculum that makes one category the organising frame of the historical narrative does not merely teach about that category; it supplies the sorting scheme.
Superordinate categories. A category containing both in-group and out-group — citizen, human being, neighbour — tends to soften the boundary. Curricula differ in whether they supply one. A book that pairs the Madina Charter with the UN human rights instrument supplies a superordinate frame; a book that constitutes the nation by a permanent religious distinction supplies its opposite.
Early category acquisition, and the difficulty of unlearning. Beliefs about social categories acquired in childhood are unusually persistent. They form before the capacity to evaluate a source is mature, become the background against which later information is interpreted, and are typically revised — where at all — by contact and experience rather than counter-argument. Age of acquisition matters independently of frequency.
Exposure is not endorsement. The distinction most commentary collapses, and the one that makes the field tractable. A person may know an account thoroughly, reproduce it fluently, and not believe it. An instrument that fails to separate what the account says from what the respondent thinks will misattribute the first to the second.
Repetition, fluency and believability. Repeated statements become easier to process, and ease of processing is routinely misread as evidence of truth. Repetition across sources is stronger than repetition within one: a claim met in History, the English reader, the Urdu reader and Islamiyat is not the position of a subject. It is the shape of the world as the school describes it, and there is no subject from which to view it.
The generation and testing effects. Material a person produces themselves is retained substantially better than material they read, and retrieving something from memory strengthens it more than re-reading does. This is among the most robust findings in the study of learning, and it is entirely indifferent to whether the material is true.
Assessment washback. What is examined governs what is taught and attended to more reliably than what a syllabus states. A curriculum’s effective content is closer to its examination than to its table of contents.
Authority cues in childhood. Children discount sources less than adults do, and a textbook carries the compounded authority of state, school and teacher. Asking who says this, and how do they know is a skill; it is taught or it is not.
Moral disengagement. Harm to an out-group becomes psychologically easier when the group is placed outside the moral community. The relevant textual property is not hostility of tone but essentialisation — whether negative attributes are presented as what a group did or as what it is. The othering lens uses exactly this test: bloc characterisation confined to pre-1947 politics scores 4; the same characterisation carried into present-tense description of a living population scores 5.
Narrative transportation. Absorption in a story reduces counter-arguing, and beliefs acquired inside a narrative are less likely to be tagged with their source. Reading a martyr’s biography is not the same cognitive operation as reading a list of facts about a war.
The hidden curriculum. What a school teaches by its structure — who speaks, who is asked, what may be questioned, what is rewarded — often outlasts what it teaches by its content.
Attitudes, beliefs and behaviour are three things. They correlate loosely, and any claim running from a textbook to a behaviour passes through two joints, each of which leaks.
3.1 The spine: reading a proposition versus being examined on it
A textbook proposition can have four independent properties. It can be available — present in the book. It can be repeated. It can be sequenced — arriving at a particular age and order. And it can be demanded — the pupil is required to produce it, or something entailing it, in a form that will be marked.
Only the fourth converts receipt into production, and only the fourth attaches a reward and a penalty. It is also the property that engages the most mechanisms at once: generation, retrieval, washback and authority. A passage practises reading. An exercise requiring a pupil to construct reasons for a conclusion supplied by the board practises the construction of reasons for conclusions supplied by authority — several hundred times across five years.
In this corpus the four properties demonstrably come apart. Punjab’s History 7, p. 32, contains the board’s only rescue-frame claim about Hindu India — lower-caste Hindus “lived a miserable and socially segregated life” and were “impressed by the equality, caste-less system and social justice of Muslims” — and no exercise, activity, review item or teacher note in the volume attaches to it. Against that, twenty distinct assessed items across all four boards require a pupil to state something about the moral condition of pre-Islamic Arabia, one of them inside a printed board model paper, and zero assessed items anywhere require a pupil to state that pre-Islamic or Hindu India was barbarous.
The sharpest single specimen of the demand form sits on a page that admits its own uncertainty. Punjab’s History 6, p. 30, states in its narrative: “We do not know why this civilization came to an end, but a number of possibilities have been suggested” — and then, in an activity box on the same page, instructs: “Justify that Indus Valley people did not learn the warfare nor developed their trade, and therefore, were easily defeated by Aryans.” Availability predicts recognition. Demand predicts production. They are not the same outcome and should not share an instrument.
4. Region-by-region predictions
Every numbered item is a prediction. Each specifies a dimension, a direction and a comparison. Each is followed by what would falsify it. All assume a study stratifying by province, school stream, cohort year, urban/rural, sex and socio-economic status — without which none is testable.
Two families of outcome measure are distinguished, and the distinction does most of the work:
- Reproduction measures — can the respondent produce the account, in the textbook’s frame? Coded from open-ended response, not from agreement scales.
- Endorsement measures — does the respondent hold it? Agreement, behavioural intention, allocation tasks, and where possible indirect measures that do not cue the school-answer.
The mechanism predicts larger provincial gaps on reproduction than on endorsement. That is counter-intuitive and it is the programme’s most discriminating prediction, because the naive “textbooks indoctrinate” hypothesis predicts the opposite ordering.
4.1 Punjab
P1. Punjab-schooled leavers of the public stream will produce the closed national narrative more fluently and completely than Sindh- or KP-schooled leavers: more externalised causation for 1971, fewer named Pakistani actors at fault, more providential framing, higher spontaneous use of the “enemy” register for India. The gap should be large on reproduction and materially smaller on endorsement. Falsified by: Punjab leavers reproducing at rates indistinguishable from KP leavers, whose books say otherwise on the same events; or an endorsement gap as large as the reproduction gap, which would indicate a process other than assessment-driven production.
P2. Punjab leavers will show higher affective valuation of military sacrifice than Sindh leavers — higher agreement that dying for the country is the highest service, higher recall of Nishan-e-Haider names, higher ranking of the army among trusted institutions. Punjab is the only board teaching martyrdom as aspiration and placing it in the language readers; Sindh’s corpus contains no military martyr at all. Falsified by: no Punjab–Sindh gap; or a gap that vanishes once family military employment and district recruitment intensity are controlled. That control is essential (§7.5).
P3. Punjab leavers will show hostility toward Pakistani religious minorities that is decoupled from their hostility toward Hindus and India. Punjab is simultaneously the most hostile board on Hindus (A6 3.17) and among the most inclusive on Pakistani minorities (A6 1.50), and the two sit in consecutive volumes of one Pakistan Studies course. If the treatment does the work, the pluralism should reappear compartmentalised into the citizenship slot, never touching the historical narrative. Falsified by: a single hostility factor on which Hindus, India and Pakistani minorities load together — which would indicate a general prejudice disposition with a non-curricular source.
4.2 Sindh
P4. Sindh-schooled public-stream leavers will show higher civic knowledge and political efficacy than any other province: more able to name a rights instrument, more likely to classify protest and petition as legitimate, more likely to assert civilian authority over the military, more willing to state an evaluative judgement about a sitting government.
P4a — the sharper form. The effect should be concentrated in respondents whose classes 6 and 7 were in Sindh public schools, because that is exactly where the evaluative assessment regime sits, and should be weaker, not stronger, among those who continued to matriculation, because Sindh’s civic provision collapses at class 8 into a reprinted federal volume with no civics at all. A treatment effect that decays with additional exposure to the same board is a strange prediction, and strange in a diagnostic way: general regional-political explanations do not predict it. Falsified by: a Sindh civic advantage increasing monotonically with years of schooling; or one present equally in respondents who left after class 5; or none at all once ethnicity, Karachi residence and party environment are controlled.
P5. Sindh leavers will show a dissociation between communal hostility and doctrinal exclusivity: relatively low on India and Hindus, relatively high on propositions about the political dominance of religion and the status of other religions. Sindh’s religious saturation climbs steeply to matriculation, its Islamiyat is the corpus’s most doctrinally assertive, and its class-8 volume asserts the pre-Islamic condition to persist today in other religions. Falsified by: Sindh scoring uniformly at the pluralist end across both families of item. Uniformity would indicate a province-level disposition rather than a curriculum-shaped profile, and would weaken the whole programme rather than this prediction alone.
4.3 Khyber Pakhtunkhwa
P6 — the most theoretically consequential test. KP leavers will show high religious identity salience without correspondingly high out-group hostility. KP has the corpus’s highest religious saturation and roughly double the class-9 religious instructional load, and its lowest othering scores on Pakistani minorities (1.33) and the West (1.33), and its most open national narrative. If religiosity and hostility are separable in the treatment, they should be separable in the population. This matters beyond Pakistan: the assumption that religious saturation produces out-group hostility is widely held and rarely tested against a case where the two are dissociated by design. KP is that case. Falsified by: KP leavers showing hostility toward minorities or the West at or above Punjab levels while also showing higher religiosity — which would indicate that saturation carries hostility regardless of how out-groups are framed, an important finding against this framework.
P7 — the sharpest available test of the demand mechanism. KP’s Pakistan Studies volumes contain the frankest political history in the corpus and none of it is assessed as judgement. The prediction is a triple dissociation:
| Critical content available | Judgement demanded | Predicted outcome | |
|---|---|---|---|
| KP | high | zero | high factual recall of critical episodes; civilian-supremacy attitudes no different from Punjab |
| Punjab | low | zero | low recall; attitudes as KP |
| Sindh | moderate | present | civilian-supremacy attitudes distinct from both; recall intermediate |
If that pattern appears — KP knowing more and believing no differently, Sindh believing differently — the demand mechanism is supported in a way no content analysis could support it, because availability is held high and demand set to zero in the same books. Falsified by: KP leavers differing from Punjab leavers on civilian-supremacy attitudes in proportion to their factual recall, which would show that availability alone shifts attitudes.
P8. KP leavers will show the widest gap in the corpus between attitudes toward religious minorities (predicted relatively pluralist) and attitudes toward women’s public roles (predicted the most conservative of the four provinces) — the two scoring best and worst in the same books. Falsified by: a general conservatism factor on which gender and minority items load together.
4.4 Balochistan — now scored, and the most theoretically interesting board in the set
Balochistan runs National Curriculum 2022-23 at classes 6–9 and National Curriculum 2006 at class 10 — two curriculum generations side by side, read from the books’ own front matter.
Balochistan turns out to be the sharpest natural experiment in the corpus, for a reason that has nothing to do with its politics. It carries the highest inquiry-verb density of any board and almost no open-opinion demand. Its History 8 sets 32 evaluation prompts and zero requests for the pupil’s own view. Its Pakistan Studies 9 sets 30 and three.
That combination isolates a variable no other board isolates: the form of critical thinking, administered without its substance.
P9 — the form-without-substance prediction. Balochistan leavers will show the corpus’s largest gap between two capacities that are usually correlated: high fluency in constructing supporting argument for a supplied conclusion, and low rate of reaching an unsupplied conclusion. Asked to justify a proposition, they should outperform leavers of the other three boards. Asked what they themselves conclude from the same evidence, they should underperform. Falsified by: the two capacities correlating in Balochistan at the same rate as elsewhere; or by Balochistan leavers showing no argumentative-fluency advantage at all, which would indicate the prompts were never actually worked.
P10 — the decoupling prediction, and the cleanest test of the demand mechanism in this document. Balochistan’s justify-that template is applied mechanically, and this is documented rather than inferred: the same volume that treats the Treaty of Sèvres as a Muslim grievance also instructs pupils to “Justify that the Treaty of Sevres is based on fair justice”. A pupil who spends a year generating reasons for whatever proposition the template drops into the slot — including ones the book itself contradicts — has practised advocacy on assignment, not belief.
Predicted: Balochistan pupils will show lower coherence between the arguments they can construct and the positions they endorse than pupils elsewhere. If this holds, it is evidence that assessment form transfers independently of content — and, importantly, it is an argument for the null on belief. It would mean the justify-that template produces skilled advocates rather than convinced nationalists, which is a materially different social outcome and a much less alarming one. Falsified by: Balochistan pupils endorsing required-to-justify propositions at rates comparable to merely-taught ones; or by fluency and endorsement being as tightly coupled there as in Punjab.
P10a — the subject-boundary prediction, newly available. Scoring showed the loaded template is not spread evenly: Balochistan’s geography, science and English books set genuinely open tasks (“debate on the reasons why oceans are…”, “evaluate the reasons for low” yields), while history and Pakistan Studies set justify-that on contested propositions. Within a single Balochistan respondent, therefore, integrative complexity should be higher on economic and environmental questions than on historical-political ones. This is a within-subject test requiring no cross-province comparison and no assumption that provinces are exchangeable, which makes it the most robust prediction in this document. Falsified by: flat integrative complexity across topic domains within Balochistan respondents; or by the same domain gap appearing equally in Punjab, where the pedagogy does not differ by subject in this way.
P10b — the self-criticism prediction. Balochistan’s History 8 is the most internally critical book found in the current corpus: it states that “Muslims themselves were responsible for this decline”, calls followers of the Hijrat movement “simple-minded”, and rejects collective blame for 1857. Its Pakistan Studies 9, one year later, closes the narrative again. A Balochistan pupil meets a narrowing between class 8 and class 9, where pupils elsewhere meet a consistent register. Predicted: Balochistan leavers will show greater within-person variance in narrative closure — more willing to attribute internal fault to historical Muslim actors, less willing on the Pakistan Movement specifically — than leavers of boards whose register is uniform. Falsified by: Balochistan leavers showing the same closure on both periods, which would indicate the class-9 book overwrites the class-8 one rather than coexisting with it.
Prior condition, unchanged and severe. Balochistan has the lowest school participation and completion in the country and the most severe conflict exposure. The dosage and selection problems of §6.1 and §7.1 are worse there than anywhere. It remains possible that no Balochistan sample can be assembled that would let any of the above be tested, and eight Balochistan books could not be found at all, breaking the class 6-to-10 sequence for English and Islamiyat. The predictions are now grounded in scores rather than in a demand profile alone; the sampling obstacle is untouched by that improvement.
4.5 The control dimensions — where the programme predicts nothing
P11. On the corpus-wide constants the four provinces should not differ: association of scientific achievement with men; the “Hindu teachers” account of 1971; absence of any account of Partition violence; absence of any economic critique of colonialism; the collapse of inquiry at matriculation. The whole programme is falsified by: large systematic provincial differences on these control dimensions in the absence of provincial differences on the dimensions where curricula diverge. That pattern would establish that whatever produces provincial attitude differences, it is not the books.
4A. Personality-level predictions
Everything above predicts attitudes — what a leaver thinks about 1971, about minorities, about the army. This section asks a harder and more contestable question: whether a decade of this treatment leaves marks at the level of personality structure rather than opinion content. Attitudes are cheap to change and cheap to fake. Structure is neither.
Three warnings before any of it.
First, the ages are exactly right, which is why the question is worth asking at all. Classes 6 to 10 span roughly ages eleven to sixteen. In Erikson’s scheme that is the whole of identity versus role confusion — the stage at which a person assembles a durable account of who they are and what they belong to. Whatever a school system does in those five years, it does it to a person in the middle of constructing themselves. That is not a claim about Pakistan; it is why adolescence is studied at all.
Second, this is the section most likely to be wrong. Attitude predictions rest on the observation that examinable propositions are compelled. Personality predictions require an additional and much larger inferential step: that compelled cognitive operations, repeated for years, generalise into dispositions. That step is plausible, it is not established in this population, and nothing in this corpus can establish it.
Third, the psychoanalytic material is flagged separately in §4A.7 and held to a different standard, because psychoanalytic constructs are notoriously able to absorb any result. Each is given an operationalisation and a falsifier or it is not stated.
4A.1 Identity foreclosure — the central prediction
Mechanism. Marcia’s operationalisation of Erikson distinguishes four identity statuses by two variables: whether a person has explored alternatives, and whether they have committed. Commitment without exploration is foreclosure — an identity adopted whole from authority, typically stable, typically defended, and characteristically brittle when challenged.
A curriculum that supplies conclusions, requires their reproduction under examination, and sets no task requiring the pupil to reach an unsupplied conclusion is, structurally, a foreclosure-producing technology in the ideological domain. It offers commitment and removes exploration. That is the definition.
PP1 — the domain-specificity prediction, and the sharpest personality test available. Public-stream leavers of all four boards will show elevated foreclosed identity status relative to achieved status — but only in the ideological and political domains, and not in the occupational, interpersonal or lifestyle domains, where the curriculum sets no comparable demand.
The domain-specificity is what makes this a curricular prediction rather than a cultural one. A society that produces foreclosure through family structure, religious authority and economic constraint should produce it across all domains at once. A school system that forecloses only the political domain should produce a jagged profile.
Instrument: the Extended Objective Measure of Ego Identity Status (EOM-EIS-2), scored per domain rather than aggregated — the aggregation is what would hide the effect. Falsified by: foreclosure elevated uniformly across all domains, which would indicate a general cultural pattern with no curricular signature; or by no elevation in the political domain at all.
PP1a — the gradient. If PP1 holds, the political-domain foreclosure gap should be largest in Punjab (highest narrative closure, near-zero open demand in Pakistan Studies) and smallest in Sindh (the only board that examines civic judgement). Balochistan should sit oddly: high argumentative fluency with foreclosed commitments, the profile P10 predicts. Falsified by: the ordering failing to track assessment demand, or tracking content saturation instead — which would relocate the mechanism from demand to exposure and undercut §3.1.
4A.2 Need for cognitive closure
Mechanism. Kruglanski’s construct describes a preference for a definite answer over ambiguity, with two phases: seizing on an early answer and freezing against revision. The justify-that template trains precisely this sequence — the answer arrives first, and the pupil’s work is to consolidate rather than to test it.
PP2. Public-stream leavers will score higher on need for closure than a matched private-stream or madrassa comparison group, with the effect concentrated in the closed-mindedness and preference for order subscales and absent from the decisiveness subscale.
That split matters. Decisiveness is a general disposition toward acting on judgement, and there is no reason curriculum would raise it. Closed-mindedness is specifically a reluctance to entertain alternatives, which is what a decade of pre-supplied conclusions would train. Instrument: Webster–Kruglanski Need for Closure Scale, subscale-scored. Falsified by: uniform elevation across subscales including decisiveness — a general trait, not a trained one.
4A.3 Ingroup glorification versus ingroup attachment
This is the distinction on which the whole “what you sow” argument turns, and it is routinely collapsed in public debate about textbooks.
Roccas and colleagues separate two things usually called patriotism. Attachment is love of one’s country and a sense of belonging to it. Glorification is the belief that one’s group is superior and that its symbols and authorities deserve deference. The two are empirically separable, and the consequential finding is that glorification, not attachment, predicts denial of ingroup wrongdoing and reduced willingness to acknowledge harm done by one’s own group.
The corpus speaks directly to this. Patriotism appears in roughly equal volume across the boards — that was one of the earlier lens findings. What differs is its kind: Punjab’s is martial and sacrificial, Sindh’s is civic and heritage-based, KP’s is concentrated in one subject.
PP3. The four boards will differ substantially on glorification and not meaningfully on attachment. Punjab leavers highest on glorification; Sindh leavers similar to Punjab on attachment but materially lower on glorification. Instrument: Roccas et al. glorification/attachment scales. Falsified by: the two moving together across provinces — which would mean the curricula differ in how much patriotism they teach rather than in what kind, and would make the entire “kind of patriotism” finding decorative.
PP3a — the consequence that makes PP3 worth testing. If glorification tracks A5 militarism, then willingness to acknowledge Pakistani state wrongdoing — in 1971, in Balochistan, toward minorities — should be predicted by glorification score and not by attachment score, within province as well as between. A patriot who loves the country without believing it superior should acknowledge harm at much higher rates. Falsified by: acknowledgement tracking attachment, or tracking neither.
4A.4 Moral foundations
Mechanism. Haidt’s framework separates individualising foundations (care, fairness) from binding foundations (loyalty, authority, sanctity). Curricula that foreground group membership, deference to authority and the sacredness of national and religious symbols load the binding foundations.
PP4. Public-stream leavers will show elevated binding-to-individualising ratios, with the gap tracking A5 and A6 by province, and with sanctity elevated specifically in the boards whose religious content sits outside religious-studies subjects — the A1 measure — rather than in the boards with the most Islamiyat hours.
The second clause is the interesting one, because it predicts that where religion is placed matters more than how much of it there is. Islamiyat is expected and bounded; religion inside a geography book is unbounded and frames the world. Instrument: Moral Foundations Questionnaire. Falsified by: sanctity tracking total religious instruction hours rather than A1 placement.
4A.5 Political self-efficacy
Mechanism. Bandura’s efficacy is built by mastery experience — doing a thing successfully. A pupil repeatedly asked to evaluate a policy, take a position, justify a view of their own is being given rehearsal in political judgement. A pupil asked when nationalisation was announced is not.
PP5. Internal political efficacy — the belief that one is competent to understand and judge politics — will be highest among Sindh leavers, the only board that examines civic judgement, and lowest among KP leavers, whose Pakistan Studies volumes set zero evaluative demand across both media.
Critically, external efficacy — the belief that the system responds — should show no such pattern, because nothing in any curriculum addresses it. Instrument: standard internal/external political efficacy items, reported separately. Falsified by: the two efficacies moving together, which would indicate a general political-culture effect rather than a rehearsal effect.
4A.6 Integrative complexity, and the within-person test
Mechanism. Integrative complexity measures whether a person’s reasoning differentiates multiple dimensions of an issue and integrates them. It is scored from free text, is not self-reported, and is therefore much harder to fake than any scale above.
PP6 — the strongest design in this document, because it needs no comparison group at all. Within a single Balochistan respondent, integrative complexity on economic and environmental questions will exceed integrative complexity on historical and political questions. The same person, the same sitting, two topic domains that their curriculum treated with opposite pedagogy — open debate in geography, justify-that in history.
No cross-province comparison. No assumption that provinces are exchangeable. No selection problem, because each respondent is their own control. Falsified by: flat complexity across domains within respondents; or by the same domain gap appearing in Punjab, where the pedagogy does not split by subject in this way.
4A.7 The psychoanalytic material, held to a stricter standard
Psychoanalytic constructs are included because they name something the cognitive measures above do not, and excluded from the main programme because they are the easiest thing in this document to confirm accidentally. Each is given a behavioural operationalisation or it is not stated at all.
Splitting, and the capacity for ambivalence. In Klein’s account, immature representation keeps good and bad objects separate; maturity is the capacity to hold both in one figure. The corpus’s national narratives are, structurally, split objects: founders without flaws, adversaries without merits. This project found almost no ambivalent portrayal of any national figure in any board.
PP7 — the ambivalence test. Asked to state one criticism of a revered national figure and one merit of a historical adversary, public-stream leavers will produce fewer of both, and with longer latency, than a matched comparison group. The prediction is specifically about capacity to generate, not about endorsement — a respondent may hold the criticism and still find it unavailable. Falsified by: ready production of ambivalent content; or by the effect appearing on adversaries but not on ingroup figures, which would indicate ordinary group loyalty rather than a representational limit.
Collective shame and narrative closure. Where a national defeat is framed as external betrayal rather than internal failure, the framing does work that a defence does: it makes the event narratable without shame. 1971 is the corpus’s test case, and the “Hindu teachers” explanation survives in book after book across two decades of editions.
PP8 — the asymmetric recall test. In free recall of 1971, public-stream leavers will produce the conspiracy elements at high rates and the surrender — the fact of it, its date, its scale — at markedly lower rates than the conspiracy elements, and lower than their recall of comparably detailed material from the same syllabus. The prediction is about a within-topic asymmetry, not about general forgetting. Falsified by: symmetric recall of both elements; or by an equivalent asymmetry appearing for a neutral topic of similar syllabus prominence, which would indicate ordinary salience effects.
What is deliberately not claimed. No prediction is offered here about superego formation, identification with authority figures, or unconscious aggression, because this project can construct no operationalisation of them that a contrary result could break. They may well be the right description. They are not, in this document, a testable one, and stating them anyway would be the exact failure mode this section exists to avoid.
4A.8 What would sink the personality programme entirely
The attitude predictions in §4 can fail individually. The personality programme has a single load-bearing assumption, and one result would end it:
If public-stream leavers show elevated foreclosure, closure need and low complexity uniformly across all domains — political, occupational, interpersonal alike — then the effect is not curricular. It is a general feature of an authority-dense upbringing, of which school is one input among many, and the curriculum is a symptom rather than a cause.
That is the null this section should be tested against, and it is entirely plausible. Everything in §4A is designed so that the pattern rather than the level carries the claim: jagged profiles implicate the curriculum, flat profiles exonerate it.
5. The vintage test — a sharper design than cross-province comparison
The corpus’s most useful structural fact for this purpose has nothing to do with geography.
Four curriculum generations are in print simultaneously. Punjab runs a uniform post-2022 series and has since begun reissuing class 10 under a further edition label dated 2026. Sindh has not adopted the Single National Curriculum at all — Social Studies 6–7 on the Sindh Curriculum of 2014/2015, Social Studies 8 an early-2000s federal reprint, Pakistan Studies a 2004 approval still in print. KP is predominantly National Curriculum 2006 with 2010–2018 approvals reprinted through 2025–26, with an SNC-era exception at class 8 covering six subjects. Balochistan is on a 2022–23 federal curriculum printed as a single-year pilot.
The edition audit established that this is structural, not an artefact of collection: the corpus holds the newest edition of every book in the matrix, and the recent-looking year stamps on KP books are print runs, not editions — its Pakistan Studies 9 is a 2010-approved book printed for academic year 2024–25.
This creates a within-province design sharper than the cross-province one, because it holds constant almost everything §7 raises: ethnicity, language, conflict exposure, regional political history, family structure, media environment, labour market, and the population itself.
The design. Compare Punjab public-stream cohorts who completed classes 6–10 wholly before the SNC rollout against cohorts who completed them wholly after, using Sindh and KP cohorts over the same calendar window as controls — those provinces having reprinted rather than revised. That is a difference-in-differences with a natural control group and a plausible discontinuity in time.
The direction makes it diagnostic. Punjab’s revision moved the national narrative upward on closure. Vintage-matched pairs establish it: Punjab History 8 (SNC 2022) against KP History 8 (2022–23) scores A4 5 against 3; Punjab Pakistan Studies 9 (2023) against KP Pakistan Studies 9 (AY 2024–25) again 5 against 3; while the boards’ older books converge at A4 ≈ 3 with no gap. It is Punjab’s recent revision that diverges. Punjab runs the newest books in the corpus and produces the most closed narrative in it.
P12. Younger Punjab public-stream cohorts will show more closed narrative reproduction than older Punjab cohorts, while Sindh and KP cohorts show no corresponding shift over the same window. This runs against any generic account of generational liberalisation, against the assumption that internet exposure erodes official narratives, and against the assumption that curriculum reform moves in one direction. That is why it is worth testing.
P13. The same design applies to gender, with an unusually specific target. The class-9 collapse in female representation generalises across five of nine subject strands and inverts in four — and all four inversions are post-2020 revisions. The collapse is therefore not a pedagogical decision about adolescence; it is what happens when nobody revised the class-9 book. Prediction: cohorts meeting revised class-9 material will show measurably different gender-role attitudes from cohorts meeting unrevised material in the same province, and the difference will be confined to the revised strands.
Falsified by: no cohort discontinuity in Punjab across the SNC boundary; or a discontinuity of the same size in Sindh and KP, which would identify a national cause rather than a curricular one.
One serious caveat, developed in §6.4. If the SNC was a political programme rather than a pedagogical one, the 2022 boundary is not a clean instrument: it coincides with whatever else that programme did, and with the political conditions that produced it. The design is the best available. It is not clean.
6. Interacting variables — dosage, authority, activation, politics
These are not confounds to be conceded. They are part of the causal structure, and three of the four make the argument sharper.
6.1 Dosage is unequal, and unequal in a way that tracks the regions
The treatment is not administered evenly. School participation, completion and literacy differ sharply across the four provinces — with Balochistan lowest on most measures and Punjab highest — and differ again by sex and by rural versus urban residence, with girls in rural areas of the lower-participation provinces receiving least. The exact figures need proper sourcing and none is asserted here; the pattern and its direction are well established, and that is what the argument requires.
The consequence is analytical, not merely cautionary. Curricular exposure is itself a region-differentiated variable. A curriculum that reaches most children in one province and a minority in another is not the same treatment, whatever the books say. Every prediction in §4 must therefore be conditioned on it, and in some places the prediction should be weaker precisely because fewer children complete the course.
P14. Provincial attitude gaps attributable to curriculum should scale with completion: largest where the relevant cohort actually finished secondary schooling in the public stream, attenuated where it did not. Concretely, Balochistan’s predicted gaps (P9, P10) should be the smallest and noisiest of the four despite its distinctive assessment regime, and the male–female gap in curricular effect should be largest in the lowest-participation provinces. Falsified by: curricular gaps of equal size in high- and low-completion provinces, or an absence of any dose–response relation between years of public-stream schooling and frame reproduction. Absence of a dose–response relation is the single most damaging possible result for this programme, because a treatment with no dose–response is not behaving like a treatment.
6.2 Where alternative sources are thin, the textbook’s relative authority rises
For a child in a low-income rural household with few books, limited connectivity and a teacher who follows the text closely, the textbook may be the only sustained formal account of history they encounter. In an urban middle-class household with heavy media consumption, English-medium private schooling and adults who contest what the school says, the textbook is one voice among many and may be discounted heavily.
This predicts an effect that is larger in exactly the settings where measurement is hardest, and it converts a nuisance variable into a moderator with a sign.
P15 — a moderation prediction, testable within a single province. The association between curricular content and attitudes will be strongest among rural, low-income, low-media-access, public-stream respondents and weakest among urban, higher-income, high-media-access respondents. Because it is a within-province prediction it survives every objection about provinces not being exchangeable, and it is the cheapest strong test in this document. Falsified by: a flat or reversed moderation — curricular association equal across media-access strata, or stronger among the heavily mediated. Either would indicate that whatever produces the association is not the textbook’s informational monopoly.
6.3 Media and social media as activators rather than competitors
The weak hypothesis is that media outweighs textbooks. The stronger one is that schooling supplies the categories and media later activates them. A frame learned at eleven and reproduced in an examination at fourteen is available to be triggered by a news cycle at twenty-five. Where a textbook frame and a media frame align they compound; where they conflict, the textbook’s advantage is that it arrived first, from an authority, and was examined — the conditions under which early-acquired category beliefs are most resistant to revision.
This is a hypothesis, and it specifies evidence.
P16. Frame reproduction should be larger under activating conditions than at baseline — after a relevant news cycle, or under experimental priming — and the size of the activation effect should track curricular exposure. A respondent who never met the frame in school should be less activatable by the same cue than one who was examined on it. Falsified by: activation effects of equal size regardless of schooling history, which would show that media supplies the frame as well as the trigger.
P17 — the Urdu-medium alignment prediction. Pakistan’s social-media environment is large and substantially Urdu-language, and the corpus’s most ideologically loaded material sits disproportionately in the Urdu-medium books — Punjab’s and KP’s Islamiyat at every grade, KP’s History and Geography, KP’s Urdu reader with its two consecutive martyrdom lessons at class 7. Where the language of the school frame and the language of the media environment coincide, compounding should be greatest. Prediction: frames delivered in Urdu-medium subjects should show stronger later activation than frames delivered only in English-medium subjects, controlling for content. This prediction has an uncomfortable property that should be stated rather than hidden: the strand where compounding is predicted to be strongest is the strand this project cannot read. Nothing in the Urdu-medium corpus is quoted anywhere in this study, and P17 therefore rests on structural inference about which subjects carry which content, not on wording.
6.4 Politics — curriculum as a political artefact before a pedagogical one
This may be the most useful of the four, and it reframes the vintage finding.
The pattern of divergence does not look like administrative drift. Punjab adopted the Single National
Curriculum comprehensively and has begun a further reissue at class 10. Sindh did not adopt it at
all: the edition audit found zero occurrences of SNC, Single National Curriculum, National
Curriculum or NCP across the board’s entire public catalogue, whose book records expose only
Class, Medium, Year and Format and carry no curriculum field at all — consistent with a public
refusal, though the audit is careful to note it rests on portal metadata rather than on imprint
pages. KP adopted partially: six subjects at one grade, under NOC numbers literally carrying the
string SNC-2022, while everything else it prints remains 2006-curriculum. Balochistan prints a
federal curriculum as a single-year pilot and hosts a federal scheme as its governing document.
Four provinces, four distinct postures toward a federal curriculum initiative. If curriculum choice is an index of a province’s relationship with the centre, then curriculum is a political artefact before it is a pedagogical one, and the “what you sow” argument must reckon with who is sowing and why.
P18 — a hypothesis to be tested against the political record, not asserted from the corpus. Curricular revisions should track changes in centre–province alignment: adoption where provincial and federal control coincide, refusal or partial adoption where they do not, and the direction of textual change should follow the political programme rather than any pedagogical rationale. The corpus supplies suggestive support — Punjab’s revision moved closure upward, which no pedagogical account predicts — but the political record itself has not been examined here, and no party-specific claim is made. Falsified by: revisions that do not align with political shifts; by adoption postures that cut across centre–province relations; or by textual changes explicable on curricular-reform grounds alone.
The cost to §5. If P18 is right, the SNC boundary is a political discontinuity, not a purely curricular one. Any cohort effect found across it is jointly attributable to the books and to whatever else the political programme changed. The design remains the best available because Sindh and KP did not revise; but it is a quasi-experiment with a politically endogenous treatment, and it should be reported that way.
7. Confounds, and the conditional null
7.1 Selection into schooling
Only a fraction of Pakistani children complete secondary education, and attrition is not random: it varies by province, sex, rurality, income and conflict exposure, in ways that correlate with the very attitudes a survey would measure. A study of school-leavers is a study of survivors. If one province retains a smaller, more urban, wealthier, more male fraction to matriculation than another, a cross-province attitude comparison is partly a comparison of who was still in the room. The direction of that bias cannot be assumed; it must be modelled, and modelling it needs enrolment and completion data disaggregated further than most published sources go.
7.2 Streams, and the fact that this corpus covers one of them
Pakistani schooling is not one system. Public board schools, low-cost private schools, elite private schools preparing pupils for foreign examinations, and madrassas of several distinct traditions differ enormously in what they teach and whom they enrol. This corpus covers public-board textbooks only. Low-cost private schools often use commercial books; elite private schools meet the board syllabus, if at all, through other materials. Streams are unevenly distributed across provinces and between urban and rural areas. Any cross-province comparison not stratified by stream is measuring stream composition.
7.3 Teachers deviate, in both directions, and nobody has measured how much
A teacher may drill recall past anything the book requires, may open a question the book closes, or may skip the ideology chapter for time. Teacher absenteeism, class size, qualification, and the gap between the language of the book and the language actually spoken in class are all substantial and all unobserved. The corpus measures the printed treatment; the delivered treatment is a different quantity and may not correlate strongly with it.
7.4 Ethnicity, language, conflict and regional political history
The four provinces differ in ways that would produce large attitude differences with no curricular cause — and the differences point in the same directions the curricular hypothesis predicts.
- KP’s population has lived through sustained militancy, military operations and mass displacement. Direct experience of what an army does in one’s own districts is an obvious alternative explanation for any KP–Punjab militarism gap, and it predicts the same sign.
- Punjab is demographically dominant and supplies a disproportionate share of military recruitment. Family connection to service predicts higher army valuation, and again the same sign.
- Sindh has a long tradition of nationalist and left political organisation, a distinct party system, and in Karachi the country’s most cosmopolitan urban population. Higher civic engagement is exactly what that predicts, with no reference to a textbook.
- Balochistan has an ongoing insurgency, a history of state coercion, and the weakest educational infrastructure in the country.
Every one of these mimics the curricular prediction. The within-province designs — the cohort test of §5 and the moderation test of §6.2 — exist because they are the only available ways to break the tie.
7.5 Non-random assignment, and reverse causation
Nobody assigned children to provinces. Worse, the arrow may run backwards: a board writes books in a political environment, staffed by people from it, under a government answerable to it. Punjab’s books may be more closed because Punjabi public opinion is, rather than the reverse. A cross-sectional correlation between curriculum content and population attitudes is what both causal directions predict. Only a temporal discontinuity can begin to separate them, and §6.4 has just argued that this particular discontinuity is itself political.
7.6 The null hypothesis, in its strong form
The proposition that textbooks matter far less than commentary assumes is not a rhetorical concession. It is live, respectable, and in one of its forms predicted by this document’s own central mechanism.
The dosage is small and unverified. A few hundred pages over five years, delivered inconsistently, in classrooms where the exercise sets may never be reached. The corpus measures print, not dose.
Didactic exposure is a weak instrument for durable attitude change. Attitudes acquired by being told things, in a low-involvement setting, by a source the recipient did not choose, tend to be shallow and to decay. Deep attitudes come from identity, social environment and experience.
The population supplies its own counter-evidence. Every dissenting Pakistani historian, journalist, lawyer and educationist was schooled in a version of these books. The informant whose recollection prompted this project’s moral-rescue thread was schooled in Punjab between 1989 and 1994 and emerged able to characterise and criticise what he had been taught. So was whoever wrote, and whoever approved, the Sindh teacher’s note quoted in §8.
And the strongest form of the null follows from the demand mechanism itself. If what a regime of recall, comprehension and justify-that transmits is a capacity — fluency in constructing reasons for supplied conclusions — rather than a conviction, the expected outcome is a population that can produce the official account on demand and does not particularly believe it. That is the reproduction/endorsement dissociation stated as a null result on belief: textbooks consequential for how people argue, inconsequential for what they hold. P10 is the direct test.
The sharper null, and the one this document actually adopts as its rival hypothesis, is not “textbooks don’t matter” but “textbooks matter conditionally — on dosage, on the density of competing sources, and on whether later media reinforces or contests the frame.” That version predicts where effects should be largest (P14, P15, P16) and is therefore easier to falsify than either the alarmist or the dismissive reading. The honest position is that the null on endorsement is most likely true, the null on available frames and fluency least likely, and the null on behaviour genuinely unknown.
8. The intergenerational argument, with its unobserved steps named
The “what you sow” question can be argued from structure. It cannot be argued from psychology, and the difference must stay visible.
The structural argument. An assessment regime is a selection mechanism. It ranks pupils, and the ranking determines who proceeds — to higher classes, universities, the civil service, teaching posts, the boards themselves. A regime composed predominantly of recall, comprehension and justification-of-supplied-conclusions ranks pupils by fidelity in reproducing given conclusions and fluency in constructing reasons for them. It does not rank them by capacity to reach a conclusion the syllabus does not hold, because it never elicits that capacity and cannot observe it. Whatever distribution of that capacity exists in a cohort, the examination is blind to it — neither rewarding nor penalising it, which across a career of examinations amounts to selecting against it relative to the traits the system does reward.
The pupils so ranked become, in some proportion, the next cohort’s teachers, examiners, curriculum officers and textbook authors. The federal curriculum directives of the 1980s and 1990s — including the instruction that the Ideology of Pakistan “be presented as an accepted reality, and be never subjected to discussion or dispute” — were written by people who had themselves been schooled and examined. So were the authors of the current books.
The unobserved steps. The argument requires four links, and this project observes one.
- Printed demand exists. Observed, page-cited, in four boards.
- Printed demand becomes set and marked demand. Unobserved. No paper that was actually sat has been consulted; no marking scheme; no script.
- Marked demand shapes pupils’ capacities and dispositions. Unobserved here, though the underlying learning mechanisms are well established independently.
- Those capacities shape the composition of the personnel who write and mark the next cycle. Unobserved, and the hardest link, since it runs through labour markets, patronage, provincial recruitment and political appointment, each with its own logic — and §6.4 has argued that the fourth link is political before it is educational.
What may be claimed. That a regime of this composition does not select for evaluative capacity; that the personnel staffing the next cycle are drawn from those it does select; and that this is a plausible reproduction mechanism.
What may not be claimed, and is not claimed here. That any individual teacher, examiner or author holds any particular view. That pupils schooled this way are incapable of evaluation — the corpus contains its own refutation. That the regime determines outcomes rather than shaping incentives. Or that one mechanism explains an educational culture in which examination boards, teacher training, class size, language policy and political economy all operate at once.
And the mechanism is not sealed. Sindh’s Social Studies 6 teaches historiography as a practice: pupils record their own accounts of a shared day, compare them in groups, and are asked how the accounts differ. The teacher’s note attached to that exercise directs the teacher to help pupils understand “that history is always incomplete”, that “histories always involve interpretation”, and to ask them “what if the version told by students of Class X, was declared to be the only ‘true’ and officially accepted story?” (Sindh, Social Studies 6, p. 11). Set against the 1995 directive that the ideology may never be disputed, the two documents are answering each other — and the answering one is currently in classrooms.
That page is the strongest single reason to state the intergenerational argument as a tendency rather than a trajectory. A system that reliably reproduced itself could not have produced it.
9. How this could be tested
A curriculum-keyed item bank. The essential first step, and it does not exist. Items whose content maps to specific page-cited propositions in specific books, so that “treatment” is a measurable quantity for a given respondent rather than a province label. Every item in paired form: a reproduction version (“according to what you were taught, why did East Pakistan separate?”) and an endorsement version (“what do you think caused it?”). Without that pairing a study cannot distinguish exposure from belief.
Stratified attitude surveys of school-leavers, stratified by province, school stream, cohort year, urban/rural, sex, medium of instruction actually experienced, and — per §6.1 and §6.2 — completed years of schooling and media access. Stream and dosage stratification are not optional.
Cohort comparison across curriculum changes (§5), requiring accurate rollout dates by district and subject, and benefiting from any respondent who can be asked which book they used.
Matched comparisons across vintages — siblings differing by a few years across a revision boundary; classmates in adjacent districts where rollout differed. Within-family comparison removes an entire class of confound at the cost of power.
Examination-script analysis. The highest-value unobserved dataset in this subject. The corpus observes printed demand, not realised demand. Archived scripts and marking schemes from one board for one year would show which items were actually set, what pupils wrote, and what markers rewarded. A demand that is printed and never set is not a treatment.
Classroom observation and teacher interview, to measure deviation in both directions and to establish whether exercise sets are reached at all. These books routinely carry more exercise material than any class could complete.
Longitudinal panel. Following one cohort from class 6 to class 10 with repeated measurement is the only design that shows change rather than difference. Nothing else substitutes for it.
Indirect and implicit measures, because respondents in a school-adjacent context give school-adjacent answers: category-association tasks, allocation decisions, third-person framings.
Priming and activation studies (§6.3), to test whether curricular exposure moderates the size of a media cue’s effect.
Triangulation with existing survey series. Large attitude surveys of Pakistani youth exist and have been run repeatedly by domestic and international organisations. A serious version of this programme should locate them, obtain microdata where possible, and check whether provincial patterns already visible in them match or contradict the predictions above. No finding or figure from any such survey is cited in this document, because none has been verified by this project.
10. Limits
No mental state is observed anywhere in this project. A pupil may write the keyed answer without holding it, hold it without having been taught it, or have been taught its opposite by a teacher. Reading is not determined by text and neither is answering.
Demand is observed; uptake is not. Whether an exercise is set, marked, or reached at all in a school year is unknown.
A large part of the treatment is invisible. Quotation is restricted to English-medium files. Punjab and KP teach Islamiyat entirely in Urdu at classes 6–10; KP additionally teaches History and Geography in Urdu; Balochistan’s History 6 and Islamiyat volumes are Urdu. Sindh is the only board teaching Islamiyat in English and is therefore the most visible board to this method — so any finding that Sindh carries more of something must be read against that. This bears directly on P17, whose predicted effect is largest in exactly the material that cannot be read.
Absence claims are bounded by extraction quality. Nastaliq OCR here is theme-grade: it recovers structure, proper nouns and dates, not wording, and breaks ligatures unpredictably. In one KP file demonstrably containing a chapter of Sufi hagiography, a high-frequency honorific returns one hit and a high-frequency name returns none. A zero token count in an Urdu file is not evidence of absence. Urdu absences are reported as “not established”, never “absent”.
Vintage confounds every board comparison, and Punjab cannot be vintage-matched against Sindh or KP in most subjects because the matching books do not exist. The honest form of most board statements is not “Punjab is more X than Sindh” but “Punjab’s current books are more X than the books Sindh is currently printing” — which is, as it happens, the policy-relevant comparison, since it is what pupils receive this year.
Balochistan is now scored on all seven axes, and §2 and §4.4 have been revised accordingly. Three limitations survive that revision and bear on every Balochistan prediction here: eight of its books could not be found at all, breaking the class 6-to-10 sequence for English and Islamiyat; twelve of its twenty-eight volumes are Urdu-medium and were read at theme grade only; and its classes 6–9 books are pilot editions printed for a single academic year, so the treatment they represent may not persist. A prediction about Balochistan is a prediction about a pilot, on an incomplete corpus, in the province with the country’s lowest school completion.
The psychological mechanisms are described, not measured. No study is cited, no effect size given, no claim made about the magnitude of any mechanism in this population. Anyone wishing to attach magnitudes must go to the primary literatures; this document deliberately does not, because a plausible-looking fabricated citation would be more damaging than an unadorned description.
The socio-economic, media and political variables in §6 are argued, not measured. No participation figure, literacy rate, media-penetration statistic or election result is asserted anywhere. The patterns invoked are described directionally and flagged as requiring proper sourcing. P18 in particular is a hypothesis about the political record, and the political record has not been examined here; no party-specific claim is made.
Every prediction in this document is untested. None has been checked against any survey, dataset or observation. They are stated in falsifiable form so that they can be checked, and should be regarded as hypotheses of unknown quality until they are.
The central inference is comparative, not causal. Even a perfect cross-province study establishes covariation. The cohort design of §5 is the strongest available approach to causation and remains a quasi-experiment with a politically endogenous treatment.
11. What would make this programme worth running
Score Balochistan. Twenty-four catalogued records with empty score objects — and four further books on disk that the catalogue does not yet list — are the largest single gap in the comparative design. A fourth scored board gives the natural experiment its fourth arm.
Get one board’s examination scripts. Printed demand is the strongest variable this project has, and it is a proxy. A single archive of sat papers, marking schemes and scripts from one board for one year would convert the spine of this argument from an inference about print into an observation about practice — or refute it.
Run the within-province tests first. The moderation test of §6.2 is cheap, needs no cross-province comparison, and survives every objection about provinces not being exchangeable. The cohort test of §5 is more expensive and more perishable: Punjab’s revision boundary currently sits between two provinces that did not revise, and that configuration will not last. Once Sindh or KP revises, the control group is gone, and the sharpest test available to this subject goes with it.
All corpus quotations are from English-medium books, cited by board, title, class and page as recorded
in the project’s extraction, where page numbers are PDF page indices except where marked f. for printed
folio. Scored figures are from the synthesis, scored against the rubric;
assessment-demand figures and quotations from [article assessment demand](/studies/article-assessment-demand/); othering, civics,
gender and history quotations from the corresponding lens study files; thread findings from
[thread moral rescue](/studies/thread-moral-rescue/) and [thread iconoclasm](/studies/thread-iconoclasm/); vintage and adoption findings from
the edition audit; catalogue facts from analysis/books.json. Emphasis inside quotations is
the analyst’s, not the textbooks’. No psychological study, survey result or statistic external to this
project is cited anywhere in this document, and none should be inferred.