Solutions
2.1
- V: looks (fine), see (the case), focus → 3. Note "on paper it looks fine" is coded visual: the process described is apprehension by looking, not a fact about paper.
- A: tone → 1. "Walked me through it" is not auditory despite being about speech; the predicate is locomotive.
- K: feel (settled), walked me through, turning it over, rough (edge), put my finger on, slides (away) → 6. "Settled" is coded kinesthetic as a somatic-affective state; a defensible alternative is U, and the rule you write down matters more than which side you pick.
- U: clean (the numbers) → 1. "Clean" is cross-modal and carries no committed sensory content here.
Totals: V = 3, A = 1, K = 6, U = 1; classified N = 10.
Boundary rule stated: a predicate is coded by the modality of the process it names, not by the modality of the object it happens to be attached to. This is what excludes "walked me through it" from A and keeps "looks fine" in V.
2.2
(a) p̂_V = 30/80 = 0.375; p̂_A = 18/80 = 0.225; p̂_K = 32/80 = 0.400. (Check: 0.375 + 0.225 + 0.400 = 1.000 ✓)
(b) SE = √(p̂_K(1 − p̂_K)/N) = √(0.400 × 0.600 / 80) = √(0.24/80) = √0.003 = 0.05477
(c) 0.400 ± 1.96(0.05477) = 0.400 ± 0.10735 → [0.293, 0.507]
The interval spans from below the equiprobable value of 0.333 to above half. Eighty predicates does not resolve this speaker's kinesthetic proportion to anything useful.
2.3
Expected count per cell: E = 80/3 = 26.667.
V: (30 − 26.667)² / 26.667 = 11.109 / 26.667 = 0.41660
A: (18 − 26.667)² / 26.667 = 75.117 / 26.667 = 2.81690
K: (32 − 26.667)² / 26.667 = 28.441 / 26.667 = 1.06654
χ² = 0.41660 + 2.81690 + 1.06654 = 4.3000
df = 3 − 1 = 2. Using p = exp(−χ²/2):
p = exp(−2.1500) = 0.1165
Not significant at α = 0.05. Even against the naive null this speaker does not separate from equiprobability.
2.4
Expected counts: E_V = 80(0.48) = 38.40; E_A = 80(0.19) = 15.20; E_K = 80(0.33) = 26.40. (Check: 38.40 + 15.20 + 26.40 = 80.00 ✓)
V: (30 − 38.40)² / 38.40 = 70.56 / 38.40 = 1.83750
A: (18 − 15.20)² / 15.20 = 7.84 / 15.20 = 0.51579
K: (32 − 26.40)² / 26.40 = 31.36 / 26.40 = 1.18788
χ² = 1.83750 + 0.51579 + 1.18788 = 3.5412
p = exp(−1.7706) = 0.1703
Also not significant — and notice that χ² fell relative to 2.3, from 4.30 to 3.54. This speaker is closer to their peers than to a uniform distribution.
Conclusion: neither test licenses classifying this speaker into a type. The comparison between them adds something the first test alone conceals: the speaker's departure from uniformity, such as it is, is the departure that everyone discussing this topic shows. Whatever variation exists here belongs to the topic and the language, not to the individual.
2.5
(a) N = (1.96)²(0.25)/(0.04)² = 3.8416 × 0.25 / 0.0016 = 0.9604 / 0.0016 = 600.25 → 601 classified predicates
(b) words = 601 / 0.023 = 26,130 words; minutes = 26,130 / 150 = 174 minutes (about 2 hours 54 minutes of continuous speech from one person)
(c) Because p(1 − p) is maximised at p = 0.5, where it equals 0.25. With no prior information, using p = 0.5 gives the largest required N and therefore guarantees the margin of error will be no worse than 0.04 whatever the true proportion turns out to be. It is the conservative choice, and the only defensible one when the quantity you are sizing for is the quantity you are trying to estimate.
2.6
Design effect:
D = 1 + (m̄ − 1)ρ = 1 + (9 − 1)(0.20) = 1 + 1.60 = 2.60
Effective standard error:
SE_eff = 0.05477 × √2.60 = 0.05477 × 1.61245 = 0.08832
Corrected interval:
0.400 ± 1.96(0.08832) = 0.400 ± 0.17311 → [0.227, 0.573]
Corrected sample size for Problem 2.5(a):
N_eff = 601 × 2.60 = 1,563 classified predicates
≈ 1,563 / 0.023 = 67,960 words ≈ 453 minutes ≈ 7.5 hours of continuous speech
Independence was doing an enormous amount of unearned work. Once you concede that a speaker inside a visual description tends to stay there for a few clauses — which is obviously true and is in fact the phenomenon the chapter is about — the cost of person-level measurement rises past any practical horizon.
2.7
(a) x̄ = 400/5 = 80. ȳ = (990 + 1580 + 2210 + 2790 + 3430)/5 = 11,000/5 = 2200.
x − x̄: −80, −40, 0, +40, +80
y − ȳ: −1210, −620, +10, +590, +1230
Sxy = (−80)(−1210) + (−40)(−620) + (0)(10) + (40)(590) + (80)(1230)
= 96,800 + 24,800 + 0 + 23,600 + 98,400 = 243,600
Sxx = 6,400 + 1,600 + 0 + 1,600 + 6,400 = 16,000
b = 243,600 / 16,000 = 15.225 ms/deg
a = 2200 − 15.225(80) = 2200 − 1218 = 982 ms
(b) rate = 1000 / 15.225 = 65.7 deg/s
(c) The two slopes differ by 16.05 − 15.225 = 0.825 ms/deg, a ratio of 1.054 — about five percent. The intercepts differ by 1004 − 982 = 22 ms, roughly two percent of a one-second baseline. Rotating a figure in the picture plane and rotating it in depth cost essentially the same. This is a substantive finding, not a null: a mechanism that manipulated a flat retinal image would find picture-plane rotation far cheaper, since it is a simple planar transform of what the eye received, while depth rotation requires recovering structure the image does not contain. Equal cost implies the representation being rotated is three-dimensional in both cases — the participant is turning a model of the object, not a picture of it.
2.8
(a) x̄ = 80. ȳ = (820 + 835 + 810 + 845 + 830)/5 = 4140/5 = 828.
y − ȳ: −8, +7, −18, +17, +2
Sxy = (−80)(−8) + (−40)(7) + (0)(−18) + (40)(17) + (80)(2)
= 640 − 280 + 0 + 680 + 160 = 1,200
Sxx = 16,000
b = 1,200 / 16,000 = 0.075 ms/deg
a = 828 − 0.075(80) = 828 − 6 = 822 ms
Over the full 160° range the fitted line rises by 0.075 × 160 = 12 ms, against a baseline of 822 ms.
(b) Fitted values: 822.0, 825.0, 828.0, 831.0, 834.0. Residuals: −2.0, +10.0, −18.0, +14.0, −4.0 (sum = 0 ✓).
SSE = 4 + 100 + 324 + 196 + 16 = 640
s² = 640/3 = 213.33, s = 14.605
SE(b) = 14.605 / √16,000 = 14.605 / 126.491 = 0.11546
t = 0.075 / 0.11546 = 0.6496, df = 3
|t| = 0.65 < 3.182, so the slope does not differ from zero. The 95% interval is 0.075 ± 3.182(0.11546) = 0.075 ± 0.367 → [−0.292, +0.442] ms/deg, which comfortably includes zero and excludes anything near 16.
(c) Because everything except the question has been held fixed. Same participants, same stimuli, same angles, same eyes, same fingers, same keys. The only thing that changed is what the participant was asked to find out — and the slope collapsed from 16.05 to 0.075 ms/deg, a factor of over two hundred.
This is the chapter's central distinction, measured. Modality-specific processing is unambiguously real: the 16 ms/deg is not an artifact, because the same people in the same chair produce 0.075 ms/deg when the question stops being spatial. But the variable that controls it is the task, not the person. A trait account has to explain how the same participant's "preferred system" changed two-hundredfold in the time it took to read a different instruction. It cannot, and that is why this null result — a null about a slope, not about an effect — does more damage to the preferred-representational-system claim than the entire outcome literature.
2.9
(a) ρ_max = √(r_XX · r_YY) = √(0.20 × 0.80) = √0.16 = 0.40
(b) ρ_TU = r_XY / √(r_XX · r_YY) = 0.09 / 0.40 = 0.225
(c) The correction divides by an estimated quantity, so it multiplies the sampling error of r_XY by 1/0.40 = 2.5 as well as the point estimate. If the observed r_XY = 0.09 came from a small sample, its own confidence interval might run from −0.10 to +0.28; disattenuating gives a corrected interval of roughly −0.25 to +0.70, which is not an estimate of anything. More fundamentally, r_XX = 0.20 means that 80% of the observed variance in X is error — and when the construct X is defined as a stable trait, that fact is a finding about the construct, not a nuisance to be algebraically removed. Disattenuation asks what the correlation would be if the trait were stable. The trait is not stable. Correcting for it is answering a question about a world we do not live in.
2.10
(a) z_r = ½ ln[(1 + 0.12)/(1 − 0.12)] = ½ ln(1.12/0.88) = ½ ln(1.272727) = ½(0.241162) = 0.120581
(b) Required: z_r √(n − 3) ≥ 1.9600 + 0.8416 = 2.8016
n − 3 ≥ (2.8016 / 0.120581)² = (23.234)² = 539.8
n ≥ 542.8 → n = 543 participants
(c) At n = 45:
z_r √(n − 3) = 0.120581 × √42 = 0.120581 × 6.48074 = 0.78145
power = Φ(0.78145 − 1.9600) = Φ(−1.17855) ≈ 0.119 → about 12%
(d) A study with 12% power fails to reject roughly seven times out of eight when the effect is exactly as large as you specified, so a null result from it carries almost no evidence against the effect — it is close to what the effect predicts. To conclude absence you need either adequate power against your smallest effect of interest, or an equivalence test with a stated equivalence bound, and n = 45 supports neither.
2.11
(a)
P(C) = (0.65)(0.30) + (0.40)(0.70) = 0.195 + 0.280 = 0.475
P(V | C) = 0.195 / 0.475 = 0.41053
LR⁺ = sensitivity / (1 − specificity) = 0.65 / 0.40 = 1.625
Check by odds: prior odds = 0.30/0.70 = 0.42857; posterior odds = 0.42857 × 1.625 = 0.69643; posterior probability = 0.69643/1.69643 = 0.41053 ✓
The prior was 0.30 and the posterior is 0.41. Observing the cue moves you eleven points — and leaves you at "probably not."
(b) Also needed: P(not C) = 1 − 0.475 = 0.525, and
P(V | not C) = (0.35)(0.30) / 0.525 = 0.105 / 0.525 = 0.200
Entropies:
H(0.30) = 0.30(1.736966) + 0.70(0.514573) = 0.521090 + 0.360201 = 0.881291 bits
H(0.41053) = 0.41053(1.284467) + 0.58947(0.762555) = 0.527307 + 0.449524 = 0.976831 bits
H(0.20) = 0.20(2.321928) + 0.80(0.321928) = 0.464386 + 0.257542 = 0.721928 bits
E[H] = 0.475(0.976831) + 0.525(0.721928) = 0.463995 + 0.379012 = 0.843007 bits
I = 0.881291 − 0.843007 = 0.038284 bits
I / H(prior) = 0.038284 / 0.881291 = 0.0434 → 4.3%
Roughly four hundredths of a bit. Even granting sensitivity and specificity well above what any test of these claims has produced, the cue removes about one twenty-third of the uncertainty it is supposed to resolve.
(c) k = 6 × 4 = 24 tests.
P(at least one) = 1 − (0.95)^24
ln(0.95) = −0.0512933; 24 × (−0.0512933) = −1.2310
e^(−1.2310) = 0.29199
P = 1 − 0.29199 = 0.708
For P ≥ 0.90:
1 − (0.95)^k ≥ 0.90 ⟺ (0.95)^k ≤ 0.10
k ≥ ln(0.10) / ln(0.95) = (−2.302585)/(−0.0512933) = 44.89 → k = 45 tests
At 24 cells, a completely empty chart yields at least one apparent confirmation about 71% of the time. This is the arithmetic of how a false claim accumulates a literature of supportive anecdotes without any of the anecdotes being fabricated.
2.12
(a) The two predictions.
The task-modality hypothesis says performance is governed by whether the task's representational demand collides with the response channel's demand. It therefore predicts a crossover interaction between task modality and response modality, with no requirement of any main effect at all: spatial-imagery tasks suffer under manual/pointing responses (both compete for spatial resources) and run well with spoken responses; verbal tasks show the reverse. In ANOVA terms it loads on the task × response interaction, and the interaction should be large relative to everything else.
The trait-PRS hypothesis says each person has a dominant channel. It therefore predicts a person × task-modality interaction — visuals do better on the imagery task, auditories on the verbal one — and, critically, it predicts that this interaction has substantial variance, is stable across sessions, and that person's channel assignment replicates on retest. It loads on the person × task variance component.
These are distinguishable because they make opposite predictions about what happens when you change the task for a fixed person. The task-modality account predicts the same person swings hugely; the trait account predicts each person carries their advantage from cell to cell. There is no reading of the data on which both large effects appear, and the design forces the question.
(b) Variance decomposition.
total variance = σ²_person + σ²_task + σ²_error = 0.4 + 8.6 + 3.0 = 12.0
proportion attributable to person = 0.4 / 12.0 = 0.0333 → 3.3%
Task modality accounts for 8.6/12.0 = 71.7%; error for 25.0%. The task explains roughly twenty-one times as much variance as the person does.
(c) Reliability of a person-level score.
R(1) = (1)(0.4) / [(1)(0.4) + 3.0] = 0.4 / 3.4 = 0.1176
R(10) = (10)(0.4) / [(10)(0.4) + 3.0] = 4.0 / 7.0 = 0.5714
R(40) = (40)(0.4) / [(40)(0.4) + 3.0] = 16.0 / 19.0 = 0.8421
Solving R(k) ≥ 0.80:
k(0.4) / [k(0.4) + 3.0] ≥ 0.80
0.4k ≥ 0.8(0.4k) + 0.8(3.0)
0.4k − 0.32k ≥ 2.4
0.08k ≥ 2.4
k ≥ 30
Thirty independent task observations per person to measure their preferred system to a reliability of 0.80. A single observation — which is what a practitioner listening to a client for a few minutes has — yields R = 0.118, indistinguishable from the κ ≈ 0.18 of Worked Example 2.3, and arrived at from a completely different direction.
(d) The verdict.
The preferred-representational-system claim was not disproved in the way a false existence claim is disproved; a small, stable person-level component is present, at roughly three percent of variance, and there is no evidence it is zero. What was disproved is the magnitude the practice required — the person factor is dwarfed by the task factor by more than twenty to one, so anything you could learn about a person from listening to them is swamped, in the very next sentence, by what they are talking about. Measuring the trait accurately enough to act on it costs thirty structured observations per person, which is more than the resulting advantage is worth even if you were willing to pay it — and this is why the honest replacement for "match their system" is not "match nothing" but the easier, better-supported and entirely free instruction: match the sentence in front of you.
That instruction needs no typology, no chart, no two-week retest, and no research programme to rescue it, because it makes a claim about an utterance rather than a person — and utterances, unlike people, wear their modality on the surface where anyone can hear it.