Skip to content

Psychology · Ch 2 — Methods of Enquiry in Psychology

Psychological Testing

2.5

Psychological Testing

Assessing individual differences has been one of psychology's central concerns since the discipline began, and psychologists have accordingly built specialised tests for assessing characteristics such as intelligence, aptitude, personality, interest, attitudes, values, and educational achievement. These tests are put to varied uses, personnel selection, job placement, training, career guidance, and clinical diagnosis, across settings as different as schools, guidance clinics, industries and the armed forces. A test itself is made up of a number of items (questions) along with their possible responses, each item tied to the particular characteristic the test is meant to measure; the characteristic being tested must be defined clearly and unambiguously, with every item genuinely relevant to that characteristic and to no other, and a test is often designed for a specific age group and may or may not carry a fixed time limit. Formally, a psychological test is a standardised and objective instrument used to assess where an individual stands relative to others on some mental or behavioural characteristic, and two words in that definition matter enormously. Objectivity means that if two or more researchers administer the same test to the same group of people, they should arrive at more or less the same scores for each person; this requires items to be worded so they mean the same thing to every reader, and requires the instructions for administering and scoring the test to be spelled out clearly and followed consistently. Constructing a usable test is a systematic process built around three technical properties. Reliability refers to how consistent a person's scores are across two different occasions or measurements: test-retest reliability is found by administering the same test to the same group twice, some time apart, and correlating the two sets of scores; split-half reliability instead checks internal consistency within a single administration, typically by dividing the test into odd-numbered and even-numbered items and correlating the two halves against each other, since items drawn from the same underlying domain should correlate well with one another. Validity asks the more fundamental question of whether the test actually measures what it claims to measure; a mathematics achievement test, for instance, needs to be shown to measure mathematical ability and not, say, language proficiency. Finally, a test becomes standardised once norms have been established for it; norms are the normal or average level of performance of a large, representative group of test-takers …