Almost every survey-based thesis contains the same sentence: “the instrument was validated and found reliable.” Supervisors read it dozens of times a year, and it often stands for very little: a Cronbach’s alpha copied from another study, a questionnaire shown to one lecturer, no pretest. This guide is a checklist for the supervisor, programme director or methods office that wants students to do the work behind the sentence. It sets out six steps, each with an owner and the artefact it produces, and it ends with a model supervisor rubric and a fictional worked example in which every number can be checked by hand.
What students confuse: validity, reliability and the statistic that stands in for both
Reliability asks whether an instrument gives consistent results. Validity asks whether it measures what the thesis says it measures. The two are related but not interchangeable, and the most common error supervisors see is treating one reported statistic as evidence for both. Taber’s review of how Cronbach’s alpha is used in science education research, published in Research in Science Education 48(6), 1273–1296, found that authors often quote alpha values without adequate explanation, that interpretations of an acceptable threshold vary widely, that alpha can look acceptable despite recognised instrument problems, and that it is sometimes misused to claim unidimensionality. It concludes that a high alpha provides limited reliability evidence. The original source for the statistic is Cronbach’s 1951 paper, “Coefficient alpha and the internal structure of tests”, Psychometrika, 16(3), 297–334.
The practical message for supervision is that a number is not a method. Students should be taught to report what was done, with whom and what the result does and does not show.
The checklist at a glance
| Step | Owner | Artefact produced |
|---|---|---|
| 1. Define the construct and its dimensions | Student, checked by supervisor | One-page construct definition with cited source |
| 2. Check the instrument’s origin and permission | Student | Instrument record: authors, year, version, licence or permission |
| 3. Judge content with an expert panel | Student convenes; supervisor approves panel | Panel brief, rating sheet, summary table |
| 4. Pretest with a small sample from the target population | Student | Pretest log: wording changes, timings, missing data |
| 5. Compute reliability on the thesis’s own sample | Student, reviewed by methods office | Reliability table with the coefficient and its caveats |
| 6. Write the methods paragraph | Student; supervisor signs off | Methods text that states procedure, not conclusions |

Step 1: define the construct and its dimensions
Nothing downstream works if the construct is vague. Require a one-page definition that names the construct, cites the source of the definition, lists the dimensions the thesis will measure and says what is out of scope. This is the same discipline described in the guide to an alignment and operationalization review, and it prevents the situation in which an instrument measures an adjacent concept.
Step 2: check the instrument’s origin and permission
If the student is using an existing instrument, the record should include authors, year, version, the population it was developed on and the route to permission. A faculty can compare how it governs such records with the approach in the piece on a validated teacher self-efficacy instrument bank. Any change to the original, including dropped or reworded items and translation, must be listed, because the adapted instrument is a new instrument and its evidence has to be gathered again.
Step 3: judge content with an expert panel
Content validity is usually evidenced by asking subject experts to judge whether items represent the construct. One widely used method is Lawshe’s content validity ratio, from C. H. Lawshe, “A quantitative approach to content validity”, Personnel Psychology, 28, 563–575 (1975). Each panellist rates every item as “essential”, “useful, but not essential” or “not necessary”. The ratio is CVR = (ne − N/2) / (N/2), where ne is the number of panellists rating the item essential and N is the panel size. It ranges from −1 to +1, and the published critical values depend on the panel size: .99 for five to seven panellists, .75 for eight, .78 for nine, .62 for ten, .42 for twenty and .29 for forty.
Supervisors should ask students to document who the experts are and why they qualify, to give them a written brief, to collect ratings independently and to keep the rating sheet. Panels of five to seven can only retain items on which almost everyone agrees, which is a practical reason to recruit eight or ten.
A fictional worked example
The following is illustrative only: no real study is described. A panel of eight experts rates a candidate item for a fictional postgraduate supervision-satisfaction scale. Seven rate it essential.
| Quantity | Value | Working |
|---|---|---|
| Panel size N | 8 | |
| Rated essential, ne | 7 | |
| CVR | 0.75 | (7 − 4) / 4 |
| Critical value for N = 8 | 0.75 | From Lawshe’s table |
| Decision | Retain | CVR meets the critical value |
Had only six of the eight rated the item essential, the ratio would be (6 − 4) / 4 = 0.50, below the critical value, and the item would be revised or dropped. A student should present such calculations in an appendix and say how experts’ comments changed the wording.
Step 4: pretest with a small sample from the target population
A pretest tests the practicalities that experts cannot judge: whether respondents understand the wording, how long completion takes, which items are skipped and whether the response format works. The pretest should sample from the population the thesis will study, not from classmates. The student should keep a log recording every wording change and why, and should not pool pretest responses with the main sample unless the instrument was unchanged afterwards, and the thesis says so.
Step 5: compute reliability on the thesis’s own sample
Reliability is a property of scores from a sample, not of an instrument in the abstract. Published coefficients are background; the thesis must report its own. The formula for Cronbach’s alpha is α = k/(k−1) × (1 − Σσi² / σt²), where k is the number of items, σi² the variance of each item and σt² the variance of the total score.
A fictional illustration: a five-item scale has item variances summing to 6.0 and a total-score variance of 15.0. Alpha is 5/4 × (1 − 6.0/15.0) = 1.25 × 0.6 = 0.75. The arithmetic is simple; the interpretation is not. Taber’s review is the reason a supervisor should ask three further questions: is the scale meant to measure one dimension, has dimensionality been checked, and is the threshold the student cites justified in the thesis rather than simply asserted? Other reliability evidence, such as repeating the instrument with the same respondents after an interval, answers a different question about stability and needs its own design and sample.

Step 6: write the methods paragraph
The methods text should describe procedures and leave interpretation to the results. A model paragraph a supervisor can offer: the construct and its source; the instrument with authors, year and version; any adaptation; the expert panel with its size and qualifications; the content validity method and what was changed as a result; the pretest sample and changes; the reliability coefficient computed on the full sample with its limits. Each element maps to an artefact above, so the methods chapter is a summary of records the student already holds.
A supervisor rubric for instrument validation
- Weak: “The questionnaire was validated by my supervisor and has an alpha of 0.8 in the literature.”
- Acceptable: names the panel, the rating method and the pretest, and reports an alpha computed on the thesis sample.
- Strong: adds a statement about dimensionality, a justified threshold, a list of adaptations and an appendix with the rating sheet and the pretest log.
A hypothesis-and-variable review at proposal stage catches many problems before they reach the instrument. The method is set out in the guide to assessing hypothesis and variable quality, and the way a faculty can standardise reporting of survey instruments is illustrated in the companion piece on student engagement survey data.
Errors supervisors see most often, and the question that exposes each
- Alpha copied from the manual. Ask: which respondents produced this number? If the answer is “the original authors’ sample”, the thesis has no reliability evidence of its own.
- One expert, one afternoon. Ask: who else judged the items, and where is the rating sheet?
- Subscales ignored. A scale with several dimensions needs a coefficient per dimension as well as, where justified, for the total. Ask: how many dimensions does the construct have, and what did you do about that?
- Items dropped to raise alpha. Removing items can lift the coefficient while narrowing what the scale covers. Ask: what did the dropped item represent in the construct definition?
- Translation without testing. A translated scale is a new instrument. Ask: what evidence shows the translated items mean the same thing to the new respondents?
- Validity claimed from reliability. Ask: what is the separate evidence that the scores mean what the thesis says they mean?
Fitting the checklist into one semester
A thesis timetable rarely leaves room for open-ended instrument work, so the checklist works best when the faculty fixes the sequence in advance. A workable pattern for a one-semester proposal-to-data schedule is to require the construct page and instrument record with the proposal, the panel brief and rating sheet before ethics submission, the pretest log before the main data collection opens and the reliability table before the results chapter is drafted. Each deliverable is short, and each is something the supervisor can read in minutes. The aim is to make the evidence a series of small, dated artefacts, not a single chapter the student assembles in a panic at the end.
Where Tesify fits
Deciding what counts as adequate validity and reliability evidence is a faculty and supervisor decision. Once a student has a validated instrument and an approved design, Tesify helps students structure and organise their thesis while they write 100% of the work themselves: more than 9,000 students have written over 15,000 chapters with Tesify. See how Tesify supports a thesis cohort.
Frequently asked questions
What is the difference between validity and reliability in a thesis?
Reliability concerns the consistency of scores. Validity concerns whether the instrument measures what the thesis claims it measures. A reliable instrument can still be invalid, so the thesis needs separate evidence for each.
Is a high Cronbach’s alpha enough?
No. Taber’s review in science education research found that alpha is often quoted without explanation, that acceptable thresholds vary, and that a high value provides limited reliability evidence and can be misused to claim unidimensionality.
Should students quote the alpha from the original study?
Only as background. Reliability describes scores from a particular sample, so the thesis should report a coefficient computed from its own data.
What is Lawshe’s content validity ratio?
It is a method from Lawshe’s 1975 article in Personnel Psychology. Experts rate each item as essential, useful but not essential, or not necessary, and the ratio is (ne − N/2) divided by N/2, where ne is the number rating the item essential.
How many experts should a panel have?
Lawshe’s critical values depend on panel size: .99 for five to seven panellists, .75 for eight, .78 for nine, .62 for ten, .42 for twenty and .29 for forty. Very small panels leave little room for disagreement.
What is a pretest for?
It checks wording, timing, skipped items and response format with respondents from the target population. It does not replace expert judgement or reliability analysis.
Does changing an item mean the instrument must be re-validated?
Yes. Dropped, reworded or translated items create an adapted instrument, so the thesis must list the changes and gather its own validity and reliability evidence.
Can a supervisor serve as the only expert?
That is weak evidence. A documented panel with stated qualifications and independent ratings is stronger, and the rating sheet should be kept as an appendix.
What should the methods chapter report?
The construct and source, the instrument and any adaptation, the panel and method, the pretest and its changes, and the reliability coefficient computed on the thesis sample with its limits.
Who should check the reliability analysis?
A methods office or second reader with statistics training, because interpreting alpha well needs checks on dimensionality and thresholds that students often skip.
