An assessment can be easy to administer, professionally presented and marked consistently, yet still fail to measure what it is supposed to measure. For example, a driving test that mainly examines a learner’s ability to memorise traffic regulations may provide limited evidence of whether the learner can control a vehicle safely on the road. This is why assessment quality cannot be judged by appearance or convenience alone.
Two of the most important ideas in assessment and evaluation are validity and reliability. Validity concerns whether an assessment supports an appropriate interpretation and use of results. Reliability concerns the consistency of the results. Understanding both helps teachers, trainers, employers, researchers and programme managers make better decisions about learning, performance and competence.
What Is Validity?
Validity is the extent to which evidence supports the interpretation and use of assessment results for a particular purpose. In simpler terms, an assessment is valid when it provides suitable evidence for the question being asked.
Validity is not an all-or-nothing label attached permanently to a test. An assessment may be valid for one purpose but unsuitable for another. A short multiple-choice quiz may be useful for checking recall of key terms, but it may not provide sufficient evidence of a learner’s ability to design a project, conduct an experiment or communicate professionally.
For instance, an entrepreneurship course may aim to develop the ability to identify customer needs, test a business idea, manage resources and explain a viable business model. If the final assessment only asks learners to define entrepreneurship terms, it measures some relevant knowledge but provides weak evidence of practical entrepreneurial competence.
What Is Reliability?
Reliability refers to the consistency, stability and dependability of assessment results. A reliable assessment produces reasonably similar results when the same performance is judged under similar conditions.
Reliability can be considered in several ways. Would the learner receive a similar result if the assessment were repeated? Would two trained assessors award similar marks for the same performance? Would different sections of a test produce reasonably consistent evidence of the same skill or knowledge area?
Reliability does not mean that an assessment is automatically fair or meaningful. A weighing scale that consistently adds five kilograms to every measurement may be reliable but inaccurate. Similarly, an assessment may produce stable results while measuring the wrong capability.
The Difference Between Validity and Reliability
The central distinction is this:
- Validity asks: Are we measuring the intended knowledge, skill, attitude or capability?
- Reliability asks: Are the results consistent enough to support dependable decisions?
Consider a written examination for a practical plumbing course. If the examination mainly rewards advanced reading ability rather than plumbing competence, its validity may be limited. If markers award very different scores to identical answers because the marking guidance is unclear, its reliability is also weak.
Validity and reliability are related. Inconsistent results make it difficult to defend an assessment interpretation. However, consistency alone is not enough. A test can be highly reliable and still lack validity if it consistently measures an irrelevant or incomplete characteristic.
A useful assessment must be both dependable enough to produce consistent evidence and appropriate enough to support the decision being made.
Main Types of Validity Evidence
Content validity
Content validity concerns whether an assessment adequately represents the knowledge and skills included in the learning outcomes or performance requirements. It is particularly important when an assessment is intended to cover a syllabus, occupational standard or training programme.
Suppose a course in financial literacy teaches budgeting, saving, responsible borrowing, risk management and interpreting financial statements. An assessment consisting only of questions about saving would not represent the full content adequately, even if every question were well written.
A table of specifications or assessment blueprint can help. It maps learning outcomes against content areas and cognitive levels, helping the assessor decide how much emphasis each area should receive.
Construct validity
Construct validity concerns whether an assessment genuinely measures the underlying ability or characteristic it claims to measure. A construct might be critical thinking, reading comprehension, mathematical reasoning, customer-service competence or leadership.
Construct-irrelevant factors can distort results. For example, an assessment intended to measure scientific reasoning may require unnecessarily complex language. Learners with weaker language proficiency could perform poorly because of the wording rather than their reasoning ability. In contrast, an assessment may omit an important part of the construct, such as practical application, and therefore provide an incomplete picture.
Criterion-related validity
Criterion-related validity considers how well assessment results relate to an appropriate external criterion. For example, scores from a carefully designed typing assessment might be compared with independently observed typing performance in a workplace.
This type of evidence can be useful when an assessment is expected to predict future performance or agree with another credible measure. The external criterion must itself be relevant and trustworthy; a poor comparison standard will not establish useful evidence.
Face validity
Face validity refers to whether an assessment appears, on the surface, to measure what it claims to measure. Learners and other stakeholders may be more willing to engage with an assessment that looks relevant and professionally designed.
Face validity can support confidence and acceptance, but it is not strong evidence by itself. An assessment may look appropriate while failing to sample the necessary knowledge or skills. Professional appearance should therefore be treated as a starting point, not proof of quality.
Important Forms of Reliability
Test–retest reliability
Test–retest reliability examines whether results remain reasonably stable when the same assessment is administered to the same people on two occasions. It is most suitable when the measured characteristic is expected to remain stable between the two administrations.
Large changes may indicate that learners remembered the questions, received additional teaching, misunderstood the instructions on the first occasion or were affected by fatigue, anxiety or other conditions. Repeating an assessment too soon can therefore produce misleading evidence.
Inter-rater reliability
Inter-rater reliability concerns the level of agreement between assessors. It is especially important for essays, presentations, portfolios, interviews, demonstrations and other performance assessments.
Two assessors may disagree because the criteria are vague, the performance is open to interpretation or the assessors have different standards. Clear criteria, performance-level descriptions, assessor training and moderation can improve consistency.
Intra-rater reliability
Intra-rater reliability concerns the consistency of one assessor’s judgements over time. A marker may apply standards more strictly at the beginning of a marking session and more generously later, or may be influenced by fatigue and familiarity with particular learners.
Using a marking guide, taking regular breaks and reviewing a sample of previously marked work can help an assessor maintain a stable standard.
Internal consistency
Internal consistency concerns whether items intended to measure the same area produce reasonably consistent evidence. For example, a set of questions designed to assess basic algebra should work together rather than measure unrelated skills.
Very high internal consistency is not always desirable. If every question is almost identical, the assessment may be narrow and fail to cover important aspects of the learning outcome. Good assessment design balances consistency with adequate breadth.
Factors That Can Reduce Reliability
Several practical conditions can make results less consistent:
- Unclear instructions that different learners interpret differently.
- Questions containing ambiguous, culturally unfamiliar or unnecessarily difficult language.
- Too few items to sample the learning outcome adequately.
- Uncontrolled differences in time limits, equipment, room conditions or access arrangements.
- Inconsistent marking decisions or incomplete scoring guidance.
- Assessor fatigue, bias, expectations or lack of training.
- Technical problems during online assessments.
- Assessment anxiety or temporary circumstances unrelated to the intended capability.
Reliability does not require every assessment condition to be identical in every situation. It requires the conditions to be sufficiently controlled and transparent for differences in results to reflect meaningful differences in performance rather than avoidable inconsistencies.
Validity, Reliability and Fairness
Validity and reliability should be considered alongside fairness and practicality. An assessment may be technically sound but unfair if it gives some learners an advantage unrelated to the intended learning outcome. Accessibility is part of good assessment design, not an optional addition.
For example, if the objective is to assess knowledge of Kenyan agricultural practices, using examples from familiar local farming contexts may improve relevance. However, if the objective is to assess accounting calculations, unnecessarily unfamiliar cultural references should not create an additional barrier. Similarly, a learner with a disability may require an appropriate adjustment, such as additional time or assistive technology, provided that the adjustment does not remove the essential skill being assessed.
Practicality also matters. A detailed individual interview may provide rich evidence but may be difficult to administer for several hundred learners. A balanced assessment system may therefore combine written questions, practical tasks, observation, discussion and project work rather than relying on one instrument.
How to Improve Assessment Validity and Reliability
- Start with clear learning outcomes. State exactly what learners should know, understand or be able to do. Use observable verbs such as analyse, calculate, construct, demonstrate, evaluate or design.
- Match the method to the outcome. Use selected-response questions for some knowledge checks, but use demonstrations, projects or simulations when learners must show practical performance.
- Construct an assessment blueprint. Map each item or task to a learning outcome, content area and level of thinking. This helps prevent over-assessing easy topics while neglecting important capabilities.
- Write clear instructions and questions. Avoid unnecessary complexity, double negatives, hidden assumptions and clues that reveal the answer. Check that each question has a clear purpose.
- Use enough evidence. A single question or brief observation is rarely sufficient for a significant decision. Sample important outcomes across more than one task or occasion where appropriate.
- Develop a rubric or marking scheme. Describe what successful performance looks like and distinguish between levels of quality. Criteria should relate directly to the learning outcomes.
- Moderate assessment judgements. Ask assessors to compare sample responses or performances, discuss differences and agree how the criteria should be applied.
- Pilot the assessment. Try the questions or tasks with a small group, where feasible, to identify confusing wording, unrealistic time limits and gaps in coverage.
- Review results and learner feedback. Look for items that almost everyone gets wrong, items that do not distinguish between levels of performance or tasks that produce unexpected patterns. Investigate before changing scores or drawing conclusions.
A Practical Example: Assessing Customer-Service Skills
Imagine that a professional training provider wants to assess customer-service competence. The learning outcomes require participants to listen actively, identify customer needs, communicate respectfully, resolve routine problems and record information accurately.
A multiple-choice test could assess knowledge of service procedures, but it would provide limited evidence of actual communication. A stronger assessment might combine a short knowledge quiz with a role-play, an observation checklist and a written case response.
The role-play improves validity because it allows the learner to demonstrate behaviour similar to the target performance. Reliability can be improved by giving all learners the same scenario structure, using the same time limit and applying a detailed rubric. If several assessors are involved, they should review sample performances together before assessing independently.
Even this assessment should be interpreted carefully. One role-play may not represent every customer situation. A second scenario or workplace observation could provide additional evidence, particularly when the result will be used for employment or certification.
Applying This in Practice
Before approving an assessment, ask the following questions:
- What decision will be made from the result?
- Which knowledge, skills or capabilities must the assessment measure?
- Does the assessment sample all important learning outcomes?
- Could language, technology, prior experience or the assessment setting affect performance in unintended ways?
- Would another assessment method provide stronger evidence?
- Are the instructions, criteria and marking process clear enough for different assessors to work consistently?
- Is there enough evidence to justify the importance of the decision?
Keep records of assessment specifications, marking guides, moderation discussions and review decisions. In a school, training centre, university or workplace, this documentation makes quality assurance more systematic and helps future assessors improve the assessment rather than repeating the same weaknesses.
For learners, understanding these principles can also be useful. A test score is evidence, not a complete description of ability. Learners should ask what the assessment actually measured, whether the conditions were fair and what additional evidence may be needed to show practical competence.
Key Takeaways
- Validity concerns whether an assessment supports an appropriate interpretation of what learners know or can do.
- Reliability concerns the consistency of results across items, occasions, assessors and conditions.
- A reliable assessment can still be invalid if it consistently measures the wrong capability or only a narrow part of it.
- Assessment methods should match the learning outcomes: practical skills require opportunities for practical demonstration.
- Clear criteria, assessor training, moderation and suitable assessment conditions improve reliability.
- Fairness, accessibility and practicality should be considered alongside validity and reliability.
No comments yet.