Understanding Research Data

Understanding Research Data

Research data is the evidence used to answer questions, test ideas and support decisions. Learn how to distinguish data types, assess quality, organise information, protect participants and choose suitable approaches for analysis.

Research data is the information collected or used to investigate a question. It may appear as survey responses, interview transcripts, financial records, photographs, laboratory measurements, observations, documents or digital activity. Understanding what data means, how it is produced and what its limitations are is essential for students, professionals, entrepreneurs and anyone making evidence-informed decisions.

Good research is not simply about gathering a large amount of information. It is about collecting relevant and trustworthy evidence in a systematic way, then interpreting it carefully. A small, well-designed dataset can be more useful than a large collection of incomplete or poorly measured information.

What Research Data Means

Research data consists of observations or records that help answer a research question. For example, a business studying customer satisfaction might collect ratings, written comments, purchase frequency and information about waiting times. A public health researcher might examine clinic records, interview responses and observations of health-service use.

Data becomes meaningful when it is connected to a clear purpose. The same information can serve different research questions. The number of customers visiting a shop may help measure demand, compare locations or assess the effect of a marketing campaign. Before collecting data, a researcher should therefore ask:

  • What question am I trying to answer?
  • What information would provide a credible answer?
  • Who or what should the data represent?
  • How will the information be collected, stored and analysed?

Main Types of Research Data

Quantitative data

Quantitative data is expressed in numbers or can be represented numerically. Examples include age, income, test scores, sales volume, distance, temperature and the number of completed applications. It is useful for identifying patterns, comparing groups and examining relationships between measurable factors.

Quantitative data may be discrete or continuous. Discrete data consists of countable values, such as the number of employees in a firm or the number of books borrowed. Continuous data can take many values within a range, such as time, weight or temperature.

Qualitative data

Qualitative data describes experiences, meanings, opinions, behaviours and social contexts. It often appears as interview transcripts, field notes, open-ended survey answers, diaries, photographs or documents. It can explain why something happens and how people understand their circumstances.

For example, a survey may show that customers are dissatisfied with a delivery service. Interviews or open-ended responses may reveal that the main problem is not delivery time alone, but unclear communication when delays occur. Qualitative data adds detail and context that numerical measures may miss.

Mixed-methods data

Mixed-methods research combines quantitative and qualitative approaches in one study. A vocational training centre could measure course completion rates and also interview learners who completed or left the programme. The numerical results show the scale of the issue, while the interviews help explain the reasons behind it.

Combining methods does not automatically make research stronger. Each method must have a clear role, and the researcher must explain how the different forms of evidence relate to one another.

Primary and Secondary Data

Primary data is collected directly for the current research project. Examples include conducting interviews, administering questionnaires, taking measurements or observing a process. Primary data can be closely matched to the research question, but collecting it may require time, money, trained staff and careful ethical planning.

Secondary data was collected previously by another person or organisation, or for a different purpose. Examples include government reports, company records, published research, administrative databases, historical documents and publicly available datasets. Secondary data can be efficient to use, but the researcher must examine how it was collected, whether it is current and whether it genuinely fits the new question.

Suppose an entrepreneur wants to estimate demand for a mobile food service. Existing information about population, transport patterns or household spending may provide useful background. However, direct customer research may still be needed to understand preferred meals, price expectations and ordering habits. Using both sources can provide a more balanced picture.

Understanding Variables

A variable is a characteristic that can take different values. In a study of employee training, variables might include training hours, assessment scores, job role and productivity. The outcome a researcher wants to explain or measure is often called the dependent variable. A factor expected to influence that outcome is commonly called an independent variable.

For example, if a researcher examines whether revision time is related to examination performance, revision time may be treated as the independent variable and the examination score as the dependent variable. This description does not prove that revision time causes higher scores. Other factors, such as prior knowledge, teaching quality or access to learning materials, may also matter.

Some variables are categorical rather than numerical. Employment type, county, product category and payment method are examples. Categorical data can still be analysed, but it must not be treated as though the categories have a meaningful numerical distance between them.

Levels of Measurement

Understanding how a variable is measured helps determine which summaries and analytical methods are appropriate.

  • Nominal: categories with no inherent order, such as department, religion or type of device.
  • Ordinal: ordered categories where the distance between categories is not necessarily equal, such as poor, fair, good and excellent.
  • Interval: numerical values with equal intervals but no meaningful absolute zero in the measurement system.
  • Ratio: numerical values with equal intervals and a meaningful zero, such as income, weight, distance or number of items sold.

A common mistake is to assume that every number supports every mathematical operation. A customer satisfaction scale from one to five has ordered categories, but the difference between one and two may not represent exactly the same change as the difference between four and five. The measurement design should guide the analysis.

What Makes Research Data Useful?

Validity

Validity concerns whether the data measures what it is intended to measure. If a researcher wants to assess financial literacy using only the number of bank accounts someone has, the measure may not be valid because account ownership does not necessarily show understanding of financial concepts.

Reliability

Reliability concerns consistency. A reliable measurement procedure produces similar results when the underlying situation has not changed. A weighing scale that gives a different reading each time an object is measured may be unreliable.

Accuracy and completeness

Accurate data is recorded correctly. Complete data contains the information needed for the analysis, including relevant responses and properly documented missing values. Researchers should distinguish between a genuine zero, a question that does not apply and a question that was left unanswered.

Timeliness and relevance

Data can be accurate but no longer suitable for a particular decision. A business using several-year-old information about customer preferences may overlook changes in technology, prices or competition. Relevance depends on the research question, the population and the period being studied.

Planning Data Collection

Data quality begins before the first response is collected. Start by defining the research problem and identifying the population of interest. The population might be all learners in a college, customers who bought a product during a specific period or households in a selected area.

A sample is a smaller group selected from that population. Sampling decisions affect how confidently findings can be applied beyond the people or cases directly studied. A sample chosen only from the easiest people to reach may exclude important perspectives. For instance, an online survey about digital services may under-represent people with limited internet access.

Next, select a collection method. Questionnaires can gather standardised responses from many participants. Interviews can explore personal experiences in depth. Observation can reveal what people do rather than what they say they do. Existing records may be efficient, but their original definitions and collection procedures need to be understood.

Questions should be clear, neutral and specific. Avoid leading questions such as, “How helpful was our excellent service?” A more balanced question would ask, “How would you rate the service you received?” Avoid double-barrelled questions that ask about two issues at once, such as, “How satisfied are you with the price and quality?” A respondent may have different views about each.

Preparing and Organising Data

Raw data often requires preparation before analysis. This may include checking for duplicate records, correcting obvious entry errors, standardising dates and labels, and documenting missing values. Keep an original, unaltered copy of the dataset before making changes.

A data dictionary is a useful tool. It records what each variable means, the units used, permitted values and how missing information is represented. For example, a variable labelled monthly_sales should specify whether it is measured in Kenyan shillings, whether returns are deducted and which month each value represents.

Organise files consistently and use meaningful names. Protect personal information by limiting access, using secure storage and removing identifying details where they are not needed. Data management is not merely administrative work: poor organisation can lead to incorrect analysis and make research impossible to verify.

Interpreting Quantitative Data

Initial analysis usually begins with descriptive statistics. Frequencies show how often categories or values occur. Percentages make groups easier to compare when their sizes differ. Measures of central tendency include the mean, median and mode.

The mean is calculated by adding values and dividing by their number. It can be strongly affected by unusually high or low values. The median is the middle value when observations are ordered and may better represent a skewed distribution, such as household income. The mode is the most frequently occurring value or category.

Measures of spread, such as the range and standard deviation, indicate how much values differ from one another. Averages without information about variation can be misleading. Two groups may have the same average score while one group has results clustered closely together and the other has widely differing results.

Researchers may also examine relationships between variables. A correlation indicates that two variables vary together, but it does not by itself establish causation. If sales rise during a period when advertising increases, other factors such as seasonal demand, a competitor leaving the market or a price change may also explain the result.

Interpreting Qualitative Data

Qualitative analysis involves carefully examining text, images or observations to identify patterns and meanings. A researcher may begin by reading the material several times, making notes and assigning codes to relevant sections. Codes can describe topics such as “cost concerns”, “trust in providers” or “difficulty accessing transport”. Related codes may then be grouped into broader themes.

Good qualitative analysis uses evidence from the data rather than selecting only comments that support an existing opinion. The researcher should consider different or contradictory views and explain how interpretations were developed. Reflexivity is also important: researchers should recognise how their own assumptions, position or relationship with participants may influence the study.

Ethics and Responsible Use

Research involving people requires respect, honesty and protection from avoidable harm. Participants should receive understandable information about the study, what participation involves and how their data will be used. Where appropriate, they should provide informed consent and be free to decline or withdraw according to the study arrangements.

Confidentiality means handling information so that unauthorised people cannot access it. Anonymity means that a participant cannot reasonably be identified from the data. These are not always the same. A small community, unusual job title or combination of details may identify someone even when their name has been removed.

Researchers should collect only information that is necessary, avoid deceptive practices unless properly justified and approved, and report findings honestly. Do not delete inconvenient results simply because they do not support the preferred explanation. If limitations, errors or missing data affect the findings, state this clearly.

Applying This in Practice

Imagine a small Kenyan agribusiness wants to understand why repeat orders have fallen. A disciplined data plan could follow these steps:

  1. Define the question: determine whether the study concerns product quality, pricing, delivery, communication or another factor.
  2. Review existing records: examine order dates, repeat purchases, complaints, refunds and delivery times.
  3. Identify information gaps: records may show what happened but not why customers chose not to return.
  4. Collect new data: use a short, neutral questionnaire and a small number of interviews with different customer groups.
  5. Check quality: remove duplicate responses, document missing answers and compare reported delivery problems with operational records.
  6. Analyse carefully: summarise response patterns, identify recurring themes and consider alternative explanations.
  7. Make a proportionate decision: test a specific improvement, such as clearer delivery updates, and monitor results rather than claiming that one change has solved every problem.

This process demonstrates the link between a research question, suitable data, careful analysis and a justified decision. It also shows why data should not be collected simply because it is available. Every item should serve a clear analytical purpose.

Key Takeaways

  • Research data is evidence collected or used to answer a specific question; its value depends on relevance and quality, not volume alone.
  • Quantitative data measures amounts and patterns, while qualitative data explains experiences, meanings and context.
  • Primary data is collected for the current study, whereas secondary data already exists and must be checked for suitability.
  • Validity, reliability, accuracy, completeness and timeliness are essential when judging data quality.
  • Clear variables, appropriate sampling and neutral questions improve the credibility of data collection.
  • Responsible researchers protect participants, document their methods and report limitations honestly.

Comments

Learner discussion on this EduHub resource.

No comments yet.