Appearance
Research Methodology — Glossary
| Term | Meaning |
|---|---|
| Abstract | A short summary of a paper: background, methods, results, and conclusion |
| Anonymity | Participant identity is never linked to the data |
| Appraisal | Critical evaluation of the quality of a source (used in reviewing) |
| Assumption | A condition required for a statistical method to hold |
| Attrition | Loss of participants during a study (can cause attrition bias) |
| Audit trail | A complete record of how data were handled, for verification |
| Bayes' theorem | Formula to update a probability given new evidence |
| Bayesian inference | Updating prior beliefs with data to obtain a posterior |
| Beneficence | Ethical duty to maximise benefits and minimise harms |
| Bias | Systematic error that distorts study results |
| Blinding | Concealing group allocation from participants, staff, or analysts |
| CRediT taxonomy | Standard contributor roles (conceptualisation, methodology, etc.) |
| Cohort study | Follows a group over time to see who develops an outcome |
| Collinearity | Near-linear dependence among predictor variables (multicollinearity) |
| Comparator | The control or alternative treatment against which a test is measured |
| Completeness | Proportion of expected data that is actually present |
| Comprehensiveness | Thoroughness of a search; how few studies are missed |
| Confidence interval | A range that, with stated confidence, contains the true parameter |
| Confounder | A variable causally linked to both exposure and outcome |
| Construct validity | Whether a measure captures the theoretical construct intended |
| Content validity | Whether an instrument samples the full domain of interest |
| Convenience sampling | Non-probability sampling of easy-to-reach participants |
| Correlation | Strength and direction of a linear relationship between two variables |
| Cross-over trial | Each participant receives multiple treatments in random order |
| Cross-sectional study | Measures exposure and outcome at a single point in time |
| Data cleaning | Checking and correcting data errors/inconsistencies before analysis |
| Data dredging | Testing many hypotheses on the same data (inflates false positives) |
| Data imputation | Filling in missing values with estimated values |
| Data saturation | Point in qualitative work where no new themes emerge |
| Descriptive study | A study describing characteristics of a population |
| Descriptive statistics | Summaries of data (counts, mean, SD, etc.) |
| Degrees of freedom | The number of independent pieces of information in a statistic |
| Dichotomous | A variable taking only two possible values (yes/no) |
| Discrete variable | A numeric variable taking countable values |
| Distribution | The set of values a variable can take and how often |
| Dummy variable | A 0/1 variable encoding a category in a regression |
| Effect size | The magnitude of a difference or relationship (not just p-values) |
| Effect modification | A variable whose effect on the outcome differs by level |
| Eligibility criteria | Rules defining who may participate in a study |
| Empirical | Based on observation or experiment rather than pure theory |
| Ethics committee | IRB/REC body reviewing protocol and participant protections |
| Ethnography | Qualitative study of culture through fieldwork and observation |
| Evaluation | Assessing the value, merit, or performance of something |
| Evidence-based | Decisions supported by the best available evidence |
| Experimental study | A study that assigns a treatment to infer causation |
| External validity | Whether findings generalise beyond the study conditions |
| F-distribution | The probability distribution used for variance-ratio statistics |
| False discovery rate | Expected proportion of false positives among significant results |
| FAIR data | Data that is Findable, Accessible, Interoperable, Reusable |
| Family-wise error rate | Probability of at least one Type I error across many tests |
| Fishbone diagram | An Ishikawa cause-and-effect diagram for root-cause analysis |
| Five whys | An iterative root-cause method of asking "why?" repeatedly |
| G*Power | Software for statistical power and sample-size analysis |
| Gap (literature) | What is not yet known, which the current study addresses |
| Grounded theory | A qualitative method building theory directly from data |
| Grey literature | Unpublished/non-commercial works (reports, theses, preprints) |
| H0 | The null hypothesis (no effect or no difference) |
| H1 | The alternative (research) hypothesis |
| Hazard ratio | A ratio of hazard rates (used in survival analysis) |
| Heterogeneity | Variation in results across studies or groups |
| Histogram | A bar chart of frequency for continuous data (bars touch) |
| H-index | An author-level metric: h papers each cited at least h times |
| Human subjects | People participating in research (or their identifiable data) |
| Hypothesis | A testable statement about the relationship between variables |
| Inclusion criteria | Characteristics required for a participant to join a study |
| Informed consent | Voluntary agreement after understanding purpose, risks, benefit |
| Instrument | A tool for measuring or observing a variable |
| Instrumentation | The process or tools used to collect/measure data |
| Integrity | Honesty and truthfulness in researching and reporting |
| Inter-rater reliability | Agreement between two or more raters (κ, ICC, %) |
| Internal consistency | How well items of a scale hang together (Cronbach's α) |
| Internal validity | Confidence that the observed effect is a true causal effect |
| Interpretivism | Epistemology that meaning is interpreted, not measured objectively |
| Intention-to-treat | Analysis including all randomised participants in original groups |
| Interquartile range | The middle 50% of data (Q3 − Q1); robust spread measure |
| Intervention | The treatment or exposure under study |
| Interview | A data-collection method using open-ended questions |
| Kappa (κ) | Statistic for agreement beyond chance between raters |
| Keyword | A term indexing/searching literature |
| Literature review | A systematic survey of published work on a topic |
| Mean | The arithmetic average; sensitive to outliers |
| Measurement | Assigning numbers or labels to attributes by defined rules |
| Median | The middle value of ordered data; robust to outliers |
| Method | A planned way of addressing a research question |
| Mode | The most frequently occurring value |
| Model | An idealised description of how data were generated |
| Multiple imputation | Creating several datasets by imputing missing values |
| Nominal scale | Categorical labels with no order (e.g., blood type) |
| Normal distribution | A symmetric bell-shaped distribution (mean, SD) |
| Null hypothesis | The "no effect" statement tested against the data |
| Observational study | A study observing exposure/outcome without assigning treatment |
| Odds ratio | Ratio of the odds of an event in two groups |
| Operational definition | A precise, concrete procedure defining how a variable is measured |
| Outlier | A value far from the rest; may be error or a true extreme |
| p-value | Probability of the data (or more extreme) if the null hypothesis is true |
| Parameter | A numerical characteristic of a population |
| Pearson correlation | A coefficient (r) ranging −1 to +1 for linear association |
| Pilot study | A small run of the method to refine procedures before the main study |
| Plagiarism | Using others' words or ideas without credit |
| Population | The entire group about which you want to generalise |
| Power | The probability of detecting a true effect (1 − β) |
| Pre-registration | Publicly registering the analysis plan before collecting data |
| Precision | Closeness of repeated estimates (low variance means high precision) |
| Primary data | Data collected firsthand for the current study |
| Primary source | The original report of an event or study |
| Probability | A number between 0 and 1 quantifying uncertainty |
| Propensity score | Estimated probability of treatment, used to balance confounders |
| Protocol | A written plan for a study |
| Publication bias | Tendency to publish significant results more readily |
| Quantitative | Data expressed numerically, amenable to statistical analysis |
| Quartile | A value dividing ranked data into four equal groups |
| Randomisation | Assigning units to groups by a chance mechanism |
| Range | The difference between the largest and smallest values |
| Relative risk | (Risk ratio) ratio of risk in exposed vs unexposed groups |
| Reliability | Consistency of a measure across time, items, or raters |
| Replication | Repeating a study to check whether findings hold |
| Reproducibility | Whether the same analysis on the same data gives the same result |
| Residual | The difference between an observed and a predicted value |
| Risk of bias | The risk that errors influenced a study's results |
| Sampling | Selecting a subset of the population |
| Sampling bias | Systematic error from how participants are selected |
| Sampling distribution | The probability distribution of a statistic over all samples |
| Sampling error | Natural variability between a sample statistic and the population |
| Sampling frame | A list used for sampling (should match the population) |
| Scale | Any measuring instrument; also a scale of measurement (NOIR) |
| Scatter plot | A plot of paired observations to show a relationship |
| Secondary source | A source interpreting or summarising primary sources |
| Significance level | The threshold α for rejecting H₀ (usually 0.05) |
| Skew | Asymmetry of a distribution (positive or negative) |
| Snowball sampling | A chain-referral method for hidden populations |
| Standard deviation | The average distance of values from the mean |
| Standard error | The standard deviation of a sampling distribution |
| Statistical significance | A result unlikely to be due to chance (small p-value) |
| Statistic | A number describing a sample characteristic |
| Stratum | A subgroup used in stratified sampling (plural strata) |
| Systematic review | A rigorous, protocol-driven literature review |
| T-distribution | The bell-shaped distribution used for small-sample inference |
| t-test | A test comparing means (one-sample, paired, or two-sample) |
| Test-retest reliability | Whether a measure gives the same result when repeated |
| Triangulation | Using multiple methods/data sources to study one question |
| Type I error | Rejecting a true null hypothesis (a false positive) |
| Type II error | Failing to reject a false null hypothesis (a false negative) |
| Uncertainty | The inherent unknownness in observations and estimates |
| Variable | A characteristic that can take different values |
| Variance | The average squared deviation from the mean |
| Visual analogue scale | A line respondents mark to rate an attribute (e.g., pain) |