COSMIN (COnsensus-based Standards for the selection of health Measurement INstruments) is an initiative led by an international multidisciplinary group, founded in 2005 in response to the lack of clarity and uniformity regarding the terminology used in the study of measurement properties of assessment tools. The group was founded and is led by Lidwine B. (Wieneke) Mokkink and Caroline B. Terwee. The team comprises researchers with expertise in epidemiology, psychometrics, medicine, qualitative research, and healthcare.

The primary objective of COSMIN is to improve the quality of studies on measurement properties through the development of a clear and concrete methodology that allows for the standardisation of results. The initiative is equally committed to developing theoretical and practical tools to promote the appropriate selection and use of assessment instruments in both clinical practice and research.

Initially, the methodology was developed for tools designed to collect data directly reported by patients (PROMs: patient-reported outcome measures). It has since been adapted to cover clinician-reported outcome measures (ClinROMs), performance-based outcome measures (PerFOMs), and laboratory values. For the sake of terminological simplicity, the term PROM will be used throughout both COSMIN documents and this entry. In this context, it is essential to introduce the concept of a construct, understood as an abstract theoretical entity that is not directly observable but is inferred from measurable indicators. Constructs represent complex dimensions of human experience — such as quality of life, participation, or performance — and form the conceptual foundation upon which measurement instruments are designed. In the case of PROMs, the construct is defined from the individual’s perspective, meaning that the instrument does not aim to capture an external objective reality, but rather the subjective perception of the person regarding their health status, functioning, or experience.

COSMIN taxonomy of measurement properties

One of the key milestones of this initiative was the unification of terminology and definitions relating to measurement properties. Between 2006 and 2007, an international consensus was reached through a Delphi study, in which agreement was achieved on a set of measurement properties — reliability, validity, and responsiveness — that should be considered in the development and evaluation of assessment tools. As shown in the figure, the consensus establishes that the quality of measurement instruments should be evaluated on the basis of nine fundamental measurement properties.

Reliability is the degree to which a PROM produces consistent results, taking into account the margin of random error inherent in any measurement in relation to the total variability of the measure. It comprises the following properties:

  • Test-retest reliability (Box 6. Reliability): the degree of stability of an assessment tool’s scores when administered on two separate occasions under similar conditions. The interval between measurements is typically two weeks, but should be adjusted according to the construct of interest, seeking a balance between minimising recall effects and preventing genuine changes in the construct being measured.
  • Inter-rater reliability (Box 6. Reliability): the degree of agreement in the scores of an assessment tool between different raters when administering the same instrument to the same subjects under equivalent conditions.
  • Fiabilidad intra-evaluador (intra-rater) (Box 6. Reliability): grado de acuerdo en las puntuaciones de la herramienta realizada por el mismo evaluador en dos o más mediciones repetidas de la misma herramienta.
  • Internal consistency (Box 4. Internal consistency): the degree to which the items within a scale are interrelated and measure the same construct.
  • Measurement error (Box 7. Measurement error): the variability in scores attributable to imprecision inherent in the assessment tool or the measurement process, which does not reflect genuine changes in the construct being measured.

Validity is the degree to which a PROM adequately measures the construct for which it was designed. It is important to note that validity is not an inherent property of the instrument itself, but rather of the interpretations drawn from its scores within a specific context. A PROM may therefore be valid for a particular population, purpose, or setting, but not necessarily for others. It comprises the following properties:

  • Content validity (Box 2. Content validity): the degree to which the content of a PROM adequately reflects the construct it intends to measure, in terms of its relevance, comprehensiveness, and comprehensibility for the target population. For example, if a PROM is designed to measure pain, its items should focus on the pain experience and should not include aspects unrelated to the construct — such as mobility upon waking — unless these form an explicit part of the theoretical definition of the construct.
  • Criterion validity (Box 8. Criterion validity): the degree to which the scores of a PROM correlate with those obtained using a gold standard, where one exists. COSMIN establishes that, in the case of PROMs, a gold standard can only be considered as such when a measure exists that captures the same construct more completely and precisely than the instrument under evaluation. In this context, a longer version of the same instrument — with a greater number of items, broader construct coverage, and superior measurement properties — may serve as the reference standard, provided it is adequately validated and conceptually aligned. For example, the 12-item Short Form (SF-12) is a generic patient-reported outcome measure designed to assess perceived health-related quality of life. It evaluates this construct across two broad dimensions: the physical component and the mental component. It was developed as an abbreviated version of the 36-item Short Form (SF-36), with the aim of reducing respondent burden while maintaining adequate capacity to estimate health-related quality of life. In this case, the SF-36 could be used as the gold standard for evaluating this property of the SF-12.
  • Construct validity (Box 9. Hypotheses testing for construct validity): the degree to which a PROM measures the construct it intends to assess, considering the appropriateness of its dimensional structure, the fulfilment of theoretical hypotheses, and the equivalence of the instrument when adapted to other languages or cultures. It is assessed through:
    • Hypothesis testing: verification of whether PROM scores behave in accordance with the expected relationships with other variables or other assessment tools.
    • Structural validity (Box 3. Structural validity): the degree to which the dimensional structure of the instrument — whether unidimensional or multidimensional — is consistent with the theoretical model of the construct it intends to measure and with the organisation of the items representing it.
    • Cross-cultural validity / measurement invariance (Box 5. Cross-cultural validity / Measurement invariance): the degree to which a PROM adapted to another language or culture retains the properties of the original instrument.

Responsiveness (Box 10. Responsiveness) is the ability of a PROM to detect changes in the construct of interest that are clinically meaningful.

Additionally, the COSMIN taxonomy of measurement property definitions includes interpret-ability. This is not a measurement property in itself, but rather a characteristic of the assessment instrument that indicates the clinical meaning of the scores obtained — that is, the qualitative and clinical significance of the results.

Por último, es importante conocer que, según el consenso de COSMIN, la propiedad de medida más importante es la validez de contenido, ya que nos permitirá conocer la relevancia, exhaustividad y comprensibilidad de la herramienta. Seguidamente, se considera muy importante la estructura interna, la cual incluye validez estructural, consistencia interna y validez transcultural.
Finally, it is important to note that, according to the COSMIN consensus, the most important measurement property is content validity, as it allows us to determine the relevance, comprehensiveness, and comprehensibility of the instrument. Internal structure — encompassing structural validity, internal consistency, and cross-cultural validity — is considered the next most important domain.

Other COSMIN resources:

  • Guide for conducting Systematic Reviews (SRs) of measurement instruments. A resource based on the combined use of COSMIN methodology and the PRISMA guidelines, given that this type of SR differs from standard reviews of clinical trials. The guide sets out eight steps to be followed, covering the search process, eligibility criteria and screening filters, data extraction procedures, and criteria for evaluating each included PROM and its measurement properties. It also includes example tables and writing recommendations for this type of SR.
  • Database for conducting SRs. Used alongside the guide above to facilitate data extraction. Types of systematic reviews of measurement instruments:
    • Quality of a single instrument: all available articles and measurement properties for one PROM are analysed.
    • Quality of multiple instruments: all available articles and measurement properties for several instruments sharing the same construct and population are analysed.
    • Quality of all available validated instruments: all available articles and measurement properties for all existing instruments sharing the same construct and population are analysed, with the aim of selecting the most appropriate one.
    • Quality of all available instruments: the analysis is conducted without specifying the construct of interest within a given population.
  • Methodological quality checklist. A checklist for evaluating the methodological quality of a study or PROM. It can be used when designing a study, drafting an article, assessing the risk of bias of a PROM, or conducting systematic reviews, among other applications. Using this checklist ensures that a study meets the recommended standards.
  • PubMed search filters. Two search filters are provided for identifying studies on the evaluation of measurement properties in PubMed: a high-sensitivity filter for retrieving PROM studies, and a more precise filter that requires less abstract screening, albeit at a slightly higher risk of missing relevant studies.
  • Guide for selecting instruments for a Core Outcome Set (COS). A joint initiative between COSMIN and Core Outcome Measures in Effectiveness Trials (COMET), providing guidance on selecting PROMs when developing a COS — a minimum set of outcomes that should be measured in clinical trials within a specific population. Once the COS is well defined, the guide offers recommendations on which constructs to measure and how to do so, with a focus on using reliable and valid instruments within the relevant context..
  • Find the right tool. A section of the COSMIN website that addresses common questions arising before starting a research project or intervention — such as what to measure and which assessment tools are available for a given population. Each question links to a dedicated page offering options and resources to help identify the most appropriate instrument for each situation.

In addition to the resources listed above, working groups have been established across Europe, North America, and Oceania to contribute to the growth of COSMIN. In Spain, the group is coordinated by Silvia Lahuerta Martín and Clara Amat Fernández, who manage the monthly sessions for Spanish-speaking participants. If you are interested in joining the COSMIN Measurement Properties Study Group, you can register via the following form: https://forms.office.com/e/9P8BLnwgHp.

In conclusion, having standards such as those provided by COSMIN is essential to ensure that the assessment instruments used in everyday practice are appropriate and fit for purpose. In research, COSMIN offers guidance on how to conduct adaptation studies and evaluate measurement properties, providing unified terminology and consistent statistical criteria. As a practical contribution to Occupational Therapy and other healthcare professions, it offers assurance that the instruments being used are stable, adapted, and valid for the target population — enabling the generation of comparable, clinically meaningful results that support evidence-based decision-making throughout treatment.

Empar Casaña
Occupational Therapist, Master's in Occupational Therapy in Neurology. FPU Pre-doctoral Researcher in the PhD Program in Public Health, Medical and Surgical Sciences. Contracted Researcher at InTeO.

 

 

 

 

Mª Paula Noce
Occupational Therapist, Master's in Neurological Occupational Therapy and Master's in Public Health.Pre-doctoral Researcher in the PhD Program in Public Health and Medical and Surgical Sciences. Collaborator at InTeO.

Translate by: Beatriz Espinosa
Occupational Therapist, holding a Master's Degree in Intervention for Functional Diversity in Childhood and specialized training in child development and mental health. University degree officially recognized in Australia.

Casaña Escriche, E., & Noce, P. (2026, Abril, 13). COSMIN: estándares para la evaluación de propiedades de medida de las herramientas de evaluación. PublicaTO – Habilidades Científicas en Terapia Ocupacional de InTeO. https://hacto.umh.es/en/2026/07/27/cosmin-standards-for-the-evaluation-of-measurement-properties-of-assessment-tools/

This work is licensed under a Creative Commons License Licencia Creative Commons