BIBLIOMETRICS · RESEARCH METRICS · CITATION GRAPHS · h-INDEX · JOURNAL METRICS · ALTMETRICS · RESPONSIBLE RESEARCH ASSESSMENT
How Bibliometrics and Responsible Research Assessment Work
At 4:17 in the afternoon, a research committee can reduce twenty years of work to three numbers. Papers. Citations. h-index. The numbers are clean enough to fit in a column. The career behind them is not.
A mathematician may publish rarely and be cited slowly. A biomedical scientist may work inside teams of hundreds. A computer scientist may place major work in conferences. A historian may spend ten years writing a book whose influence arrives through teaching, archives and public argument rather than journal citations. A software engineer may build a research tool used by thousands of scientists without becoming first author on the papers that depend on it.
Metrics are compressed stories. The danger begins when we forget what was compressed.
Bibliometrics studies patterns in publications, references, citations, authors, journals, institutions and scholarly networks. Research metrics turn some of those patterns into indicators: citation counts, h-index, journal-level measures, field-normalised indicators, usage measures and attention signals. Responsible research assessment asks a harder question: when, if ever, should those indicators influence decisions about people, projects, institutions and knowledge?
The missing piece in most explanations is the whole machine. Citation counts do not fall from the sky. They are produced by databases that ingest records, match references, resolve identities, define time windows and decide what counts as a work. Metrics then sit on top of those data choices. Committees sit on top of the metrics. Incentives sit on top of the committees. Behaviour changes in response. The measurement system enters the system it is measuring.
The short answer
RESEARCH OUTPUT → METADATA → REFERENCE LIST → DATABASE INGESTION → IDENTITY RESOLUTION → CITATION GRAPH → METRIC DEFINITION → FIELD / TIME / CAREER CONTEXT → INDICATOR → HUMAN INTERPRETATION → ASSESSMENT DECISION → INCENTIVES → RESEARCHER BEHAVIOUR → NEW OUTPUTS → MEASUREMENT SYSTEM CHANGES AGAIN
1. Bibliometrics begins with documents, not scores
Before there is an h-index, there are research objects: articles, books, conference papers, reviews, datasets, software, preprints, protocols and other outputs.
Bibliometric systems decide which of these objects they index, how they describe them and how they connect them. A metric can only measure what its database can see.
2. A citation is a relationship between works
When one publication references another, a citation relationship can be represented as a directed edge in a graph.
PAPER A → cites → PAPER B PAPER C → cites → PAPER B PAPER D → cites → PAPER B PAPER B cited_by_count = 3 if all three links are successfully indexed and matched.
That final sentence matters. Citation counts depend on successful extraction and matching, not only on what authors wrote.
3. Reference matching is an identity problem
A reference may contain a DOI, title, author, year, journal, volume, issue and pages. Some references are complete. Others contain errors. Old literature may have no DOI at all.
Modern scholarly graphs match references through identifiers where possible and bibliographic inference where necessary. The weaker the metadata, the more uncertain the match.
4. Crossref makes citation links part of open scholarly infrastructure
Crossref encourages publishers to deposit reference lists with DOI metadata so citations can be linked, discovered and reused by other systems.
In May 2026, Crossref reported that its metadata corpus had passed two billion citation links. The scale is extraordinary, but the more important idea is architectural: citation data are becoming a reusable research graph rather than a private list at the end of a PDF.
5. Metadata quality determines metric quality
If author names are incomplete, dates wrong, references missing or titles malformed, downstream citation matching weakens.
Crossref’s current bibliographic guidance emphasises complete and accurate contributors, titles, dates and identifiers because those fields support citation matching and discovery.
6. The database is part of the metric
A researcher can have different citation counts and different h-index values in Web of Science, Scopus, Google Scholar or OpenAlex because the databases index different material and resolve records differently.
Therefore a metric without a named data source is incomplete.
7. Coverage is never neutral
Databases differ in journal coverage, conference proceedings, books, languages, regions, historical depth and document types.
A metric can look precise while quietly inheriting the worldview of the database beneath it.
8. Books and conferences expose field differences
Many humanities disciplines rely heavily on books. Computer science often values conference proceedings. Some biomedical fields communicate primarily through journal articles.
A journal-centric database will represent these fields unevenly even before any metric is calculated.
9. Language coverage can become impact bias
Research published in globally dominant languages is often easier for international databases and citation communities to discover.
Local-language scholarship may have substantial social or policy impact while appearing weak in internationally oriented citation databases.
10. Regional relevance and global citation are not the same thing
A study on a specific public-health intervention in Singapore may be highly useful to local practitioners without attracting thousands of global citations.
Responsible assessment asks whether the metric matches the mission.
11. Citation count measures citation attention
The simplest bibliometric indicator is the number of indexed works that cite a publication.
That sounds like “impact”, but citation is narrower. It records one kind of scholarly attention inside a particular database.
12. Citations do not all mean agreement
A paper may be cited because later researchers build on it, criticise it, correct it, replicate it, use its method or mention it as historical background.
A citation count usually does not encode the reason for citation.
13. Negative citation is still citation
A flawed but influential paper can accumulate citations precisely because researchers keep explaining why it is flawed.
High citation therefore means influence on scholarly attention, not automatic endorsement.
14. Review articles have structural citation advantages
Reviews summarise broad areas and are frequently used as convenient gateways to a literature.
Comparing the citation count of a review article with a narrow methods paper as though they had equal citation opportunity can mislead.
15. Older papers have had more time to accumulate citations
A five-year-old paper and a five-month-old paper are not competing under equal exposure time.
Time windows are therefore part of responsible interpretation.
16. Citation half-lives differ by field
Some fields move quickly and cite recent work heavily. Others build slowly and continue citing work decades later.
Short assessment windows systematically favour faster citation cultures.
17. Self-citation can be legitimate
Researchers often build on their own earlier methods, datasets or theories. Citing those works can be necessary for provenance.
The integrity question is whether self-citation is relevant or strategically inflated.
18. Citation cartels are a system-level failure
Groups of authors or journals can coordinate unnecessary citations to raise metrics.
When a metric becomes valuable enough, behaviour can be reorganised around the metric rather than the underlying research quality.
19. Goodhart’s Law belongs in every metrics discussion
The familiar formulation is that when a measure becomes a target, it can cease to be a good measure.
Publication counts encourage more papers. Citation targets encourage citation strategies. Journal-metric targets encourage venue gaming. The measurement changes the behaviour it was designed to observe.
20. The h-index compresses productivity and citation impact
Jorge Hirsch proposed the h-index in 2005. A researcher has an h-index of h when h of their papers have at least h citations each.
Citation counts by paper: 52, 31, 17, 10, 8, 5, 3, 2 At least 5 papers have ≥5 citations. But 6 papers do not have ≥6 citations. h-index = 5
The attraction is obvious: one number rewards both sustained output and sustained citation.
21. The h-index ignores the tail above h
A researcher with five papers cited five times each can have the same h-index as a researcher whose top five papers have thousands of citations if both have only five papers above the threshold.
Once a paper is safely inside the h-core, additional citations do not raise h until another paper crosses the next threshold.
22. The h-index grows with career length
A senior researcher has had more years to publish and accumulate citations.
Comparing h-index across career stages without adjustment can reward age rather than current research quality.
23. The h-index is field dependent
Fields with large communities and dense citation practices tend to produce higher h-indices than small or slower-moving fields.
Cross-field h-index rankings therefore need extreme caution.
24. The h-index inherits database coverage
If one database indexes more conference papers or books, the same person may receive a higher h-index there.
Always report the data source and date with the h-index.
25. i10-index is even simpler
The i10-index counts how many works have received at least ten citations.
Its simplicity is useful for quick summaries but gives little context about field, age, contribution or citation distribution.
26. OpenAlex exposes metrics while warning about rigorous use
Current OpenAlex documentation includes convenience indicators such as h-index, i10-index and two-year mean citedness for entities, while explicitly recommending that rigorous analyses compute metrics from the underlying works.
That is a useful design principle: summary metrics should remain auditable back to the records that generated them.
27. Journal metrics measure journals, not individual articles
This sounds obvious and is violated constantly.
A journal-level citation average describes a publication venue. It does not establish that every article in that journal has the same quality or impact.
28. The Journal Impact Factor has a specific calculation logic
The Journal Impact Factor is a journal-level indicator calculated from citations in a defined year to recent content and a denominator of specified citable items from the preceding publication window.
The exact data and inclusion rules belong to Clarivate’s Journal Citation Reports methodology, so responsible use should refer to the current JCR definition rather than informal approximations.
29. JIF distributions are skewed
A small number of highly cited papers can raise a journal average while many individual articles receive far fewer citations.
This is one reason journal averages are weak proxies for the value of an individual paper.
30. DORA rejects journal prestige as a substitute for research quality
The San Francisco Declaration on Research Assessment argues against using journal-based metrics such as the Journal Impact Factor as surrogate measures of the quality of individual research articles or the contributions of individual researchers.
DORA does not say numbers are forbidden. Its current guidance emphasises responsible, contextual use.
31. DORA’s five principles are operational
- Be clear about what is being assessed.
- Be transparent about data and methods.
- Be specific about what the indicator can represent.
- Be contextual about discipline, career stage and purpose.
- Be fair about bias, access and consequences.
These principles turn “use metrics responsibly” from a slogan into an assessment workflow.
32. CiteScore is another journal-level metric
CiteScore uses Scopus data and a defined citation window to measure average citations per peer-reviewed document under Elsevier’s methodology.
It should be interpreted as a journal metric with its own database coverage and rules, not as a universal quality number.
33. SNIP attempts field context
Source Normalized Impact per Paper adjusts journal citation impact using the citation potential of its subject field.
The deeper principle is more important than the brand name: raw citation averages are hard to compare across fields with different citation densities.
34. SJR uses network prestige
SCImago Journal Rank weights citations according to the prestige of the citing source rather than treating every citation as equivalent.
This moves bibliometrics from counting edges to modelling the structure of the citation network.
35. Eigenfactor uses another network approach
Eigenfactor-style measures analyse journal influence through citation networks and attempt to account for the importance of where citations originate.
Different network metrics encode different assumptions about what influence means.
36. There is no metric without a model
Every indicator chooses a denominator, time window, database, document type, normalisation method and aggregation rule.
Metrics look like observations but contain design decisions.
37. Field normalisation asks: compared with what?
A raw citation count becomes more interpretable when compared with other work of similar field, age and document type.
Field-normalised indicators try to answer whether a paper is cited more or less than an appropriate reference set.
38. The reference set is the hidden heart of normalisation
Which papers count as the same field? Which publication years? Which document types? Which database?
Two normalised metrics can disagree because they define the expected baseline differently.
39. Interdisciplinary work challenges field boundaries
A paper spanning medicine, computing and ethics may not fit one citation culture.
Forcing interdisciplinary work into a single field baseline can introduce artefacts precisely where the research crosses boundaries.
40. Percentiles can be easier to interpret than averages
A statement such as “this paper is in the top 10% of citations among comparable papers in its field and year” can be more robust to skewed citation distributions than comparing raw averages.
But the quality of the comparison still depends on the reference set.
41. Highly cited does not mean methodologically sound
Research can become highly cited because it is fashionable, controversial, flawed, useful as a method, or central to a large literature.
Bibliometric impact and research quality overlap imperfectly.
42. Quality requires reading evidence
Study design, methods, data integrity, reproducibility, argument, originality and contribution cannot be reduced reliably to citation counts alone.
Metrics can help locate unusual patterns. Expert judgement remains necessary to interpret them.
43. CoARA makes qualitative judgement primary
The Agreement on Reforming Research Assessment sets a shared direction in which assessment is based primarily on qualitative judgement, with peer review central and quantitative indicators used responsibly in support.
This reverses a common bad workflow in which committees begin with numbers and use narrative merely to justify the ranking produced by the spreadsheet.
44. The Leiden Manifesto offers ten principles
The Leiden Manifesto for Research Metrics was created because research evaluation had become increasingly driven by data rather than expert judgement.
Its ten principles include supporting qualitative expert assessment, measuring performance against institutional missions, protecting locally relevant research, accounting for field variation, keeping data and analysis open, allowing verification, respecting career differences and regularly reviewing indicators.
45. Mission should precede measurement
A teaching university, national laboratory, public-health institute and humanities centre have different missions.
Using the same metric dashboard for all of them silently defines one mission as universal.
46. Research assessment has multiple levels
- Individual article.
- Researcher.
- Research group.
- Department.
- Institution.
- Journal.
- Country or region.
- Research field.
A metric valid at one level can become invalid when moved to another.
47. Ecological fallacy appears in research metrics too
A high-impact journal does not imply every article is high impact. A highly cited institution does not imply every researcher is highly cited.
Group averages should not be assigned mechanically to individuals.
48. Researcher assessment needs contribution context
A paper with one thousand authors cannot be interpreted by publication count alone.
Contribution statements, author roles and team science become essential as collaboration scales.
49. Hyperauthorship changes counting
Large physics, genomics and clinical collaborations can produce papers with hundreds or thousands of authors.
Full counting gives every author one publication. Fractional counting divides credit. Contribution-based approaches use role information. Each choice answers a different question.
50. First-author conventions are not universal
Some fields use first author for primary contribution and last author for senior leadership. Others use alphabetical order. Large collaborations may publish consortium authorship.
Metrics that assume one authorship convention can misread another discipline completely.
51. Career interruptions affect output trajectories
Parental leave, clinical duties, disability, caregiving, migration, military service, institutional disruption and other interruptions affect publication and citation accumulation.
Responsible assessment should not treat uninterrupted output as the only normal career.
52. Early-career researchers are structurally disadvantaged by cumulative metrics
h-index and total citations reward accumulated time.
For early-career assessment, recent contributions, methods, independence, trajectory, open practices and qualitative evidence may be more informative.
53. Institutional prestige can enter the citation loop
Researchers at famous institutions may receive more visibility, invitations and citations partly because of existing prestige.
Metrics can therefore amplify prior advantage rather than measure contribution from a neutral starting line.
54. Matthew effects create cumulative advantage
Recognition attracts recognition. Highly cited researchers become easier to discover, receive more invitations and may attract more collaborators and resources.
Bibliometric inequality can be both outcome and cause.
55. Gender and demographic bias can be embedded in the graph
Citation practices reflect scholarly communities, including their historical inequalities.
A metric can reproduce those patterns while appearing impersonal because the bias occurred upstream in opportunity, visibility, collaboration and citation behaviour.
56. Geography affects visibility
Research from globally central institutions is often more visible in international networks than equally useful research from less connected regions.
Responsible bibliometrics should distinguish global visibility from intrinsic merit and local value.
57. Open access can alter citation opportunity
When more readers can access a paper without a subscription barrier, the opportunity for reading and citation can increase.
But open access status is entangled with field, funder, journal and author choices, so citation differences should not be interpreted as a simple causal law without careful study.
58. Search ranking also shapes citation
Researchers often cite what they can find. Discovery systems rank results.
Visibility in search engines, databases and recommendation systems can influence future citation, creating feedback between discovery and metrics.
59. Altmetrics measure attention outside traditional citations
Altmetrics can include signals from news coverage, social media, policy documents, online reference managers, blogs or other digital attention sources depending on the provider.
They can reveal forms of attention that accumulate faster or outside formal scholarly publishing.
60. Attention is not impact and impact is not quality
A viral paper may be discussed because it is surprising, controversial or wrong. A highly useful technical standard may receive little public attention.
Altmetrics should be interpreted as attention signals with context, not as replacement truth scores.
61. Policy citations can reveal societal pathways
When research is cited in government guidance, clinical recommendations or standards, the influence pathway differs from scholarly citation.
This can be especially important for applied research whose purpose is action rather than disciplinary prestige.
62. Patent citations represent another knowledge route
Research cited in patents may indicate technological relevance, although patent citation practices have their own legal and procedural dynamics.
No single downstream signal should be treated as universal impact.
63. Downloads measure use opportunity
Views and downloads can show that people reached the research object.
They do not show whether the work was read carefully, believed, reused or acted upon.
64. Usage metrics are platform dependent
A paper mirrored across a publisher site, repository and preprint server can split usage across platforms.
Comparing raw download counts without understanding platform scope can be misleading.
65. Research software needs different evidence
Software can have enormous scientific value while receiving fewer traditional citations than papers.
Usage, dependencies, releases, community adoption, documentation, software citations and maintenance can supplement article metrics.
66. Research data need different evidence too
Datasets can be reused across studies and deserve persistent identification and citation.
Crossref and DataCite have increasingly emphasised metadata relationships between publications and datasets as part of a richer research graph.
67. Data citations are becoming more visible
In 2026 Crossref introduced a dedicated beta data-citation endpoint that surfaces links between scholarly works and identified datasets.
This is a sign of where bibliometrics is moving: from counting articles toward mapping a network of research objects.
68. The research graph is larger than papers
PERSON ↔ ORCID ↔ INSTITUTION ↔ ROR ↔ GRANT ↔ PROJECT ↔ PAPER ↔ DATASET ↔ SOFTWARE ↔ PREPRINT ↔ PEER REVIEW ↔ CORRECTION ↔ RETRACTION ↔ POLICY / PATENT / GUIDELINE
Future research assessment will increasingly operate on this graph rather than on publication counts alone.
69. Retractions create a metric integrity problem
A retracted paper can continue receiving citations after retraction.
A raw citation count that ignores publication status can reward attention to unreliable research without showing the reason for that attention.
70. Citation databases need status metadata
Correction, retraction, withdrawal and expression-of-concern metadata help downstream systems interpret whether a citation points to a current or compromised publication state.
Crossref’s 2026 public data work has increasingly incorporated research-integrity relationships, including Crossmark and Retraction Watch metadata.
71. A metric should preserve a receipt
A responsible metric should be reproducible from a documented query or dataset.
METRIC RECEIPT Object assessed: Researcher / Paper / Journal / Unit Purpose: Promotion / Discovery / Benchmarking / Review Database: named source Snapshot date: YYYY-MM-DD Coverage: document types / years / fields Deduplication rule: stated Self-citation rule: stated Retraction rule: stated Field normalisation: stated Metric formula: stated Known limitations: stated Human interpretation: recorded
Without a receipt, the number cannot be audited properly.
72. Metric freshness matters
Citation counts change continuously. Database coverage is corrected. Author profiles are merged. Retractions are added.
A metric is a snapshot, not a permanent property of the researcher.
73. Report the date
“h-index 34” is incomplete.
“h-index 34 in OpenAlex, calculated 10 September 2026 under this work-resolution rule” is much more defensible.
74. Data cleaning changes rankings
Duplicate works, misattributed authors and missing references can alter citation counts materially.
Before ranking people, clean the identity graph.
75. Author disambiguation is one of the hardest problems
Many researchers share names. One researcher may publish under initials, changed surnames, transliterations or multiple affiliations.
Persistent identifiers such as ORCID help, but historic records still require probabilistic matching.
76. Institutional disambiguation matters too
Universities rename faculties, merge centres and appear under abbreviations.
Organisation identifiers such as ROR help make affiliation data more stable across time.
77. Rankings amplify small data errors
A missing publication may barely affect a descriptive profile but change a boundary ranking or promotion threshold.
The more consequential the decision, the more rigorous the data verification should be.
78. Thresholds create cliff effects
If a policy requires h-index ≥20, a researcher with 19 can be treated categorically differently from one with 20 despite almost identical records.
Continuous indicators become arbitrary gates when converted into hard thresholds without justification.
79. League tables change institutional behaviour
Universities may recruit, publish or restructure strategically to improve ranking indicators.
Once a ranking influences money and prestige, it becomes part of the institutional incentive system.
80. Responsible metrics require anti-gaming design
- Use multiple indicators rather than one target.
- Keep formulas transparent.
- Inspect unusual citation patterns.
- Separate discovery metrics from evaluation metrics.
- Use qualitative review.
- Audit for field and career bias.
- Review the indicator periodically.
No design removes gaming entirely. The objective is to make manipulation less rewarding than genuine contribution.
81. Metrics can help discovery without deciding worth
Citation networks are excellent for finding influential papers, related authors, emerging clusters and historical pathways.
The same metric may be useful for discovery and inappropriate for promotion.
82. Bibliometrics can map a field
Co-citation, bibliographic coupling, co-authorship and keyword networks can reveal intellectual communities and research fronts.
These maps are models of recorded scholarly relationships, not perfect maps of knowledge itself.
83. Co-citation links works that are cited together
If later papers repeatedly cite A and B together, the two works may occupy a shared intellectual neighbourhood.
Co-citation patterns can reveal conceptual communities even when A and B never cited one another.
84. Bibliographic coupling links works with shared references
If papers C and D cite many of the same earlier works, they may be drawing from a similar knowledge base.
This can be useful for mapping recent work before it has accumulated many incoming citations.
85. Collaboration networks reveal structure, not contribution quality
Co-authorship graphs can identify hubs, institutions and cross-border collaborations.
Being centrally connected does not automatically mean producing better research.
86. Bibliometrics can study science itself
The “science of science” uses publication and citation data to examine careers, collaboration, novelty, diffusion and inequality.
This is a legitimate research field, but it faces the same measurement problem as any other science: observable database traces are not identical to the full underlying phenomenon.
87. Singapore uses research metrics in real institutional workflows
NTU Library’s research services include citation reporting for promotion and tenure, research-impact advisory work, ORCID support, altmetrics guidance and repository services.
This is a useful local model because metrics sit alongside identity, open access, research data and archiving rather than being treated as an isolated score.
88. Singapore’s research mission is broader than citation prestige
Singapore’s RIE system places substantial emphasis on scientific excellence, national priorities, translation, industry relevance and public value.
A responsible national research assessment system therefore needs evidence of contribution and impact that can extend beyond journal citation counts.
89. Metrics can support grant review, but should not pre-decide it
Past productivity and influence can provide context for feasibility and track record.
But high past citation counts do not prove that a new proposal is methodologically strong, strategically relevant or likely to produce value.
90. Promotion assessment should recognise diverse outputs
Datasets, software, standards, patents, clinical practice, policy work, public scholarship, teaching and mentorship can be central to a researcher’s contribution.
CoARA and DORA both push assessment toward broader evidence rather than publication prestige alone.
91. Narrative CVs solve one problem and create another
Narrative CVs let researchers explain contributions that do not fit citation metrics.
They also reward writing skill and can create new forms of strategic presentation. Responsible assessment combines narrative evidence with verification rather than assuming prose is bias-free.
92. Peer review has bias too
Replacing every metric with expert judgement does not eliminate subjectivity.
Experts can be influenced by prestige, networks, disciplinary fashion and implicit bias. The goal is not numbers versus humans. It is transparent, accountable combination.
93. The strongest model is mixed evidence
QUALITATIVE REVIEW + CONTEXTUAL METRICS + CONTRIBUTION EVIDENCE + OPEN RESEARCH PRACTICES + MISSION ALIGNMENT + CAREER CONTEXT + DATA QUALITY CHECK = MORE DEFENSIBLE ASSESSMENT
This is slower than sorting a spreadsheet. It is also closer to the actual decision being made.
94. AI can automate bibliometric mapping
AI can cluster research topics, identify citation pathways, reconcile author identities, summarise portfolios and detect unusual network patterns.
These tools can make large research graphs more navigable.
95. AI can also automate old biases at greater scale
If an evaluation model learns from historical promotion outcomes, publication prestige and citation counts, it may reproduce prior institutional preferences while presenting them as prediction.
Automation can hide value judgements inside statistical weights.
96. LLMs can generate persuasive assessment narratives
A language model can turn a publication list into a fluent summary of “impact”.
Fluency is dangerous when the model has not verified citation data, contribution roles, retraction status or field context.
97. AI-era assessment needs an extension of responsible-metrics principles
Research published in 2026 has begun extending the Leiden Manifesto’s logic to large language models and automated research evaluation.
The core insight is durable: automated tools should support expert judgement, disclose limitations and remain auditable rather than become opaque authorities.
98. AI should return to the underlying works
If an AI says a researcher is “high impact”, the user should be able to inspect the works, citation graph, field baseline and evidence supporting that statement.
Responsible AI assessment needs a return path from conclusion to source.
99. AI should distinguish metric from inference
OBSERVED: 4,820 citations in database X on date Y CALCULATED: field-normalised indicator Z using method M INFERRED: “research has broad scholarly influence” JUDGED: “candidate demonstrates research excellence” Each layer requires different evidence.
The system should not collapse all four layers into one confident sentence.
100. Search engines are becoming invisible assessment infrastructure
Committees, reviewers and researchers increasingly discover work through algorithmic ranking.
What appears on the first screen can shape what gets read, cited and eventually measured.
101. Recommendation systems create feedback loops
Highly cited work is recommended because it is highly cited, then receives more visibility and potentially more citations.
Metrics can become partly self-reinforcing through discovery algorithms.
102. Open metadata can reduce black-box dependence
Crossref and OpenAlex make large portions of scholarly metadata openly reusable, allowing institutions and researchers to audit and reproduce analyses without relying entirely on proprietary dashboards.
Open infrastructure does not guarantee perfect data. It improves inspectability.
103. Proprietary databases still provide valuable curation
Commercial databases can offer carefully curated coverage, specialised indicators, interfaces and institutional analytics.
Responsible assessment should understand the strengths and boundaries of the chosen source rather than assume open is automatically superior or proprietary automatically authoritative.
104. Reproducible assessment needs versioned data
If a committee cannot reconstruct the dataset used for an assessment six months later, disputes become difficult to resolve.
Snapshot dates, exported records and calculation code should be preserved for consequential decisions where appropriate.
105. Metrics need correction pathways
An author profile may merge two people. A paper may be duplicated. A retraction may be missing. A grant may be assigned to the wrong institution.
Assessment systems should allow researchers to challenge data errors before irreversible decisions are made.
106. Appeal is part of metric governance
When quantitative evidence materially affects hiring, funding or promotion, affected researchers should know which data were used and have a route to correct factual errors.
Transparency without correction is incomplete accountability.
107. A metric dashboard should show uncertainty
Many dashboards display exact numbers without showing missing coverage, identity uncertainty or field-definition choices.
Where uncertainty is material, the interface should expose it rather than imply false precision.
108. Rank intervals can be more honest than exact rank
If small metadata changes can move an institution from rank 47 to 39, the exact number may overstate stability.
Confidence ranges, sensitivity analysis or multiple indicators can communicate robustness better than a single ordinal rank.
109. Sensitivity analysis asks whether the result survives reasonable alternatives
Does the assessment change if self-citations are excluded? If another database is used? If the field normalisation changes? If the time window is five rather than three years?
A robust conclusion should not depend entirely on one arbitrary parameter.
110. Metrics are most useful when they answer narrow questions
“How often has this paper been cited in this database?” is a narrow, answerable question.
“Who is the best scientist?” is not one metric question. It hides multiple values and purposes.
111. Define the decision before choosing the metric
DECISION → WHAT VALUE MATTERS? → WHAT EVIDENCE REPRESENTS THAT VALUE? → WHICH METRIC, IF ANY, ADDS INFORMATION? → WHAT CONTEXT IS REQUIRED? → WHAT BIAS CAN ENTER? → WHAT QUALITATIVE REVIEW IS NEEDED? → HOW CAN THE RESULT BE CHALLENGED?
This sequence prevents the common mistake of collecting whatever numbers are easy and then pretending they were the right numbers.
112. Do not use metrics because the dashboard has them
Availability creates temptation. A platform may display h-index, citation totals, journal percentiles and collaboration scores because they are easy to compute.
The decision-maker still has to justify why each indicator belongs in the assessment.
113. Responsible research assessment is a governance discipline
The problem is not solved by replacing one metric with a better metric.
Institutions need policies, reviewer training, transparent criteria, data-quality checks, conflict management, appeals and periodic review of incentives.
114. Metrics should be reviewed because systems evolve
Publishing models change. Open access grows. AI changes discovery. New output types emerge. Databases expand. Researchers learn to optimise old indicators.
An indicator that was reasonable ten years ago may become distortive today.
115. The 2026 environment is materially different
Clarivate’s 2026 Journal Citation Reports now cover more than twenty-two thousand journals across hundreds of categories. Crossref’s open graph contains billions of citation links. OpenAlex provides large-scale open scholarly data. CoARA and DORA continue pushing assessment reform. AI systems can generate portfolio evaluations in seconds.
Scale has increased faster than our ability to mistake scale for wisdom should be allowed to increase.
116. A practical researcher checklist
- Know which databases index your field well.
- Maintain ORCID and institutional identity records.
- Check your author profiles for duplicates and omissions.
- Report metric source and date.
- Do not compare raw metrics across unrelated fields.
- Use narrative evidence to explain contribution.
- Document software, data, policy and teaching outputs where relevant.
- Avoid strategic self-citation or citation exchange.
- Check retraction and correction status of highly cited work you rely on.
- Preserve evidence supporting claimed impact.
117. A practical evaluator checklist
- Define the purpose of assessment first.
- Use qualitative judgement as the main decision process.
- Select metrics only when they answer a specific question.
- Name the data source and snapshot date.
- Check field, career-stage and output-type context.
- Do not use journal metrics as article or researcher quality proxies.
- Inspect contribution roles in team science.
- Use multiple evidence types.
- Test sensitivity to database and parameter choices.
- Allow factual corrections and appeals.
- Review incentives created by the assessment system.
118. A practical institution checklist
- Publish assessment criteria before evaluation.
- Train reviewers in responsible metrics.
- Align indicators with institutional mission.
- Audit databases for disciplinary and demographic coverage.
- Record calculation methods.
- Protect locally relevant and non-journal research.
- Recognise data, software, policy, teaching and mentorship where mission-relevant.
- Avoid hard metric thresholds without evidence.
- Monitor gaming and perverse incentives.
- Version assessment policies.
- Audit AI-assisted evaluation systems.
- Provide a data-correction path.
119. Failure modes
| Failure | What breaks |
|---|---|
| Citation count = quality | Attention is mistaken for methodological merit. |
| Journal metric = article quality | Venue average is assigned to an individual work. |
| h-index without database/date | The number cannot be reproduced. |
| Cross-field raw comparison | Different citation cultures are treated as equivalent. |
| Metric becomes target | Behaviour reorganises around gaming the indicator. |
| Publication count ignores contribution | Team science is misread. |
| Altmetric attention = societal benefit | Visibility is confused with impact. |
| Retraction status ignored | Unreliable papers retain positive metric weight without context. |
| AI evaluates from prestige proxies | Historical bias is automated. |
| No appeal or data correction | Database errors become career decisions. |
| Dashboard chosen before mission | Available numbers define institutional values by accident. |
120. The deeper model: research metrics are a control system
Metrics do not merely describe research. Once connected to money, hiring, promotion and prestige, they help control research behaviour.
MEASURE → RANK → REWARD → RESEARCHER ADAPTS → OUTPUT CHANGES → CITATION NETWORK CHANGES → METRIC DISTRIBUTION CHANGES → INSTITUTION ADJUSTS MEASURE
This feedback loop is why responsible assessment matters. A badly designed metric does not merely misdescribe yesterday. It can reshape tomorrow.
121. The Wintour V1.0 rule: every number must return to its evidence
Under Wintour V1.0, a metric is never allowed to become an orphaned assertion.
NUMBER → FORMULA → DATASET → RECORDS → IDENTIFIERS → SOURCE WORKS → PUBLICATION STATE → HUMAN INTERPRETATION → DECISION RECEIPT
If the return path breaks, confidence should fall.
The purpose of research assessment is not to discover the cleanest number. It is to make the fairest possible judgement about work that matters.
122. Why bibliometrics matters to civilisation
Civilisation produces more knowledge than any individual can read. Metrics emerged partly because institutions need ways to navigate that abundance: which papers shaped a field, where collaboration is growing, which research programmes are visible, how knowledge moves across borders.
That compression is useful. It becomes dangerous when the compression is mistaken for the thing itself.
A society that rewards the wrong research signals can slowly change what researchers choose to study, how they publish, whom they cite and which forms of contribution survive. Responsible bibliometrics is therefore not a technical footnote to science. It is part of the governance of knowledge.
Current authority routes
- DORA: Guidance on the Responsible Use of Quantitative Indicators in Research Assessment
- DORA: Research Assessment FAQs
- CoARA: Agreement on Reforming Research Assessment
- Leiden Manifesto for Research Metrics
- Clarivate: Journal Citation Reports 2026
- Crossref: Two Billion Citation Links
- Crossref: 2026 Public Data File
- Crossref & DataCite: Metadata for Research Integrity
- OpenAlex: Common Attributes and Bibliometric Indicators
- OpenAlex: Citations and Reference Matching
- Hirsch: The Original h-index Paper
- NTU Library: Research Services and Research Impact Support
Continue the Archives and Publishing series
- How Citations, References and Scholarly Linking Work
- How Scholarly Publishing and Peer Review Work
- How Research Integrity and Publication Ethics Work
- How Open Access Publishing Works
- How Preprints and Research Repositories Work
- How Publishing Metadata and Distribution Work
- Wintour House | The eduKate Publishing House
Publication control: Wintour House V1.0 · eduKate Publishing · Rainbolt discovery pass · CivDJ synthesis · evidence, metadata, attribution, context, bias, model-limit, correction, freshness and archive gates.
World Return: The next time a research ranking gives you one clean number, ask for the dirty machinery beneath it: which records, which database, which field, which time window, which formula, which missing outputs, which incentives—and which human judgement still has to be made.