INTRODUCTION: THE BIRTH OF IDEOMETRICS – AN ATTEMPT TO UNIFY THE SCIENCE OF IDEAS
Throughout human history, the capacity to generate, evaluate, and prioritise ideas has been at the heart of individual creativity, collective progress, and civilisational development [1]. Whether the context was philosophical debate in ancient Athens, medical diagnosis in Renaissance Italy, engineering problem-solving during the Industrial Revolution, or policymaking in modern democracies, the underlying process remained the same: faced with multiple competing ideas or possibilities, humans have always needed to decide which are most worthy of pursuit [2].
What differs across epochs, however, is the sophistication of the methods and processes used for such decisions. Traditional cultures emphasised ancestral traditions and folklore, the experiences of elders, intuition, revelation, and spiritual enlightenment [3,4]. Alongside these approaches, many informal and formal techniques have been devised to aid idea generation and application (Table 1) [5–88]. They span across all human activities in their attempts to offer structured and replicable approaches to generating, evaluating, and prioritising ideas, aiming to assist societies in allocating scarce time, funding, and attention to the most promising plans.
Table 1. The landscape of approaches, methods, and tools that have been historically used to generate, evaluate, and prioritise ideas
| Idea generation methods (approaches used to produce novel ideas, concepts, or hypotheses) | Individual and cognitive techniques | Stream of consciousness and freewriting (William James, 1890 [5]; Peter Elbow, 1973 [6]) |
|---|---|---|
| Theory of Inventive Problem Solving – TRIZ (Genrich Altshuller, 1946 [7]) | ||
| Morphological analysis (Fritz Zwicky, 1957 [8]) | ||
| Lateral thinking (Edward de Bono, 1967 [9]) | ||
| SCAMPER technique (Robert F. Eberle, 1971 [10]) | ||
| Mind mapping (Tony Buzan, 1974, 1993 [11,12]) | ||
| Heuristics and biases framework (Amos Tversky and Daniel Kahneman, 1974 [13]) | ||
| Six Thinking Hats (Edward de Bono, 1985 [14]) | ||
| Group-based and social approaches | Brainstorming (Alex F. Osborn, 1942, 1953 [15,16]) | |
| Focus groups and in-depth interviews (Robert K. Merton, Marjorie Fiske, Patricia L. Kendall, 1956 [17]) | ||
| Delphi technique (Norman Dalkey and Olaf Helmer, RAND Corporation, 1963 [18]) | ||
| Nominal Group Technique (Andre L. Delbecq and Andrew H. Van de Ven, 1971 [19]) | ||
| World Café (Juanita Brown and David Isaacs, 1995 [20]) | ||
| Open Space Technology (Harrison Owen, 1997 [21]) | ||
| InnoCentive (Alpheus Bingham, Aaron Schacht, and Dwayne Spradlin, 1998, 2011 [22]) | ||
| James Lind Alliance (Nick Partridge and John Scadding, 2004 [23], with Iain Chalmers) | ||
| Child Health and Nutrition Research Initiative – the CHNRI method (Igor Rudan, 2006, 2008 [24,25]) | ||
| IdeaScale (Vivek Bhaskaran and Rob Hoehn, 2009 [26]) | ||
| Design and innovation frameworks | Human-Centered Design (John E. Arnold, 1958 [27]; Donald A. Norman, 1988 [28]; IDEO.org, 2009 [29]) | |
| Design Thinking (Herbert Simon, 1969 [30]; Tim Brown, 2008 [31]) | ||
| Agile ideation sprints (Ken Schwaber and Jeff Sutherland, 1995 [32]) | ||
| Hackathons (John Gage and OpenBSD, 1999 [33]) | ||
| Lean startup methodology (Eric Ries, 2011 [34]) | ||
| Computational and AI-driven methods | Genetic algorithms and evolutionary computation (John Holland, 1975 [35]) | |
| Automated hypothesis generation (Don R. Swanson, 1986 [36]) | ||
| Generative adversarial networks for idea synthesis (Ian Goodfellow, 2014 [37]) | ||
| Large language models for ideation support (Ashish Vaswani et al., 2017 [38], OpenAI, 2019 [39]) | ||
| Idea evaluation methods (techniques to judge the quality, feasibility, novelty, or value of proposed ideas) | Expert-based evaluation | Peer review (Henry Oldenburg, 1665 [40]) |
| Modified Delphi for scoring (Norman Dalkey and Olaf Helmer, RAND, 1963 [18]) | ||
| Expert panels and consensus conferences (National Institutes of Health, 1977, 1990 [41]) | ||
| Analytical Hierarchy Process (Thomas L. Saaty, 1977,1980 [42,43]) | ||
| Quantitative assessment metrics | Cost-benefit and cost-effectiveness analysis (Abbé de Saint-Pierre, 1708 [44]; Burton Weisbrod, 1960s [125]; Ezra J. Mishan, 1971 [45]) | |
| Net present value and internal rate of return (Irving Fisher, 1907 [46]; Joel Dean, 1951 [47]) | ||
| Patent metrics (US Patent & TM Office, 1947; Adam Jaffe, Manuel Trajtenberg, Bronwyn Hall, 1993 [47]) | ||
| Bibliometric and scientometric indices (Eugene Garfield, 1955 [48] and 1972 [49]; Jorge Hirsch, 2005 [50]) | ||
| Technology Readiness Levels (Stanley Sadin, John C. Mankins and NASA, 1970s, 1995 [51]) | ||
| Scoring models and criteria-based frameworks | Multi-criteria decision analysis (MCDA) (Harold W. Kuhn and Albert W. Tucker, 1951 [49]) | |
| Strengths, Weaknesses, Opportunities, and Threats (SWOT) analysis (Albert S. Humphrey, 1960s [52]) | ||
| Weighted scoring models (Stanley Zionts, 1979 [53]) | ||
| Pugh Matrix – Decision-Matrix Method (Stuart Pugh, 1980s, 1990 [54]) | ||
| Feasibility-desirability-viability framework (Tim Brown, 2009 [31]) | ||
| Crowd-based assessment | Wisdom of the crowd techniques (Francis Galton, 1907 [55]) | |
| Prediction markets (Robin Hanson, 1980s, 1990 [56]) | ||
| James Lind Alliance (Nick Partridge and John Scadding, 2004 [23], with Iain Chalmers) | ||
| Child Health and Nutrition Research Initiative method (Igor Rudan, 2006, 2008 [24,25]) | ||
| Social media engagement metrics as proxies for idea traction (Jason Priem, 2010 [57]) | ||
| Scientific and philosophical validity tests | Logical consistency and deductive reasoning (Aristotle, 4th century BC [58,59]) | |
| Empirical testability, replicability, and falsifiability (Francis Bacon, 1620 [60]; Karl Popper, 1934 [61]) | ||
| Paradigm shift (Thomas Kuhn, 1962 [62]) | ||
| Idea prioritisation methods (methods select the most promising ideas for action, investment, or further study) | Structured decision-making frameworks | Paired comparison methods (Louis L. Thurstone, 1927 [63]) |
| Multi-voting and dot-voting (group facilitation practices, 1950s–1960s [64,65]) | ||
| Delphi with ranking rounds (Norman Dalkey and Olaf Helmer, RAND, 1963 [18]) | ||
| Nominal Group Technique with voting (Andre L. Delbecq and Andrew H. Van de Ven, 1971 [19]) | ||
| Analytic Hierarchy Process (Thomas L. Saaty, 1977,1980 [42,43]) | ||
| Priority-setting frameworks in health and science | RAND/UCLA Appropriateness Method (RAND Corporation and UCLA clinicians, 1980s, 2001 [66]) | |
| Essential National Health Research framework (COHRED, 1990 [67]) | ||
| GRADE methodology with Evidence to Decision frameworks (GRADE Working Group, 2000 [68]) | ||
| James Lind Alliance Partnerships (Nick Partridge and John Scadding, 2004 [23], with Iain Chalmers) | ||
| Combined Approach Matrix (Abdul Ghaffar, Andres de Francisco, Stephen Matlin, 2004; [69]) | ||
| Child Health and Nutrition Research Initiative method (Igor Rudan, 2006, 2008 [24,25]) | ||
| Portfolio and pipeline management tools | R&D portfolio matrices (Bruce Henderson and Boston Consulting Group, 1970 [71]) | |
| Real Options Analysis (Stewart C. Myers, 1977 [70]) | ||
| Product roadmapping and prioritisation grids (Robert Phaal and colleagues, 1970s to 2000s, 2004 [72]) | ||
| Stage-Gate Model (Robert G. Cooper, 1980s, 1990 [73]) | ||
| AI-driven prioritisation tools | Knowledge graphs and semantic similarity clustering (Allan M. Collins and M. Ross Quillian, 1960s [74]) | |
| Reinforcement learning-based portfolio optimisation (Richard S. Sutton and Andrew G. Barto, 1998 [75]) | ||
| Automated priority setting via large language models (Peige Song and Igor Rudan, ISoGH, 2024 [76]) | ||
| Participatory and democratic prioritisation | Cross-cutting philosophical and meta-theoretical approaches (Paul Feyerabend, 1975 [77], and others) | |
| Citizen juries and deliberative democracy forums (Ned Crosby, 1974 [78]; James Fishkin, 1991 [79]) | ||
| Public consultations and e-surveys with weighting (Stephen Sedley, 1985 [80]) | ||
| Participatory budgeting (Tarso Genro and Raul Pont, 1989 [81–83]) | ||
| Approaches to prioritising ideas beyond specific techniques | Occam’s Razor (William of Ockham, 1323-1328 [84,85]) | |
| Bayesian inference (Thomas Bayes, 1763 [86]) | ||
| Dialectical method (Georg Wilhelm Friedrich Hegel, 1807 [87]) | ||
| Epistemic humility and pluralism (John Stuart Mill, 1859 [88], and others – from Socrates [218] to Paul Feyerabend [77]) |
Yet despite their ubiquity and centrality across all domains – from science, medicine, and philosophy, to business, governance, and the arts – these methods and processes have rarely been treated as components of a unified scientific field; instead, they have been dispersed across disciplines, each evolving its own tools, terminology, and criteria. Historically, their evolution was driven by local polymaths, rather than reductionist approaches, and their uptake and persistence likely depended on a degree of serendipity in surrounding contexts and circumstances. In isolation from each other, psychologists, for example, developed creativity techniques, while engineers constructed optimisation models, political scientists refined deliberative forums, economists introduced cost-effectiveness analysis, information theorists examined knowledge flows, and funders in global health used the Child Health and Nutrition Research Initiative (CHNRI) method. Yet, all these approaches were essentially trying to achieve the same goal in their own area of interest, using the same basic approach that led them from generating, to evaluating, and then prioritising ideas.
As a result, many of the underlying intellectual challenges, such as how to quantify the potential of an idea, or assess its moral worth or validity, balance expert versus public input, or prioritise in the face of uncertainty, have been approached in parallel, but have not yet been synthesised. This fragmentation is precisely what motivates the emergence of a new integrative approach that we have named ideometrics: the science of generating, evaluating, and prioritising ideas. Ideometrics is not intended as a reductive or technocratic pursuit, but rather as a transdisciplinary framework that recognises the diversity of human knowledge systems, while seeking coherence in how we work with ideas across them. Just as bibliometrics unified the study of scientific publications [89] and econometrics structured quantitative reasoning in economics [90], ideometrics aspires to provide a systematic lens through which the life cycle of ideas – conception, scrutiny, and their elevation and spread – can be understood and improved.
The philosophical foundations of this project were laid in three earlier conceptual papers by one of the co-authors (IR). The first one proposed that the human brain might be understood not merely as a rational processor or associative network of information gathered by the senses, but as a sensor of ideas itself – an evolved organ that continuously perceives ideas or generates, evaluates, and prioritises them based on the core criteria of attractiveness (i.e. a subjective, emotional component), feasibility (i.e. an objective, rational component), and potential future impact (i.e. a foresight based on an appropriate understanding and valuing of the information available about the context, which includes the brain’s ability to model time, space, and events) [1]. The second paper explored the value of information – a concept traditionally rooted in economics and game theory – as a foundational principle for evaluation of ideas within the context of reality, along with the dangers of disinformation for perception of ideas [91]. It argued that information influences the brain’s sense of ideas through the dimensions of relevance, credibility, and leverage, thus having the capacity to reduce uncertainty, improve prediction, and inform future actions in a physical world. This logic applies equally to scientific hypotheses, business strategies, and ethical frameworks. The third paper focused on what makes science successful and what is necessary for a field of science to originate, thrive, and eventually start to fulfil its mission [92].
Building on those earlier contributions, this work takes a more practical turn: we attempt to provide the first comprehensive synthesis of the many methods that humans have developed and used across numerous disciplines to generate, evaluate, and prioritise ideas. In so doing, we fully acknowledge the futility of trying to identify every tool, framework, or heuristic ever devised as the intellectual landscape is too vast, and many methods are proprietary, undocumented, or culturally specific. Also, historically, those in power have often and continue to resort to various forms of propaganda or violence to ensure that only their ideas thrive, while those of their opponents are censored and silenced [93,94]. We do not think of such approaches to prioritising ideas as scientifically structured or grounded, but we do consider their impact below. Nonetheless, by tracing the evolution of major techniques, spanning from the classical heuristics of Aristotle and Bacon [95] to contemporary systems powered by artificial intelligence (AI), we have sought to create the first map of the landscape of ideas, presenting it in a format that is amendable to future updates and revisions.
As anticipated, most of these methods originated in the social sciences and humanities, often grounded in qualitative reasoning, normative deliberation, and participatory judgment. Techniques like the dialectical method, epistemic humility and pluralism, the Delphi method, focus groups, citizen juries, public consultations, and cross-cutting philosophical and meta-theoretical approaches were developed not merely to reach technical decisions, but to navigate value pluralism, stakeholder inclusion, and legitimacy in contested spaces. Others, like peer review and consensus conferences, aimed to institutionalise epistemic standards [96]. However, in recent decades, the landscape has been transformed by the advent of quantitative metrics, digital tools, and AI. Citation counts, h-indices, and journal impact factors introduced the notion that influence of the proposed ideas within science could be numerically tracked. Cost-benefit and cost-effectiveness analyses formalised trade-offs in economic and policy domains. Crowdsourcing platforms, real-time feedback systems, and online participatory tools allowed for idea generation and ranking at unprecedented scale and speed. Most recently, large language models (LLMs), knowledge graphs, and reinforcement learning (RL) algorithms have begun to automate parts of the idea evaluation and prioritisation pipeline, raising fundamental questions about the future of human creativity, judgment, and responsibility [97].
We believe that this convergence between traditional humanistic methods and emerging machine-assisted systems demands reflection on the science of ideas. Ideometrics is our attempt to provide such a structure: the framework we propose divides the field into three principal domains:
- Idea generation, encompassing classical creativity techniques, collaborative design practices, and computational ideation tools, from brainstorming and the Theory of Inventive Problem Solving (Teoriya Resheniya Izobretatelskikh Zadach (TRIZ) in Russian), and CHNRI’s expert crowdsourcing, to mind mapping, hackathons, and LLMs;
- Idea evaluation, covering both expert-driven and data-driven methods, including peer review, Delphi panels, Analytic Hierarchy Process (AHP), citation and patent metrics, cost-effectiveness analysis, and foundational principles like falsifiability and testability;
- Idea prioritisation, including structured decision tools such as the Stage-Gate Model, Grading of Recommendations Assessment, Development and Evaluation (GRADE), portfolio matrices, as well as participatory and philosophical approaches like participatory budgeting, citizen deliberation, and epistemic humility.
Within each of these three fundamental domains of ideometrics we attempt to trace the historical origins, intellectual foundations, key contributors, methodological innovations, and current applications of each major approach. We also explore how different traditions overlap, diverge, or can be integrated with one another, particularly as digital and AI-based methods gain prominence. Importantly, we view this work not as a static catalogue, but merely as the first iteration of a living framework. As new methods emerge and existing ones are unearthed or gain more visibility in the literature, whether through technological innovation, social experimentation, or philosophical reflection, this structure should allow for their systematic incorporation and comparative analysis. We anticipate that, as fields such as science policy, global health, digital governance, and AI ethics continue to grapple with how to prioritise among competing visions, ideometrics will prove increasingly relevant.
Clearly, at present, dozens of methods for generating, evaluating, and prioritising ideas are used across diverse decision contexts, such as science, policy, technology development, innovation, and governance. However, these methods have evolved independently, they use different terminologies, rest on different epistemological assumptions, and have rarely been examined as manifestations of a shared underlying process. As a result, decision-makers face a fragmented methodological landscape with limited guidance on how to choose appropriate tools, how methods relate to one another, or how their outputs can be compared across domains. This fragmentation leads to several practical challenges, such as inconsistent standards for evaluating ideas across disciplines; poor comparability between prioritisation exercises conducted with different methods; limited cross-fertilisation (because tools developed in one field remain unknown or underused in others); inefficiencies and duplications (as new fields reinvent evaluation techniques without awareness of historical precedents); and a lack of cumulative science (because the methods are not studied within a unified framework). Our goal is thus not simply to introduce a new term, but to provide a conceptual home for the study of how ideas come into existence, are scrutinised, and are ultimately selected. By describing the common structural logic underlying these methods and offering a framework for comparing them, ideometrics addresses a previously unmet need: a systematic, cross-disciplinary science of ideas that is rigorous, transparent, and cumulative. To expand on this introduction, Box 1 explains what makes ideometrics a scientific approach, Box 2 explains how ideometrics differs from existing fields of science, and Box 3 considers what makes ideometrics applicable and relevant.
Box 1. What qualifies ideometrics as a field of science?
For Ideometrics to qualify as a scientific field, analogous to econometrics, bibliometrics, psychometrics, or decision sciences, it should satisfy recognised criteria of scientific legitimacy: falsifiability, testability, methodological rigour, predictive value, and the capacity to support a cumulative research programme. Here, we argue that ideometrics meets all these requirements. We also outline the empirical, conceptual, and methodological foundations that distinguish ideometrics as a science, rather than merely an integrative framework.
- Falsifiability: Ideometrics generates empirically refutable claims
Ideometrics enables explicit hypotheses about how different idea-processing methods behave. These hypotheses are not rhetorical; they are empirically falsifiable. Examples include: different MCDA techniques (e.g. AHP vs. MAUT vs. PROMETHEE) will yield systematically different priority rankings under defined conditions; expert-crowdsourced scoring systems, such as CHNRI, will produce more reproducible rankings than unstructured expert panels; participatory approaches will elevate systematically different types of ideas than expert-driven approaches, controlling for context; hybrid human–machine evaluators (LLMs + experts) will outperform human-only evaluators in consistency or robustness. All of these hypotheses and claims can be scientifically tested and potentially refuted using experiments, simulations, and real-world decision problems. - Testability: Idea-processing methods can be empirically evaluated
Ideometrics provides a structure for rigorous empirical assessment of: inter-method agreement (correlations, concordance coefficients); sensitivity to weighting assumptions or scoring scales; reproducibility of decisions under repeated evaluation; robustness to cognitive biases or framing effects; performance under uncertainty; efficiency and time-cost trade-offs; impact on downstream outcomes (e.g. quality of funded research, performance of selected interventions). Methods can be compared on identical decision tasks, using controlled experiments, algorithmic simulations, or retrospective analysis of historical prioritisation exercises. This should provide a strong empirical foundation. - Methodological rigour: Structured mapping, standardisation, and mathematical characterisation
Ideometrics introduces rigorous methodological infrastructure across previously fragmented practices: formal mapping of method structures (criteria, aggregation rules, weighting logic); quantitative characterisation of assumptions (compensability, independence, sensitivity); standardised reporting frameworks to improve transparency and comparability; replicable scoring systems where different evaluators can produce comparable outputs, with agreement statistics computed; cross-method benchmarking on shared, well-defined decision problems; integration with algorithmic methods (LLMs, genetic algorithms, reinforcement learning) to assess scalability and computational properties; opportunities for experimental manipulation of method parameters to understand causal drivers of divergence. This rigour allows ideometrics to function as a systematic empirical discipline. - Predictive value: ideometrics supports testable predictions about idea performance
A scientific field must generate hypotheses that can be tested prospectively. Ideometrics enables predictions such as: Which evaluation approaches produce the most stable results under high uncertainty? Which prioritisation frameworks best identify high-impact research proposals or technologies? How do idea portfolios evolve under different prioritisation rules? How do human-machine hybrid evaluators perform relative to human-only or machine-only systems? Which methods are most resilient to groupthink, status bias, or institutional incentives? All these predictions can be validated in real-world settings such as science funding, global health research and development, product innovation pipelines, and policy decision-making. - A cumulative research programme: Data, experiments, and theoretical refinement
Like any mature scientific field, ideometrics enables accumulation of evidence and progressive refinement of theory. This includes: shared datasets of evaluation and prioritisation tasks across domains; controlled experiments comparing method performance on identical problems; longitudinal studies examining outcomes of decisions made using different frameworks; meta-analyses identifying which factors most influence method divergence; integration of human cognition and AI, testing how hybrid evaluations improve idea selection; development of validated performance metrics (stability, reproducibility, bias-resistance, efficiency, interpretability); formal theory-building linking cognitive, organisational, behavioural, and computational dimensions of idea generation, evaluation and prioritisation. This cumulative programme differentiates ideometrics from merely being descriptive or integrative.
We believe that ideometrics qualifies as a scientific field because it offers: falsifiable hypotheses; testable predictions with empirical evaluability; methodological rigour; predictive power; and a cumulative research agenda. Taken together, these elements show that ideometrics is not simply a synthesis of existing ideas, concepts, and frameworks, but a new scientific discipline capable of producing generalisable, empirically grounded knowledge about how humans, institutions, and artificial intelligence models generate, evaluate, and prioritise ideas across domains.
Box 2. How does ideometrics differ from existing fields of science?
Although numerous disciplines analyse aspects of how ideas emerge, develop, or are selected, none currently conceptualise the entire lifecycle of ideas as a unified object of scientific study. Ideometrics differs from adjacent fields in three fundamental ways: its scope, level of integration, and methodological ambition.
- Scope: the entire lifecycle of ideas
Decision sciences, innovation studies, behavioural economics, philosophy of science, management studies, and knowledge management each focus on particular stages or contexts of idea development. For example, decision sciences focus primarily on evaluation and choice under uncertainty; innovation studies concern idea development within economic and organisational contexts; knowledge management studies how information and expertise are stored, shared, and transferred; philosophy of science examines justification, logic, and theory selection; and behavioural sciences investigate cognitive biases and heuristics. Ideometrics differs fundamentally in scope, because it treats the entire lifecycle of ideas, i.e. generation, evaluation, and prioritisation, as a single, unified, generalisable cognitive and organisational process that recurs across scientific, technological, organisational, policy, and creative domains. This holistic scope is not present in any existing field – there is no discipline that encompasses all three stages as a single object of scientific study. - Integrative level: a meta-scientific framework
Most disciplines analyse ideas within their own disciplinary contexts (e.g. hypotheses in science, products in innovation, interventions in global health). Ideometrics operates one level above, at a meta-level, by identifying common structures that hold across contexts, independent of domain. It offers a comparative architecture, a cross-domain taxonomy, shared metrics, and a unifying conceptual language. This cross-context integration does not currently exist in any field. Therefore, the primary contribution of ideometrics is to provide a unifying conceptual architecture that connects methods which evolved independently, revealing them as domain-specific solutions of a shared underlying process. - Methodology: a comparative science of idea processes
While many fields use tools for evaluating or prioritising ideas, none systematically study how these tools compare, why their outputs differ, or which methods are best suited to which contexts. There is currently no field of science that maps the relationships among evaluation and prioritisation methods, studies their assumptions and properties, develops standards for transparent reporting, experimentally tests performance across decision contexts, or integrates human and machine contributions within shared frameworks. Ideometrics proposes to fill this methodological gap by establishing a scientific field for the comparative study, validation, and refinement of idea-processing methods, regardless of their origin discipline.
Therefore, ideometrics differs from existing fields not merely by synthesising disparate traditions, but by providing a distinct conceptual object (‘the idea lifecycle’), a higher level of integration, and a methodological agenda aimed at building the first integrative, cumulative, cross-domain science of ideas.
Box 3. What makes ideometrics applicable and relevant?
The term ‘idea’ encompasses an expansive concept, and different decision contexts will generate, evaluate and prioritise different types of ideas. Preserving, studying and understanding this diversity of ideas is one of the main motivations for proposing ideometrics as an integrative scientific field. Although the terminology varies across disciplines, the underlying cognitive and procedural structure of moving from idea generation, to evaluation, to prioritisation is fundamentally similar.
This foundational paper develops a unifying framework flexible enough to accommodate this diversity, rather than to impose premature categorisations that might artificially constrain the field. At this point, we can outline some of the major decision contexts in which ideometrics operates, along with the typical types of ideas encountered:
- Scientific research: hypotheses, study designs, measurement strategies, novel conceptual models;
- Technology and engineering: design concepts, prototypes, algorithmic architectures, technological pathways;
- Business and innovation: product concepts, market strategies, organisational processes, investment opportunities;
- Policy and governance: regulatory options, programme designs, institutional reforms, resource-allocation scenarios;
- Global health and development: intervention models, funding strategies, system-strengthening approaches;
- Creative and humanistic fields: artistic concepts, narrative structures, aesthetic frameworks;
Across these varied contexts, ideas differ substantially in content, scope, and evaluative criteria. Yet the structural questions, such as: What are the possible options? How should they be assessed? Which should be prioritised? – they remain consistent. This structural unity is precisely what ideometrics seeks to define and systematise, to assist decision making by providing the alternative. In research institutes, creative agencies, government ministries, private companies and local communities, instead of governing based on intuition and instinct, a scientific alternative could be used that would reassure the leaders in the likelihood that the ideas they prioritise will eventually succeed in reaching the stated aims.
We could not identify any prior attempts to systematically list methods according to their aim to generate, evaluate, and then prioritise ideas. Given a highly multidisciplinary and dispersed nature of the reviewed information, in addition to searching the standard academic databases (mainly Google Scholar), we also used general Google searches. We selected these based on a clear reference that was the point of origin for each method (listed as references in Table 1). This work represents a preliminary, still imperfect and likely incomplete, but essential contribution toward unifying the study of ideas as structured phenomena, subject to scientific inquiry. It has both a descriptive and a normative component: we believe that by bringing together centuries of fragmented knowledge into a coherent whole, we can better equip individuals, organisations, and societies to navigate the abundance of new ideas in the 21st century, and to focus their attention on those most likely to have a positive impact on humanity’s present and future.
IDEA GENERATION METHODS
Intuitive approaches: individual and cognitive techniques
Stream of consciousness and freewriting
Freewriting and stream-of-consciousness writing are two of the most enduring techniques for unlocking creativity. They share a commitment to spontaneity and unfiltered thought. Although often mentioned together, they emerged from distinct intellectual traditions and have served different purposes. The philosophical foundations of stream-of-consciousness writing can be traced to the American psychologist and philosopher William James. In his landmark work in 1890, ‘The Principles of Psychology’ [5], James introduced the concept of the ‘stream of thought’, later popularised as the ‘stream of consciousness’. He described the human mind as a flowing, continuous succession of ideas and sensations: an ever-shifting experience, rather than a static sequence of mental events. His insights, especially the emphasis on nonlinear and introspective awareness, proved profoundly influential in shaping 20th-century literature. Although James did not formally propose stream-of-consciousness writing as a literary technique, his concept of ‘consciousness as a continuous flow’ inspired literary modernists such as James Joyce, Virginia Woolf, and William Faulkner, who adopted his idea to explore the complexities of human conscious thought in their novels. The style was first formalised in modern literature in Édouard Dujardin’s 1887 novel ‘Les Lauriers sont coupés’ [98], but reached broader recognition with James Joyce’s ‘Ulysses’ in 1922 [99]. These authors developed stream-of-consciousness into a literary device meant to simulate the internal monologue and fragmented experience of consciousness, rather than to generate ideas per se.
Freewriting, by contrast, emerged in the 1970s as a pedagogical tool. Peter Elbow, in his influential 1973 book ‘Writing Without Teachers’ [6], introduced it as a method for overcoming internal censorship and accessing subconscious thought. The technique involves writing continuously for 10–15 minutes without editing, pausing, or planning. This allows thoughts to flow in language with minimal inhibition. While Elbow was inspired by earlier educational theories and perhaps indirectly by Jamesian introspection, he was the first to formalise freewriting as a systematic approach to idea generation in the classroom. The genealogy of these methods reflects a fascinating convergence: Jamesian psychology led to modernist literary experimentation, and then to Elbow’s pedagogical formalisation. While stream-of-consciousness writing aimed to depict the raw texture of thought, freewriting sought to liberate the creative process itself.
Theory of Inventive Problem Solving
The Theory of Inventive Problem Solving, acronymised as TRIZ from its Russian name Teoriya Resheniya Izobretatelskikh Zadach, was developed by Soviet navy engineer and patent examiner Genrich S. Altshuller. Beginning his work in 1946, Altshuller analysed thousands of patents in search of recurring patterns in inventive solutions. This analysis led to the formulation of a structured methodology for creativity in engineering and innovation. Although imprisoned during Stalin’s purges from 1950 to 1954, Altshuller resumed his research upon release. He introduced TRIZ formally in 1956 [7] and continued to refine it over the next several decades. ‘Creativity as an Exact Science’, his major work published in English in 1984, outlined the core components of TRIZ, including dozens of inventive principles, engineering parameters, the contradiction matrix, and patterns of technological evolution [100]. The TRIZ provides a systematic approach to overcoming technical contradictions without compromise; by abstracting problems and applying universal solution patterns derived from successful inventions, it seeks to enable innovation through structured logic rather than trial and error.
Morphological analysis
Morphological analysis is a method of generating creative solutions by systematically exploring all permutations of relevant variables in a structured framework. It was introduced during the 1940s and 1950s by Fritz Zwicky, a Swiss astrophysicist working at Caltech. It was an effort to make invention more routine within aeronautics. Zwicky’s work culminated in two landmark publications: ‘Morphological Astronomy’ in 1957 [8] and ‘Discovery, Invention, Research – Through the Morphological Approach’ in 1969 [101]. His method involved identifying key parameters of a problem, listing all possible values or states for each parameter, creating a multi-dimensional matrix (a ‘Zwicky box’), and exploring combinations to uncover novel, feasible configurations. It exhaustively maps the solution space and the landscape of potential configurations to encourage idea generation through recombination, thus revealing previously overlooked solutions. It has proven especially useful in complex problem structuring, scenario planning, and exploratory design [101].
Lateral thinking
Lateral thinking, a term that entered popular culture for deliberate creativity, was coined in 1967 by Edward de Bono, a Maltese physician, psychologist, creativity theorist and Oxford University’s Rhodes scholar. Unlike traditional ‘vertical’ thinking, which proceeds logically and sequentially, lateral thinking seeks to disrupt habitual mental patterns to provoke fresh insights. De Bono introduced this concept in his book ‘The Use of Lateral Thinking’ in 1967 [9]. He outlined several practical techniques, including breaking out of dominant mental patterns, using deliberate techniques to generate new perspectives and ideas, encouraging discontinuous thinking rather than sequential logic, and challenging assumptions and established frames. Unlike brainstorming or logical deduction, which we discuss later in the text, lateral thinking provokes unexpected connections and reversals to enable creative insight, escaping conventional logic and discovering unexpected solutions. By breaking assumptions and encouraging discontinuities in thought, lateral thinking formalised the idea that creativity could be taught and systematised. It led to innovative training programmes in education, strategy, and organisational problem-solving.
SCAMPER technique
The Substitute – Combine – Adapt – Modify (or Magnify) – Put to another use – Eliminate – Reverse (or Rearrange), (SCAMPER) technique is a structured ideation tool developed by Robert F. Eberle, an American creativity educator and promoter, and introduced in his book ‘SCAMPER: Games for Imagination Development’ in 1971 [10]. Originally aimed at helping children learn how to think creatively, it was quickly adopted in business and product development for its practical utility. Each of the seven prompts that form its acronymised name encourages the user to manipulate an existing idea in a specific way, enabling systematic variation and creative expansion. SCAMPER evolved from earlier brainstorming methods introduced by Alex Osborn (see details on the Delphi technique below), but stands out for its operational simplicity and memorable structure.
Mind mapping
Mind mapping was introduced by British psychologist and author Tony Buzan in the early 1970s. It gained broader attention through his BBC television series ‘Use Your Head’ [11], and was further developed in ‘The Mind Map Book’ in 1993 [12]. Mind maps begin with a central idea and radiate outward in branches, using keywords, images, colours, and spatial layout to mirror the brain’s associative thinking. This nonlinear format facilitates ideation, improves memory, and enhances understanding by engaging both hemispheres of the brain. Mind mapping is particularly useful in brainstorming, planning, studying, and visual problem-solving. Its popularity stems from its intuitive format and cognitive alignment with how people naturally organise information.
Heuristics and Biases Framework
Although not designed specifically for ideation, the Heuristics and Biases Framework, developed by Israeli psychologists Daniel Kahneman and Amos Tversky at the Hebrew University in Jerusalem, has significantly influenced how we understand idea formation and evaluation. First articulated in their 1974 article ‘Judgment Under Uncertainty: Heuristics and Biases’ published in ‘Science’ [13], the framework identified mental shortcuts such as representativeness, availability, and anchoring, noting that, while facilitating fast decision-making, they often lead to systematic cognitive errors or biases. This body of work helped explain why some ideas gain traction and others are prematurely dismissed, laying the foundation for debiasing strategies, such as ‘devil’s advocacy’, red teaming, and structured decision protocols, that seek to enhance ideation processes by counteracting intuitive blind spots. The collaboration between Kahneman and Twersky is credited to have established behavioural economics; Kahneman later moved to Princeton University, USA, and received the Nobel Prize in Economics in 2002 for integrating insights from psychological research into economic science, especially concerning human judgment and decision-making under uncertainty. His lifetime work has been summarised in a popular book ‘Thinking, Fast and Slow’ in 2011 [102].
Six Thinking Hats
Also developed by Edward de Bono, the Six Thinking Hats technique was introduced in the eponymous 1985 book [14] to structure group thinking and ideation. Each ‘hat’ represents a distinct cognitive role or mode and is accordingly given its own colour: neutral facts and data (white); emotions and intuition (red); caution and critical judgment (black); optimism and benefits (yellow); creativity and alternatives (green); process and meta-thinking (blue). By synchronising group attention on one thinking mode at a time, this method reduces conflict, enhances focus, and promotes constructive ideation. The green hat is explicitly used for creative generation, but all hats contribute to well-rounded idea development.
Participatory, collaborative methods and surveys for idea generation: group-based and social approaches
Brainstorming
Brainstorming was introduced by advertising executive Alex F. Osborn in the late 1930s to address the lack of innovation in team settings at his advertising agency, Batten, Barton, Durstine & Osborn. He first described the technique briefly in his book ‘How to “Think Up”’ in 1942 [15] and expanded it fully in ‘Applied Imagination’ in 1953 [16]. Building on the ideas of others, such as Lewin’s ‘field theory’, Osborn’s method emphasised four cardinal rules: defer judgment, strive for quantity, encourage wild ideas, and combine and improve ideas. These principles were intended to foster psychological safety, stimulate divergent thinking and creative collaboration, and break habitual patterns of thought in group settings. Despite later critiques questioning its comparative effectiveness in groups versus individuals, brainstorming laid the foundation for later creativity frameworks, including SCAMPER, TRIZ, lateral thinking, and design thinking. It remains a widely used starting point for ideation in education, business, and research.
Focus groups and in-depth interviews
Group interviewing techniques, a precursor to focus groups, were used in the 1920s for developing survey questionnaires and other social research purposes. Focused, in-depth interviews emerged as a systematic method for eliciting group insights in the early 1940s, introduced by sociologist Robert K. Merton in collaboration with Patricia L. Kendall and Marjorie Fiske. Working for the US Office of War Information during World War II, Merton sought to understand audience responses to radio broadcasts and propaganda through what he termed ‘focused interviews’. These interviews emphasised participant interaction, spontaneous discussion, and subjective reactions to specific stimuli. The foundational publication, ‘The Focused Interview: A Manual of Problems and Procedures’ (1956) codified this method [17]. It introduced the principles of non-directive moderation, guided but open-ended questioning, and the analytical value of group dynamics. Although the term ‘focus group’ gained popularity and recognition later, popularised by Ernest Dichter [103] and other scholars, Merton’s methodological contribution laid the groundwork for its widespread adoption. Initially rooted in social research, focus groups expanded into market research, healthcare, public policy, and innovation by the latter half of the 20th century. Their ability to generate and explore perceptions, emotions, and preferences made them especially useful in idea generation, concept testing, co-creation workshops, and participatory design. By the 1980s, works such as David L. Morgan’s ‘Focus Groups as Qualitative Research’ helped formalise the approach in academic literature [104]. Today, focus groups remain a cornerstone of qualitative inquiry and idea generation across diverse fields, valued for their capacity to surface detailed and unexpected insights through structured, yet flexible dialogue.
Delphi technique
The Delphi method was developed in the 1950s by Norman Dalkey and Olaf Helmer at the RAND Corporation as a tool for structured expert consultation, initially aimed at military and technological forecasting during the Cold War. Its foundational principles were iterative surveys of respondents for their proposed ideas, anonymity, controlled feedback, and statistical synthesis. They were all designed to minimise bias from group dynamics while enabling ‘collective intelligence’. The technique was first formally introduced to the academic community in the paper ‘An Experimental Application of the Delphi Method to the Use of Experts’ in 1963, which outlined its procedures and logic [18]. Experts respond to a series of questionnaires in multiple rounds, with anonymity to reduce dominance bias. Controlled feedback is provided after each round, with responses aggregated and refined until convergence towards consensus is achieved. Though originally intended for forecasting, the Delphi method quickly found applications in strategic planning, health research prioritisation, foresight exercises, and multidisciplinary consensus-building [18]. Its rigorous, yet flexible structure makes it especially suitable for contexts where empirical data are limited and expert judgment is essential. It became widely used in trying to find consensus, which also meakes it relevant to evaluating and prioritising ideas (see later).
Nominal Group Technique
The Nominal Group Technique (NGT) was introduced in the late 1960s by organisational theorists Andre L. Delbecq, Andrew H. Van de Ven, and later David H. Gustafson at the University of Wisconsin-Madison. It aims to balance individual ideation with structured group decision-making, encouraging equal participation and minimising conformity pressure and the influence of dominant personalities or group dynamics on idea generation, making it especially useful for diverse groups. Participants first generate ideas silently and independently, after which they share them without discussion in a round-robin format. After clarification, they privately vote or rank ideas to establish group priorities. This structured flow ensures balanced participation and clear prioritisation (see later). The technique was described in their article ‘A Group Process Model for Problem Identification and Program Planning’ in 1971 [19], with a definitive guide and distinction from brainstorming and Delphi processes in the book ‘Group Techniques for Program Planning’ in 1975 [105]. The NGT has since become a popular tool in public health, strategic planning, and participatory research settings.
World Café
The World Café method was introduced by Juanita Brown and David Isaacs in 1995. While designing strategic dialogue forums with senior leaders, they came up with the World Café as a way to harness collective intelligence through structured conversation. Inspired by an impromptu café-style dialogue at real-life meetings, they formalised the method into a participatory approach for large-scale idea generation. The process simulates a café environment with small tables, rotating participants, and a ‘table host’ who captures the evolving dialogue. Over successive rounds, ideas are cross-pollinated, deepened, and synthesised. A concluding ‘harvest’ session captures emerging themes. Brown and Isaacs’ 2005 book, ‘The World Café: Shaping our Futures through Conversations that Matter’, details the principles of the design, offers implementation guidance, and showcases real-world applications in education, business, and public policy [20].
Open Space Technology
Open Space Technology was introduced in the mid-1980s by Harrison Owen, an episcopal priest whose academic training focused on the nature and function of myth, ritual, and culture. It is a self-organising facilitation method for addressing complex issues. Inspired by the informal creativity observed during conference coffee breaks, Owen created a format in which participants build their own agendas around a central theme. Key features include an open marketplace of ideas, voluntary participation, and flexible session structure guided by principles such as ‘whoever comes is the right people’ and ‘the law of two feet’ (encouraging people to move between sessions based on interest). Open Space Technology was first implemented in 1985 and formally described in Owen’s book ‘Open Space Technology: A User’s Guide’ in 1997 [21]. It has since been adopted in change management, crisis response, multi-stakeholder engagement, and innovation workshops.
InnoCentive and other ‘open innovation’ platforms
InnoCentive, a platform for ‘open innovation’, was conceived by Alpheus Bingham and Aaron Schacht in 1998 at Eli Lilly and Company. The initial idea evolved into ‘Molecule.com’ and then ‘BountyChem’ before becoming InnoCentive. Bingham, then a senior executive at Eli Lilly, helped create InnoCentive as an external innovation platform to tackle research and development (R&D) bottlenecks. The company was launched in 2001 and became a spin-out in 2005. It sought to crowdsource solutions to challenges in R&D by connecting organisations with global problem-solvers. The platform invites ‘seekers’ to post technical challenges, which are tackled by a distributed network of ‘solvers’, with winning solutions receiving cash prizes. The approach was elaborated in the book ‘The Open Innovation Marketplace’ in 2011, co-authored by Alpheus Bingham and Dwayne Spradlin [22], which explained the logic of challenge-driven innovation. InnoCentive helped catalyse the broader open innovation movement, inspiring similar platforms in data science (Kaggle), global development (XPRIZE) and space technology (NASA’s Centrifuge).
The James Lind Alliance priority setting partnerships
The James Lind Alliance (JLA) was initiated in 2004 by Sir Iain Chalmers, a British physician, health services researcher, and one of the founders of the Cochrane Collaboration, with subsequent support from the funding bodies such as the UK Medical Research Council (MRC) and the National Institute for Health and Care Research (NIHR). It was first described by its co-founders, Sir Nick Partridge and John Scadding, in an article published in ‘The Lancet’ in 2004 [23]. The goal of this non-profit initiative is to promote equitable and transparent research prioritisation by bringing together patients, carers, and clinicians in structured collaborations. These priority setting partnerships (PSPs) aim to identify and rank unanswered questions in healthcare that are most relevant to those directly affected. The initiative was named after James Lind, the 18th-century Scottish naval surgeon credited with conducting one of the first controlled clinical trials on scurvy. The PSPs perform a series of steps: collection of uncertainties, verification against existing research, stakeholder surveys, and consensus workshops using voting-based techniques, often adapting the NGT (see later). The final outcome is a list of ‘top 10’ priorities that informs research funding, guideline development, and policy planning.
The JLA PSPs generate research ideas in the first step through a broad stakeholder survey, often referred to as the open call for the ‘initial survey of uncertainties’. Then, patients, carers, and clinicians submit questions or uncertainties about a particular health condition or topic that they believe should be answered by research. Ideas are collected through an online or paper-based survey disseminated via charities, professional bodies, clinics, social media, etc. The survey is intentionally open-ended, and respondents are not prompted with predefined questions. The method ensures that research ideas come directly from those most affected, and not from academics or policymakers, helping to surface real-world uncertainties that may not be apparent to researchers alone. The raw responses are then reviewed and interpreted by the PSP team, duplicates merged, out-of-scope entries removed, and submissions rewritten into researchable summary questions using the patient/population, intervention, comparison, and outcome (PICO) format where possible. Each summary question is checked against existing evidence, such as systematic reviews, to ensure it is a true uncertainty. This bottom-up ideation approach ensures the prioritised research agenda is relevant, meaningful, and co-produced by those who live with or treat the condition.
The Child Health and Nutrition Research Initiative method
The Child Health and Nutrition Research Initiative (CHNRI) method was initially developed in 2006 by Igor Rudan from the University of Edinburgh, a medical doctor with interests in anthropology, human genetics, epidemiology and global health economics. He was contracted as a consultant of the CHNRI of the Global Forum for Health Research in Geneva, Switzerland, by Professors Shams El Arifeen and Robert E. Black, who co-authored the first version of the CHNRI methodology’s conceptual framework [24]. Detailed guidelines for implementation were then published by Rudan and colleagues in 2008 [25]. Originally focused on child health research in low- and middle-income countries (LMICs), the method is credited for achieving several conceptual advances, such as providing a way of generating hundreds of research ideas in a systematic and transparent way. In short, many leading experts in a given field propose ideas for research, which are then grouped into: those that can better describe the problem; those that can optimise the way in which current resources and interventions are addressing the problem; those that can further develop and improve the available interventions, to make them more cost-effective and/or equitable; and those that could discover entirely novel interventions. When applied to the field of health research, this ‘four D’ framework – ‘description, delivery, development and discovery’ – translates into epidemiological research, health systems and policy research, research on improving the existing health interventions, and research to discover new health interventions. From those four fundamental instruments of health research, hundreds of proposed research ideas would then be categorised into broad ‘research avenues’, narrower ‘research options’, and specific ‘research questions’, accounting for the ‘depth and breadth’ of each proposed idea. This method also enables the analysis of the level of saturation of the proposed spectrum of ideas: it can trace the increase in probability that any further ideas would have already been suggested previously by contributing experts. Since 2008, the CHNRI method has been widely adopted and applied by the World Health Organization (WHO), United Nations Children’s Fund, national governments, academic collaborations, and major global funders across diverse domains of health and development research. Its conceptual advances make it flexible and applicable to almost any problem in any area of science. We will review its approach to evaluating and prioritising ideas later in the text.
IdeaScale
IdeaScale, launched in 2008–09 by Vivek Bhaskaran and Rob Hoehn, is a cloud-based software company that licenses the eponymous platform for crowdsourcing ideas and managing innovation workflows [26,106]. It gained immediate traction when 23 US federal agencies adopted it under President Obama’s Open Government Initiative. The platform enables users to submit, vote on, comment, and refine ideas in a structured interface. A configurable workflow guides submissions from inception to evaluation and implementation, promoting collective intelligence in both public and corporate settings. While lacking a single key academic publication, IdeaScale’s success is documented in industry case studies and civil technology literature. It remains a leading software-as-a-service solution for participatory innovation in innovation management.
Generating ideas for creation: design and innovation frameworks
Human-centered design
Human-centred design (HCD) is an approach to innovation that begins with understanding the needs, behaviours, and experiences of the people for whom solutions are being developed. While the roots of the concept extend across disciplines, including engineering, psychology, anthropology, and the arts, the formalisation of HCD began in the mid-20th century. One of the earliest proponents of the approach was Professor John E. Arnold, who taught creative engineering at MIT and later at Stanford, where he helped establish the design programme in 1958 [27]. Arnold’s work emphasised that engineering and product development should centre on human needs, setting the stage for the philosophy that would later mature into HCD. The concept evolved significantly in the late 20th century, gaining traction in design, technology, and public service. Cognitive scientist Donald A. Norman played an important role in establishing the psychological foundations of user-centred design, most notably through his book ‘The Psychology of Everyday Things’ in 1988, which highlighted usability, intuitive interfaces, and feedback systems [28]. HCD was further popularised by the design firm IDEO, founded by David Kelley, which helped institutionalise the methodology in business and development sectors. The key elements of the process are inspiration, through deeply understanding people’s lives and challenges; then, ideation, through generating and refining ideas collaboratively; and implementation, through prototyping, testing, and iterating based on real-world feedback. The collaboration between IDEO and Stanford’s d.school – officially Hasso Plattner Institute of Design – catalysed a global movement in human-centred innovation. A widely used field manual, The ‘Field Guide to Human-Centered Design’ by IDEO.org in 2015 [29], later made the methodology accessible to social innovators, non-governmental organisations, and policy practitioners.
Design thinking
Design thinking is a human-centred, iterative process for creative problem-solving, distinguished by its blend of analytical rigour and empathy-driven insight. Its intellectual roots trace back to Herbert Simon’s ‘The Sciences of the Artificial’ in 1969 [30], where he described design as a generic cognitive strategy for addressing ill-structured problems. Robert McKim’s ‘Experiences in Visual Thinking’ in 1973 [107] further emphasised creativity and visual reasoning in engineering design. Yet it was in the 1990s and 2000s that the method was crystallised and disseminated globally, primarily through the work of David Kelley, founder of IDEO and Stanford’s d.school. IDEO’s CEO Tim Brown articulated and popularised the philosophy in his influential book ‘Change by Design’ in 2009 [31]. It helped formalise design thinking as a structured method involving five nonlinear stages: empathise, i.e. deeply understand users’ needs; then, define, by framing the core problem; follow by ideating, through brainstorming creative solutions; develop a prototype, thus building tangible representations; and, finally, test, gather feedback and refine. The process emphasises human-centeredness, rapid prototyping and iteration, and cross-disciplinary collaboration. Design thinking has been widely adopted across industries, from business and education to healthcare and international development, often overlapping with or incorporating HCD.
Agile ideation sprints
Agile ideation sprints are short, structured cycles for collaboratively generating and validating ideas. They emerged from the Agile software development movement, particularly Scrum methodology, and were later formalised for design and innovation contexts by Google Ventures in the form of the ‘design sprint’. Jeff Sutherland and Ken Schwaber formalised Scrum in the context of the Agile movement in 1995, emphasising iterative progress, team collaboration, and flexibility [32]. Although early sprints focused on delivering working software, they also fostered idea evolution during planning and review sessions. The ‘Manifesto for Agile Software Development’, presented in 2001, introduced principles such as rapid iteration, collaboration, and working solutions over documentation [108]. The concept of a dedicated ‘design sprint’ was later introduced by Jake Knapp at Google Ventures in the early 2010s. Documented in the book ‘Sprint’ in 2016, this five-day process guides teams from problem framing to a tested prototype [109]. It combines principles from design thinking, user research, and Agile workflows, making it accessible beyond software teams. The typical sprint includes mapping the problem, sketching ideas, deciding on a solution, prototyping, and testing with real users – all in a compressed timeframe [109].
Hackathons
Hackathons are intense, time-bounded events where individuals or teams collaborate to generate, develop, and prototype new ideas, often in the domains of software, hardware, or policy innovation. The term ‘hackathon’ originates from the words ‘hack’ and ‘marathon’, where ‘hack’ is used in the sense of exploratory programming (and not a reference to breaching digital security) [33]. It was independently coined in June 1999 during two separate events. The first was a cryptographic coding sprint organised by Theo de Raadt and the OpenBSD community in Calgary, Canada. Soon after, Sun Microsystems hosted the JavaOne Hackathon, where developers worked for 24 hours to build Java applications, marking the first branded corporate hackathon [33]. Initially focused on software engineering, hackathons quickly spread to fields like health innovation, civic technology, education, and sustainability. Their hallmark is rapid problem-solving under tight constraints, culminating in demo pitches to peers or judges. While no single publication introduced hackathons academically, key studies have since analysed them as engines of innovation. Gerard Briscoe and Catherine Mulligan’s paper ‘Digital Innovation: The Hackathon Phenomenon’ in 2014 explored their role in digital ecosystems [110], Lilly Irani examined how they shape entrepreneurial culture in 2015 [111], while Thomas James Lodato and Carl DiSalvo studied them as vehicles for participatory and issue-driven design in 2016 [112].
Lean startup methodology
The lean startup methodology provides a systematic framework for developing ideas and innovations through iterative testing, rapid prototyping, and validated learning. Introduced by entrepreneur Eric Ries in the late 2000s, it draws from lean manufacturing, Agile software development, and Steve Blank’s customer development model [113]. Ries began promoting the framework through blog posts and talks around 2008, eventually formalising it in his bestselling book ‘The Lean Startup’ in 2011 [34]. The core cycle of build-measure-learn encourages innovators to test ideas or hypotheses via minimum viable products, then launch quickly, measure performance, and learn from real user behaviour, use learning to pivot or persevere, validate learning, measure progress by evidence that an idea works in the real world, track meaningful learning and progress through metrics tied to business hypotheses, and make structured course corrections based on feedback and data. It challenges traditional notions of success based on features or funding, instead measuring progress through evidence of what works. Concepts like ‘innovation accounting’ and ‘pivoting’ became part of the startup and corporate lexicon, with major adoption by companies like General Electric, Intuit, and government innovation programmes.
Approaches to ideation in the digital age: computational and AI-driven methods
Genetic algorithms and evolutionary computation
Genetic algorithms (GAs) and the broader domain of evolutionary computation apply principles of biological evolution to computational problem-solving and idea generation. This field was pioneered by John H. Holland, professor of psychology, electrical engineering, and computer science at the University of Michigan in the 1960s. He initially developed these concepts while studying cellular automata and then tried to apply principles of biological evolution to computer systems. Holland formalised the theoretical foundations of GAs in his landmark work ‘Adaptation in Natural and Artificial Systems’ in 1975 [35]. He encoded potential solutions as ‘chromosomes’/‘strings’, and then used mechanisms like crossover and mutation to evolve and optimise them over successive ‘generations’. These solutions are evaluated using a fitness function and refined through selection, crossover, and mutation, closely mimicking Darwinian processes of natural selection. Key components of GAs include: chromosomes (encoded representations of candidate solutions; fitness function (evaluates the performance of each solution); selection (favours better-performing solutions for reproduction); crossover (combines parts of parent solutions to form offspring); mutation (introduces variation to explore new possibilities); generations (repeat cycles of selection, recombination, and mutation over time). GAs have become a foundational tool in evolutionary computation, used across optimisation, design, computational creativity, engineering design, logistics, neural architecture search in deep learning, robotics and game AI, and artistic generation.
Automated hypothesis generation
The concept of automated hypothesis generation, where machines identify novel, testable scientific ideas, originated in the 1980s and has since evolved into a powerful class of tools for discovery. The foundational idea was articulated by Don R. Swanson, a medical informatician at the University of Chicago, in 1986 through his ‘literature-based discovery’ and the ‘theory of undiscovered public knowledge’ [36]. Swanson demonstrated that connections could be inferred between disjoint literatures, for example, linking fish oil and Raynaud’s syndrome via blood viscosity, by identifying intermediate ‘B-terms’. This logic became the basis for platforms such as Arrowsmith, co-developed by Swanson and Neil Smalheiser [114], and later tools like LION LBD [115], IBM Watson Discovery [116], and Semantic Scholar [117]. These tools automate the identification of ‘B-terms’ linking disconnected ‘A’ and ‘C’ domains (e.g. A → B → C) and thus propose novel hypotheses, particularly in biomedicine and materials science. Recent advances leverage AI and graph-based embedding to generate hypotheses in neuroscience, physics, and chemistry.
Generative adversarial networks for idea synthesis
Generative adversarial networks (GANs) were introduced by Ian Goodfellow, an American computer scientist, engineer, and executive, and a pioneer of artificial neural networks and deep learning. He worked as a research scientist at Google DeepMind and Google Brain, was the director of machine learning at Apple, and was one of the first employees at OpenAI. He developed GANs in 2014 with his colleagues at the Université de Montréal [37], thus revolutionising the landscape of machine creativity. GANs consist of two neural networks, a ‘generator’ and a ‘discriminator’, that compete in a zero-sum game. The generator produces synthetic outputs, while the discriminator attempts to distinguish real from fake data. Through this adversarial training process, the generator learns to create outputs that are increasingly indistinguishable from real data. This architecture enables the synthesis of realistic and novel content across a range of domains, including art, design, language, science, and drug discovery. It laid the groundwork for a vast field of research and application, with a large potential that is difficult to fully grasp at this time. While GANs generate novel content (e.g. images, text), their primary conceptual contribution is the adversarial training framework itself, which translates into a model for generating novel ideas.
LLMs for ideation support
LLMs have become powerful tools for supporting ideation through natural language. Built upon the transformer architecture introduced by Ashish Vaswani and colleagues in 2017 [38], LLMs evolved from statistical language models into general-purpose reasoning and creativity engines. OpenAI’s Generative Pre-trained Transformer (GPT) series was the first to demonstrate emergent ideation capabilities at scale. GPT-2 revealed early signs of generalisation, creativity, and ideation abilities with zero-shot and few-shot learning in 2019 [39]. Soon after, GPT-3 introduced robust few-shot learning, reasoning, synthesis, and creative, domain-crossing ideation in 2020 [118]. These models could generate hypotheses, expand on prompts, synthesise analogies, and co-create with human users in diverse fields, becoming ‘collaborative ideation tools’ or ‘co-creative support systems’ in fiction, design, business, science, policy, and other fields. LLMs learned these capabilities not through hardcoded rules, but through massive-scale pretraining on Internet-scale corpora, enabling them to recombine concepts in novel ways – ‘ideating’ not by explicit programming, but through latent concept combination and abstraction. Therefore, LLMs now routinely assist in research ideation, scientific writing, product design, fiction, strategy, and many other tasks, constituting a new era of ‘machine-augmented creativity’. As of 2024, models like GPT-4, Claude, Gemini, DeepSeek, and open-source systems such as LLaMA and Mistral continue to push the frontiers of AI-supported creativity – again, with seemingly vast potential for ideation that is difficult to predict at this time.
IDEA EVALUATION METHODS
Methods for evaluating ideas based on knowledge: expert-based evaluation
Peer review
The practice of peer review, which consists of soliciting judgments from experts in subject matter to assess the validity, originality, and relevance of research ideas, has a long history. Informal scholarly critique dates back centuries, with critiques of ideas through lectures, books, papers and other means – a useful example being Risalah, one of the most famous books of Islamic jurisprudence. However, the structured, editorial peer review system as practiced today is a relatively modern institutional development. The earliest known examples of peer-like evaluation emerged in the 17th century with the Royal Society of London and its journal ‘Philosophical Transactions’, introduced it in 1665 under the editorship of Henry Oldenburg [40]. A similar practice has been associated with the ‘Journal des Sçavans’ in France.
Oldenburg is often regarded as an early pioneer of the peer review model due to his habit of seeking informal feedback on submissions from other scholars. However, this process was inconsistent, highly discretionary, and lacked transparency. Throughout the 19th century, prominent journals such as ‘The Lancet’ and ‘Nature’ relied largely on editorial judgment rather than external review. The contemporary model of anonymous or named external experts rigorously evaluating submitted manuscripts only emerged in the mid-20th century. This transformation was catalysed by several post-World War II developments, including a surge in scientific output and the expansion of publicly funded research via institutions such as the National Institutes of Health (NIH) and the National Science Foundation. These trends created an urgent need for systematic, fair, and expert-driven evaluation mechanisms to allocate limited research funds and publication space. One of the earliest pioneers in formalising peer review was the ‘Journal of the American Medical Association’ (JAMA), which began implementing structured review processes in the 1940s. However, even leading journals like the ‘New England Journal of Medicine’ did not adopt external peer review until 1976. A landmark analysis of the historical development of editorial peer review was provided by John C. Burnham in JAMA in 1990 [119]. This paper traces how editorial judgment gave way to systematic external scrutiny, documenting the social and institutional forces that shaped the peer review model as a gatekeeping mechanism in modern science.
Modified Delphi for scoring
The origins of the Delphi method at the RAND corporation were described in the previous text. It has proved to be a helpful structured technique for eliciting expert judgment, particularly well-suited to situations with incomplete data, high uncertainty, or multidisciplinary inputs. It was initially a part of classified military forecasting projects during the Cold War, so although it was used since 1950s, it became publicly known in 1963 [18]. The key innovation was a systematic approach to achieving expert consensus through multiple rounds of anonymous input, controlled feedback, and statistical aggregation of opinions. Here, we focus on the way that it has been used to evaluate proposed ideas. This was done through iterative rounds of input with controlled feedback, and then statistical convergence toward consensus and quantifiable scoring of subjective judgments. The Delphi method uses median or mean scores for any quantitative items, supporting it by frequency distributions or histograms to show variability among the scorers and thematic analysis for any qualitative comments. This aggregation represents the collective judgment of the panel in the first round, which is then followed by feedback to experts, showing how their responses compare to the group’s central tendency, how much disagreement or convergence exists, and any justifications or rationales offered by other participants. This step allows participants to re-evaluate their own positions in light of the group’s thinking. Then, experts are asked to reassess their initial responses and re-rate the ideas across two to four rounds, typically using either Likert scales (e.g. 1–9 for importance, feasibility, and impact) or ranking or scoring of competing ideas. The aim is to move toward consensus over successive rounds, though some level of dissent is acceptable and even encouraged to identify uncertainty or alternative scenarios; namely, consensus does not necessarily mean unanimity, but rather convergence or a stable majority, which is better reflected in the use of interquartile ranges. By the final round, ideas that have achieved high median ratings, narrow interquartile ranges that indicate consensus, and stable ratings across rounds are considered robust, high-priority, or widely supported. Conversely, ideas with low scores or wide disagreement may be deprioritised or marked for further study. Therefore, evaluation of ideas in the Delphi process is, in essence, both quantitative (through aggregated expert ratings, medians, and convergence indicators), qualitative (through open comments, rationales, and recorded justifications), iterative (refining judgments through exposure to anonymous peer perspectives), and systematic (using structured questionnaires and statistical summaries). Although originally designed for forecasting technological developments and national security scenarios, the Delphi method was quickly adapted for use in healthcare, education, and research prioritisation, particularly where empirical data were limited or absent. Its core strength lied in its ability to combine subjective expert opinions and insights with procedural rigour and replicability [18].
Expert panels and consensus conferences
Structured expert panels and consensus conferences represent another classical approach to evaluating and prioritising research ideas, particularly in contexts where high-stakes decisions must be made in the absence of definitive empirical evidence. These methods gained prominence during the 1970s and 1980s, especially in the field of health policy. The NIH Consensus Development Program, launched in 1977, institutionalised the practice of consensus development conferences [41], which brought together multidisciplinary panels of experts to review available evidence, hear stakeholder input, and arrive at public consensus statements on key medical or research issues. The Delphi technique and other consensus development techniques, such as the NGT, differed in their structure and anonymity from consensus development conferences, but they all aimed to facilitate reliable judgments among diverse experts. These methods became standard in organisations such as the United Nations (e.g. its Millennium Development Goals), NIH, the WHO, the Agency for Healthcare Research and Quality (AHRQ), and various national research councils. Much of the foundational documentation on consensus conferences resides in institutional reports, highlighting the role of expert panel methods in transforming subjective expertise into defensible, consensus-based research recommendations [120].
Analytical Hierarchy Process
The Analytical Hierarchy Process (AHP) is a mathematical and psychological framework for structured decision-making under complex, multi-criteria conditions. Developed by Thomas L. Saaty, a professor at the Joseph M. Katz Graduate School of Business at the University of Pittsburgh in the 1970s, the AHP allows for transparent prioritisation by decomposing decisions into hierarchical elements and comparing them pairwise [42]. The key components of the method’s contribution to evaluating ideas are decomposition of complex problems into hierarchical levels, their pairwise comparisons using a scale from 1 to 9, weighting of alternatives using eigenvalue computation and consistency ratio to assess the reliability of judgments. The AHP was further elaborated in Saaty’s book ‘The Analytic Hierarchy Process: Planning, Priority Setting, Resource Allocation’ in 1980 [43]. Originally applied to engineering and military strategy, the AHP was soon adopted in public policy, health research, and strategic planning. Its ability to make trade-offs explicit made it particularly attractive for research funding allocation, technology assessment, and national science agenda setting, with adoption by institutions such as the National Science Foundation and the European Commission’s Framework Programmes. The AHP became one of the first structured, quantitative tools applied to evaluate and prioritise research ideas based on multiple weighted criteria, such as scientific merit, feasibility, societal impact, and cost-effectiveness. Its quantitative approach and transparency allowed it to be combined with Delphi panels or newer frameworks like the CHNRI method, especially when multiple stakeholders must weigh diverse criteria. We should also explain that AHP is one of the earliest, most influential, and most widely applied multi-criteria decision analysis (MCDA) techniques, and that it is presented separately from MCDA only to illustrate its historical and practical significance, not to imply that it lies outside the MCDA category.
Crunching the numbers: quantitative assessment metrics
Cost-benefit and cost-effectiveness analysis
Two additional tools that have profoundly influenced how ideas and projects are evaluated, particularly in public policy and health economics, are cost-benefit analysis (CBA) and cost-effectiveness analysis (CEA). They share roots in welfare economics, but were formalised in distinct institutional and disciplinary contexts. The roots of CBA can be traced back to early economic thinkers like Abbé de Saint-Pierre in France in 1708, who conducted one of the earliest cost-benefit analyses, specifically focusing on the utility of road improvements, while in 19th-century France, Jules Dupuit introduced the concept of consumer’s surplus that established its economic basis [121]. Alfred Marshall, a prominent British economist, focused on supply and demand in his book ‘Principles of Economics’ published in 1890 [122], which laid the foundation for neoclassical economics. Relevant to CBA, he also refers to the ‘Green Book’, a concept of economic progress and social welfare that includes access to open spaces and recreational facilities as crucial components of a thriving society. He believed that economic progress should not solely be measured by material wealth, but also by improvements in human health, education, and overall quality of life. However, it was not before the 1930s that the principle of CBA was proposed in the USA and formalised for public-sector application through the Flood Control Act of 1936, which mandated that the benefits of federal infrastructure projects must exceed their costs, as a key principle on how to evaluate ideas in this space [123]. The concept was then spread through US Army Corps of Engineers and US federal agencies. Scholars like Otto Eckstein and Ezra J. Mishan provided theoretical foundations, with Mishan’s 1971 textbook becoming a standard reference [124].
CEA emerged slightly later, first in the US Department of Defense and then in healthcare contexts, where the key challenge was that benefits are difficult to monetise. Economists like Burton Weisbrod in the 1960s [125] and later Weinstein and Stason in the 1970s/1980s [51] helped define CEA’s methodological framework and laid the groundwork for using quality-adjusted life years in health evaluations. Today, both CBA and CEA are central to frameworks such as health technology assessment and are widely used by organisations like the National Institute for Health and Care Excellence (NICE) in the UK, WHO-CHOICE, and global health funders.
Net present value and internal rate of return
When evaluating new ideas in the realm of finance and investment, two metrics stand out as foundational tools for assessing their feasibility and profitability: net present value (NPV) and internal rate of return (IRR). The NPV was formalised by the economist Irving Fisher in his 1907 book ‘The Rate of Interest’ [126]. Fisher was an American economist, statistician, inventor, and progressive social campaigner, educated and working at Yale University. He introduced the concept of discounted cash flow and the time value of money, forming the basis of modern investment analysis, which he expanded upon in his 1930 follow-up ‘The Theory of Interest’ [127].
The IRR, while grounded in similar logic, became prominent in the 1950s through the work of Joel Dean, an American economist and one of the founders of business economics, who worked at several universities in the US. He introduced the IRR to business decision-making in his book ‘Capital Budgeting’ in 1951 [128]. The roots of IRR can be traced to at least two distinct origins. Eugen von Böhm-Bawerk, an economist and federal minister of finance of Austria, discussed maximising net cash flow for each invested currency unit in his book ‘Positive Theorie des Kapitales’ in 1889 [44]. John Maynard Keynes, an English economist and philosopher whose ideas fundamentally changed the theory and practice of macroeconomics and the economic policies of governments, had discussed similar concepts. He used the term ‘marginal efficiency of capital’ in his ‘General Theory’ in 1936, linking IRR to investment expectations [129]. These tools remain core to capital budgeting, used extensively in both private-sector finance and public-sector project appraisal. They also complement methods like CEA and CBA in multi-criteria decision-making.
Patent metrics: citations and the originality index
Patent analysis offers another window into the evaluation of novel ideas, particularly in the domains of innovation economics and technology development. While Eugene Garfield – an American linguist and information scientist who is often considered the father of scientometrics – focused mainly on scientific literature, his influence permeated the development of patent metrics. Patent examiners at the US Patent and Trademark Office began using citation cards in 1947, which were an early foundation for metrics such as forward citation counts, originality, and generality indexes. Building on Garfield’s foundational work in scientific citation indexing [130], researchers in the 1980s and 1990s began to apply similar techniques to patents. The aim was to measure the novelty or originality of an invention based on its patent citations, which became widely recognised as indicators of technological impact. This was promoted by researchers such as Adam Jaffe, Manuel Trajtenberg, and Bronwyn Hall. In 1993, Jaffe and colleagues demonstrated that the number of times a patent is cited by subsequent patents, i.e. forward citations, correlates with its technological significance and influence [45]. In a subsequent 2001 working paper, Hall, Jaffe, and Trajtenberg introduced two additional patent metrics: the originality index, which measures how diverse the sources of a patent’s knowledge base are, and the generality index, which captures how widely a patent is cited across technological fields [131]. These metrics, important for the evaluation of novel ideas in innovation, are now embedded in patent analytics platforms and are routinely used by economists, venture capitalists, and technology strategists. It is also worth noting that they are embedded in databases such as the National Bureau of Economic Research Patent Citations Data File, the US Patent and Trademark Office, European Patent Office, and the Organisation for Economic Co-operation and Development (OECD) Regional Patent Database. They are also used in commercial tools such as Derwent, IFI CLAIMS, and Google Patents.
Citation metrics, impact factor, and h-index
The assessment of scientific influence and research productivity has evolved significantly over the past century, largely through the introduction of citation-based metrics. These tools, developed independently across several decades, have shaped the way ideas, research impact, and scholarly productivity are evaluated in academia, publishing, and funding landscapes. The foundations of citation analysis were laid by Eugene Garfield. In his seminal 1955 paper, ‘Citation Indexes for Science: A New Dimension in Documentation’, Garfield introduced the idea that scholarly influence could be traced through citations [130]. This laid the groundwork for the Science Citation Index (SCI), which was launched in the 1960s and allowed for systematic tracking of idea diffusion through scientific literature. Garfield later established the Institute for Scientific Information (ISI), which played a central role in the development of citation metrics. Garfield, together with Irving H. Sher, later introduced the journal impact factor (JIF), which was formally published in ‘Science’ in 1972 [132]. This metric evaluates a journal’s average citation frequency within a defined window, initially two years, but more recently five years, and has become a widely used – and debated – proxy for journal quality and scientific prestige. Garfield devised the JIF as a tool to primarily assist librarians in evaluating journal quality and influence. Half a century after Garfield’s work, Jorge E. Hirsch, a physicist at the University of California, San Diego, introduced the h-index in 2005 [46]. This metric aimed to capture both the productivity and the citation impact of an individual researcher’s body of work, whereby a scholar has an h-index of the largest number h such that their h articles have received at least h citations each. The metric gained rapid traction in academic evaluations due to its intuitive appeal and simplicity, while providing a useful measure of the impact of an individual’s pursued research ideas within the research community. These three citation-based metrics – the citation count, the JIF, and the h-index – are now embedded in research assessment frameworks, promotion and tenure decisions, grant reviews, and institutional rankings. However, their limitations have also sparked movements such as the San Francisco Declaration on Research Assessment and the Leiden Manifesto, advocating for more nuanced and responsible approaches to evaluating scientific contributions [133].
Technology Readiness Levels
The Technology Readiness Levels (TRLs) framework was pioneered by NASA in the 1970s to evaluate the maturity of technologies intended for space exploration. Originally developed by Stanley Sadin, then director of NASA’s Advanced Projects Office, the system allowed decision-makers to assess how far a particular technology had advanced along its development path [47]. The earliest internal use of TRLs occurred between 1974 and 1977. The framework was refined and standardised in subsequent years, culminating in a pivotal 1995 white paper by John C. Mankins titled ‘Technology Readiness Levels’, which became the definitive articulation of the nine-level TRL scale [134]. This structured approach ranges from TRL 1, representing basic research, to TRL 9, indicating fully operational and proven systems. Initially a tool for NASA, TRLs were later adopted by a range of public sector agencies, including the US Department of Defense, Department of Energy, and the European Space Agency. The European Commission integrated the framework into its Horizon 2020 and Horizon Europe funding programmes. Today, TRLs are used broadly across industries including aerospace, health technology, energy, and manufacturing, to evaluate the progress of an idea in technology from its early conceptual progress to application. The nine levels that were used to evaluate ideas were: TRL 1, in which basic principles were observed; TRL 2, in which technology concept was formulated; TRL 3, with experimental proof of concept; TRL 4, with laboratory validation of components; TRL 5, with validation in relevant environment; TRL 6, with system/subsystem demonstration; TRL 7, with a prototype demonstration in operational field; TRL 8, with the actual system completed and qualified; TRL 9, with system proven in real-world operational use [134,135]. The TRL system has become a global benchmark for technology maturity assessment, guiding investment strategies, R&D funding, and innovation management.
Value-driven approaches: scoring models and criteria-based frameworks
Multi-Criteria Decision Analysis
Multi-criteria decision analysis (MCDA), or multi-criteria decision making (MCDM), is a formalised family of methods designed to assist in decisions involving multiple, often conflicting, criteria. Its intellectual foundations trace back to Enlightenment thinkers like Benjamin Franklin and his ‘moral or prudential algebra’ in 1772, although it was primarily a method for comparing two alternatives [48]. However, the field was shaped into its modern form in the mid-20th century through advances in mathematics, decision theory, operations research, and utility modelling. Research by Harold W. Kuhn and Albert W. Tucker, mathematicians from Princeton University, laid the groundwork for MCDA in 1951 [49]. The work by Howard Raiffa and Robert Schlaifer at Harvard Business School was instrumental in shaping decision analysis in its current form [50]. Bernard Roy pioneered one of the first MCDA methods, Elimination and Choice Expressing Reality, in France in 1968 [136]. Another approach that falls into this family of methods is Thomas Saaty’s AHP, launched in the USA in the 1970s and mentioned in earlier text [42,43]. Simultaneously, Keeney and Raiffa’s 1976 treatise ‘Decisions with Multiple Objectives’ laid the theoretical groundwork for Multi-Attribute Utility Theory. This was considered a key progress that highlighted the relevance of multi-attribute utility theory to practical decision analysis, so it continues to underpin many MCDA tools today [137]. In 1979, Stanley Zionts’ article, ‘MCDM – If not a Roman Numeral, then What?’, contributed to popularising the acronym MCDM, as well as MCDA [53]. In early 1980s, the joint work of Stan Zionts and Jyrki Wallenius on interactive multi-objective linear programming was influential [138]. Various MCDA methods have emerged and have been applied to a wider range of complex decision problems from the 1980s until today, with the approach evolving to encompass diverse approaches – more than 40 of them are listed on the MCDA’s Wikipedia page [139].
The MCDA is useful in evaluating ideas on how best to address challenges where decisions involve multiple, often conflicting objectives. It provides a framework for structuring complex problems, assessing alternatives, and evaluating preferences. It evolved over time, building upon earlier ideas and expanding its scope to address increasingly complex decision-making scenarios. Today, it has become a broad family of methods used to assess and rank competing ideas, interventions, or alternatives based on multiple conflicting criteria in various fields, including public decision-making, resource allocation, project evaluation, public health energy policy, infrastructure planning, and technology forecasting. Organisations such as the WHO, NICE (UK), and the European Commission routinely apply MCDA in funding, priority-setting, and evaluation.
Strengths, Weaknesses, Opportunities, and Threats (SWOT) analysis
Strengths, Weaknesses, Opportunities, and Threats (SWOT) analysis, perhaps one of the most widely used strategic evaluation tools in business, policy, and project planning, was developed in the 1960s at the Stanford Research Institute, under the leadership of business and management consultant Albert S. Humphrey. Originally coined as the Satisfactory, Opportunity, Fault, and Threat (SOFT) approach, the method was later refined into the now-familiar SWOT framework [140]. Although no single academic paper formalised the method at the time, Humphrey’s approach gained traction through internal corporate planning seminars and practitioner literature. Its aim was to help Fortune 500 companies improve strategic planning by better understanding internal capabilities and external market conditions. Its core innovation was precisely in bridging internal factors (i.e. strengths and weaknesses) with external conditions (i.e. opportunities and threats), enabling organisations to craft strategies that were both self-aware and environmentally responsive. Therefore, SWOT enabled the evaluation of longer-term strategic ideas for the future development of companies. A retrospective account by Humphrey, published posthumously in the Stanford Research Institute Alumni Newsletter in 2005, recounts how the technique emerged during efforts to improve long-range planning within Fortune 500 companies [52]. A parallel lineage of thought could also be found in the influential Harvard Business School text ‘Business Policy: Text and Cases’ by Learned and colleagues in 1965 [141], which laid the groundwork for internal and external analysis in strategic thinking, although it did not yet use the SWOT acronym. Other notable frameworks for strategic idea evaluation include PESTEL and Porter’s Five Forces (see later in the section on idea prioritisation methods).
Weighted scoring models
Weighted scoring models, also known as MCDA tools, emerged in parallel with many MCDA methods. They emerged during the mid-20th century through the confluence of operations research, decision theory, and systems engineering. These models provided a structured way to evaluate competing alternatives against multiple weighted criteria, especially when trade-offs were necessary and stakes were high. They were formally introduced by Stanley Zionts in 1979 [53], whose mathematical model aimed to aid evaluation processes in decision-making by comparing and ranking alternatives based on various criteria. Pivotal progress in applications in decision theory and value-focused thinking was contributed by Howard Raiffa and Ralph L. Keeney from Harvard University’s International Institute for Applied Systems Analysis in 1976 [50], while the roots of application in operations research reach back to the work by C. West Churchman, Russell Ackoff, and E. Leonard Arnoff from the Wharton School at the University of Pennsylvania in 1957 [142]. The latter three authors are often credited for introducing systems thinking and the idea of assigning weights to criteria in complex evaluations. Stanley Zionts and Barry Boehm helped bring these methods into applied domains, such as engineering, software evaluation, defence procurement, and project evaluation in technology settings, with Boehm’s ‘Software Engineering Economics’ in 1981 explicitly incorporating weighted scoring models into technology decision-making frameworks [143]. A typical structure of a weighted scoring model involves: defining decision alternatives; selecting criteria relevant to evaluation; assigning weights to each criterion based on its importance; scoring each alternative against each criterion; calculating the weighted sum of scores for each alternative; ranking alternatives based on total weighted score. These kinds of models became popular in project selection, product evaluation, and R&D portfolio analysis [143,144].
Pugh Matrix (Decision-Matrix Method)
The Pugh Matrix – also known as the Decision-Matrix Method – was introduced by Stuart Pugh, British mechanical engineer and Professor of Design at the University of Strathclyde in the late 1980s. It is a systematic tool for evaluating ideas in engineering and product development by comparing alternatives in design concepts, ultimately leading to the selection of the best option. It enabled teams to evaluate multiple options against a set of predefined criteria, but relative to a baseline. Pugh formalised this approach in his 1990 book ‘Total Design: Integrated Methods for Successful Product Engineering’ [54]. His method emphasises comparative, rather than absolute assessment, using simple symbols (+, –, S) to indicate whether an option performs better, worse, or the same as a reference solution. Pugh’s approach helps teams avoid premature convergence on a single idea and fosters objective dialogue in innovation processes. The eventual tally scores guide decisions, but discussion and insight can matter more than raw totals. This makes the matrix particularly useful in engineering, manufacturing, healthcare, and R&D portfolio selection. Avoiding early overcommitment to a single idea and encouraging iterative refinement through team-based analysis are this method’s strengths, finding its application in the Six Sigma (particularly in the Define, Measure, Analyse, Improve and Control framework), lean product development, and innovation portfolio management approaches (see later).
Design thinking: Feasibility-Desirability-Viability (FDV) framework
The feasibility-desirability-viability (FDV) framework for evaluation of ideas in design was popularised by the global design and innovation consultancy IDEO in the late 1990s and early 2000s. Although no formal academic origin underpins this triad, it emerged organically from IDEO’s human-centred design practice and design thinking (described in the section on idea generation methods). It was institutionalised through its teaching at Stanford’s d.school, formally the Hasso Plattner Institute of Design. The framework assists teams in evaluating innovation ideas through three essential lenses: feasibility (can it be built with current technology and skills?), desirability (do people want it?), and viability (is it economically and organisationally sustainable?). Thus, the framework emphasises balancing what is desirable from a user’s perspective, what is technologically feasible, and what is economically viable. Tim Brown’s 2009 book ‘Change by Design’ [31] serves as the primary reference for FDV, integrating insights from earlier systems thinkers like Buckminster Fuller [145] and Herbert Simon [30], as well as management theorists such as Peter Drucker [146], Alexander Osterwalder and Roger Martin [147]. While conceptually simple, FDV’s strength lies in forcing innovation teams to balance user needs, technical constraints, and financial realities when evaluating novel ideas in design.
Democratic approaches: crowd-based assessment
Wisdom of the crowd techniques
The concept of the ‘wisdom of the crowd’ evaluates ideas by non-expert voting. It is rooted in the notion that collective judgments can outperform those of individuals. While this approach has deep intellectual roots, it was popularised in its modern form by journalist James Surowiecki in his 2004 book ‘The Wisdom of Crowds’ [148]. Surowiecki outlined the conditions under which crowd-based judgments tend to be superior: diversity of opinion, independence of individuals, decentralisation of decision-making, and the existence of an aggregation mechanism. The principle was first empirically observed in 1907 by Francis Galton, an English polymath and the early pioneer of behavioural genetics [55]. Galton analysed 787 guesses by fairgoers estimating the weight of an ox. The median estimate was remarkably close to the actual weight, suggesting that the aggregate judgment of a diverse group could be quite accurate. Galton’s findings were published in ‘Nature’ under the title ‘Vox Populi’ [55]. In contemporary contexts, the evaluation of various ideas using the ‘wisdom of the crowd’ is operationalised through platforms and processes such as ranked voting, upvoting systems, prediction markets, and participatory idea selection tools. It has found applications in innovation challenges, participatory budgeting, open policymaking, and deliberative democracy. The ‘wisdom of the crowd’ principle led to the development of crowdsourced idea evaluation platforms, including voting-based idea platforms such as IdeaScale, Google Moderator, or Dell IdeaStorm, and crowdsourced research funding and prioritisation such as OpenIDEO and Foldit. These platforms invite broad participation and aggregate rankings to identify promising ideas [149]. Voting mechanisms may include simple upvotes/downvotes, ranked-choice ballots, dotmocracy, or pairwise comparisons – each serving as an aggregation mechanism that aligns with Surowiecki’s criteria for collective intelligence [149].
Prediction markets
Prediction markets, also known as ‘information markets’ or ‘idea futures’, offer a formalised mechanism for aggregating collective judgments. They allow participants to trade contracts based on the likelihood of future events, effectively translating beliefs into prices that reflect the perceived probability of outcomes. The conceptual foundation of prediction markets was laid by Robin Hanson, an economist and polymath at George Mason University and formerly Caltech/NASA, in the late 1980s. In his 1990 article ‘Could Gambling Save Science?’, Hanson proposed ‘idea futures’ as a method for forecasting scientific outcomes [56]. His later work further refined the mathematical underpinnings of prediction markets, including combinatorial betting and market scoring rules [150]. One of the earliest real-world implementations was the Iowa Electronic Markets, launched in 1988 at the University of Iowa, which originally focused on political forecasting and aimed to forecast US election outcomes [151]. Economists James Wolfers and Eric Zitzewitz contributed empirical validations of prediction markets’ accuracy across domains [152]. Prediction markets have since been adapted to assess R&D portfolios, evaluate the viability of new technologies, and forecast policy impacts. Companies like Google, HP Laboratories, Eli Lilly, and Microsoft have experimented with internal prediction markets to inform strategic decisions [153].
The James Lind Alliance (JLA) priority setting partnerships (PSPs)
The JLA follows a transparent, inclusive, and stepwise process to evaluate proposed research ideas, which are typically framed as ‘uncertainties’. After they are gathered from patients, carers, and clinicians, they undergo a structured evaluation and prioritisation process. In the first step, the PSP performs data collation and theming: all submissions, which may number in the hundreds or even thousands, are collated and subjected to a qualitative data synthesis. This results in grouping of similar or duplicate questions, removing out-of-scope entries, and merging overlapping questions into broader, clearly phrased, indicative questions, as well as performing thematic analysis and deductive coding. In the second step, each proposed research idea is checked against the existing evidence to determine whether it is still unanswered, and those answered are then removed. In the third step, all validated research uncertainties are presented in an interim prioritisation survey, which is distributed again to patients, carers and clinicians. They are asked to select the questions they consider most important. Responses are quantitatively analysed, often by calculating frequency of selection, sometimes stratified by stakeholder group to preserve balance between professional and public voices. The top 25–30 most frequently selected questions typically proceed to the final prioritisation workshop – a face-to-face or virtual consensus workshop, facilitated using NGTs: mixed small group discussions that bring together patients, carers and clinicians, followed by iterative ranking rounds, where each group discusses and re-orders the questions. In a final plenary session, all groups’ rankings are combined, debated, and agreed upon. The outcome is a ‘top 10’ list of jointly agreed research priorities, reflecting a consensus across all stakeholder perspectives. Through this approach, the key principles that guide evaluation of proposed research ideas are equity of voice for all stakeholders, transparency of the process, grounding in existing evidence, and consensus-building, where deliberative methods are used to arrive at shared priorities [23,154,155].
The CHNRI method
The approach to evaluating proposed research ideas in the CHNRI method was initially described in the CHNRI method’s guidelines for implementation in 2008 [25]. Further refinements were detailed by Igor Rudan in 2016 in the fourth paper of the series that updated CHNRI method and explained the method’s key conceptual advances [156], and later in his book ‘Measuring ideas: The CHNRI method’ published in 2022 and co-edited with Sachiyo Yoshida from the WHO, Kerri Wazny from the University of Edinburgh, and Simon Cousens from the London School of Hygiene and Tropical Medicine [157]. In the first step, the method sought to define the existing context within a chosen field of health research and identify a small set of criteria that would recognise one research idea as better than another. These criteria could have been, for example, the likelihood that the proposed research idea would be answerable; that it would lead to effective, deliverable, affordable and/or sustainable health intervention; that it could reduce a large portion of the existing disease burden; and that its effects would be equitable. In the next step, dozens of the leading experts in the chosen field would be invited to propose several of their best research ideas. Then, a consolidated list of research ideas is developed after merging similar ideas or removing duplicate ones. In the third step, all experts ‘score’ the ideas by assessing their likelihood of satisfying each criterion by simply answering ‘yes’ or ‘no’. They are also allowed to use an ‘informed maybe’, or to simply leave the scoring field blank if they did not have enough knowledge to make this judgment. The preference of ‘yes’/‘no’ answer over larger scales with several increasing options prevents regression of the scores to the mean value, but ‘maybe’ is still allowed if it is well-informed, while the ‘blank’ option serves to minimise the noise if the scorers are uninformed. As a result, the collective opinion of dozens of leading experts in the field of health research about several hundred systematically gathered research ideas are scored through this expert-sourcing, leading to a simple table that shows the ‘collective optimism’ of the experts towards how the ideas would satisfy each of the criteria. All scores are intuitive, as they range between 0–100%, and there are statistical tools that allow computing agreement statistics for each idea, as well as 95% confidence intervals for all scores. This scoring approach offers several advantages over other priority-setting methods: it ensures transparency, as all evaluations follow a clear structure; minimises personal biases by relying on a structured survey process rather than subjective discussion; provides a systematic, reproducible method with well-defined outcomes; and generates intuitive, quantitative scores that can be used for ranking and funding decisions [156,157].
Social media engagement metrics as proxies for idea traction
With the advent of Web 2.0 and the rise of participatory digital platforms, social media engagement metrics, such as likes, shares, comments, retweets, upvotes, and follower growth, began to serve as informal proxies for gauging the public traction of ideas, content, or innovations [158]. While not initially designed for this purpose, such metrics became increasingly influential in decision-making within marketing, science communication, and innovation ecosystems. The period between 2004 and 2009 was marked by the early emergence of social platforms, such as Facebook, Twitter (now X), and Reddit, while the formalisation of engagement metrics as indicators of idea diffusion began in the 2010s. Jason Priem, a PhD student at the University of North Carolina at Chapel Hill, and colleagues initiated the Altmetrics movement with their Altmetrics Manifesto in 2010 [57]. They proposed that online interactions, i.e. tweets, blog posts, bookmarks, and other similar expressions on the internet, could complement citation counts in capturing scholarly impact. In parallel, marketing science contributed robust empirical analyses. Katie Delahaye Paine’s book ‘Measure What Matters: Online Tools for Understanding Customers, Social Media, Engagement, and Key Relationships’ in 2011 is often cited as one of the key references in this vast new field [159]. Paine is a prominent figure in the field of communications and measurement, and her book focuses on using data to understand and improve public relations, social media, and communication strategies. At the Wharton School, professors of marketing Jonah Berger and Katherine Milkman published a study in 2012 entitled ‘What Makes Online Content Go Viral?’, in which they demonstrated that emotional resonance, novelty, and social currency drive engagement and diffusion, making these signals useful for predicting the reach of new ideas [160]. Professor Igor Rudan and his PhD student with a background in music and media industries, Iain H. Campbell, from the University of Edinburgh, experimented with a series of videos themed on global health in 2017–19 and used multiple media outlets to explore what themes cause virality among which audience and on which platform [161,162]. In 2018, Anatoli Colicev, a scholar in marketing at Bocconi University and his coworkers developed specific social media metrics, including engagement, to study the impact of social media on brand awareness, purchase intention, and customer satisfaction [163]. Engagement metrics are now routinely used in many areas – marketing analytics, research dissemination analytics, technology adoption, innovation trend forecasting, public health messaging and campaign evaluation, and early-stage evaluation of product or policy ideas. However, limitations to their use remain, because ‘engagement’ is a complex construct with various dimensions, such as likes, comments, shares, or time spent on content, that all need to be considered. Engagement data can also be manipulated by bots or coordinated campaigns, so it may not correlate with long-term impact or quality. Researchers have urged caution, emphasising the importance of linking engagement to tangible outcomes [162,163].
Historically useful principles: scientific and philosophical validity tests
At the end of this section on the methods used for evaluating ideas, we remind of historically useful principles that do not strictly represent specific, well-developed methodologies, but that have been used throughout history towards a shared goal, so they underlie many of the listed methods.
Logical consistency and deductive reasoning
The principles of logical consistency and deductive reasoning represent some of the earliest formal tools for evaluating the basic validity of ideas. These concepts were systematically developed by the Greek philosopher Aristotle in the 4th century BCE. Widely regarded as the ‘father of logic’, Aristotle’s work laid the foundation for much of Western rational inquiry, influencing disciplines ranging from philosophy and mathematics to modern science. In his treatise ‘Prior Analytics’ from around 350 BC, Aristotle introduced syllogistic logic, a formal system in which conclusions are derived logically from two premises. It emphasised that a valid argument, or syllogism, must be logically consistent, for example: ‘All humans are mortal. Socrates is a human. Therefore, Socrates is mortal’. Such Socratic reasoning exemplifies deductive logic, where the truth of the conclusion necessarily follows from the truth of the premises [58]. Aristotle also emphasised logical consistency, notably in ‘Metaphysics (Book IV)’, where he formulated the law of non-contradiction: ‘It is impossible for the same thing to both belong and not belong to the same thing, in the same respect, at the same time’. His methods included reductio ad impossibile – testing an argument by assuming the opposite and showing it leads to a contradiction. These ideas became the basis for assessing coherence and truth in any argumentation [59]. Aristotle’s logical system profoundly influenced the development of Western thought and remained influential for centuries, shaping how people reasoned about ideas and assessed arguments. The principles of logical consistency and deductive reasoning for assessing ideas are among the oldest formal methods in intellectual history. They influenced medieval scholasticism (e.g. the legacy of Thomas Aquinas), modern logic and set theory (e.g. the many works of Gottlob Frege, Bertrand Russell, and Kurt Gödel), and the scientific method, where deductive reasoning is used to test hypotheses for internal coherence.
Empirical testability, replicability, and falsifiability
The criteria of empirical testability, replicability, and falsifiability are central to the modern scientific method and the evaluation of scientific hypotheses. These principles assert that for an idea, hypothesis, or theory to be considered scientific, it must be subject to empirical observation and capable of being tested under controlled conditions. Furthermore, the results must be replicable by independent researchers using the same methods. Early advocates of empirical inquiry, most notably Francis Bacon, philosopher and former Lord High Chancellor of Great Britain, emphasised observation and induction in his ‘Novum Organum’ in 1620 [60]. His work was followed by David Hume, who raised concerns about induction and causation, leading to the need for reproducibility of scientific discoveries in his book ‘A Treatise of Human Nature’ in 1739 [164]; then, Claude Bernard, who emphasised the need for experimental verification in physiology in 1865 [165]; and John Stuart Mill, who included the methods of agreement and difference, anticipating modern experimental design, in his book ‘A System of Logic’ in 1843 [166]. However, it was Karl Popper, an Austrian philosopher of science, later based in the UK, who gave these criteria their most influential formal articulation. In ‘The Logic of Scientific Discovery’ (originally ‘Logik der Forschung’, 1934) [61], Popper proposed the key principles for evaluation of scientific ideas: empirical testability, i.e. that an idea must be observable and testable; replicability, i.e. that independent tests must yield consistent results; and falsifiability, i.e. that it must be possible to disprove the idea. These criteria became central to scientific method, hypothesis testing, clinical trials, reproducibility crisis discussions, and the philosophy of science. Popper considered falsifiability the defining criterion of science. He argued that science progresses not through verification, but through conjecture and refutation. Therefore, scientific theories must be structured in a way that they can, in principle, be proven false, and bold hypotheses must expose themselves to possible disproof. For a theory to be scientific, it must predict what should not happen and then be testable in a way that observations can potentially contradict it. Popper also highlighted the importance of replicability as a safeguard against bias, error, or coincidence, demanding that independent researchers should be able to reproduce results under the same conditions. Still, he underscored that confirming a theory through repeated observation is inherently weaker than designing tests to falsify it [61]. The criteria of testability, replicability, and falsifiability offered a logical solution to the philosophical problem of demarcation: how to distinguish between science and non-science. A theory is scientific only if it makes specific, uncertain predictions that exclude specific outcomes. If no conceivable observation could contradict the theory, then it is unfalsifiable, which makes it unscientific. Popper cited Einstein’s theory of relativity as a model of falsifiability, since it made precise predictions that could be tested and refuted. In contrast, he viewed Freud’s psychoanalysis and Marxist historicism as pseudoscientific because they could accommodate any outcome and therefore resisted falsification.
Paradigm shifts
In contrast to Popper’s focus on individual hypotheses and their testability, Thomas S. Kuhn, an American historian and philosopher of science who worked at several US universities – Harvard, Berkeley, Princeton and MIT – offered a sociological and historical perspective on how scientific ideas evolve. In his groundbreaking book ‘The Structure of Scientific Revolutions’ in 1962, Kuhn introduced the concept of a paradigm shift – a radical, collective reorientation in the basic assumptions, methods, and questions within a scientific community [62]. According to Kuhn, science proceeds through long periods of traditional research, during which researchers solve puzzles within an established paradigm. Over time, however, anomalies accumulate and generate an increased amount of data that the prevailing theory cannot explain. Eventually, a crisis emerges, culminating in a scientific revolution and the adoption of a new paradigm. This new paradigm is often incommensurable with the old. His concept of ‘incommensurability’ suggests that different paradigms may be fundamentally incompatible and cannot even be directly compared. Examples cited by Kuhn include the transition from Ptolemaic to Copernican astronomy, from Newtonian physics to Einstein’s relativity, from classical chemistry to atomic theory, and from classical genetics to molecular biology. Therefore, Kuhn argued that scientific progress is not a linear accumulation of facts; instead, it occurs through revolutionary shifts between different paradigms. His work challenged the traditional view of science as a purely objective and cumulative process, emphasising the role of social and historical context in shaping scientific knowledge, professional consensus, and sociocultural factors. Essentially, against Popper’s argument that ideas should only be judged by falsifiability or logical consistency, he noticed that they are also assessed within the context of dominant paradigms, which define what counts as evidence, what questions are worth asking, and what methods are valid. Thus, a paradigm shift changes the entire assessment framework, where once rejected concepts can become foundational, and vice versa.
IDEA PRIORITISATION METHODS
Rational decisions through structure: structured decision-making frameworks
Paired comparison methods
Paired comparison methods enable decision-makers to prioritise ideas by comparing these two at a time. They have a long history in psychology, decision theory, and preference elicitation. Louis L. Thurstone, the American psychologist from the University of Chicago and pioneer of psychometrics, is credited with introducing them formally in the 1920s as part of psychometric scaling. In his 1927 paper ‘A Law of Comparative Judgment’ [63], he demonstrated that this cognitively simple, yet powerful approach can be useful in investigating a wide range of psychological attributes, such as ‘seriousness of crime’. His work allowed for the derivation of quantitative preference scales from simple ordinal choices. By presenting respondents with pairs of alternatives and asking them to choose the preferred one, paired comparison methods create a matrix of preferences from which overall rankings can be statistically inferred. Thurstone introduced the statistical model behind paired comparisons and derived interval-scale values for preferences based on how often one item is chosen over another. These models, such as Thurstone’s and later the Bradley-Terry [167] and Elo models [168], are now foundational in decision science and preference elicitation. Paired comparison remains particularly valuable when stakeholders are uncomfortable assigning numerical weights, but can reliably express ordinal preferences. Modern frameworks like AHP and certain Delphi adaptations often integrate paired comparisons as a core or optional feature. Paired comparison methods are used to prioritise health research (e.g. JLA, Delphi hybrid methods), educational learning needs, ideas in marketing, product development, and military and intelligence analysis, and for crowdsourced prioritisation such as voting systems.
Multi-voting and dot voting
Multi-voting, also known as ‘dot voting’, ‘dotmocracy’, or ‘sticker voting’, is a simple and highly adaptable method used to identify priorities in group settings. It likely originated informally in the 1950s and 1960s in industrial and educational group facilitation and grew out of group facilitation practices. It was then popularised in the 1970s and 1980s by practitioners of total quality management, Six Sigma, facilitation training programmes, and later by design thinking and Agile teams. Because of this, it has no single credited inventor. It later became a standard tool in design thinking, quality management, and Agile project management. In multi-voting exercises, participants are given a fixed number of ‘votes’. These can be either dots, stickers, checkmarks, or online tokens. Then, the participants allocate their votes among competing ideas that are visually displayed to them – usually on flipcharts, whiteboards, or similar surfaces. The votes are then tallied to generate a prioritised list based on collective preference. This method’s appeal lies in its simplicity, speed, and inclusiveness. It is now widely featured in facilitation manuals, such as Bens’ Facilitating with Ease [64], continuous improvement guides like Brassard’s and Ritter’s The Memory Jogger II [65], and IDEO’s and Stanford d.school’s design toolkits that were mentioned earlier.
Nominal group technique with voting (NGT)
The NGT, a structured, face-to-face method for idea generation and prioritisation using voting, was introduced in the late 1960s and early 1970s by Andre L. Delbecq, Andrew H. Van de Ven, and David H. Gustafson, American management scholars working at the universities of Wisconsin and Minnesota. The NGT was designed as a structured method for group brainstorming and idea prioritisation that ensured inclusive participation and minimised the dominance of strong, outspoken personalities in face-to-face settings, which was a common weakness of many other methods in use at the time. The approach combined individual idea generation with structured group discussion and prioritisation by anonymous voting, as well as facilitated consensus-building while preserving individual creativity. It was first explained in their book ‘A Group Process Model for Problem Identification and Program Planning’ in 1971 [19], followed by another book that compared it to Delphi processes in 1975 [105]. These two books provided a comprehensive guide to the technique, including its application for consensus-building and action planning. The NGT outlines a five-step process. In silent idea generation, individuals silently generate ideas in writing; then, in round-robin sharing, each member shares one idea at a time, recorded without debate; in the clarification stage, the group discusses each idea for understanding, but not evaluation; finally, the anonymous voting or ranking takes place, in which participants independently and anonymously vote or rank ideas, often using a point allocation method. The results are then aggregated so that votes are tallied to produce a prioritised list. This method preserves individual creativity, encourages equal input, and produces quantifiable prioritisation. [19,105]. The NGT has been widely used in health research priority-setting, clinical guideline development, educational planning and curriculum design, as well as community and strategic planning in government and nonprofit sectors, including the JLA and the NIHR UK.
Delphi process with ranking rounds
The elements of generation and evaluation of ideas within the Delphi process have already been described in previous text [18,105]. Here, we focus on its approach to prioritisation of ideas. In the Delphi method, once expert ratings are collected for each idea, facilitators compute summary statistics, such as the mean, median, interquartile range, or degree of agreement in order to quantify how each item was evaluated. This aggregated information is then shared back with the participants in anonymised form, allowing them to see how their views align or diverge from those of their peers. In the subsequent round – typically the third and beyond if needed – participants are invited to review their previous ratings in light of this group summary and may revise their scores. This controlled feedback loop serves to highlight outlier opinions, surface emerging consensus, and encourage convergence toward a collective judgment. When stability in responses across rounds indicates that sufficient consensus has been achieved, the final results are compiled into a ranked list of ideas. This final prioritisation can be based on median or mean scores, or through weighted scoring when some criteria are deemed more important. In more complex cases, composite scoring systems may be used to combine multiple evaluation criteria into a unified index. Some Delphi exercises conclude at this stage, while others proceed to classify the ranked ideas into categorical tiers, such as identifying the top five priorities, or assigning levels of urgency (e.g. high, medium, low), or implementation horizons (short-, medium-, or long-term). To ensure rigour and consistency, many Delphi protocols define consensus thresholds in advance. They may require at least 70% of participants to rate an item 7 or higher on a 1–9 scale or expect an interquartile range of no more than 2 to indicate sufficient agreement. These predefined rules help determine which ideas are ultimately retained, discarded, or considered most important. Several methodological features enable this structured ranking and prioritisation process. Anonymity protects participants from dominance bias; iteration allows progressive refinement of views; controlled feedback fosters peer learning; quantitative scoring supports systematic evaluation; and statistical aggregation enables transparent consensus measurement. In cases where the goal is to narrow down a shortlist of priorities, e.g. in funding or policy, a final round may include a forced ranking exercise or a points-allocation system in which participants distribute a fixed number of points across the most promising ideas.
Analytic hierarchy process (AHP)
The elements of generation and evaluation of ideas within the AHP, developed by Thomas L. Saaty [42,43], the American mathematician and operations researcher from the University of Pittsburgh, have already been described earlier in this paper. In this section, we reflect on its approach to prioritisation of ideas. Saaty sought a rigorous, yet intuitive framework for making structured judgments in the face of competing alternatives, and for selecting rational choices in the face of multiple, conflicting objectives. He broke complex decisions down into a hierarchy of goals, criteria, sub-criteria, and alternatives. Once the users of AHP have evaluated the relative importance of elements through pairwise comparisons, AHP generates quantitative weights via eigenvalue calculations. Then, it computes a ‘priority vector’ using eigenvalue methods. A consistency ratio is also calculated to check the reliability of these judgments. The AHP has since been applied extensively in fields ranging from global health research and public policy to corporate strategy and engineering design.
Addressing health and development: priority-setting frameworks in health sciences
RAND/UCLA Appropriateness Method
The RAND/UCLA Appropriateness Method (RAM) was developed in the mid-1980s by researchers at RAND Corporation in collaboration with clinicians at the University of California, Los Angeles to assess the appropriateness of medical interventions and surgical procedures. It was partly motivated by the concerns about overuse, underuse, and variability of the delivered interventions in clinical practice. The RAM combines a systematic review of the literature with expert panel judgment, where experts rate medical procedures using a scale of 1 to 9, considering both the evidence base and clinical scenarios. Ratings are analysed for consensus and disagreement, helping to identify and prioritise procedures that are appropriate, inappropriate, or uncertain. The key reference, ‘RAND/UCLA Appropriateness Method User’s Manual’, was published in 2001 following years of experience [66], formally codifying the method and providing step-by-step guidance for its implementation. Since its inception, RAM has been widely applied in healthcare quality improvement, technology assessment, and coverage decisions. It represents a hybrid between empirical data use and expert deliberation.
Essential National Health Research framework
The Essential National Health Research (ENHR) framework was introduced in 1990 by the Commission on Health Research for Development, which later evolved into the Council on Health Research for Development. The framework was articulated in the landmark report ‘Health Research: Essential Link to Equity in Development’, published in 1990 on behalf of the Commission [67]. The ENHR sought to reorient national research systems in LMICs toward equity-driven, needs-based, and country-led priorities. It emphasised integrating research with health system goals, involving diverse stakeholders, and supporting national capacity for implementation. The ENHR became a foundational model for aligning research with development and equity goals and promoting research as a public good, rather than a purely academic enterprise. Although it was received with great enthusiasm by the governments in LMICs due to those welcome aims, the uptake and practical application of this framework have been relatively modest to date. In the paper by Sachiyo Yoshida from the WHO, based on her analysis of the usage of different methodological approaches in the PubMed database, the ENHR was used in only 0.6% of all the papers that attempted to set health research priorities between 2001 and 2014 [169]. This is likely because, although ENHR does well to point to the general, broad research aims and goals for national health systems in terms of country leadership, increased focus on meeting the needs, and improved equity in the population, it is not prescriptive and decisive in picking priorities, as many other methods are, and does not end with clear lists and ranks of ideas, but rather with broad and general recommendations. Although these are most often very useful and point the policy makers to the right direction, the output lacking in specificity can later be subject to differences in interpretation among the stakeholders
GRADE methodology and Evidence to Decision frameworks
The GRADE methodology was initiated by the GRADE Working Group in 2000 as a global collaboration to improve the clarity, transparency, and reliability of clinical recommendations [68]. It now gathers over 500 scientists, clinicians, methodologists, and other experts dedicated to a transparent and systematic approach to assessing evidence and developing recommendations. It has since become the dominant approach used by guideline developers, including WHO, NICE, and Cochrane. The GRADE approach distinguishes the certainty of evidence (high, moderate, low, very low) from the strength of recommendations, incorporating factors such as benefits, harms, values, preferences, and resource use. The Evidence to Decision (EtD) frameworks specifically are used to make well-informed healthcare choices by systematically evaluating evidence and making transparent decisions. They were developed during the DECIDE project (2011–15) [170]. These frameworks make explicit how evidence is translated into recommendations by offering structured domains such as equity, feasibility, acceptability, and cost-effectiveness. Over subsequent years, iterative methods that include literature review, stakeholder feedback, and user testing were used to produce EtD frameworks tailored to different decision-making contexts: clinical recommendations, coverage decisions, and health system/public health decisions, making this a priority-setting methodology.
James Lind Alliance Priority Setting Partnerships (JLA PSPs)
The elements of generation and evaluation of ideas in the JLA PSPs process have already been described above [23,154]. Here, we focus on its approach to prioritisation of ideas. Once the uncertainties are identified, they are then organised into longlists and subsequently shortlisted through stakeholder surveys where participants score questions based on their perceived importance. The final stage of prioritisation is carried out in face-to-face consensus workshops using structured, voting-based techniques – typically adaptations of the NGT (see earlier text), where diverse stakeholders deliberate and rank the shortlisted questions. This inclusive, democratic process culminates in a final ‘top 10’ list of agreed research priorities, which serves as a guidance tool for funders, policymakers, and research organisations. These ‘top 10’ lists are widely recognised for their legitimacy and relevance, having informed funding calls and guideline development in multiple clinical areas. More than a hundred PSPs have been completed across various diseases and care settings, establishing the JLA as a globally influential model for equitable and evidence-based research prioritisation. Its impact has extended internationally, shaping practices adopted by institutions such as the WHO, Cochrane, the UK’s NIHR and MRC, and others. Key documentation of the JLA’s methods includes the evolving ‘JLA Guidebook’ [154] and a landmark article that provides a conceptual and empirical foundation for reducing waste in research through inclusive priority setting [171]. The legacy of the JLA lies in its pioneering approach to shared decision-making in research agenda setting, creating a replicable model for integrating lived experience into the production of evidence and ensuring that scientific inquiry is more responsive, democratic, and impactful.
Combined Approach Matrix
The Combined Approach Matrix (CAM) was introduced by the Global Forum for Health Research in 2004 as a practical tool for health research priority setting, particularly in LMICs. Developed under the leadership of Stephen Matlin, an international expert in global health innovation, through technical leadership by Abdul Ghaffar and contribution from Andres de Francisco, both global health and development experts [69], CAM aimed to address the so-called ‘10/90 gap’. This was the observation that less than 10% of global health research resources targeted conditions affecting 90% of the world’s population [172]. The CAM integrates multiple dimensions in prioritising health research, aiming to improve the process in which scientists discuss and decide on funding priorities based on their own views and knowledge. It was piloted in countries such as Mexico, Tanzania, and Burkina Faso, where it helped to shift attention to locally owned research agendas. It was useful for systematic classification, organisation, and presentation of the large body of information needed at different stages of the priority setting process [172,173]. As a result, decisions of committees in the three countries were based on information, rather than personal knowledge and judgment. CAM incorporates ‘economic’ dimension along one axis, and ‘institutional’ along the other, thus covering the determinants of health at the population level. Components of the economic dimension are ‘disease burden’, its ‘determinants’, ‘present level of knowledge’, ‘cost and effectiveness’, and ‘resource flows’. Components of the institutional dimension are ‘the individual, household and community’, ‘health ministry and other health institutions’, ‘sectors other than health’, and ‘macro-economic policies’. The CAM can be applied at the level of disease, risk factor, group, or condition, and also at local, national, or international levels [69,172,173]. In terms of uptake, the above-mentioned paper by Sachiyo Yoshida from the WHO showed that CAM was used in only 1.8% of all the papers that attempted to set health research priorities between 2001 and 2014 [169]. The likely reason is that it did not offer an algorithm or system for ranking or discriminating between the competing investment options. Therefore, in the absence of reliable information, which is common in countries that are in need of tools like the CAM, most of the decisions were still based on discussions and agreements within the panels of experts. However, the CAM laid methodological foundations for further work of the Global Forum for Health Research on addressing health research priority setting. The later framework, the CHNRI method, which also arose from the Global Forum for Health Research in 2006–08, built on the CAM and gradually received global uptake and widespread use in global health research prioritisation.
The CHNRI method
Building on the CAM and several other previous attempts to prioritise health research ideas, the CHNRI method proposed a more structured approach. It has led to more than 200 published exercises based on the CHNRI method to date, with some beginning to expand to areas beyond health research. As a result of the scoring process described earlier, each proposed research idea receives an intermediate score for every criterion that is used for its evaluation. Each score assesses how well a research idea satisfies a specific criterion – e.g. answerability, effectiveness, deliverability, equity, and others. The intermediate scores are computed by averaging all ‘non-blank’ responses (i.e. ‘1’, ‘0’, or ‘0.5’ points), thus ensuring that missing responses do not distort the calculations. All intermediate scores thus become percentages, providing a standardised measure of the collective opinion of many leading experts in the field. At this stage, the management team may decide to place weights or apply thresholds on selected criteria, based on the input from many stakeholders. The overall research priority score (RPS) can then be calculated as the weighted mean of all intermediate scores after the ideas that did not meet the thresholds are excluded. The final RPS can then be used for ranking, and the uncertainty of the final RPS for each research idea can be statistically evaluated using bootstrapping-generated confidence intervals.
This approach to prioritisation offers several advantages. It ensures transparency, minimises biases, and eliminates undue influences among participants. It provides a systematic, reproducible approach with simple, well-defined, and intuitive outcomes that can be used for funding decisions [76,157]. It also assesses the level of agreement between the scorers, thus exposing the areas of greatest controversy and providing an opportunity for targeted discussions on priorities after the completion of the process. The level of agreement can be assessed by a simple measure of average expert agreement (AEA), but also by Kappa statistics, or more recently, an improved version of the AEA score based on information theory, defined as the exponential of the negative entropy. Entropy is a widely used information criterion to quantify uncertainty; in this case, higher entropy implies greater uncertainty and less agreement.
The CHNRI method can also assess the internal structure of the scorers through studying the diversity of their responses and identifying clustering. Various methods of hierarchical clustering can be used to this end. One of the most recent advances of the CHNRI method was its adoption of AI and LLMs. In one of the recent exercises, the output of the LLM, based on its training on human collective knowledge, was compared with the results obtained through human collective opinion of the leading experts [76]. The results showed both striking similarities and interesting differences. This development of LLMs may assist us in better understanding the differences between human and AI-based prioritisation, representing an early example where humans and AI can supplement each other towards a desirable outcome. The CHNRI method also includes two approaches to linking priorities either with specific decisions regarding using a new or optimising an existing funding portfolio. To ensure that research investments remain cost-effective and aligned with current needs, the methodology allows for periodic re-evaluation and refinement [76,157].
Prioritising in private sector: portfolio and pipeline management tools
Real options analysis (ROA)
Real options analysis (ROA) is a method for prioritising innovation investments under uncertainty. Just as a financial option has value, so too does the flexibility to delay, expand, scale, or abandon a real-world project or idea under uncertainty. A tool that takes this into account would assist in prioritising strategic ideas based not only on expected value, but also on the value of waiting or adapting decisions over time. It would, therefore, capture option value in decisions where uncertainty and learning are significant, especially in strategic planning, innovation, R&D, pharmaceuticals, and rapid technology development. The approach extends financial options theory to real-world assets. It emphasises the value of flexibility in decision-making, offering a valuation framework that quantifies the strategic value of flexibility in uncertain environments. The concept was introduced in 1977 by Stewart C. Myers, Professor of Financial Economics at the MIT Sloan School of Management. He argued that traditional NPV methods undervalue investments by ignoring managerial flexibility [70]. He defined ‘real options’ as ‘opportunities to purchase real assets on possibly favourable terms’. Myers’ work focused on the idea that firms can be viewed as having both ‘real assets’ and ‘real options’, where real options represent opportunities to acquire assets on potentially advantageous terms. This approach allowed for the valuation of investments with embedded flexibility, like the ability to expand, contract, or abandon a project based on evolving information, which are not adequately captured by traditional discounted cash flow methods. Thus, it became a tool that allowed for the prioritisation of strategic decisions. The approach was further formalised by Avinash K. Dixit and Robert S. Pindyck, professors of economics at Princeton and MIT universities, in their 1994 book ‘Investment under Uncertainty’, which provided a rigorous economic foundation for the theory [174]. Lenos Trigeorgis, professor of strategy and finance at Durham University Business School, then expanded its practical applications, especially for strategic decision-making in R&D, energy, and public policy [175].
Research and development (R&D) portfolio matrices: Boston Consulting Group (BCG) and GE-McKinsey
R&D portfolio matrices, such as the Boston Consulting Group (BCG) Matrix and the GE-McKinsey Matrix, were developed in the 1970s as strategic tools to prioritise project or product ideas, R&D ideas, or business units. Their prioritisation was based on key dimensions like ‘market attractiveness’ and ‘internal capabilities’. These matrices were then applied in strategic management and later adapted to research portfolio decision-making in sectors like pharmaceuticals, technology, and global health, which are heavily dependent on R&D. Portfolio matrices guided strategic investment decisions and became popular in the 1970s. They have since been adapted to manage research, development, and innovation portfolios.
The BCG Matrix, also known as the growth-share matrix, was introduced in 1970 by Bruce Henderson [71] to help companies allocate resources across a portfolio of businesses or products based on market growth rate and relative market share. It classifies business units or projects into four quadrants – ‘Stars’, ‘Cash Cows’, ‘Question Marks’, and ‘Dogs’ – based on two axes – market growth rate (external opportunity) and relative market share (internal strength). It helps decision-makers allocate resources effectively by identifying which units to invest in, maintain, or divest.
The GE-McKinsey Matrix was developed in the early 1970s through a collaboration between General Electric and McKinsey & Company [176]. Its purpose was to evaluate business units based on industry attractiveness and competitive strength, providing a more detailed approach than the BCG Matrix. It introduced greater complexity with nine boxes instead of four, evaluating business units based on industry attractiveness and business strength. By the 1980s, both matrices were adapted for use in research-heavy fields such as pharmaceutical development, technology planning, and global health research funding. They remain influential in identifying which R&D initiatives to scale, monitor, or close. Comparatively, the BCG Matrix is a 2 × 2 grid with axes of market growth versus market share, while the GE-McKinsey Matrix is 3 × 3 grid with axes of industry attractiveness versus business strength [177].
Product road-mapping and prioritisation grids
Two influential tools for prioritising innovative ideas in dynamic environments are product roadmaps and prioritisation grids. They evolved over time through contributions from product development, technology management, and Agile software communities. The concepts were formalised and popularised between the 1970s and early 2000s through academic, industry, and consulting literature. Though developed independently, both help organisations to align innovation initiatives with their strategic goals. Product roadmapping evolved in technology management circles during the 1970s and 1980s. It was academically formalised by Robert Phaal, Clare Farrukh, and David Probert at the University of Cambridge’s Institute for Manufacturing [72]. Its purpose is to align product and technology development with strategic goals across time horizons. It is a visual timeline that links market drivers, product goals, and enabling technologies across time horizons. It helps coordinate innovation efforts and communicate priorities to diverse stakeholders. Prioritisation grids visually compare ideas along two or more criteria, such as impact vs feasibility or effort vs value. Their purpose is to compare and rank ideas based on two or more criteria (e.g. Impact vs Feasibility). They are rooted in quality management, design thinking, and Agile product development. Notable contributors to these tools include Noriaki Kano, who introduced the Kano Model to categorise features based on customer satisfaction in 1984 [178]. The Kano model is a set of guidelines and techniques used to categorise and prioritise customer needs, guide product development and improve customer satisfaction. Then, IDEO and Stanford d.school popularised effort-impact matrices in innovation workshops. The Lean Startup movement and Agile sprint planning, which integrated prioritisation grids into software product backlogs. Both tools are now widely used in global health strategy, public sector foresight, and corporate innovation planning.
Stage-Gate Model
The Stage-Gate Model, also known as the Phase-Gate Model, was developed in the 1980s by Robert G. Cooper based on empirical studies of high-performing innovation projects at companies such as DuPont, Exxon, and United Technologies. The model was formally introduced in Cooper’s seminal 1990 publication, ‘Stage-Gate Systems: A New Tool for Managing New Products’, published in ‘Business Horizons’ [73]. The model structures the innovation process into a series of discrete stages, each focused on specific activities such as idea generation, concept testing, development, and commercialisation. These stages are separated by decision ‘gates’, where cross-functional teams assess whether the project should advance, be modified, or terminated. Each gate evaluates the project against pre-defined criteria such as strategic alignment, market potential, and technical feasibility. The Stage-Gate Model remains a cornerstone of structured innovation, especially in sectors such as consumer goods, pharmaceuticals, and industrial R&D.
Modern approaches in digital age: AI-driven prioritisation tools
Knowledge graphs and semantic similarity clustering
Knowledge graphs and semantic similarity clustering have emerged as powerful computational tools for prioritising research ideas by systematically mapping relationships among concepts and identifying high-impact clusters for further exploration. Rooted in early work on semantic networks from the 1960s, such as the models of Allan M. Collins and M. Ross Quillian [74], American cognitive scientists from Northwestern University and BBN Technologies. These approaches have matured significantly over the past two decades through advancements in knowledge representation, ontology development, and natural language processing. A knowledge graph is a structured representation of entities (like diseases, genes, or interventions) and the semantic relationships that connect them. It enables researchers to visualise and traverse complex idea landscapes, helping identify conceptual clusters, detect research gaps, and inform strategic funding. The concept gained widespread attention in 2012 when Google formally introduced its ‘Knowledge Graph’ to improve semantic search [179], but the biomedical domain had already laid foundational work with the Unified Medical Language System developed by the US National Library of Medicine in the 1990s [180]. This ontology-driven framework enabled early applications of semantic clustering in health research. Semantic similarity clustering, a complementary approach, uses algorithms to group ideas or terms based on ontological proximity or vector-space similarity, often using tools like MeSH, the Unified Medical Language System, and embedding models derived from deep learning. This technique has been used for deduplicating proposed research questions, thematic mapping of large literature corpora, and identifying emergent clusters that may merit prioritisation. Since 2015, researchers like Percha and Altman have applied these tools to structure biomedical knowledge extracted from unstructured text using text mining and semantic networks for biomedical idea mapping and hypothesis generation [181]. Moreover, Himmelstein and colleagues demonstrated how graph traversal techniques can prioritise drug repurposing candidates based on the connectivity of evidence [182]. Large-scale initiatives, such as IBM Watson Health [183], Semantic Scholar from Allen Institute for AI [184], Elsevier’s research mapping systems [185], Microsoft Academic Graph [186], and the NIH’s NCATS Biomedical Data Translator Consortium [187], have operationalised these tools to support automated hypothesis generation, research portfolio analysis, and strategic foresight. From around 2018, these methods began being explicitly applied to idea prioritisation, especially in biomedical sciences where the volume and complexity of ideas are immense. Unlike traditional deliberative or expert-driven methods, knowledge graph-based prioritisation is data-driven, scalable, and less prone to individual bias. It enables automated and reproducible discovery of promising research areas by identifying semantically rich clusters of ideas that align with knowledge gaps or interdisciplinary convergence. Other key publications defining this field include those by Ehrlinger and Wöß in 2016 on formal definitions of knowledge graphs [188] and Wang and colleagues in 2017 on graph embeddings to compute semantic similarity between concepts and enable clustering [189]. Together, these works have laid the groundwork for a new class of AI-enabled methods for research prioritisation that are now integral to modern scientific discovery and funding intelligence systems [190,191].
Reinforcement learning-based portfolio optimisation
RL-based portfolio optimisation is a relatively recent advancement at the intersection of AI and decision science. It applies RL techniques to dynamically manage and prioritise investment portfolios, including R&D, financial, and portfolios of diverse investment ideas within uncertain, dynamic, and rapidly evolving environments. In this new approach to prioritisation, a reinforcement agent learns to take actions in an environment to maximise cumulative reward over time, optimising a portfolio to allocate resources (e.g. capital, research funding, investment ideas) across options to maximise expected value. RL has increasingly been applied to the optimisation of research, innovation, and investment portfolios. While the concept has no single originator, it builds on the convergence of foundational work in RL theory, portfolio management, and deep learning. The theoretical underpinnings were established by Richard S. Sutton and Andrew G. Barto, whose 1998 textbook ‘Reinforcement Learning: An Introduction’ remains the standard reference [75]. The first applications of RL to portfolio optimisation were pioneered by Moody and Saffell in 2001 [192], who used policy-gradient methods for financial trading. By the 2010s, researchers such as Jiang and colleagues applied deep RL to dynamic asset allocation problems, extending the method’s relevance to complex, nonlinear decision spaces, with a key reference published in 2017 [193]. OpenAI and Google’s DeepMind then extended RL to large-scale planning in science and complex decision environments in gaming with examples such as AlphaZero [194]. RL introduced adaptive learning agents that allocate resources or prioritise ideas dynamically based on observed feedback and evolving environments. They enabled self-improving, data-driven prioritisation under uncertainty, making it especially valuable for innovation portfolios where outcomes are difficult to predict and decisions must be staged or revisited. Such agents have emerging use in R&D portfolio management, drug development pipelines, global health prioritisation, and AI-led scientific discovery.
Automated priority setting via LLMs
The application of LLMs such as ChatGPT, Claude, Gemini, and DeepSeek to automated research prioritisation is a frontier development that began to emerge between 2023 and 2024. These models, built on billions of parameters and trained on vast amounts of textual sources, are now capable of reasoning across criteria and ranking ideas at scale, which was previously the domain of expert panels and consensus methods. The first peer-reviewed demonstration of LLMs applied to research idea scoring and its direct comparison to human expert group’s scoring based on the CHNRI method was published in 2024 by global health experts Peige Song from Zhejiang University, Igor Rudan from the University of Edinburgh, and their colleagues from the International Society of Global Health (ISoGH) [76]. They compared the top research priorities that the ISoGH’s global health experts generated using the CHNRI method to those generated by ChatGPT for the challenge of pandemic preparedness. It was an example of how human collective opinion can be compared to human collective knowledge that was used for AI LLM’s training, with the outcomes indicating the similarities and differences between humans and AI. This approach opened the door for AI-led research foresight, but also rapid priority setting by funders and public health agencies in emergency situations. The approach is highly scalable and cost-effective, allowing the rapid evaluation of a very large number of ideas, particularly useful in time-sensitive or resource-constrained settings [76,190]. It has the potential to complement and even augment traditional frameworks like CHNRI, Delphi panels, or MCDA. It is expected that other early adopters may be the Chan Zuckerberg Initiative with their use of AI for scientific discovery [195], MIT and Stanford with their work in hypothesis generation [196], and Semantic Scholar integrations at the Allen Institute for AI [197].
Engaging society in deliberation: participatory and democratic prioritisation
Cross-cutting philosophical and meta-theoretical approaches
Cross-cutting philosophical and meta-theoretical approaches for the prioritisation of ideas are not attributed to a single inventor, group or publication. They represent a meta-level evolution of thought across epistemology, philosophy of science, and decision theory. These approaches aim to critically assess, compare, and synthesise various idea-evaluation frameworks by exploring their underlying assumptions, value systems, and conceptual boundaries. Therefore, beyond procedural methods described in the rest of this paper, a rich body of philosophical and meta-theoretical work has shaped how we can think today about the prioritisation of ideas, critically examining the epistemological, ethical, and institutional foundations of prioritisation frameworks. These approaches offer insight into how knowledge, values, and power interact in decision-making, encompassing diverse intellectual traditions. They aim to develop meta-criteria, such as robustness, fairness, scalability, and inclusiveness, for choosing among evaluation frameworks. They can examine the limits of quantification, value-ladenness of science, and disciplinary worldviews in prioritisation. They also attempt to reconcile qualitative and quantitative paradigms in evidence and decision-making. The representatives of this line of thought are philosophers of science such as Thomas Kuhn, Imre Lakatos, Helen Longino, Paul Feyerabend, and Nancy Cartwright, and science policy scholars such as Helga Nowotny, Sheila Jasanoff, Silvio O. Funtowicz, and Jerome Ravetz. Their foundational ideas were presented between the 1960s and 1990s and applied to idea prioritisation from the 1990s onward, especially in science, research programmes, and meta-framework evaluations. Thomas Kuhn introduced the notion of paradigm shifts and the incommensurability of scientific worldviews, pointing to the issue of paradigm-dependence in how ideas are evaluated and prioritised [62]. Imre Lakatos and Paul Feyerabend challenged the idea of singular rational methods, advocating for methodological pluralism. In the 1970s, Imre Lakatos developed the idea of research programmes as evolving systems of ideas, challenging static hypothesis testing [198]. Paul Feyerabend advocated for epistemological anarchism, critiquing rigid methodological prioritisation [77]. The biopsychosocial model, a cross-cutting philosophical and meta-theoretical approach for prioritising ideas, was developed by George Engel in 1977 [199]. Helen Longino, in 1990s, emphasised the social nature of knowledge production and contextual empiricism – the role of communal norms and social epistemology in scientific credibility of ideas, and not just individual rationality [200]. Funtowicz and Ravetz pioneered post-normal science in 1993, arguing for participatory decision-making when ‘facts are uncertain, values in dispute, stakes high, and decisions urgent’ [201]. Sheila Jasanoff, Helga Nowotny, and others extended this thinking into science and technology governance, exploring issues of legitimacy, reflexivity, and public engagement [202]. Cartwright emphasised the contextuality and partiality of methods in prioritising ideas [203]. All of these approaches can help assess the assumptions and values embedded in prioritisation frameworks; examine the balance between qualitative and quantitative reasoning; address the limits of objectivity and the need for epistemic humility; and navigate ethical trade-offs in global health (e.g. disability-adjusted life years versus equity), AI governance, and innovation policy. They are applied in global health (e.g. ethical debates around disability-adjusted life, cost-effectiveness, and epistemic justice); technology governance (e.g. responsible research and innovation, anticipatory governance); research funding, AI and foresight (e.g. meta-analyses of value systems, ethics, and biases in algorithmic prioritisation). These works continue to influence frameworks such as the CHNRI, Delphi, MCDA, and (more recently) LLM-based prioritisation tools by reminding us that the way we evaluate ideas is always shaped by deeper worldviews, institutional interests, and contested values. These four clusters of ideation prioritisation methods – participatory budgeting, citizen juries, public consultations with weighting, and cross-cutting meta-theories – illustrate the growing sophistication of participatory and philosophical approaches in shaping what ideas societies pursue. They vary in structure, epistemic assumptions, scalability, and inclusiveness, but all contribute essential tools for navigating complexity, uncertainty, and value pluralism in modern decision-making.
Citizen juries and deliberative democracy forums
Citizen juries and deliberative forums emerged in the 1970s and 1980s as methods to integrate reasoned public deliberation into policy and research decisions. These approaches were rooted in democratic theory and the ideal of a more informed and reflective public discourse. They were pioneered by Ned Crosby, initially in his Ph.D. work at the University of Minnesota and then through his work at the Jefferson Center, while the first implementation in 1974 was on healthcare policy [78]. In citizen juries, a group of randomly selected citizens are given the information and time to discuss a particular issue and make recommendations. They usually bring together 12–24 persons, who receive balanced information, engage with experts, deliberate over multiple days, and ultimately make policy recommendations. This process provides a counterbalance to both elite-driven decision-making and uninformed mass opinion.
Parallel efforts in deliberative democracy were advanced by thinkers such as John Dryzek [204], who offered philosophical grounding for these practices in their work, shaping the theory behind deliberative forums. The concept of deliberative democracy, which emphasises the importance of reasoned discussion and debate in decision-making, has roots in ancient Greece. A deliberative forum is a larger, often one-off or multi-day event, where diverse lay participants discuss societal challenges with structured facilitation. Crosby’s work, particularly with citizen juries, is seen as a practical application of deliberative democratic principles, and deliberative forums are seen as its extension. As the next step, the deliberative poll combines public opinion polling with deliberation, with respondents surveyed before and after exposure to expert input and discussion. James Fishkin later developed Deliberative Polling® in the 1980s and 1990s as a scalable alternative to small juries, combining opinion polling with structured deliberation to assess shifts in public views after informed discussion [79,205]. He justified it as a democratic innovation to balance public opinion and expert knowledge and presented evidence of deliberative polling effects in real-world decision-making. Citizen juries, deliberative forums and deliberative polls have been used in various contexts, including policy development, candidate evaluation, and community problem-solving, demonstrating their potential to engage citizens in meaningful ways. These approaches have found wide application in climate adaptation funding, technology ethics, end-of-life care, and AI governance. They exemplify how randomly selected laypersons, when given time, information, and facilitation, can produce thoughtful, legitimate prioritisation of complex issues.
Public consultations and e-surveys with weighting
Public consultations and e-surveys with weighting are participatory tools that have matured alongside digital technologies and institutional demands for evidence-informed policymaking. Rather than being the product of a single innovator, these methods have evolved from legal, political, and technological traditions, especially gaining traction from the 1990s onward. The legal foundations for fair consultation are rooted in the legal concept of ‘legitimate expectation’, where the idea of weighting in public consultations were laid by cases such as Council for Civil Service Unions versus Minister for the Civil Service AC 374 (‘GCHQ’). The Gunning Principles, established by Stephen Sedley QC in 1985 in the R versus Borough of Brent ex parte Gunning case [80], provide a framework for fair and lawful public consultations: timely initiation, transparency of intent, adequate response periods, and serious consideration of inputs. E-surveys, particularly e-mail surveys, emerged in parallel, with early adoption in the 1980s. Kiesler and Sproull published one of the first analyses of email surveys in ‘Public Opinion Quarterly’ in 1986 [206]. As internet access expanded, digital surveys became essential tools for gathering structured feedback, with weighting mechanisms introduced to differentiate stakeholder influence based on expertise, role, or representativeness. Pioneering institutions that implemented public consultations include OECD in 1990s and 2000s [207], promoting public consultation and participation in governance, and as a formal tool for evidence-based policymaking; then, UK NICE, with integrated consultation and stakeholder weighting in health technology appraisal process in 2004 [208]; also, the JLA [23,154] and the CHNRI [24,25], which developed equal-weight or weighted e-surveys for research priority-setting; and the European Commission, which adopted structured e-consultations in science policy using structured questionnaires with weighted responses [209]. Applications range from health research funding (e.g. NIHR, WHO) and policy design (OECD strategies) to science and innovation roadmapping (e.g. Horizon Europe). Tools like Pol.is and IdeaScale exemplify how digital platforms now integrate statistical weighting, multi-criteria decision analysis, and feedback loops into large-scale idea prioritisation exercises.
Participatory budgeting
Participatory budgeting (PB) is one of the most influential and widely replicated models for citizen-led ideas for allocation of public resources. It was first introduced in 1989 in Porto Alegre, Brazil, when the Workers’ Party took the office. Developed as a response to growing demands for transparency, equity, and democratic participation, PB allows community members to directly propose, discuss, and vote on how public funds are allocated. The Porto Alegre model emerged during Brazil’s transition from dictatorship to democracy, serving as a mechanism to rebuild trust in public institutions and empower marginalised communities. It has been recognised for its success in prioritising the needs of the poorest areas, improving public services, strengthening governance, and increasing citizen participation. It represented a citizen-led process to generate, discuss, and vote on ideas for public funding, redistributing power, and enhancing transparency, and has influenced deliberative democracy, community-led innovation, and bottom-up research prioritisation frameworks worldwide. Key figures in its development include municipal leaders like Tarso Genro and Raul Pont, along with scholars such as Boaventura de Sousa Santos [81], Gianpaolo Baiocchi [82], and Brian Wampler [83], who later provided robust academic analyses of its mechanisms and outcomes. At its core, PB is structured around a deliberative process: citizens generate ideas, debate them in local assemblies, and vote on proposals, which are then implemented using municipal funds. It has been applied across various domains, such as urban infrastructure, education, health, and youth programmes. PB’s success in Porto Alegre inspired more than 7000 replications worldwide, from Paris and New York to Nairobi and Seoul. It has evolved into digital formats, such as Decidim and Consul platforms [210,211], and school-level budgeting initiatives. Its global legacy lies in demonstrating that inclusive, deliberative models can redistribute power, foster equity, and enhance governance legitimacy.
Foundational philosophical heuristics and reasoning frameworks: approaches to prioritising ideas beyond specific techniques
At the end of this section on the methods used for prioritising ideas, we remind yet again of historically useful principles that do not strictly represent specific, well-developed methodologies, but that still serve as foundational philosophical heuristics that continue to shape how we compare and prioritise competing ideas. These approaches – Occam’s Razor, Bayesian inference, the dialectical method, and epistemic humility – are not decision-making tools in the narrow sense, but rather enduring intellectual frameworks that guide reasoning under uncertainty, complexity, and value plurality. Their influence spans science, medicine, policy, and AI, embedding centuries of epistemological development into contemporary idea evaluation.
Occam’s Razor
Occam’s Razor (also spelled Ockham’s Razor), one of the most enduring heuristics in Western intellectual history, prioritises explanatory simplicity: ‘Prefer simpler ideas when explanations are equal’. Although the principle predates him, it was William of Ockham, a 14th-century English Franciscan friar and philosopher, who popularised its application in logic and metaphysics. He articulated the core principle circa 1323–28 during his major theological and philosophical works, especially in Avignon and Munich. He advocated that ‘entities should not be multiplied beyond necessity’, a phrase later paraphrased to summarise the principle. While Ockham never used the term ‘Occam’s Razor’ himself, the idea is evident throughout his major works, particularly ‘Summa Logicae’ (circa 1323) and ‘Quodlibetal Questions’ (circa 1328) [84,85]. It introduced foundational principles of logic and parsimony in reasoning, foundational for both evaluating and prioritising ideas based on simplicity and explanatory power. Ockham frequently applied these principles in theological and logical debates. Occam’s Razor became integrated into scientific method, logic, AI model selection, medicine, and all disciplines that confront competing hypotheses or explanations. The term ‘Occam’s Razor’ was coined retrospectively, gaining currency in the 17th century through philosophers like Johannes Clauberg and Libert Froidmont, and was later formalised by Sir William Hamilton, a renowned Scottish philosopher and professor at the University of Edinburgh, in 1852 [212]. The principle, as it is applied today, suggests that among competing hypotheses, the one making the fewest assumptions should be selected, unless additional complexity is clearly warranted. It underpins model selection in scientific research, medical differential diagnostics, and machine learning, where it appears in formalised metrics such as the Akaike information criterion [213] and Bayesian information criterion [214].
Bayesian inference
Bayesian inference provides a formal mechanism for updating beliefs about hypotheses or models that underlie research ideas in light of new evidence. Though named after Thomas Bayes, English Presbyterian minister and mathematician, who outlined a specific probabilistic problem in his 1763 posthumous essay, the method was given modern mathematical form by Pierre-Simon Laplace in the late 18th and early 19th centuries. Bayes’ original work, ‘An Essay Towards Solving a Problem in the Doctrine of Chances’, was edited and published by his friend Richard Price in the ‘Philosophical Transactions of the Royal Society’ [86]. This publication introduced what we now call Bayes’ Theorem, which allows updating the probability of a hypothesis in light of new evidence. Laplace, French polymath who influenced the fields of physics, astronomy, mathematics, engineering, statistics and philosophy, extended Bayes’ ideas into a general framework for probabilistic reasoning, applying it to astronomy, population statistics, and political decision-making [215].
Bayesian inference calculates the probability of a hypothesis (posterior) by combining its prior probability with the likelihood of observing current evidence. It inherently incorporates a penalty for overly complex hypotheses, which is a formal echo of Occam’s Razor. Importantly, it can be used to rank competing ideas based on both plausibility and fit with data. Today, it underlies model selection, decision theory, policy forecasting, health technology assessment, AI safety frameworks, and many more areas of progress. Its conceptual elegance lies in combining subjective beliefs with objective evidence through formal probability.
Dialectical method – thesis, antithesis, synthesis
The dialectical method – resolving tension between opposing ideas through synthesis – originated in ancient philosophical debates, but was systematised in modern form by German idealists, particularly Johann Gottlieb Fichte and Georg Wilhelm Friedrich Hegel. Though Hegel did not use the terms ‘thesis – antithesis – synthesis’ explicitly, this triadic model was later used to describe the dynamic movement of ideas in his philosophy, especially in his book ‘Phenomenology of Spirit’ in 1807 [87]. The method involves an initial idea (thesis), its negation or contradiction (antithesis), and their resolution in a higher-order synthesis, which usually incorporates the elements of both. This process is not merely cyclical but developmental: each synthesis becomes a new thesis, advancing understanding through iterative contradiction and integration. Heinrich Moritz Chalybäus was the first to formalise this triadic structure in 1837, retroactively applying it to Hegel’s work [216]. The dialectical method influenced prioritisation of ideas in critical theory, scenario planning, and design thinking, particularly wherever synthesis of conflicting perspectives is essential. While it was not originally formulated for prioritising ideas in a modern evaluative sense, it has been used for evaluating competing perspectives, identifying contradictions, and advancing knowledge – making it relevant to meta-theoretical prioritisation in philosophy, politics, and science.
Epistemic humility and pluralism – valuing diverse viewpoints in evaluation
The principles of epistemic humility and pluralism challenge the idea that knowledge is objective, singular, or complete. Rooted in ancient Greek philosophy, particularly in Socratic dialogue, they found renewed expression in modern political philosophy, feminist epistemology, decolonial theory, science and technology studies, social sciences and post-modern approaches. Socrates’ assertion of his own ignorance (Plato’s Apology) provided a foundation for epistemic humility [217]. In the 19th century, John Stuart Mill argued in 1859, in his book ‘On Liberty’, that exposure to dissenting views strengthens knowledge [88]. In the 20th century, Paul Feyerabend’s ‘Against Method’ (1975) rejected rigid scientific orthodoxy, advocating instead for methodological pluralism and epistemological anarchism: ‘Anything goes’ [77]. Further contributions came from many more recent thinkers, who all argued that diverse perspectives, including those from lay, indigenous, and marginalised groups, improve the legitimacy, robustness, and ethical grounding of idea evaluation. Today, these principles are embedded in participatory research models like CHNRI’s weighting and setting thresholds by diverse shareholders, or JLA partnerships. The heuristics and meta-frameworks explored in this section, which included parsimony, probabilistic updating, dialectical synthesis, and epistemic humility, offer essential cognitive and ethical guidance in the face of competing ideas. Though they emerged from different epochs and disciplines, they converge on a common premise: the process of prioritisation is not merely technical, but also philosophical. By combining these timeless principles with contemporary participatory methods, we can foster more reasoned, inclusive, and reflexive approaches to idea evaluation in science, policy, and innovation.
DISCUSSION
Toward a unified science of ideas
This exercise of systematically identifying and structuring the methods that were historically used by humans to generate, evaluate, and prioritise ideas has revealed several insights. Despite their disciplinary diversity, these methods exhibit striking structural, functional, and philosophical parallels. Their fundamental aim has remained the same: how to make sense of competing possibilities and decide which ideas are worth acting upon? This rational approach opposes an alternative where individuals with strong beliefs, and often vested interests, hold and enforce their ideas upon others, while finding ways to legitimise them in the process.
When viewed through the ideometrics lens, several recurring patterns emerge across methods, regardless of their disciplinary origin, epistemic orientation, or technical complexity, uncovering interesting parallels across disparate traditions. A number of methods, especially the more contemporary ones, seem implicitly or explicitly to follow a layered process described in the brain’s ‘sense of ideas’ concept: generation, evaluation, and prioritisation of ideas [1]. For example, brainstorming generates divergent options, Delphi panels evaluate them through expert feedback, and AHP ranks them using quantitative pairwise comparisons. The CHNRI method follows a similar triad: generation through open solicitation, evaluation using expert scoring, and prioritisation by aggregated weighted scores.
There is also a persistent tension across methods stemming from the balancing of subjective human judgment with the aspiration for objectivity. Peer review, Delphi, and NGT depend on subjective expert or stakeholder input, while TRLs, citation metrics, and MCDA aim to quantify value. Although trained on inherently subjective human content, AI-based systems like LLMs and reinforcement learning introduce a new variant: machine-generated outputs grounded in massive amounts of data. This interplay suggests that the tools of ideometrics should not aspire to remove subjectivity. Instead, they need to structure and balance it transparently.
Recursive refinement is another common theme, as many techniques rely on iterative feedback loops. Delphi rounds, Agile sprints, design thinking cycles, and machine learning training epochs all reflect recursive logic. They aim not for immediate certainty, but for progressive approximation. This emphasis on refinement over resolution implies that the systems within ideometrics are most effective when treated as dynamic, evolving processes rather than static decision tools.
Across the methods of ideometrics, the act of transparent aggregation seems to be the ‘core function’. Whether scores, votes, arguments, or preferences are in question, aggregation is the pivotal mechanism that transforms individual judgments into collective decisions. The MCDA aggregates weighted scores; prediction markets aggregate probabilistic bets; knowledge graphs aggregate semantic relations. This central role of aggregation suggests that future tools in ideometrics should continue to innovate around how inputs are synthesised, how transparency is preserved, and how differing values are represented.
Finally, each method encodes implicit assumptions about what counts as a ‘good’ idea, based on some embedded values and norms: whether it is logical coherence (e.g. deductive reasoning), novelty (e.g. patent originality), societal need (e.g. JLA PSPs), stakeholder consensus (e.g. deliberative democracy), or impact on equity (e.g. CHNRI). Philosophical heuristics such as Occam’s Razor or epistemic humility make these values explicit, while technical tools like AHP or CHNRI embed them in scoring rubrics. Recognising these normative assumptions is crucial, because ideometrics is never ‘just technical’. It needs to be appreciated that it is ethical, can be subjective, and even political in some contexts.
What is still missing?
On the note of the role of politics in this field of science, our current work placed insufficient attention to the role of political authority and financial power on prioritising ideas. It is difficult to scientifically address political decisions that seem to conflict rational reasoning, but still have large impact on the ideas that are prioritised. This is especially true when such political decisions have wide public support based on either insufficient information, or even disinformation [91]. Furthermore, funding decisions on future strategy by the main funding agencies can also be political and have impact on prioritisation of ideas for the entire scientific community. Michel Foucault and several other thinkers through history have addressed these themes profoundly in their highly influential work [94], so we saw it as beyond the scope of our current efforts to define and position ideometrics.
We already explained in the introductory section that we were unable to grasp all the approaches to generating, evaluating, and prioritising ideas that may exist. There are certainly approaches to priority setting that may be quite well known, but we managed to miss these for several possible reasons. Language barrier might be one – we could not really assess the tools that likely exist in many local settings world-wide and are described in the literature in local languages.
Another substantial limitation is that our work approached this field from a largely Western and rationalist perspective. We mentioned in the opening paragraphs of this work that, in more traditional cultures, the emphasis in choosing best ideas could often be based on ancestral traditions and folklore, the experiences of elders, intuition, revelation and spiritual enlightenment [3,4]. We did not attempt to enter that vast space and try to highlight the most enduring and/or promising approaches.
It is possible that, at times, our work could give impression that it is mixing approaches that are too different in their specificity, from the rather sublime debate that characterised the ancient Greece and its philosophers to the prescriptive structures of the CHNRI or JLA methods. In trying to be systematic in a large task that we undertook, we left together the methods that were at disposal thousands of years ago and the modern, AI-based tools, although times and options have surely changed over that period.
There is also an area within ideometrics where evidence-informed policymaking needs to be based on the information that is emerging in real time, through iterative rounds, as we saw during the COVID-19 pandemic [218,219]. Within this specific context, the evolution of so-called ‘rapid reviews’ and ‘living reviews’ [220,221], and the role of systematic reviews in general, as a methodologically sophisticated approach to summarising the empirical evidence-base [222], deserve special attention within the field of ideometrics. This work describes the context within which the ideas are prioritised. Therefore, it can mitigate the detrimental role of disinformation and attenuate the impact of biased or skewed perceptions. This is why they should be elevated to a formal research and development priority, as methodological integration of evidence synthesis into ideometric prioritisation, ensuring their role is actively studied.
In priority-setting processes that involve many individuals, conflicts of interest can significantly influence outcomes, potentially distorting the collective judgment. Individuals with financial, professional, or personal stakes in particular outcomes may consciously or unconsciously advocate for priorities that align with their interests rather than the broader public good. When such conflicts are not disclosed or adequately managed, they can bias the ranking of research topics, divert resources away from areas of greatest need, and undermine the credibility of the entire process. This risk increases in multi-stakeholder exercises where participants come from diverse sectors, including academia, industry, government, and civil society. To safeguard the integrity and legitimacy of priority-setting exercises, it is essential to require transparent disclosure of all potential conflicts of interest, adopt independent facilitation, and apply structured consensus-building methods that dilute individual influence and promote collective reasoning. Clear governance mechanisms are key to ensuring fairness, accountability, and trust in the final outcomes [223].
There is another area well-worth exploring and developing within ideometrics: the one of empirical assessment whether certain ideas are indeed ‘better’ in practice. Namely, researchers can track real-world implementation and outcomes of competing ideas over time. Comparative evaluations, such as randomised controlled trials, natural experiments, or pilot programmes, can measure the effectiveness, cost-efficiency, and scalability of ideas in different contexts. Outcome indicators, like health impact, uptake, economic return, or social acceptance, provide measurable benchmarks. Longitudinal studies and impact evaluations can further reveal which ideas deliver sustained benefits. By linking idea selection to empirical feedback loops, we can iteratively refine priority-setting processes and identify which types of ideas consistently lead to meaningful, evidence-backed improvements [224].
Moreover, we believe that ideometrics is needed precisely because methods that serve similar purposes can lead to different outputs. Therefore, the field requires systematic mapping, comparison, and eventually empirical evaluation of how methods behave under different decision contexts. Different techniques within the same methodological umbrella may not be interchangeable, and their outputs may diverge. Understanding the sources of divergence – mathematical, epistemological, or contextual – is one of the core scientific tasks for the emerging field of ideometrics. Future empirical research should evaluate how different MCDA approaches perform when applied to identical decision challenges, under controlled conditions, to provide clearer guidance for method selection. Therefore, ideometrics is both integrative, critical and comparative, as it recognises methodological variability as one of the central phenomena that require scientific study.
Importantly, our phrasing around the brain’s ‘sense of ideas’ and the application of the ‘value of information’ concept should not be misinterpreted as implying a literal sensory modality or a direct neuroscientific claim. In our paper, these concepts serve as conceptual analogies, used to illustrate the intuition that the human mind continually generates, evaluates, and selects among competing representations, much as it processes sensory inputs. Our argument is therefore conceptual, not neuroscientific: we introduce a high-level organising principle that unifies multiple independent traditions. The mind continually processes ‘candidate ideas’, much as the senses process competing sensory signals, and their dominance depends on an implicit assessment of attractiveness, feasibility, expected impact, and other possible criteria. In relation to this, we anticipate a rise in quantitative studies of the ‘life course’ of ideas – their birth, patrimony, spread, evolution, change, and eventual rejections and resurrections through time.
Finally, it needs to be acknowledged that some of the best ideas throughout human history, that eventually resulted in large positive impact on societies, science, technology, art, and human experience in general, came from a sequence of serendipitous events that happened seemingly by mere chance, in a very beneficial way. In fact, there were also great ideas that resulted from a gross error in rational judgement, or through entirely disinformed sequence of thought or actions, but eventually had a very positive impact [225]. That may be the ultimate challenge for ideometrics, as a field of science, to learn how to harness without unacceptable risks. This area of study, i.e. understanding the non-systematic paths to impactful ideas, will require a dedicated research priority – ‘the ideometrics of serendipity’ – alongside the others proposed in the research agenda for the future of ideometrics (see below).
From framework to field: developing the science of ideometrics
Having identified, organised, and synthesised a broad range of methods to support ideometrics, we now turn to the question: how to elevate ideometrics from a useful framework to a recognised field of scientific inquiry? The first step is empirical – ideometrics must subject itself to the very tools it has catalogued, and the first examples of such approach exist already [226,227]. We therefore propose to conduct a quantitative bibliometric analysis of the ‘footprint’ of each ideometric method in the scholarly literature. Using established citation databases such as Scopus, Web of Science, and Google Scholar, we aim to track the total number of publications referencing each method; measure citation velocity and longevity (e.g. half-life); assess disciplinary spread and cross-sectoral uptake; identify highly cited papers and thought leaders in each tradition; and evaluate their adoption in policy, health, engineering, business, and education. This analysis will not only provide an empirical baseline for future research, but also reveal which methods have achieved practical traction, which remain underutilised, and which have declined in relevance. We will then continue to monitor the landscape for any new methods that may emerge.
In parallel, we aim to introduce a set of reporting guidelines for ideometric exercises – i.e. attempts to generate, evaluate and prioritise ideas using any of the methods – analogous to the PRISMA or CONSORT statements in health research [228,229]. These guidelines would standardise the documentation of ideation exercises, covering elements such as: who participated, and how they were selected; what criteria were used for evaluation and why; which aggregation methods were employed; how transparency, replicability, and equity were ensured; and how results are intended to inform policy or decision-making. These guidelines would promote comparability of reporting across methods, improve the reproducibility of ideometric studies, open the door to meta-analyses of approaches that addressed the same challenge, and enable comparisons across domains (e.g. by comparing Delphi vs CHNRI vs LLM rankings on the same problem set). They could also form the basis of journal submission requirements for ideometric-based research papers. Box 1 summarises what we believe could eventually qualify ideometrics as a field of science.
A possible research agenda for the future of ideometrics
To establish ideometrics as a coherent and dynamic field of science, we propose the following R&D priorities:
Building a living repository of methods and use cases: we propose the creation of a curated, openly accessible digital repository that catalogues the methods of ideometrics, case studies, scoring templates, decision matrices, and user experiences. This could, for example, become a ‘Wikipedia meets GitHub’ for idea management – both archival and interactive.
- Conduct cross-method comparative trials: just as new treatments are tested against standards of care, the methods of ideometrics should be empirically compared. For example: if CHNRI, Delphi, AHP, and an LLM-based method are all applied to the same set of global health research questions, do they produce similar priorities? If not, why? What assumptions are driving the divergence? Comparative trials would reveal method sensitivity, value alignment, and practical feasibility.
- Investigate cognitive and social dynamics: as ideometrics increasingly shapes real-world decisions, we must understand its psychological and sociological dimensions. How do different stakeholder groups engage with ideometrics tools? What cognitive, social, financial, political, power-driven, and other biases persist despite formal structures? How do power dynamics influence prioritisation in participatory settings? This work would draw from behavioural science, sociology of science, and ethics.
- Integrate AI and human-centred design: our next frontier is the development of a user-friendly, transparent, and democratic software tool for ideometrics-based decision support. We are currently working on such a platform [230]. It will allow users (individuals, organisations, governments) to input ideas and associated metadata; facilitate transparent scoring across selected criteria; include participatory features (public voting, weighting, deliberation); offer AI-enhanced clustering, validation, and prioritisation tools; and provide structured outputs that are traceable, auditable, and replicable. This software will serve as both a research engine and a policy tool, bridging the gap between academic rigour and practical application. Its ultimate aim is to support evidence-informed, participatory, and adaptive decision-making across sectors.
Gaps that ideometrics is expected to fill
Although many fields debate how ideas should be generated, evaluated, or prioritised, such as decision science, innovation studies, behavioural economics, philosophy of science, evidence-based medicine, and public policy, no existing discipline has attempted to unify these debates under a single conceptual and methodological framework. The gaps that ideometrics is expected to fill are:
- A conceptual gap: Existing debates treat idea processes as discipline-specific, not universal. Decision scientists examine structured choice under uncertainty; innovation theorists analyse product design; philosophers examine logic and scientific justification; public-health researchers develop priority-setting tools; AI researchers explore computational creativity. These debates are rich, but fragmented: each field sees only its own techniques, rarely recognising that they share a common three-step cognitive structure of generating, evaluating, and prioritising ideas contributes a unifying conceptual lens that identifies this structural commonality and treats the lifecycle of ideas as a generalisable object of scientific inquiry.
- A methodological gap: There is no comparative science of idea-processing tools. Although many decision frameworks exist, no systematic effort has been made to map them across domains, compare their structures, analyse their assumptions, evaluate their convergence/divergence, study their performance on shared problems, develop reporting standards, or define cross-domain selection criteria. Ideometrics directly addresses this methodological vacuum by proposing the architecture through which such comparisons can be made.
- A policy and practice gap: Societies face unprecedented idea overload and lack tools for rigorous prioritisation. Policymaking, scientific funding, global health, technology investment, and AI governance all face the same challenge: too many ideas, too little time and capacity. Existing frameworks are siloed, incompatible, and inconsistently used. Decision-makers lack a coherent evidence base for choosing between prioritisation methods. Ideometrics positions itself as a practical science capable of informing how institutions structure expert input, integrate citizen participation, combine human and machine intelligence, evaluate trade-offs, avoid bias and disinformation, and choose appropriate methods for specific decision contexts. This explicitly addresses the policy need for systematic, transparent, and scalable approaches to idea selection under scarcity.
In summary, ideometrics is a new integrative approach that should fill a well-defined conceptual gap in how the lifecycle of ideas is understood, address a methodological gap by enabling cross-method comparison, and respond to a policy gap created by the increasing volume, complexity, and consequences of idea-based decisions in the 21st century.
Practical implementation and societal impact
For ideometrics to realise its transformative potential, it must go beyond academic theory. We envision its practical implementation across several domains, which include science and research prioritisation, public policy, corporate strategy and innovation, education and capacity building, global governance and foresight, and others. Detailed practice-oriented extensions of this work will be forthcoming, ensuring that ideometrics goes beyond this manuscript in its ambition to be implemented (see also Box 2), while this paper is intentionally focused on laying the conceptual, theoretical and normative foundations.
Other important considerations
We provided ample additional information in Online Supplementary Document. Table S1 lists the scientific field(s) of origin and application for the presented methods, while Table S2 provides comparative attributes of methods within ideometrics. Then, Box S1 addresses epistemic risks explicitly, such as over-generalisation, value-ladenness, or the limits of quantification, when unifying qualitative and quantitative traditions to pre-empt possible critiques of reductionism for ideometrics. Box S2 provides a discussion of ethical, equity, and governance implications, such as algorithmic bias in AI-based ideation, ownership of collective knowledge, and participatory fairness. Box S3 brings a schematic conceptual model illustrating the relationships and feedback loops among idea generation, evaluation, and prioritisation. Box S4 explains our position on including non-Western epistemologies. Box S5 considers the potential biases in LLM-generated ideas (e.g. hallucination or cultural skew) and their relevance to the ‘value of information’ concept. Box S6 explains LLM’s research priority setting within the CHNRI exercise. Box S7 considers the role of political authority, financial power, deliberate political propaganda and disinformation in ideometrics. Finally, Box S8 recognises and appropriately manages this paper’s reliance on the lead author’s prior work in this field.
CONCLUSION
Working on this paper, we increasingly felt like archaeologists who have discovered a field of science that has been here all along for centuries, present in virtually every line of human activity, but that no one could easily notice, let alone grasp in its wholeness, simply because it was dispersed too widely and broadly, in chunks too small and diverse to connect. Connecting it into a mosaic of ideometrics across many fields, using the simple framework to ‘generate-evaluate-prioritise ideas’, it felt like an effort in archaeology of human approaches to ideation, assessment of the new ideas, and choosing those that seemed worth following to people over times and spaces.
Ideometrics is expected to evolve along several complementary dimensions: as a theoretical paradigm, offering a unifying conceptual language for understanding the lifecycle of ideas across domains; as a methodological discipline, enabling systematic comparison, evaluation, and refinement of the diverse tools used for generating, evaluating, and prioritising ideas; and as a practical toolkit, providing structured, transparent, and context-appropriate methods that decision-makers, researchers, and innovators can apply in real-world settings. For science, it should provide a cumulative framework for improving hypothesis generation, research prioritisation, and methodological consistency; for policy, it should offer transparent, evidence-based decision processes for balancing competing ideas under uncertainty, incorporating expert and citizen perspectives; and for innovation and technology, it should improve portfolio decision-making, concept selection, and evaluation of emerging technologies, including human–AI hybrid ideation systems. In the longer term, ideometrics should become a unified, rigorous, and empirically grounded science of ideas that supports better decision-making and priority-setting across multiple sectors.
Ideometrics does not seek to mechanise human creativity or reduce thought to computation. Rather, it invites us to see the life of ideas, from their inception to their evaluation and selection, as a subject of structured inquiry. It respects the depth of ancient heuristics and the promise of modern algorithms. It embraces pluralism, transparency, and methodological humility. It recognises that ideas are not only cognitive artefacts, but also social commitments: they are shaped by who speaks, who listens, and how we all make decisions. By documenting its foundations, mapping its landscape, and outlining a roadmap for growth, we will help this work to catalyse the emergence of ideometrics as a vibrant interdisciplinary field of science. Its potential is not only intellectual, but also highly practical: a better science of ideas can lead to generating and prioritising better ideas, and better ideas can, in turn, lead to a better world.