Abstract
Over seven decades the question of how creativity works has moved from speculation into the laboratory, yet what reaches the general reader is mostly conclusions with nothing attached to them. This article tries to close that gap: in each case the experimental design is described before the result is reported. The argument runs across three layers. The cognitive layer shows that creating is not one operation but two, with different logics, and that the order of the two has a measurable effect on the originality of what is produced. The neural layer shows that those two operations rest on networks whose activity is normally inversely related. The contextual layer accounts for why people with the same underlying mechanism produce different work. A final section sets out the limits of the evidence, including a serious objection that applies to much of this literature and the incomplete answer that has been offered to it.
Keywords: creativity, creative cognition, divergent thinking, incubation, default mode network, constraint, preinventive structure, sketching
1. The problem
Before asking how something works, it helps to be clear about what is being explained.
The definition now standard in the research literature has two conditions: originality and effectiveness. Runco and Jaeger (2012) arrived at it by reviewing dozens of earlier definitions and extracting what they shared. The second condition is the one usually dropped. Something merely new, that does no work, is not creative under this definition. A random answer is original. It is not creative.
A limitation follows from that definition and should be kept in view to the end of this article. Effectiveness has to be judged, and the judge stands outside the person. Creativity is therefore not a wholly intra-individual phenomenon, and no individual measure settles it on its own. Section 7 returns to this, and by then it is no longer an abstract worry.
The scientific study of the subject is usually dated to Guilford's (1950) address to the American Psychological Association. What mattered there was not the introduction of a concept but a methodological claim: that creativity could be measured and studied like any other cognitive capacity. The concept he offered for the purpose was divergent thinking, the ability to produce many varied responses to a problem that has no single correct answer.
Against this stands the everyday explanation, which is an explanation by trait: some people are creative and some are not. The trouble with it is not that it admits individual differences, which are real. The trouble is that it explains nothing. Saying that someone is creative because they have talent is an explanation in exactly the way that saying a sleeping draught induces sleep because it possesses a dormitive property is an explanation. The name of the thing has been repeated and then offered as its own cause.
The three layers below are an attempt to fill that gap.
2. Scope and method
This is a review of existing literature, not a report of new work. Its sources are peer-reviewed papers and meta-analytic reviews from 1950 to 2018, and the order of presentation follows the logic of the argument rather than the chronology of publication.
Two limitations should be stated at the outset:
- Most of the neural literature is correlational. Causal conclusions are not available from it.
- The most widely used instrument in this field has weak predictive validity, and that objection applies to some of the very studies cited below. Section 7 opens the problem rather than hiding it. Without it, the conclusion of this article would go past its evidence.
3. The cognitive layer
This layer answers one question: when someone creates, what operations actually run? The answer is built in four steps: material, operations, timing, and the moment of solution.
3.1 Material: why some minds make more distant combinations
Mednick (1962) gave an answer whose simplicity is deceptive. To create is to bring together elements that sit far apart in memory. The further apart they are, the more novel the connection; and if the connection also does some work, that is what we call creativity.
To make the claim testable he built a task still in use. Three apparently unrelated words are given and the participant must find a fourth that links all three. The classic item is cottage, swiss, cake, and the answer is cheese. Solving it works exactly as the theory describes: a chain of associations runs out from each word, and the answer sits where the three chains meet.
On that basis Mednick attributed individual differences to what he called associative hierarchies. Given a stimulus, some people produce a few common associations and stop; that hierarchy is steep. Others continue past the common associations into less probable ones; that hierarchy is flat. The subtle point is that both groups start from the common associations. The difference is not where they begin but how far they travel before stopping.
For decades this remained a hypothesis without a quantitative measure. What changed it was the arrival of network science. Kenett, Anaki and Faust (2014) modelled semantic memory as a network whose nodes are concepts and whose edges are associative links, then compared the structural properties of that network across people with different creativity scores. The networks of more creative individuals proved richer, less rigidly structured and more interconnected. In such a network the distance between any two concepts is shorter, so moving from one to another is more probable.
Two results from this section carry forward. What was a hypothesis about behaviour in 1962 is now a property of memory structure that can be measured. And, more consequentially, that property depends on the contents of memory. The network is built from what has gone into it. A mind holding less material makes fewer combinations, and that is a limit on input, not a defect in the mechanism.
3.2 Operations: creating is two things, not one
Knowing the material does not tell us what is done with it. The most precise available answer is the model Finke, Ward and Smith (1992) set out in Creative Cognition.
The model's name, Geneplore, is built from generate and explore, and the compound carries the claim: creating is not one operation but two, with different logics.
Generation covers:
- retrieval of knowledge from memory
- association
- mental synthesis and transformation
- analogical transfer
- categorical reduction
What comes out of this phase is not a solution. The authors coined the term preinventive structure for it: unfinished forms with no settled meaning or function, but with promise in originality and usefulness.
Exploration covers:
- interpretation of those same structures
- hypothesis testing and attribute finding
- functional inference
- contextual shifting
- searching for limitations
This is where a preinventive structure acquires meaning and becomes usable.
So far the model could be one more descriptive taxonomy. What turned it into an empirical claim were the experiments Finke (1990) designed and ran over a three-year period.
The design. Participants are shown a set of simple three-dimensional visual parts: sphere, cube, cone, cylinder, wire, ring, flat plate and so on. Three parts are selected. Participants are asked to combine them mentally into a form. At this stage nothing is said about use and they are not asked to make anything useful. What they produce is the preinventive structure. Only afterwards are they given a functional category, such as toys, furniture or transport, and asked to interpret the form they have made as an object of that category. Independent judges then rate the outputs for originality and practicality.
The result that matters comes from comparing this order against its reverse. When the functional category is announced before the form is built, so that participants know from the start that they are making a toy, the outputs are rated less original. Knowing the goal in advance narrows the range of combinations, and what gets made stays close to familiar examples of the category. In this literature the phenomenon is called fixation.

The value of the finding is that it converts an order into a measurable effect. Separating generation from exploration is not, in this experiment, a stylistic preference. It is a variable, and manipulating it changes the output.
A further point follows and is usually missed: the incompleteness of a preinventive structure is a condition of its working. A structure that is finished and fixed too early leaves no room for reinterpretation, and exploration has nothing to do with it. Ambiguity, at this stage, is not a defect.
Notably, the same conclusion has been reached from an unrelated direction. Goel (1995) showed that ambiguous symbol systems such as freehand sketches preserve, precisely through their ambiguity, the possibility of reinterpretation and lateral movement between alternatives. Goldschmidt (1991) showed that an external representation itself induces new images in the designer. Two separate literatures, one result.
The present author proposes that the two can be joined: a sketch is an externalised preinventive structure. Section 8 develops the proposal and sets out its testable predictions.
3.3 Timing: testing a hundred-year-old claim
The model above says which operations run. It says nothing about the interval between them, and that question has a long history.
Wallas (1926) described the process in four stages: preparation, incubation, illumination, verification. The second, deliberately setting the problem aside, remained for decades a much-repeated and untested claim, and a large part of the popular literature on creativity was built on it.
The design of incubation experiments is simple and consistent. A participant works on a problem for a period. Then one of three things happens: they continue without a break; they fill an interval with an unrelated task; or they rest. Then they return to the problem. What is compared is the solution rate across the three conditions.
The difficulty was that individual studies disagreed. Some found an effect and some did not. Sio and Ormerod (2009) resolved the disagreement by meta-analysis, pooling the existing results and separating out the moderators. Three findings emerged:
- The incubation effect is positive. Across the pooled studies, a break raises the solution rate.
- The effect size is not uniform. Divergent thinking tasks benefit more than linguistic and visual insight tasks.
- The length of the preparation phase is the main moderator. The longer someone has struggled with the problem before the break, the stronger the incubation effect. This is the practically important result.
That third finding overturns the folk reading of the stage. In the popular account, incubation is a kind of letting go: stop working and the idea will arrive. What the data say is close to the opposite. Incubation is the yield on the work that preceded it. Stepping away from a problem you have not yet engaged with does nothing, because there is nothing running in the background.
3.4 The moment: insight has a prior state
The last question in this layer concerns what is ordinarily called the "aha" moment.
Kounios and Beeman (2014) review a long research programme in which participants solve problems and report, for each one, whether they solved it by sudden insight or by step-by-step analysis. That self-report splits the data in two, and the two sets are then compared.
Two findings are reported. At the moment the solution becomes available in the insight condition, there is a burst of gamma-band activity, around 40 Hz, over the right anterior superior temporal gyrus, a region associated with understanding metaphor and jokes and grasping the gist of a conversation. And roughly 1.5 seconds before that, alpha power rises sharply over the right occipital cortex, which is usually read as a temporary reduction of visual input so that attention can turn inward.
A third finding from the same programme matters more here. The state of the brain before the problem is presented, in the rest interval between problems, partly predicts whether the next problem will be solved by insight or by analysis.
The implication is that insight is not an event that arrives from outside. It has a prior state, and states, unlike events, respond to conditions.
3.5 What this layer says
The four sections above compose one picture. The material is the structure of semantic memory. The operations are two, and their order has a measurable effect. The interval between them helps, but only if work has been done first. And the moment of solution has a prior state.
4. The neural layer
The question here is different. It is not what operations run, but what they run on, and whether the underlying architecture explains why the work is hard.
4.1 What has to be set aside
Attributing creativity to the right hemisphere has no support in neuroimaging, and the sorting of people into "right-brained" and "left-brained" types has been rejected in large-sample studies. What is observed during creative tasks is widely distributed activity across both hemispheres, not the dominance of one side.
4.2 Three networks, and what "antagonistic" actually means
The current literature is organised around three large-scale networks:
- The default mode network. Active when attention is not on an external task. Associated with spontaneous association, mind-wandering and autobiographical recall.
- The executive control network. Associated with goal-directed attention, evaluation, and holding information in working memory.
- The salience network. Switches between the other two, determining which has control at a given moment.
The phrase "these two networks are antagonistic" is repeated constantly in general writing, usually without its meaning being given. What it means precisely is that in imaging data the activity of the default and executive networks is negatively correlated: when one rises, the other falls. The pattern recurs at rest and across many tasks.
The measure should also be clear. What these studies call functional connectivity is the correlation of the signal time series between regions. If the signal in two regions rises and falls together over time, the two are said to be coupled. The measure reports co-variation, not that one drove the other.
4.3 The findings
Beaty and colleagues (2015) scanned participants performing a divergent thinking task. The standard task in this field is the alternate uses task: a person is given an ordinary object, a brick for instance, and asked to produce as many unusual uses for it as they can in a fixed time. Responses are later scored for fluency, flexibility and originality.
What they found was that creative idea production is accompanied by coupling between the default and executive networks. Given the previous section, the significance is clear: what has been observed is a state that is not normally observed.
Temporal analysis of that coupling shows a pattern consistent with the model in section 3. Early in the task, coupling is mainly between the default and salience networks; later, between the default and executive networks. Something resembling the priority of generation over evaluation is traceable in the time series as well.
Beaty and colleagues (2018) went further, using a method whose value lies in its design. In connectome-based predictive modelling, the connectivity patterns and creativity scores of one group of participants are used to build a model, and the model is then tested on participants who played no part in building it. If it can predict the scores of the held-out participants, what was found is not a chance fit to the original data.
With 163 participants, a network associated with high creative ability was identified, composed of frontal and parietal regions across all three default, salience and executive systems. In the same study, the correlation between creative thinking ability and self-reported creative behaviour and achievement in the arts and sciences was 0.54, which counts as substantial in individual-differences research.
4.4 What this layer does not establish
The permitted reading of this section is narrower than what is usually taken from it.
These findings are correlational. They do not show that network coupling causes creativity; both could be effects of a third factor. Reverse inference is also not permitted: from the activity of a network one cannot conclude that a person is engaged in a particular cognitive process, because each network is active across many tasks.
What can be said with reasonable confidence is this: the experienced difficulty of generating and evaluating at the same time is not merely a psychological figure of speech, and it is not inconsistent with the network architecture of the brain. That is the limit of the claim.
5. The contextual layer
If the underlying mechanism is shared, the visible difference in output needs explaining. The third layer fills that gap.
Amabile's (2012) componential theory treats creativity as the product of three components:
- Domain-relevant skills: knowledge and expertise in the field.
- Creativity-relevant skills: working style and habits of thought.
- Task motivation, and in particular intrinsic motivation.
The intrinsic motivation principle holds that people are most creative when the primary driver is the interest, enjoyment and challenge of the work itself rather than an external reward.
The first component agrees with section 3.1: the range of association depends on the contents of semantic memory, and expertise is those contents. Two separate literatures, arriving from different directions at the same place.
The second factor with specific experimental support is constraint, and its result contradicts the common intuition.
The design. Haught-Tromp (2017) asked participants to write short two-line rhymes for greeting cards. One group received only a topic. The other received a topic plus a mandatory word that had to appear in the text, a constraint that makes the task logically harder. Independent judges rated the outputs.
The result. The constrained group produced more original work. The hypothesis is named after the wager in which a writer agreed to produce a book using only fifty permitted words, and the result became one of the best-selling children's books ever published.
The proposed explanation is that a constraint does not shrink the search space so much as move it: it closes the familiar routes and sends the mind into regions of the semantic network it would not otherwise have visited.
The third factor, preparation, was covered in section 3.3.
6. What the three layers agree on
The three layers came at the problem by entirely different means: behavioural analysis in the laboratory, measurement of brain signal, and the study of organisational and motivational context. That they converge lends the conclusion weight.
The cognitive layer says generation and exploration are two operations with different logics, that their order has a measurable effect, and that a preinventive structure needs its incompleteness in order to work. The neural layer says the corresponding networks are normally negatively correlated and that their order of engagement is traceable in the time series. The contextual layer says output depends on the contents of memory, on motivation, on constraint, and on the length of preparation.
The shared conclusion is that separating generation from evaluation in time is a structural requirement of the process. Premature evaluation fixes the preinventive structure before exploration can act on it, and the process halts at a stage it has not yet worked through. Finke's experiment measured exactly this.
Against that picture, the trait explanation is weak. Not because individual differences are absent, but because no variable by that name appears in any of the three layers. What appears instead is the structure of semantic memory, the pattern of network coupling, domain knowledge, motivation, and working conditions. Some of these change more easily than others, but none of them is a fixed trait.
7. Limits of the evidence
This section cannot be dropped, because without it the conclusion in section 6 goes past its evidence.
First, the field's measuring instrument has weak predictive validity. Divergent thinking tests, including the alternate uses task described in section 4.3, have failed across dozens of studies to predict real-world creativity. Zeng, Proctor and Salvendy (2011) enumerate six weaknesses:
- inadequate construct validity
- failure to assess the integrated creative process
- neglect of domain specificity
- neglect of the role of expertise
- weak predictive and ecological validity
- weak discriminant validity
Further, the correlation of these tests with creative performance is weaker over the long term than it is concurrently.
And that objection brings a sharper one with it. If the alternate uses task is a poor measure of real creativity, then what did the neural studies in section 4 measure? At best the neural correlate of performance on a laboratory task, not the neural correlate of creativity. The objection applies to much of that literature and should not be stepped over.
The partial answer offered to it is in the 2018 study itself: the correlation of 0.54 between divergent thinking ability and self-reported creative achievement in the arts and sciences. The task does relate to something outside the laboratory. It is not a complete answer, because self-report carries error of its own. The fair formulation is that the task bears some relation to real creativity, the relation is not strong, and claims in this field should stay correspondingly cautious.
Third, the definition is itself contested. The effectiveness condition depends on external judgement, and that judgement varies with context and period. Some of what counts as creative now would not have counted at another time. This limitation is intrinsic to the subject and no improvement in method removes it.
The practical upshot for the non-specialist reader is plain.
No test can tell you that you are creative. And by the same token, no test can tell you that you are not.
The limitation leaves both claims unsupported, though in practice it works mostly against the second, since the second is the one usually accepted without any test at all, on the strength of a memory.
8. A theoretical proposal: the sketch as externalised preinventive structure
Section 3.2 reached a point it did not open. Two literatures that make no reference to one another have arrived at the same result, and the shape of the overlap is exact enough to make the question of their relation a fair one.
On one side, creative cognition. Finke and colleagues (1992) hold that the generative phase produces structures that as yet have no meaning or function, and that this temporary meaninglessness is what allows the exploratory phase to read them in more than one way. A structure that acquires meaning too early passes out of exploration's reach.
On the other side, design cognition. Goel (1995) holds that ambiguity in the symbol system preserves reinterpretation in the early stages of designing, and Goldschmidt (1991) holds that a designer reads things from their own sketch that they had not thought of before drawing it.
The proposal is that these describe one thing from two places. The correspondence:
| In creative cognition | In design cognition |
|---|---|
| Preinventive structure, unfinished, no assigned function | Early sketch, ambiguous, undetailed |
| Incompleteness is a condition of its working | Ambiguity is a feature, not a defect |
| The exploratory phase interprets the structure | The designer reads the sketch and infers from it |
| Premature fixing closes exploration | Premature precise notation closes reinterpretation |
If the correspondence holds, a sketch is not a record of a preinventive structure. It is that structure, outside the head.
What the proposal does that neither literature does alone. Neither one explains why externalising works. Geneplore is silent about where a preinventive structure resides and tacitly treats it as mental. The design literature documents the effect of sketching but does not place it within a framework of creative cognition. The correspondence above fills both gaps: externalising works because it fixes incompleteness. A mental representation is rewritten on each recall and drifts toward coherence; an external one stays as unfinished as it was, and so remains open to reinterpretation on later passes.
Testable predictions. A theoretical proposal is worth what its testability is worth. At least three predictions follow:
- The ordering effect should reappear in drawing. If the functional category is announced before drawing, outputs should be less original than when the category is announced after. This is Finke's design with an external rather than a mental representation.
- Degree of ambiguity should moderate the effect. The more precise and less ambiguous the external representation, the lower the yield of the exploratory phase. In early stages, a precise tool should produce less original output than a rough one. This connects to an existing question in the design literature about the differing effects of freehand sketching and precise tools.
- The incubation effect should be stronger in the presence of an external representation. If incompleteness has been fixed outside the head, returning to the problem after a break should be more productive than when the person relies on mental recall alone.
The status of the proposal. This mapping was not found in the sources reviewed for this article and is offered here as a hypothesis, not a finding. Before any claim of priority, a systematic review of the design cognition literature is needed; it may have been stated earlier under different terminology.
9. Implications
Four practical implications follow, each resting on specific evidence:
- Separate generation from evaluation in time (sections 3.2, 4.3 and 6). Evaluation is necessary and without it work does not reach quality; what the evidence addresses is its place in the sequence, not its removal.
- Incubate after engagement, not instead of it (section 3.3). Stepping away works when enough work has preceded it.
- Add a constraint rather than remove one (section 5). The advice contradicts the common intuition but has experimental support.
- Externalise the unfinished structure (sections 3.2 and 8). If incompleteness is a condition of the structure's working, an external representation preserves it, while a mental representation is continually rewritten and loses it. Unlike the three above, this implication rests on a theoretical proposal rather than an established finding.
10. Conclusion
What is known today about the mechanism of creativity does not describe a single unified ability.
It describes a two-operation process whose raw material is the structure of semantic memory, whose order of operations has a measurable effect, which rests on neural networks that are not normally active together, and which is sensitive to external conditions including constraint, motivation and the length of preparation.
None of these components is a fixed trait. At the same time, what we hold is not enough for large claims: the field's measuring instrument has weak predictive validity and the neural evidence is correlational.
The cautious and defensible summary is this. Creativity cannot be measured well enough to be assigned to people. The conditions under which the process works better or worse are reasonably well known, and largely changeable.
References
Amabile, T. M. (2012). Componential theory of creativity (Working Paper No. 12-096). Harvard Business School.
Beaty, R. E., Benedek, M., Kaufman, S. B., & Silvia, P. J. (2015). Default and executive network coupling supports creative idea production. Scientific Reports, 5, 10964. https://doi.org/10.1038/srep10964
Beaty, R. E., Kenett, Y. N., Christensen, A. P., Rosenberg, M. D., Benedek, M., Chen, Q., Fink, A., Qiu, J., Kwapil, T. R., Kane, M. J., & Silvia, P. J. (2018). Robust prediction of individual creative ability from brain functional connectivity. Proceedings of the National Academy of Sciences, 115(5), 1087–1092. https://doi.org/10.1073/pnas.1713532115
Finke, R. A. (1990). Creative imagery: Discoveries and inventions in visualization. Lawrence Erlbaum Associates.
Finke, R. A., Ward, T. B., & Smith, S. M. (1992). Creative cognition: Theory, research, and applications. MIT Press.
Goel, V. (1995). Sketches of thought. MIT Press.
Goldschmidt, G. (1991). The dialectics of sketching. Creativity Research Journal, 4(2), 123–143. https://doi.org/10.1080/10400419109534381
Guilford, J. P. (1950). Creativity. American Psychologist, 5, 444–454. https://doi.org/10.1037/h0063487
Haught-Tromp, C. (2017). The Green Eggs and Ham hypothesis: How constraints facilitate creativity. Psychology of Aesthetics, Creativity, and the Arts, 11(1). https://doi.org/10.1037/aca0000061
Kenett, Y. N., Anaki, D., & Faust, M. (2014). Investigating the structure of semantic networks in low and high creative persons. Frontiers in Human Neuroscience, 8, 407. https://doi.org/10.3389/fnhum.2014.00407
Kounios, J., & Beeman, M. (2014). The cognitive neuroscience of insight. Annual Review of Psychology, 65, 71–93. https://doi.org/10.1146/annurev-psych-010213-115154
Mednick, S. A. (1962). The associative basis of the creative process. Psychological Review, 69(3), 220–232.
Runco, M. A., & Jaeger, G. J. (2012). The standard definition of creativity. Creativity Research Journal, 24(1), 92–96. https://doi.org/10.1080/10400419.2012.650092
Sio, U. N., & Ormerod, T. C. (2009). Does incubation enhance problem solving? A meta-analytic review. Psychological Bulletin, 135(1), 94–120. https://doi.org/10.1037/a0014212
Wallas, G. (1926). The art of thought. Jonathan Cape.
Zeng, L., Proctor, R. W., & Salvendy, G. (2011). Can traditional divergent thinking tests be trusted in measuring and predicting real-world creativity? Creativity Research Journal, 23(1), 24–37. https://doi.org/10.1080/10400419.2011.545713