Skip to content

Three-Factor Learning

A synaptic learning-rule structure in which presynaptic activity, postsynaptic state, and an additional modulatory signal jointly govern plasticity.

Version
v1 · 2026-10-04 · History
Domain-specific #
13777
Domain group
Natural Sciences
Origin domain
Neuroscience
Subdomain
Computational Neuroscience → Neuroscience
Aliases
Three-factor learning rule, Neo-Hebbian three-factor rule

Core Idea

A three-factor learning rule changes synaptic efficacy through the joint influence of (1) a presynaptic activity variable, (2) a postsynaptic activity or state variable, and (3) an additional modulatory signal. The first two describe information available at the connection; the third is extrinsic to that local pair. A rule may multiply a pre/post term by a modulator, let the modulator alter the postsynaptic state, or combine the three inputs another way. In theoretical treatments the modulator may encode reward, a prediction error, novelty or another signal; in one biological setting it can be carried by dopamine. The identity is operative three-input dependence, not one factorization, molecule or reward task.[1]

For a delayed third signal, a transient eligibility trace can retain a synapse-specific mark after the fast pre/post event. A schematic rule is \(\Delta w_{ij}\propto e_{ij}M\), where \(e_{ij}\) summarizes recent local activity at synapse \(ij\) and \(M\) is the later modulation. The equation is an illustrative family form, not a universal molecular law. Some three-factor rules can act with little or no delay and need not use an explicit decaying trace; others use different mathematical combinations.[1][2]

In delayed trace-based variants, a later broad signal can act selectively on connections that were recently eligible, addressing a part of temporal credit assignment. The three-factor identity alone does not prove that the signal identifies the correct synapses in every task, completely solves distal reward, or automatically stabilizes learning. Those results depend on the particular update rule, eligibility window when used, modulation semantics, network dynamics and objective.[2]

Structural Signature

Sig role-phrases: presynaptic variable; postsynaptic variable; distinct extrinsic modulator; joint three-input efficacy update; optional eligibility trace.

  1. Presynaptic activity: a signal available at the connection from an upstream neuron.
  2. Postsynaptic state: a downstream activity or response variable that contributes to the update alongside presynaptic information and the extrinsic modulator.
  3. Local information: the pre/post variables supply connection-specific information to the update. They need not first form an independent plasticity tendency or eligibility mark; plain correlation alone does not specify a three-factor rule.[1]
  4. Third-factor modulation: an additional extrinsic signal affects the update, whether by gating a local term, altering a postsynaptic variable, or another specified dependence. Its spatial reach and semantic content vary by model or circuit.[1]
  5. Eligibility trace, in delayed-trace variants: a temporary memory of recent local activity that can outlast a brief spike pairing until modulation arrives.
  6. Weight or efficacy change: the eventual synaptic alteration; the rule must specify how all three inputs jointly affect it, not merely measure three co-occurring quantities.

Condensed: synaptic update depending on pre-activity, post-state and an extrinsic modulator, optionally linked across time by local eligibility = three-factor synaptic learning.

What It Is Not

  • Not ordinary two-factor Hebbian learning. A rule driven only by the pre/post pair lacks the defining third modulation.
  • Not identical to spike-timing-dependent plasticity (STDP). STDP supplies one possible local timing rule; it becomes an instance of this family only when an additional factor modulates the change.[1]
  • Not necessarily dopamine or reward. These are important examples, while error, novelty or other signals can occupy the modulatory role.
  • Not necessarily a brain-wide broadcast. The third signal may have a wider reach than the local connection without being literally global or uniform.
  • Not always an explicit eligibility-trace rule. Traces are especially useful when the relevant signal is delayed; the three-factor definition itself is about joint dependence.
  • Not a universal cure for credit assignment or instability. Gating can constrain updates, but its success depends on the signal, trace timing and network design.
  • Not simply three pieces of data. All three factors must take operative roles in the plasticity rule.

Scope of Application

In theoretical synaptic plasticity, the three-factor framework collects rules that couple local pre/post information to an additional contextual or feedback variable. Frémaux and Gerstner distinguish pre- and postsynaptic factors from the modulator and discuss eligibility traces as a way to span reward delay. This is a rule family, not one mandated equation.[1]

In spiking-network learning models, Izhikevich's original dopamine-modulated STDP model offers a concrete delayed-reward example. Rapid spike-timing information is retained in a slower local plasticity state; a subsequent dopaminergic signal selectively affects recently involved connections in that model. The model shows how such credit assignment can work under its assumptions, not that all brains implement the same algorithm.[2]

In striatal experimental neuroscience, Yagishita and colleagues separately optically stimulated glutamatergic and dopaminergic inputs. Dopamine promoted dendritic-spine enlargement only when delivered 0.3 to 2 seconds after glutamatergic input. This is source-attested structural plasticity in the tested preparation, not a measured behavioral reward-learning rule or a universal eligibility window.[3][4]

Clarity

Consider a delayed, trace-based instance: several connections are active while an action is chosen, and a later favorable outcome produces a modulatory signal. If every connection changed solely because that signal arrived, there would be no local credit assignment. Here pre/post activity creates different local eligibility states; the later signal acts through those states. This illustrates one implementation of three-input dependence, not a requirement that every three-factor rule first create a mark or that the signal be perfectly informative.

Without delay, one can imagine all three variables being available together. With a delay, the local event must leave a memory long enough to overlap with the modulator. A trace that decays too quickly misses the outcome; a trace that persists indiscriminately may also credit irrelevant activity. The timing is therefore a substantive model choice, not a decorative detail.

Manages Complexity

The rule distinguishes connection-specific information from an extrinsic influence on plasticity. In one common delayed implementation, synapses register pre/post events locally while a modulatory system carries later feedback without sending a distinct error message to every connection. Other rules combine the same three inputs differently. This distinction gives a compact vocabulary for comparing biological experiments and computational models; it does not imply that one scalar encodes every relevant cause of an outcome.[1][2]

Abstract Reasoning

To test whether a proposed mechanism is three-factor learning, write down the local pre-variable, local post-variable and distinct extrinsic modulator. Show how each causally enters the efficacy-update rule; do not assume the function factors into a pre/post term multiplied by a gate. If the signal arrives later, test whether and how a temporary local state preserves relevance, and for how long. Then ask what the third signal means in that case: reward, error, novelty or something else. Finally separate a formal rule from evidence that a particular circuit realizes it.[1][3]

The diagnostic question is: Which pre- and postsynaptic variables and which extrinsic modulator jointly determine this synapse's update, and through what functional dependence?

Knowledge Transfer

The broader pattern is a local update influenced by an additional contextual variable. Marking local candidates and later authorizing them is one delayed subtype, not the definition of every three-factor rule. The domain-specific form concerns synaptic efficacy, pre/post neural activity and neuromodulatory or algorithmic plasticity variables. It should not be flattened into generic reinforcement learning, which can use very different credit-assignment mechanisms.

Examples

Delayed reward in a spiking-network model

Izhikevich's original spiking-network model links millisecond-scale spike-timing-dependent plasticity to a later dopamine-like reward signal. The author-hosted abstract specifies that slow subsequent synaptic plasticity stays dopamine-sensitive for a few seconds and that random firing during the wait does not affect the earlier STDP mark in that model. The modeled network thus gives the broad dopamine signal selective leverage over recently marked connections; it does not prove a universal biological solution to delayed reward.[2]

Mapped back: pre/post factors = near-coincident spikes at a particular connection; eligibility = model's slow local plasticity state; third factor = extracellular dopamine change after reward; outcome = selective modeled synaptic update over a few-second critical period; limit = model-specific rather than a brain-wide law.

Dopamine timing at striatal spines

Yagishita and colleagues' original experiment separately stimulated glutamatergic and dopamine inputs and found dopamine promoted spine enlargement only 0.3–2 seconds after the local glutamatergic input. The observation links prior local activation and subsequent modulation for this structural endpoint; it does not identify a universal reward-prediction error or show how a whole animal assigns behavioral credit.[3][4]

Mapped back: local factors = prior glutamatergic activation and postsynaptic spine response in tested cells; eligibility-like state = transient post-input susceptibility inferred from timing; third factor = separately stimulated dopamine input; observed output = spine enlargement only in the 0.3–2-second window; limit = structural plasticity, not directly measured learned behavior.

Ungated STDP near miss

A rule updates a synapse only from the temporal order of its presynaptic and postsynaptic spikes. It has two locally determined ingredients. Unless an additional contextual signal influences the update, it remains STDP, not a three-factor rule.

Mapped back: the absent third factor is identity-changing.

Structural Tensions

Short eligibility specificity versus delayed-reward reach. A short-lived local mark limits how many irrelevant intervening events remain eligible when a broad modulator arrives, but may have decayed before a late reward; a longer-lived mark reaches more delayed outcomes but widens the set of potentially credited local events. Izhikevich's few-second modeled dopamine-sensitive period and Yagishita's 0.3–2-second observed structural window are different settings and cannot be equated. Diagnostic: does the trace last long enough for the task's feedback delay without making unrelated synaptic events eligible?[2][4]

Structural–Framed Character

Three-factor learning is structural in the operative pre/post/modulator role conjunction, but framed in what the modulator means and whether a particular circuit instantiates the mathematical rule. Its evaluative weight differs for Izhikevich's delayed-reward model and Yagishita's spine-enlargement experiment: the first demonstrates a computational credit mechanism under assumptions, the second measures a bounded structural timing effect, not behavior. Human modeling practice chooses a trace and objective, while electrophysiology and imaging laboratories supply distinct empirical institutions and endpoints. The vocabulary travels literally from a dopamine-gated STDP model to another plasticity system only if all three causal factors govern a connection's update; calling any reward-related neural activity “three-factor learning” imports the label without that rule. Its character: a family of locally informed, context-modulated synaptic update rules with model- and circuit-dependent timing and feedback semantics.[1][2][3]

Structural Core vs. Domain Accent

The skeletal relation is an update depending on local variables and an additional contextual variable; that portable pattern belongs, if anywhere, to a separately justified learning or feedback prime. The domain-bound rule changes synaptic efficacy through pre- and postsynaptic variables plus a distinct third modulator, sometimes mediated by an eligibility trace. The named entry fails the prime bar because supervised gradient updates, generic reinforcement learning and ordinary feedback can shape changes without synapses or this three-factor dependence. The broad skeleton may travel, but the synaptic pre/post/modulator identity and any particular biological timing do not automatically travel with it.

  • Hebbian Learning: supplies a pre/post association component; the third-factor family adds a distinct extrinsic modulator.
  • Feedback: the modulator may carry later outcome information.
  • Credit Assignment: eligibility can help connect a delayed outcome to earlier local activity.

These are candidate conceptual relations, not approved DAG edges.

Neighborhood in Abstraction Space

Three-Factor Learning sits in a sparse region of the domain-specific corpus (79th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Neuronal Signaling & Plasticity (14 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

Hebbian Learning is a related local co-activity principle whose current live identity excludes an additional steering signal; it is not a strict genus of the three-factor rule. STDP is a timing-dependent form of local synaptic plasticity, which may be unmodulated or modulated. Eligibility Trace is the temporary local mark often used by delayed three-factor variants; it is not the entire rule. Reinforcement Learning is a wider computational framework and can be implemented without this synaptic update structure.

References

[1] Frémaux and Gerstner, “Neuromodulated Spike-Timing-Dependent Plasticity, and Theory of Three-Factor Learning Rules” (2016). registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i

[2] Izhikevich, “Solving the Distal Reward Problem through Linkage of STDP and Dopamine Signaling” (2007), author-hosted original article summary. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g

[3] Yagishita et al., “A critical time window for dopamine actions on the structural plasticity of dendritic spines,” Science (2014), original-study abstract; full discussion not inspected in this review. DOI: 10.1126/science.1255514. registry ↩a ↩b ↩c ↩d

[4] Yagishita et al., PubMed original-study abstract, separately stimulated glutamatergic/dopaminergic inputs and 0.3–2-second spine-enlargement window. DOI: 10.1126/science.1255514. registry ↩a ↩b ↩c