Bayesian Network¶
Represent a joint probability law by a directed acyclic graph and one conditional distribution per variable given its parents.
Core Idea¶
A Bayesian network represents a joint probability law over random variables by pairing a directed acyclic graph (DAG) with a conditional probability distribution for each variable given its graphical parents. If the nodes are \(X_1,\ldots,X_n\) and \(\mathrm{pa}(X_i)\) denotes the parents of \(X_i\), the represented joint law obeys \(P(x_1,\ldots,x_n)=\prod_i P(x_i\mid\mathrm{pa}(X_i))\). The graph is not a decorative diagram: its parent relation determines the local factors, and its Markov semantics states which conditional independences are entailed by that factorization.[1]
One model can be queried in several directions. Observed findings may update beliefs about diseases; sensor readings may update beliefs about component faults. Posterior inference is a use of the represented law, not a particular algorithm required for the representation to be a Bayesian network. Exact elimination, simulation, or compiled circuits can answer queries with different cost and approximation properties.[2][3]
Structural Signature¶
Sig role-phrases:
- Random-variable family — each node denotes a specified uncertain quantity and a value domain. Merely naming concepts as vertices does not specify a probability model.
- Acyclic directed parent structure — each variable has a parent set, and the graph has no directed cycle. Direction organizes conditional factorization; it does not automatically establish physical causation.[1]
- Local conditional laws — each node carries a normalized law \(P(X_i\mid\mathrm{pa}(X_i))\), expressed as a table, density, or other valid conditional family. A bare graph without these laws leaves the joint unspecified.[1][3]
- Global product rule — multiplying all local laws gives the represented joint distribution. If a proposed distribution does not factor this way, this DAG is not its Bayesian-network representation.[1]
- Graphical independence semantics — d-separation identifies conditional independences guaranteed by the DAG factorization. It does not say every open path forces dependence for every numerical parameterization.[1][4]
Evidence and query nodes are common ways to use a network, but a model remains a Bayesian network before an inference engine is selected or an observation entered.
What It Is Not¶
It is not every probabilistic graphical model. The live umbrella also includes undirected Markov networks and factor graphs with different factor and separation rules. A Bayesian network specifically uses directed acyclic parenthood and one conditional law per node. Nor is it a decision graph or causal diagram merely because arrows are drawn: ordinary probabilistic factorization does not by itself authorize intervention claims.[1]
It is not d-separation. D-separation is a graphical criterion applied to a DAG to test implied conditional independences; it is neither a synonym for the whole model nor a fitted inference algorithm. The frozen Wikipedia selection includes “D-separation” as a redirect to the Bayesian-network page, but that routing fact does not collapse the two technical identities.[4]
Scope of Application¶
The identity applies wherever a joint probability model can be expressed as parent-local conditionals on a DAG: expert systems, diagnostic models, uncertain sensor systems, prediction, and other probabilistic reasoning. Variables may be observed or hidden; local laws may be tabular or parametric, provided they define a coherent joint law.[1]
Shwe and Cooper's QMR-DT work describes a two-level, multiply connected medical belief network for disease and finding variables. Its complexity made most exact posterior-marginal algorithms impractical for that model, motivating likelihood-weighting approximation. This is a research model, not clinical advice or proof of reliability outside the study.[2]
NASA's ADAPT electrical-power testbed represented faults probabilistically with a Bayesian network and used observations to isolate component and sensor failures. Arithmetic circuits compiled from the network served the implementation's speed needs; compilation is not part of what makes the underlying model a Bayesian network.[3]
Clarity¶
The entry separates three claims. The graph and local laws define a factorized joint. D-separation reads independences the graph guarantees under that semantics. A posterior query conditions the joint on evidence. Confusing these layers turns a modeling assumption into an observed fact or a computational method into the model's identity.[1]
Graphical nonseparation has a subtler meaning than many summaries suggest. A d-connected pair may be dependent under compatible parameters, but a particular fitted distribution can have extra independences that the graph does not entail. The CMU source explicitly gives such a parameterized counterexample. An absent graph implication is not proof of dependence.[1]
Manages Complexity¶
Without exploitable conditional independences, a joint table over \(n\) binary variables has \(2^n-1\) free probabilities. Local factors can be much smaller when each node has few parents. The gain is conditional: a nearly complete DAG can reproduce the full chain-rule burden, and sparse representation does not guarantee cheap Inference.[1]
Shwe and Cooper's QMR-DT network illustrates the distinction. A large, multiply connected network can have a local description while exact disease-posterior computation remains difficult. Factorization manages representation; an inference procedure must still handle connectivity and the query.[2]
Abstract Reasoning¶
Define variables and the DAG, then specify a local conditional law for each node. The product rule reconstructs the joint. To assess independence, test whether the conditioning set d-separates the queried node sets. If it does, independence is implied for every distribution obeying that graph's factorization; if it does not, inspect actual parameters rather than automatically asserting dependence.[1][4]
For a three-node fork \(A\leftarrow B\rightarrow C\), the factorization \(P(B)P(A\mid B)P(C\mid B)\) entails \(A\perp C\mid B\). A collider \(A\rightarrow B\leftarrow C\) has \(P(A)P(C)P(B\mid A,C)\); conditioning on \(B\) can open a path between \(A\) and \(C\). Arrow configuration and conditioned evidence matter more than visual distance between nodes.[1]
Knowledge Transfer¶
The complete pattern transfers literally between medical and electrical-system models: replace diseases/findings with components/sensors while retaining variable nodes, directed parent sets, local laws, joint factorization, and conditional queries. Probabilities and graph must be justified afresh. A medical conditional table cannot be imported into a power system merely because both are Bayesian networks.
Dependency, factorization, and representation transfer further, but “Bayesian network” remains a probabilistic-model term. A software-dependency graph without random variables and conditional distributions is an analogy, not another instance.
Examples¶
QMR-DT medical belief network¶
Shwe and Cooper describe a two-level multiply connected network built from Quick Medical Reference knowledge and evaluate approximate disease-posterior computation on difficult diagnostic cases.[2] Mapped back: disease and finding quantities are random variables; the two-level directed connections supply the DAG; local probabilities supply conditional laws; their product defines the joint disease–finding model; the graph entails only its d-separation independences; observed findings are evidence for posterior-disease queries. Likelihood weighting is a query strategy, not a defining component.
NASA ADAPT electrical-power diagnosis¶
NASA's ADAPT work constructs Bayesian networks for an electrical-power system and uses observations for fault isolation; the NASA report describes discrete component and sensor variables with parent-conditioned tables.[3] Mapped back: component state, sensor state, and observations are variables; directed dependencies form a DAG; parent-conditioned tables specify local laws; their product is the system joint model; graphical separation expresses model-implied independences; observed readings support fault-state queries. Compiled circuits accelerate inference but do not replace the network's identity.
Structural Tensions¶
Compact locality versus missing dependence. Fewer parents make local tables smaller but assert more conditional independences. Adding every plausible link may avoid an omission at the cost of lost compression. Diagnostic: Which omitted dependencies have evidence behind them, and how do table sizes change with a proposed edge?
Graph entailment versus parameter coincidence. D-separation gives independence guarantees across the graph's factorizing family; particular local probabilities can create extra independences. Diagnostic: Is the relation forced by structure, or found only in one parameterization?[1]
Exact answer versus feasible inference. The factorized law may be exact while a desired posterior is expensive. Diagnostic: Has the query's computational cost been assessed, and if approximation is used, what convergence evidence applies?[2][3]
Probabilistic direction versus causal intervention. An arrow organizes conditional distributions, but a causal effect claim asks what happens under manipulation. Diagnostic: What additional assumptions make a graph edge causally interpretable rather than merely probabilistically useful?
Structural–Framed Character¶
Bayesian Network is mixed-structural, leaning formal: directed acyclicity and local conditional factorization have exact consequences, while a modeler selects variables, edges and conditional distributions. Its evaluative weight is not positive by definition; a coherent network can fit observations poorly or be misread as causal. It is partly human-practice-bound as a probabilistic model constructed for a question, though its implied conditional-independence statements can be tested independently of the modeler's wishes. Its institutional origin is probabilistic graphical modeling research rather than a governing convention that makes those dependencies true. Its vocabulary travel reaches medicine, diagnosis and engineering wherever the same random-variable/DAG factorization is specified; ordinary arrows on a workflow chart do not carry probabilistic semantics. Import versus recognition requires a normalized joint law represented by local parent-conditioned factors, not just a graph that looks acyclic.
Live Directed Acyclic Graph supplies a portable ordering/dependency skeleton, but is only a component, not the asserted genus. The proposed strict parent is the live domain-specific Probabilistic Graphical Model; the child's distinctive semantics live in probability. Its character: a formal probability model whose graphical structure is broadly recognizable but whose identity requires random variables, conditionals and factorization.
Structural Core vs. Domain Accent¶
The graph can be lifted, but the probability law cannot be dropped without changing the object.
What is skeletal. A directed acyclic graph organizes local dependencies so a larger structure can be reasoned about through constrained parent relations. Live Directed Acyclic Graph carries that portable component across domains. The asserted Probabilistic Graphical Model parent supplies the narrower family of joint distributions organized by graphs; it is domain-specific, not a substitute for a cross-domain prime.
What is domain-bound. Nodes must represent random variables, directed edges define parent sets, local conditional laws are normalized, and their product represents a joint probability law with the corresponding Markov/separation interpretation. Remove the probabilities and a DAG of software modules is not a Bayesian network. Variable selection, discrete versus continuous states, parameter estimation and predictive versus causal use vary. A causal reading needs additional assumptions; acyclicity alone does not confer intervention semantics.
Why this is not a prime. DAG structure travels broadly under its live prime. Bayesian networks are recognized across unlike applications only when the probabilistic factorization and independence rules genuinely hold. A generic dependency diagram imports the familiar visual language while missing the stochastic model, so the child's additional semantics do not inherit the graph prime's full breadth.
Instantiates / Related Primes¶
No strict typed parent relation is asserted in the current DAG. Independently reviewed without a defensible necessary parent selected in the current catalog; admitted unparented pending later DAG densification.
Neighborhood in Abstraction Space¶
Bayesian Network sits in a sparse region of the domain-specific corpus (76th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Formal Models & Logical Foundations (33 abstractions)
Nearest neighbors
- Bayesian Programming — 0.84
- Chow–Liu Tree — 0.83
- Exponentially Modified Gaussian Distribution — 0.83
- Gram Matrix — 0.83
- Equicontinuity — 0.83
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- An undirected Markov network or factor graph, whose factorization and separation syntax differ.
- A causal DAG without specified probabilities, or an ordinary Bayesian network falsely assumed causal.
- D-separation, the criterion for graph-implied independence rather than the entire model.
- An inference algorithm, which operates on the model but does not define it.
- A sparse drawing assumed to make every posterior query computationally easy.
References¶
[1] Carnegie Mellon University, “10-708 PGM, Lecture 2: Bayesian Networks”, “Bayesian Network: Factorization Theorem,” “I-maps,” “Local Markov assumptions,” and “D-separation criterion.” Defines parent-conditional factorization, graphical independence, and parameter-specific extra independence. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m
[2] Michael Shwe and Gregory F. Cooper, “An Empirical Analysis of Likelihood-Weighting Simulation on a Large, Multiply-Connected Belief Network”, Proceedings of the Sixth Conference on Uncertainty in Artificial Intelligence (UAI 1990), 498–508, §§2–3. Describes QMR-DT's disease/finding nodes, local laws, and inference study; this is the conference paper, distinct from the authors’ 1991 journal article. registry ↩a ↩b ↩c ↩d ↩e
[3] Ole J. Mengshoel, Mark Chavira, Keith Cascio, Scott Poll, Adnan Darwiche, and Serdar Uckun, Efficient Probabilistic Diagnostics for Electrical Power Systems, NASA/TM-2008-214589 (2008), abstract and §2.1. Official ADAPT case and Bayesian-network semantics. registry ↩a ↩b ↩c ↩d ↩e
[4] Dan Geiger, Tom S. Verma, and Judea Pearl, “d-Separation: From Theorems to Algorithms”, author-posted manuscript. Establishes the graphical criterion for network-implied independence. registry ↩a ↩b ↩c