The Giry monad (Giry 1980, following Lawvere 1962) is the monad on a category of suitable spaces which sends each suitable space to the space of suitable probability measures on .
This is one of the main examples of probability monads, and hence one of the main structures used in categorical probability.
The Giry monad is defined on the category of measurable spaces, , assigning to each measurable space the space of all probability measures on endowed with the -algebra generated by the set of all the evaluation maps
sending a probability measure to , where ranges over all the measurable sets of . The unit of the monad at , traditionally denoted , sends a point to the Dirac measure at , , while the monad-multiplication is defined by the natural transformation
given by
This makes the endofunctor into a monad, , and as such this is the Giry monad on measurable spaces, as originally defined by Lawvere 1962.
An alternative choice, convenient for analysis purposes, and introduced by Giry, is obtained by restricting the category of measurable spaces to the (full) subcategory which are those measurable spaces generated by Polish spaces, , which are separable metric spaces for which a complete metric exists. The morphisms of this category are continuous functions.
Write
for the endofunctor which sends a space, , to the space of probability measures on the Borel subsets of . is equipped with the weakest topology which makes the integration map continuous for any , a bounded, continuous, real function on .
There is a natural transformation
given by
This makes the endofunctor into a monad, and this is the Giry monad on Polish spaces.
The Kleisli morphisms of the Giry monad on Meas (and related subcategories) are Markov kernels. Therefore its Kleisli category is the category Stoch (=). It is one of the most important examples of a Markov category.
Let denote the category of algebras of the -monad which as as objects those measurable spaces for which there exists a -algebra which is an object in the Eilenberg-Moore category of the -monad. The morphisms of are those measurable functions such that constitutes an arrow in the Eilenberg-Moore category of the -monad.
If is any measurable space then the space of probability measures has a convex space structure defined pointwise: if is finite collection of probability measures on then, for every sequence with each such that , the affine sum , is also a probability measure, defined at the measurable set in by
Given any -algebra the base space has the structure of a convex space which makes the measurable function an affine (measurable) map. Moreover, morphisms of -algebras are also affine maps.
Given define the convex space structure on by
Because a -algebra must satisfy we have, for any finite sequence ,
where the last line makes use of the definition of the convex structure on . Thus, every -algebra is affine.
To prove that any map of -algebras is an affine map, we compute
Lemma 3.1 shows that any object in lies in both the categories and , where is the category of convex spaces. Similarly, any morphism in is also a morphism in both and . An example of such an object is which, as a measurable space, is the one-point compactification of the real-line with the Borel -algebra. As a convex space, has the natural convex space structures on the real-line extended by the point with the property that for all and all .
Let denote the category whose objects are convex spaces and, in addition, also posseses a measurable space structure such that all the algebraic operations of taking finite affine sums of elements, , where is the coordinate projection function, is a measurable function. Note these maps are always an affine function because of the convex space structure on product spaces. Moreover, we require that each object in must have enough affine measurable functions to coseparate the points of . The objects of are called measurable convex spaces. The morphisms of are affine measurable functions. Because is a coseparator in it follows that is a coseparator in .
The space with the pointwise convex space structure and measurable structure is a measurable convex space because the evaluation maps coseparate any two distinct probability measures in and the operation of taking affine sums is a measurable function because
Let . The operation of taking an affine sum
is a measurable function. This implies that all finite affine sum operators are measurable.
Let be a measurable set in . By the pointwise convex space structure of we have
where is the measurable affine sum on elements of the measurable convex space , and is an affine measurable function because is an affine measurable function. Since the right-hand side of the equation is measurable it follows, for every measurable set of , that
This last equation is true for all measurable sets in , and since the affine measurable functions generate the initial -algebra on it follows that the function is measurable.
The second statement follows from the observation that we can use induction on the number of components in a product space, and, assuming components, that we are free to choose parameters freely. Since every affine sum is uniquely defined by parameters the result follows.
Given any measurable space and any let denote the functional sending . The value is the expected value of the measurable function with respect to the measure . Note that the functional is (1) weakly averaging, and (2) linear. If we let denote the set of all weakly averaging linear functionals from the hom set to then we have a bijective correspondence between this set of all weakly averaging linear functionals and . The correspondence is and where the probability measure is defined on a measurable set by .
Let . Taking we obtain the space of affine measurable endomaps on .
An -generalized point of an object in is a functional satisfying, for all and all , the equation
which implies that is (1) weakly averaging, and (2) for all : . Moreover, just as in the identification of with the functional space consisting of all weakly averaging linear functionals , which uses the -linear vector space structure of the hom set and , we require the -generalized points to satisfy the linearity property that . (Generalized points are defined in Definition 8.19 of Sets for Mathematics, and several basic properties are discussed therein. Weakly averaging functions are also defined there.)
Note that if then the functional , which is the restriction of the functional to affine measurable functions, is an -generalized point of since, for all and all , , and the linearity property of .
Since we are restricting the operators to operate only on affine measurable functions, yielding the operators , we have
If is an -generalized element of then there exists a such that .
Let be an object in . We have the inclusion function which induces the restriction mapping
which is a surjective function. (The conditions on both functional spaces are identical: the elements are weakly averaging and linear.) This specifies an equivalence relation on the set defined by if and only if the restriction of those functionals are equal on the set of all affine measurable functions . Thus every -generalized point of comes from a functional , which in turn arises from the probability measure on .
We say an object in satisfies the fullness property if and only if for every the property
holds.
Define to be the full subcategory of consisting of those objects which satisfy the fullness property.
The space is an object in because, for every , we have . Trivially, the space is also an object in .
Let denote the full subcategory of consisting of the single object . The restricted Yoneda embedding functor defined (on objects) by
is a full and faithful functor. ( is the category of -linear vector spaces.)
In the category every affine measurable function is determined by its value on points . Hence to prove the fully faithful property it suffices to prove those properties on points.
Faithful: Note is the evaluation map . Let , for be two points of . If for all , then since has enough affine measurable maps to to coseparate points it follows that and is faithful functor.
Full: If is a natural transformation then, at the single component of , is a linear functional from the -linear space to the -linear space . The naturality condition requires for all and all . All told, is a weakly averaging linear functional i.e., is an -generalized point of . Now to complete the proof we employ Lemma 3.3: for some . Then, because , it satisfies the fullness property and we have, for all , the property that . So there exists an element such that . But the fullness property says . Thus there exist an such that for all from which it follows that and we conclude that is a full functor.
If is an object in then there exists a unique affine measurable function such that for all .
Let denote the inclusion functor. Let denote the slice category whose objects consist of affine measurable functions , and let denote the projection functor. For Theorem 3.4 is equivalent to saying with the projection map at component being . In other words, the inclusion functor is a codense functor. See Propositions 1 and 2, page 242 of CWM.
Consider the cone over with vertex and natural transformation components .
Since there exists a unique -morphism such that for all affine maps . It follows that for each Dirac measure that, for all in that . Since is a coseparator in it follows .
The defining property of is that it is the unique affine measurable function such that, for all , the property
holds. In the special case of or for a closed convex subset of it follows that for each fixed that . Henceforth we use the notation for the unique affine measurable function such that, for each fixed , the property
holds. In other words, by definition, is the unique point in such that the preceding equation holds.
If then the function is a G-algebra.
We need to show the following two properties: (1) for all we have , and (2).
The first property follows from the preceding corollary and remark concerning notation.
To prove the second property let and compose both sides of the required condition by an affine measurable map , and use the property that . The left-hand side of the required condition, after composition on the left hand side with , is given by
On the other hand, the right hand side of the required condition in (2), after composition with , yields
which coincides with the left hand side of the required condition. The result of the lemma now follows from the property that coseparates, i.e, the set of affine measurable functions are jointly monic on .
Let . Every affine measurable function yields a morphism of -algebras.
We have already noted, for every , that is the unique point in such that for all affine measurable functions . That equation is equivalent to the statement
But since both and are objects in it follows by Lemma 3.6 that both and are -algebras. Hence is a morphism of those algebras.
For every measurable space the space is an object in . For every measurable function the pushforward map is an affine measurable function.
By Lemma 3.2 and the paragraph preceding that lemma we know that is a measurable convex space. We need to show it also satisfies the fullness property.
First note that because we have, for every measurable set in , the property that
Since these evaluation maps are jointly monic on it follows that . Using this result, along with Lemma 3.7, we obtain the result that for and all affine maps the element which shows the fullness property is satisfied for . Hence .
The fact that is an affine function follows immediately by the pointwise convex space structure of and . The fact that it is a measurable function follows because is an endofunctor on .
By the preceding lemma we obtain the ‘’free functor’‘ , which is the Giry monad (functor) viewed as a functor into . There is also a ‘’forgetful functor’‘ which forgets the convex space structure. (We are using the terms free functor and forgetful functor before we actually show they have adjoints simply because we need to name them.)
Let . The affine measurable functions are the components of a natural transformation
Suppose and are objects in and that is an affine measurable function. We need to verify that . Applying both sides of this equation to any yields
which is a true statement. Hence the required commutativity condition holds showing is a natural transformation.
It is now easy to verify that forms an adjunction and that the induced monad is the Giry monad .
Now apply Beck’s monadicity theorem (Theorem 2.2) to prove that the comparison functor
is an equivalence. (The monadicity theorem is Theorem 1, page 147 of CWM. Although see the note on strict monadicity at monadicity theorem.) In Theorem 1, Chapter VI, page 152 CWM, MacLane provides the proof that the comparison functor between any algebraic variety and the corresponding category of algebras induced by the free functor and forgetful functor of that algebraic variety is an equivalence. Indeed, the category is one such algebra, and it is convenient to view the proof as a proof that the comparison functor is an equivalence. That free functor sends a set to the set of all formal finite affine sums, . By changing the base space from to , and using the fact from Lemma 3.2 that all the affine sum operations are measurable in rather than just set functions, that proof carries through verbatim to show that the comparison functor is an equivalence. That completes the proof that is equivalent to .
See also monads of probability, measures, and valuations.
Vladimir Voevodsky has also worked on a category theoretic treatment of probability theory, and gave few talks on this at IHES, Miami, in Moscow etc. Voevodsky had in mind applications in mathematical biology?, for example, population genetics:
…a categorical study of probability theory where “categorical” is understood in the sense of category theory. Originally, I developed this approach to probability to get a better understanding of the constructions which I had to deal with in population genetics. Later it evolved into something which seems to be also interesting from a purely mathematical point of view. On the elementary level it gives a category which is useful for the work with probabilistic constructions involving complicated combinations of stochastic processes of different types. On a more advanced level, applying in this context the old idea of a functor as a generalized object one gets a better view of the relationship between probability and the theory of (pre-)ordered topological vector spaces.
A talk in Moscow (20 Niv 2008, in Russian) can be viewed here, wmv 223.6 Mb. Abstract:
In early 60-ies Bill Lawvere defined a category whose objects are measurable spaces and morphisms are Markov kernels. I will try to show how this category allows one to think about many of the notions of probability theory in categorical terms and to connect probabilistic objects to objects of other types through various functors.
Voevodsky’s unfinished notes on categorical probability theory have been released posthumously.
Prakash Panangaden in Probabilistic Relations defines the category (stochastic relations) to have as objects sets equipped with a -field. Morphisms are conditional probability densities or stochastic kernels. So, a morphism from to is a function such that
If is a morphism from to , then from to is defined as .
Panangaden’s definition differs from Giry’s in the second clause where subprobability measures are allowed, rather than ordinary probability measures.
Panangaden emphasises that the mechanism is similar to the way that the category of relations can be constructed from the power set functor. Just as the category of relations is the Kleisli category of the powerset functor over the category of sets Set, is the Kleisli category of the functor over the category of measurable spaces and measurable functions which sends a measurable space, , to the measurable space of subprobability measures on . This functor gives rise to a monad.
What is gained by the move from probability measures to subprobability measures? One motivation seems to be to model probabilistic processes from to a coproduct . This you can iterate to form a process which looks to see where in you eventually end up. This relates to being traced.
There is a monad on , . A probability measure on is a subprobability measure on . Panangaden’s monad is a composite of Giry’s and .
measure, probability measure, pushforward measure, convex mixture
Radon monad, distribution monad, extended probabilistic powerdomain
The adjunction underlying the Giry monad was originally developed by Lawvere in 1962, prior to the full recognition of the relationship between monads and adjunctions. Although P. Huber had already shown in 1961 that every adjoint pair gives rise to a monad, it wasn’t until 1965 that the constructions of Eilenberg-Moore, and Kleisli, made the essential equivalence of both concepts manifest.
Lawvere’s construction was written up as an appendix to a proposal to the Arms Control and Disarmament Agency, set up by President Kennedy as part of the State Department to handle planning and execution of certain treaties with the Soviet Union. This appendix was intended to provide a reasonable framework for arms control verification protocols (Lawvere 20).
At that time, Lawvere was working for a “think tank” in California, and the purpose of the proposal was to provide a means for verifying compliance with limitations on nuclear weapons. In the 1980’s, Michèle Giry was collaborating with another French mathematician at that time who was also working with the French intelligence agency, and she was able to obtain a copy of the appendix. Giry then developed and extended some of the ideas in the appendix (Giry 80)
Gian-Carlo Rota had also (somehow) obtained a copy of the appendix, which ended up in the library at The American Institute of Mathematics, and only became publicly available in 2012.
From Lawvere 20:
I’d like to say that the idea of the category of probabilistic mappings, the document corresponding to that was not part of a seminar, as some of the circulations say, essentially it was the document submitted to the arms control and disarmament agency after suitable checking that the Pentagon didn’t disagree with it. Because of the fact that for arms control agencies as a side responsibility the forming of arms control agreements and part of these agreements must involve agreed upon protocols of verification. So the idea of that paper did not provide such protocols, but it purported to provide reasonable framework within which such protocol can be formulated.
The idea originates with
The original copy of the appendix is
W. Lawvere, The category of probabilistic mappings, ms. 12 pages, 1962 (Lawvere Probability 1962)
(notice that the statement of origin on p.1 is wrong.)
This idea was picked up and published in:
Michèle Giry, A categorical approach to probability theory, Categorical aspects of topology and analysis (Ottawa, Ont., 1980), pp. 68–85, Lecture Notes in Math. 915 Springer 1982 (doi:10.1007/BFb0092872)
(there are allegedly a few minor analytically incorrect points and gaps in proofs, observed by later authors).
Historical comments on the appearance of Lawvere 62 are made in
According to E. Burroni (2009), the Giry monad appears also in
The article
shows, in effect, that the Giry monad restricted to countable measurable spaces (with the discrete -algebra) yields the restricted Giry functor which has the codensity monad . This suggest that the natural numbers are ‘’sufficient’‘ in some sense. Indeed, the full subcategory of Polish spaces consisting of the single object of all natural numbers with the powerset -algebra is codense in - every continuous function is completely determined by its values on the countable dense subset of .
The article
views probability measures via double dualization, restricted to weakly averaging affine maps. A more satisfactory description of the Giry monad arises from recognizing the need for viewing them as weakly-averaging linear maps, obtained by double dualizing into , which then yields the characterization of -algebras summarized above. These ideas originally appeared as
but it was realized the method applied to all measurable spaces.
Some corrections from an earlier version of the Categorical Probability Theory article, were pointed out in
Apart from these papers, there are similar developments in
Franck van Breugel, The metric monad for probabilistic nondeterminism, features both the Lawvere/Giry monad and Panangaden’s monad.
Ernst-Erich Doberkat, Characterizing the Eilenberg-Moore algebras for a monad of stochastic relations (pdf)
Ernst-Erich Doberkat, Kleisli morphisms and randomized congruences, Journal of Pure and Applied Algebra Volume 211, Issue 3, December 2007, Pages 638-664 https://doi.org/10.1016/j.jpaa.2007.03.003
N. N. Cencov, Statistical decisions rules and optimal Inference, Translations of Math. Monographs 53, Amer. Math. Society 1982
(blog comment) Cencov’s “category of statistical decisions” coincides with Giry’s (Lawvere’s) category. I ( somebody) have the sense that Cencov discovered this category independently of Lawvere although years later.
category cafe related to Giry monad: category theoretic probability, coalgebraic modal logic
Samson Abramsky et al. Nuclear and trace ideals in tensored ∗-Categories,arxiv:math/9805102, on the representation of probability theory through monads, which looks to work Giry’s monad into a context even more closely resembling the category of relations.
There is also relation with work of Jacobs et al.
Robert Furber, Bart Jacobs, Towards a categorical account of conditional probability, arxiv:1306.0831
Bart Jacobs, Probabilities, distribution monads and convex categories, Theoretical
Computer Science 412(28) (2011) pp.3323–3336. https://doi.org/10.1016/j.tcs.2011.04.005, (preprint)
J. Culbertson and K. Sturtz use the Giry monad in their categorical approach to Bayesian reasoning and inference (both articles contain further references to the categorical approach to probability theory):
Jared Culbertson and Kirk Sturtz, A categorical foundation for Bayesian probability, Applied Cat. Struc. 2013 (preprint as arXiv:1205.1488)
Jared Culbertson and Kirk Sturtz, Bayesian machine learning via category theory, 2013 (arxiv:1312.1445)
Elisabeth Burroni, Lois distributives. Applications aux automates stochastiques, TAC 22, 2009 pp.199-221 (journal page)
where she derives stochastic automata as algebras for a suitable distributive law on the monoid and Giry monads.
B. Fong has a section on the Giry monad in his paper on Bayesian networks:
See also:
Discussion of the Giry monad extended to simplicial sets and used to characterize quantum contextuality via simplicial homotopy theory:
Cihan Okay, Sam Roberts, Stephen D. Bartlett, Robert Raussendorf, Topological proofs of contextuality in quantum mechanics, Quantum Information and Computation 17 (2017) 1135-1166 [arXiv:1701.01888, doi:10.26421/QIC17.13-14-5]
Cihan Okay, Aziz Kharoof, Selman Ipek, Simplicial quantum contextuality, Quantum 7 (2023) 1009 [arXiv:2204.06648, doi:10.22331/q-2023-05-22-1009]
Aziz Kharoof, Cihan Okay, Homotopical characterization of strongly contextual simplicial distributions on cone spaces [arXiv:2311.14111]
exposition:
Last revised on September 20, 2026 at 10:30:38. See the history of this page for a list of all contributions to it.