<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Richard Mason</title>
<link>https://richardpmason.com/</link>
<atom:link href="https://richardpmason.com/index.xml" rel="self" type="application/rss+xml"/>
<description>Notes on maths and other things.</description>
<generator>quarto-1.10.18</generator>
<lastBuildDate>Sat, 15 Aug 2026 00:00:00 GMT</lastBuildDate>
<item>
  <title>Why do we want to do approximate inference?</title>
  <dc:creator>Richard Mason</dc:creator>
  <link>https://richardpmason.com/posts/free_energy_derivation/</link>
  <description><![CDATA[ 





<p>Suppose we have an agent that makes an observation about the world <img src="https://latex.codecogs.com/png.latex?y%5Cin%5Cmathcal%7BY%7D"> and has some internal model that describes the probability of that observation <img src="https://latex.codecogs.com/png.latex?y"> conditioned on some hidden state of the world <img src="https://latex.codecogs.com/png.latex?x%5Cin%5Cmathcal%7BX%7D">, which we denote by <img src="https://latex.codecogs.com/png.latex?p(y%20%7C%20x)">, and a prior distribution on the hidden states of the world <img src="https://latex.codecogs.com/png.latex?p(x)">. Bayes’ rule states that <img src="https://latex.codecogs.com/png.latex?p(x%20%7C%20y)%20=%20%5Cfrac%7Bp(y%20%7C%20x)p(x)%7D%7Bp(y)%7D%20=%20%5Cfrac%7Bp(y%20%7C%20x)p(x)%7D%7B%5Cint%20%7Bp(y%20%7C%20x)p(x)%7Ddx%7D,"></p>
<p>so in principle the agent can calculate the probability of a hidden state of the world under its internal model by multipling the likelihood <img src="https://latex.codecogs.com/png.latex?p(y%20%7C%20x)"> by the prior <img src="https://latex.codecogs.com/png.latex?p(x)"> and normalising by the marginal <img src="https://latex.codecogs.com/png.latex?p(y)">. The challenge is that computing the marginal term requires integrating over all of the possible hidden states, which tends to be intractable unless the internal model is chosen to have a form that has a closed form integral e.g., a Gaussian.</p>
<p>The consequence is that if an Agent want to do Bayesian inference they face a choice:</p>
<ol type="1">
<li>large computational cost to do exact inference on arbitrary internal models.</li>
<li>restrict the internal model to a family that is tractable to integrate in closed form.</li>
<li>do some form of approximate inference e.g., variational inference.</li>
</ol>
<p>Depending on the compute time budget available, exact inference is likely to be infeasible, and limiting ourselves to families of distributions that have closed form integrals is quite restrictive, which leaves us with branch 3 to explore.</p>
<p>One of the approaches to approximate inference that I like is the Free Energy principle by Friston. The idea is to first rearrange Bayes’ rule to focus on the intractable marginal <img src="https://latex.codecogs.com/png.latex?p(y)">.</p>
<p><img src="https://latex.codecogs.com/png.latex?p(y)%20=%20%5Cfrac%7Bp(y%7Cx)p(x)%7D%7Bp(x%7Cy)%7D"></p>
<p>Let <img src="https://latex.codecogs.com/png.latex?q_%7B%5Ctheta%7D(x)"> be an arbitrary probability distirbution parameterised by <img src="https://latex.codecogs.com/png.latex?%5Ctheta%5Cin%5CTheta">. We have the freedom to insert this distribution into the top and bottom of the fraction on the RHS</p>
<p><img src="https://latex.codecogs.com/png.latex?p(y)%20=%20%5Cfrac%7Bp(y%7Cx)p(x)q_%7B%5Ctheta%7D(x)%7D%7Bp(x%7Cy)q_%7B%5Ctheta%7D(x)%7D."></p>
<p>Next we take logs of both sides and negate, converting the LHS into surprise</p>
<p><img src="https://latex.codecogs.com/png.latex?-%5Clog%7Bp(y)%7D%20=%20-%5Clog%7B%5Cfrac%7Bp(y%7Cx)p(x)q_%7B%5Ctheta%7D(x)%7D%7Bp(x%7Cy)q_%7B%5Ctheta%7D(x)%7D%7D."></p>
<p>Then taking the expectation of both sides over <img src="https://latex.codecogs.com/png.latex?q_%7B%5Ctheta%7D(x)"> we have</p>
<p><img src="https://latex.codecogs.com/png.latex?-%5Cmathbb%7BE%7D_%7Bq_%7B%5Ctheta%7D(x)%7D%5B%5Clog%7Bp(y)%7D%5D%20=%20-%5Cmathbb%7BE%7D_%7Bq_%7B%5Ctheta%7D(x)%7D%5Cleft%5B%5Clog%7B%5Cfrac%7Bp(y%7Cx)p(x)q_%7B%5Ctheta%7D(x)%7D%7Bp(x%7Cy)q_%7B%5Ctheta%7D(x)%7D%7D%5Cright%5D."></p>
<p>The LHS is not a function of <img src="https://latex.codecogs.com/png.latex?x"> so the integral simplifies</p>
<p><img src="https://latex.codecogs.com/png.latex?%5Cmathbb%7BE%7D_%7Bq_%7B%5Ctheta%7D(x)%7D%5B%5Clog%7Bp(y)%7D%5D%20=%20%5Cint%20q_%7B%5Ctheta%7D(x)%5Clog%7Bp(y)%7Ddx%20=%20%5Clog%7Bp(y)%7D%20%5Cint%20q_%7B%5Ctheta%7D(x)dx%20=%20%5Clog%7Bp(y)%7D."></p>
<p>On the RHS we can choose how to split the terms inside the logarithm</p>
<p><img src="https://latex.codecogs.com/png.latex?-%5Cmathbb%7BE%7D_%7Bq_%7B%5Ctheta%7D(x)%7D%5Cleft%5B%5Clog%7B%5Cfrac%7Bp(y%7Cx)p(x)q_%7B%5Ctheta%7D(x)%7D%7Bp(x%7Cy)q_%7B%5Ctheta%7D(x)%7D%7D%5Cright%5D."></p>



 ]]></description>
  <category>AI</category>
  <guid>https://richardpmason.com/posts/free_energy_derivation/</guid>
  <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Welcome</title>
  <dc:creator>Richard Mason</dc:creator>
  <link>https://richardpmason.com/posts/welcome/</link>
  <description><![CDATA[ 





<p>This is the first post on the blog. More to follow.</p>



 ]]></description>
  <category>news</category>
  <guid>https://richardpmason.com/posts/welcome/</guid>
  <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
</item>
</channel>
</rss>
