Skip to content
LESSON NOTES · 03

3. Importance sampling and reweighting

Position in the course: Lesson 3 of 10. Complete the preceding derivation and use the explained exercises to check understanding.

Prerequisites: Expectations; canonical distribution

Learning goal: explain the mathematical steps, reproduce the analytical examples, and state the conditions under which the conclusions hold.

1. Move effort toward important regions

Uniform points waste effort when most of an integrand is concentrated in a small region. Choose a normalized density \(q(x)\) with support wherever the integrand \(f(x)\) is nonzero. Multiplying and dividing by \(q\) gives

\[ I=\int f(x)dx=\mathbb E_q[f/q],\quad \operatorname{Var}_q(f/q)=\int\frac{f(x)^2}{q(x)}dx-I^2. \]

This equality requires meaningful finite integrals. A proposal that omits a narrow contributing region introduces bias that extra samples never remove. By Cauchy–Schwarz, \(\int f^2/q\geq(\int|f|)^2\); equality occurs at \(q\propto|f|\). For a nonnegative integrand this would make the estimator constant, but its normalization is the unknown integral. The optimum is a guide to designing an approximation, not a circular algorithm.

2. Worked improvement for a known integral

For \(f=x^2\) on \([0,1]\), uniform sampling has variance \(4/45\). Choose \(q=2x\). Except at the measure-zero endpoint, \(f/q=x/2\). Its mean is \(1/3\) and second moment is \(\int_0^1(2x)(x^2/4)dx=1/8\), so its variance is \(1/72\). This is a reduction by a factor \(6.4\) in variance. The choice \(q=3x^2\) would give zero variance, illustrating the ideal but already-known normalization. These values are analytical comparisons.

3. When only unnormalized weights are available

Suppose a chain samples normalized \(p_0\) and the desired density is \(p_1\propto p_0w\). Dividing numerator and denominator by the same unknown normalization gives

\[ \langle A\rangle_1=\frac{\langle wA\rangle_0}{\langle w\rangle_0},\qquad \widehat A_1=\frac{\sum_sw_sA_s}{\sum_sw_s}. \]

At different temperatures \(w_s=e^{-(\beta_1-\beta_0)E_s}\). The ratio estimator is generally biased at finite sample size even when numerator and denominator estimates are individually unbiased. The delta method expands it around their means and includes their covariance. When independent weights are nonnegative, \((\sum w)^2/\sum w^2\) is a useful weight concentration diagnostic; it is not a complete Markov-chain effective sample size and ignores missing regions.

4. Overlap is a physical requirement

If the target favors configurations never encountered at the original temperature, no histogram transformation can create them. Examine distributions of energy and order parameters, perform independent runs, and use intermediate states or expanded ensembles when overlap is poor. Evaluate exponential weights by subtracting their largest log value to avoid overflow; the common scale cancels. Numerical stability preserves information already sampled; it cannot manufacture overlap.

A control variate offers another route to variance reduction. If \(B\) has known mean \(b\), use \(A-c(B-b)\) without changing the expectation. Its variance is \(\operatorname{Var}(A)+c^2\operatorname{Var}(B)-2c\operatorname{Cov}(A,B)\). Minimizing this quadratic gives \(c=\operatorname{Cov}(A,B)/\operatorname{Var}(B)\) when the denominator is nonzero. Separate tuning samples help avoid finite-sample effects from selecting the coefficient on the reporting data.

5. Exercises with solutions

Exercise: Weights are \(1,1,1,9\). Compute the concentration diagnostic and interpret it.

Solution

It is \(12^2/(1+1+1+81)=12/7\approx1.71\). Four saved configurations supply less than two evenly weighted equivalents. Correlation or unvisited states can make the actual situation worse.

Exercise: Can an importance distribution have zero probability where \(f\ne0\)?

Solution

No, unless that region has zero contribution measure. The change-of-measure identity would otherwise omit part of the integral, not merely increase its variance.

Importance sampling and reweighting

Original teaching schematic of the mathematics or algorithm; it is not simulation or experimental data.

6. Sources and connections

Related theory: Molecular dynamics · Quantum Monte Carlo

Software connection: CP2K · Quantum ESPRESSO

These software courses provide related background on energies, orbitals, or convergence management; they do not imply that the Monte Carlo or QMC examples on this page were executed there.

Molecular dynamics · Quantum Monte Carlo · RASPA3


Previous lesson · Course overview · Next lesson