Inference, Constraints, and World Models

One Gaussian, two peaks

The fit is normalized. It still assigns probability to a valley the data avoid.

Wheel or pinch to zoom · drag a range to zoom · Shift-drag to pan · double-click to reset. Hover for values.

The 2⁻ᵈ example assumes independent half-volume restrictions. It is not a law for every rejection sampler.

Follows under the stated assumptions · Conditional needs a restriction · Does not follow has a counterexample · Judgment is empirical or a preference. Open a claim for its derivation and sources.

Density, sampling, transport

03:23One Gaussian misses two modesConditional

For equally weighted peaks at ±a, the moment-matched normal is N(0, a² + σ²): its maximum sits at x = 0, between the peaks. But a reverse-KL variational approximation may select one peak instead. Variational inference is not restricted to Gaussians and usually approximates a posterior, not a raw data histogram. Auto-Encoding Variational Bayes · video

04:44Rejection sampling loses in high dimensionsConditional

For envelope Mq ≥ f, acceptance is Z/M where Z = ∫f. If each independent coordinate occupies half the proposal’s volume, acceptance = 2⁻ᵈ. A closer proposal need not have this decay. von Neumann, “Various Techniques Used in Connection with Random Digits” (1951) describes accept/reject sampling; the envelope identity follows by integrating the accept probability. video

05:40MCMC needs no partition function; it can stick between modesConditional

In Metropolis–Hastings, α(x,y) = min(1, e−E(y)+E(x)q(x|y)/q(y|x)): the same Z cancels. A local proposal still needs many unlikely moves to cross a deep valley. Not all MCMC uses rejection, and none is universally slow. Hastings (1970) · video

06:53GFlowNets preserve flow through a construction graphFollows

For an internal node, incoming flow = outgoing flow; at a terminal x, outflow = reward R(x). Then the terminal sampling law is R(x)/ΣR. This assumes a specified graph and terminal rewards, not an arbitrary inference oracle. The original paper seeks objects whose probability is proportional to a given positive reward. Bengio et al. · video

08:21; 18:10; 29:23Energy gives ratios, not free likelihoods or free trainingConditional

p(x) = e−E(x)/Z, so p(x)/p(y) = eE(y)−E(x). A normalizable E can generate samples in principle. But maximum-likelihood training has ∂θ log p(x) = −∂θE(x) + Ep[∂θE], which still needs model expectations. Score matching and some other objectives avoid Z; that is a choice of objective, not a property of every EBM. LeCun et al., EBM tutorial · Hyvärinen, score matching · video

11:32–17:54Token likelihoods, squared error, uncertainty and priorsConditional

p(x₁,…,xₙ) = ∏ᵢp(xᵢ|x₁,…,xᵢ₋₁) by the chain rule. Under N(μθ, σ²I) with fixed σ, −log p(y|x) = ‖y−μθ(x)‖²/(2σ²) + constant. A normalized softmax is not automatically calibrated; a prior is not merely an interpolation rule. An exactly fair coin has posterior probability zero under a continuous-only bias prior: add point mass at ½ to test that precise hypothesis. Q-values can rank all actions; exploration is not exclusive to policy gradients. Kingma & Welling · Watkins & Dayan · video

19:35–28:31Diffusion, score, continuous flow: related, not identicalConditional

Noise gives pₜ and score ∇log pₜ. For dX = f dt + g dW, a probability-flow ODE has velocity v = f − ½g²∇log pₜ (scalar, state-independent g). A sufficiently regular ODE permits log-density change d log pₜ(Xₜ)/dt = −∇·v. It is this ODE—not an invertible individual stochastic noise path—that offers what Song et al. call exact likelihood computation, given the ideal field and exact integration. Finite-step diffusion exists; “infinite latents” is an analogy. Sohl-Dickstein et al. · DDPM · Song et al. · video

26:43; 55:32–1:01:47Flow matching includes diffusion paths; it does not equal diffusionFollows

For an invertible map F, pₓ(x) = pᵤ(F⁻¹(x))|det JF⁻¹(x)|. In continuous time, ∂ₜp + ∇·(pv) = 0. Flow matching regresses a chosen conditional path’s velocity without integrating the ODE during training; sampling still integrates it. The paper says its path family subsumes existing diffusion paths as specific instances—and explicitly permits non-diffusion paths. Neither ODE uniqueness nor cheap divergence follows for an arbitrary discontinuous field. Neural ODEs · Lipman et al. · Rectified Flow · video

59:49“All diffusion is KL gradient flow”Does not follow

The heat equation ∂ₜp = Δp is the Wasserstein gradient flow of entropy under specific conditions. An arbitrary learned noise schedule or reverse sampler is not therefore the gradient flow of KL to its target. That theorem has an objective and specified dynamics; the word “diffusion” does not supply them. Jordan–Kinderlehrer–Otto · Score-SDE · video

Representation learning

29:05; 33:44; 2:04“JEPA is basically diffusion”Does not follow

The video does not say this: it connects score-based energies to diffusion, then calls JEPA a broad label. The proposed equivalence also fails three independent tests. Objective: I-JEPA predicts target embeddings from context; DDPM regresses injected noise or a denoising target. Output: I-JEPA yields representations, not a density score or reverse sampler. Counterexample: constant student and teacher embeddings give zero bare predictive error for any data; a constant field cannot be the score of a two-peaked, normalizable density. Collapse-prevention mechanisms change JEPA’s training, not its objective into a diffusion SDE. LeJEPA’s Gaussian regularizer is on embeddings, not a pixel-space forward noise chain. Some hybrids can combine both; overlap is not identity. I-JEPA · LeJEPA · DDPM · video

33:44; 2:04“DINO would be called JEPA”; “world model is branding”Judgment

DINO and JEPA both learn representations without labeled targets, but their training targets are not identical. Naming families retrospectively is a taxonomy choice. Predictive dynamics predate either label; neither that fact nor the label erases distinct methods. The guest’s DINO-over-JEPA result is explicitly limited to his own tasks. DINO · I-JEPA · World Models · video

36:14–40:02Contrastive, non-contrastive, and LeJEPAConditional

SimCLR uses positive pairs and in-batch negatives; MoCo adds a queue. VICReg prevents collapse with variance/covariance penalties. I-JEPA predicts masked target features with a teacher. LeJEPA adds Sketched Isotropic Gaussian Regularization to a predictive loss. A reference distribution is not part of every JEPA, nor proof that contrastive learning is intrinsically unstable; batch size and target choice matter empirically. SimCLR · MoCo · VICReg · LeJEPA · video

40:03–44:53; 53:00“EBMs are obsolete”; “continuous media needs diffusion”Judgment

Cost and sample quality depend on the model, objective and resources. Even continuous variables factorize autoregressively: p(x₁,…,xₙ) = ∏ᵢp(xᵢ|x₁,…,xᵢ₋₁). EBM ratios, physical energies or pretraining losses may be useful without fast generation. These equations prove possibility, not which approach wins a benchmark. EBM tutorial · Flow Matching · video

44:54–52:58Search, Q-learning and SARSAFollows

AlphaGo paired a learned policy/value with Monte Carlo tree search. Branch-and-bound also searches a tree, but neither process is necessarily Bayesian inference. For a discounted MDP, Q*(s,a) = E[r + γ maxa′ Q*(s′,a′)|s,a]. SARSA bootstraps the action actually selected; Q-learning bootstraps a maximizing action. State/action values are not uncertainty intervals. AlphaGo · Watkins & Dayan · video

Deep learning theory

1:01:48–1:07:52Coarse-to-fine images; architecture stops matteringJudgment

AlphaFold 3 uses a diffusion-based structure module; some generators visibly refine coarse shapes before details. Neither the diffusion equations nor a straight flow path force that perceptual order. Vision training can show spectral bias, but the schedule, data and architecture matter. Width alone does not establish that every architecture behaves alike. AlphaFold 3 · Spectral bias · video

1:08:12–1:15:04NTK and mean field are two different limitsConditional

With the appropriate infinite-width scaling, a limiting NTK stays constant during training and function updates linearize: ∂ₜf = −K∇fL. Under a different scaling, a mean-field limit evolves a distribution of neurons and can learn features. Neither says every finite network has a convex objective, or that a billion parameters guarantee a better function. ReLU spline partitions describe local geometry, not generalization by themselves. Jacot et al., NTK · Mei et al., mean field · video

1:15:05–1:23:53Barron’s rate is dimension-free only for a function classConditional

For functions with bounded Fourier moment / Barron norm C, a width-m two-layer approximation has squared-error order C²/m without an explicit exponential in input dimension. C can itself grow with dimension or oscillation. Constant-in-most-coordinates is easy; “all high-dimensional learning escapes the curse” is not a corollary. Double descent is an observed risk curve, not a proof that every wide net generalizes. A useful theory must also survive tests on its assumptions. Barron (1993) · Belkin et al. · video

1:23:53–1:29:09A manifold plot is not a manifold proofDoes not follow

A 2D projection of activations can show a loop even when the original distribution has noise in every dimension. ReLU networks partition inputs into piecewise-affine regions; this neither proves nor refutes a low-dimensional data manifold. A modular-arithmetic circuit does not establish a circuit for every concept. And “in infinite dimensions every problem becomes convex” is simply false: f(x) = −‖x‖² is nonconvex on a Hilbert space of any dimension. Grokking circuits · video

Reward and constrained RL

1:29:09–1:40:01“Reward is enough” does not say data is freeConditional

You can define r(s,a) = 1 for the desired action and 0 otherwise. This proves a representation exists if the desired action is already known. With N indistinguishable actions and only one rewarding choice, even finding it may require O(N) trials; sparse feedback is costly. The original is a hypothesis about reward maximization, not a sample-complexity theorem. Control may identify an unknown plant; model-based RL may learn dynamics. The opposition “control has a model, RL has none” is too neat. Silver et al., Reward is Enough · video

1:35:00–1:42:15Expected rewards can hide failures and correlated shaping termsConditional

Expected cost E[C] ≤ B permits a 1% chance of cost 100B if the other 99% cost 0. Reward weights can double-count correlated outcomes and need retuning; an explicit constraint expresses a threshold instead. Neither is a deployment guarantee. The episode’s OpenAI Five appendix anecdote and “wrong lane” example are not treated as verified measurements here. A cited worst-case ReLU training difficulty also requires its missing assumptions; expected error alone cannot guarantee uniform error. OpenAI Five · CPO · video

1:42:16–1:47:56Constrained policies can need randomnessFollows

The two-action lab above proves it in one state: a budget B ∈ (0,10) is met optimally by playing risky with probability B/10. In finite discounted CMDPs, occupancy-measure constraints can be linear; a neural network’s parameter optimization need not be convex. A chance constraint limits violation probability; CVaR averages losses in a tail. CPO claims near-constraint satisfaction at each iteration under its assumptions, not a blanket 99%-safe-per-update guarantee. Altman, CMDPs · Achiam et al., CPO · video

1:45–1:55Projection, dual penalties, expert constraints and maskingConditional

A penalty L(π,λ) = Jr(π) − λ(Jc(π)−B) can enforce an expected budget at a suitable saddle point; approximate updates can oscillate. Projection is only as good as the learned feasible set. Action masking can enforce known local limits; uncertain future state constraints require dynamics or a safety monitor. Generic constraint satisfaction can be hard; that does not forbid every hard-coded safe architecture. CPO · Altman · video

1:55:45–2:00:32Abstraction needs a contract, not a universal-approximation wishConditional

A lower-level system that guarantees a bounded error for a specified disturbance set offers a usable interface; a network’s capacity to approximate a function does not guarantee training finds it or that it stays safe. The conjectures about creativity, hierarchical RL’s popularity and fuzzy logic’s usefulness are not mathematical consequences of this contract. Hornik et al., approximation · video

World models and planning

2:00:33–2:04:37One name, several prediction targetsConditional

Sensor dynamics, latent transitions and video generation can all be called “world models”; the term alone says nothing about action conditioning or controllability. Ha and Schmidhuber trained a generative model and a separate agent in its own dream environment. Dreamer learns a latent dynamics model with a policy. Whether to reserve “model-based RL” for models updated during interaction is the guest’s proposed vocabulary, not a theorem. World Models · DreamerV3 · video

2:04:38–2:08:13A perfect predictor does not select a good actionFollows

A complete transition rule T(s,a) gives each next state, not the best policy. With b actions and horizon H, exhaustive search considers bᴴ paths. Pruning, structure or a learned value function can save work; the rule alone does not. This refutes “prediction removes the need for decision-making,” but does not prove that RL specifically is necessary. World Models · MuZero (model + search) · video

2:08:14–2:14:14Rollouts, robot demonstrations, reliabilityConditional

A cheap simulator allows many trials only while its errors remain tolerable for the policy. The host’s unnamed ball-handling clip cannot establish generalization. For independent tasks with per-task success 0.9, all 20 succeed with probability 0.9²⁰ ≈ 0.12; correlated failures can be worse. Repeating a 51%-correct measurement helps only with sufficiently independent evidence—not as a generic Kalman-filter safety guarantee. Predictions about manufacturing adoption and “demo-first” methods remain opinions. DreamerV3 · video

Source: full interview, checked against the captioned 2:14:14 version. One caption gap near 1:54 is not interpreted; examples and derivations here are toy constructions, not reported measurements. A mathematical counterexample rules out a universal claim, not a particular experiment.