Inference, Constraints, and World Models
One Gaussian, two peaks
The fit is normalized. It still assigns probability to a valley the data avoid.
Wheel or pinch to zoom · drag a range to zoom · Shift-drag to pan · double-click to reset. Hover for values.
The 2⁻ᵈ example assumes independent half-volume restrictions. It is not a law for every rejection sampler.
JEPA and diffusion objectives
Try collapsing the embeddings. Then add noise to the two-peaked data. Only one procedure still has a density to denoise.
I-JEPA: predict a representation
L = ‖predictor(encoder(context)) − teacher(target)‖²
The bare predictive term can collapse. The training method needs an anti-collapse mechanism. This objective has no reverse-noise sampler.
Diffusion: estimate a density’s score
A score is ∂ₓ log pₜ(x). A Gaussian embedding regularizer, as in LeJEPA, matches a latent distribution; it does not supply a reverse-time score field. I-JEPA calls itself “a non-generative approach.” Score-SDE explicitly constructs a reverse-time sampler.
Expected cost and tail risk
Choose between a safe action (reward 0, cost 0) and a risky action (reward 10, cost 10). Spend an expected-cost budget of 4.
The optimal mixture here uses risky with probability budget / 10. No deterministic action attains that return for a budget strictly between 0 and 10. A hard per-run rule would forbid the risky action altogether.
Prediction and planning
Suppose every transition is known exactly. How many length-H action sequences must an exhaustive search consider?
3⁵ = 243 possible action sequences.
This counts brute-force trajectories, not the time required by every planner. Structure, pruning and value estimates can help; a perfect predictor alone gives none of them.
Density, sampling, transport
03:23One Gaussian misses two modesConditional
For equally weighted peaks at ±a, the moment-matched normal is N(0, a² + σ²): its maximum sits at x = 0, between the peaks. But a reverse-KL variational approximation may select one peak instead. Variational inference is not restricted to Gaussians and usually approximates a posterior, not a raw data histogram. Auto-Encoding Variational Bayes · video
04:44Rejection sampling loses in high dimensionsConditional
For envelope Mq ≥ f, acceptance is Z/M where Z = ∫f. If each independent coordinate occupies half the proposal’s volume, acceptance = 2⁻ᵈ. A closer proposal need not have this decay. von Neumann, “Various Techniques Used in Connection with Random Digits” (1951) describes accept/reject sampling; the envelope identity follows by integrating the accept probability. video
05:40MCMC needs no partition function; it can stick between modesConditional
In Metropolis–Hastings, α(x,y) = min(1, e−E(y)+E(x)q(x|y)/q(y|x)): the same Z cancels. A local proposal still needs many unlikely moves to cross a deep valley. Not all MCMC uses rejection, and none is universally slow. Hastings (1970) · video
06:53GFlowNets preserve flow through a construction graphFollows
For an internal node, incoming flow = outgoing flow; at a terminal x, outflow = reward R(x). Then the terminal sampling law is R(x)/ΣR. This assumes a specified graph and terminal rewards, not an arbitrary inference oracle. The original paper seeks objects whose probability is proportional to a given positive reward
. Bengio et al. · video
08:21; 18:10; 29:23Energy gives ratios, not free likelihoods or free trainingConditional
p(x) = e−E(x)/Z, so p(x)/p(y) = eE(y)−E(x). A normalizable E can generate samples in principle. But maximum-likelihood training has ∂θ log p(x) = −∂θE(x) + Ep[∂θE], which still needs model expectations. Score matching and some other objectives avoid Z; that is a choice of objective, not a property of every EBM. LeCun et al., EBM tutorial · Hyvärinen, score matching · video
11:32–17:54Token likelihoods, squared error, uncertainty and priorsConditional
p(x₁,…,xₙ) = ∏ᵢp(xᵢ|x₁,…,xᵢ₋₁) by the chain rule. Under N(μθ, σ²I) with fixed σ, −log p(y|x) = ‖y−μθ(x)‖²/(2σ²) + constant. A normalized softmax is not automatically calibrated; a prior is not merely an interpolation rule. An exactly fair coin has posterior probability zero under a continuous-only bias prior: add point mass at ½ to test that precise hypothesis. Q-values can rank all actions; exploration is not exclusive to policy gradients. Kingma & Welling · Watkins & Dayan · video
19:35–28:31Diffusion, score, continuous flow: related, not identicalConditional
Noise gives pₜ and score ∇log pₜ. For dX = f dt + g dW, a probability-flow ODE has velocity v = f − ½g²∇log pₜ (scalar, state-independent g). A sufficiently regular ODE permits log-density change d log pₜ(Xₜ)/dt = −∇·v. It is this ODE—not an invertible individual stochastic noise path—that offers what Song et al. call exact likelihood computation
, given the ideal field and exact integration. Finite-step diffusion exists; “infinite latents” is an analogy. Sohl-Dickstein et al. · DDPM · Song et al. · video
26:43; 55:32–1:01:47Flow matching includes diffusion paths; it does not equal diffusionFollows
For an invertible map F, pₓ(x) = pᵤ(F⁻¹(x))|det JF⁻¹(x)|. In continuous time, ∂ₜp + ∇·(pv) = 0. Flow matching regresses a chosen conditional path’s velocity without integrating the ODE during training; sampling still integrates it. The paper says its path family subsumes existing diffusion paths as specific instances
—and explicitly permits non-diffusion paths. Neither ODE uniqueness nor cheap divergence follows for an arbitrary discontinuous field. Neural ODEs · Lipman et al. · Rectified Flow · video
59:49“All diffusion is KL gradient flow”Does not follow
The heat equation ∂ₜp = Δp is the Wasserstein gradient flow of entropy under specific conditions. An arbitrary learned noise schedule or reverse sampler is not therefore the gradient flow of KL to its target. That theorem has an objective and specified dynamics; the word “diffusion” does not supply them. Jordan–Kinderlehrer–Otto · Score-SDE · video
Representation learning
29:05; 33:44; 2:04“JEPA is basically diffusion”Does not follow
The video does not say this: it connects score-based energies to diffusion, then calls JEPA a broad label. The proposed equivalence also fails three independent tests. Objective: I-JEPA predicts target embeddings from context; DDPM regresses injected noise or a denoising target. Output: I-JEPA yields representations, not a density score or reverse sampler. Counterexample: constant student and teacher embeddings give zero bare predictive error for any data; a constant field cannot be the score of a two-peaked, normalizable density. Collapse-prevention mechanisms change JEPA’s training, not its objective into a diffusion SDE. LeJEPA’s Gaussian regularizer is on embeddings, not a pixel-space forward noise chain. Some hybrids can combine both; overlap is not identity. I-JEPA · LeJEPA · DDPM · video
33:44; 2:04“DINO would be called JEPA”; “world model is branding”Judgment
DINO and JEPA both learn representations without labeled targets, but their training targets are not identical. Naming families retrospectively is a taxonomy choice. Predictive dynamics predate either label; neither that fact nor the label erases distinct methods. The guest’s DINO-over-JEPA result is explicitly limited to his own tasks. DINO · I-JEPA · World Models · video
36:14–40:02Contrastive, non-contrastive, and LeJEPAConditional
SimCLR uses positive pairs and in-batch negatives; MoCo adds a queue. VICReg prevents collapse with variance/covariance penalties. I-JEPA predicts masked target features with a teacher. LeJEPA adds Sketched Isotropic Gaussian Regularization
to a predictive loss. A reference distribution is not part of every JEPA, nor proof that contrastive learning is intrinsically unstable; batch size and target choice matter empirically. SimCLR · MoCo · VICReg · LeJEPA · video
40:03–44:53; 53:00“EBMs are obsolete”; “continuous media needs diffusion”Judgment
Cost and sample quality depend on the model, objective and resources. Even continuous variables factorize autoregressively: p(x₁,…,xₙ) = ∏ᵢp(xᵢ|x₁,…,xᵢ₋₁). EBM ratios, physical energies or pretraining losses may be useful without fast generation. These equations prove possibility, not which approach wins a benchmark. EBM tutorial · Flow Matching · video
44:54–52:58Search, Q-learning and SARSAFollows
AlphaGo paired a learned policy/value with Monte Carlo tree search. Branch-and-bound also searches a tree, but neither process is necessarily Bayesian inference. For a discounted MDP, Q*(s,a) = E[r + γ maxa′ Q*(s′,a′)|s,a]. SARSA bootstraps the action actually selected; Q-learning bootstraps a maximizing action. State/action values are not uncertainty intervals. AlphaGo · Watkins & Dayan · video
Deep learning theory
1:01:48–1:07:52Coarse-to-fine images; architecture stops matteringJudgment
AlphaFold 3 uses a diffusion-based structure module; some generators visibly refine coarse shapes before details. Neither the diffusion equations nor a straight flow path force that perceptual order. Vision training can show spectral bias, but the schedule, data and architecture matter. Width alone does not establish that every architecture behaves alike. AlphaFold 3 · Spectral bias · video
1:08:12–1:15:04NTK and mean field are two different limitsConditional
With the appropriate infinite-width scaling, a limiting NTK stays constant during training
and function updates linearize: ∂ₜf = −K∇fL. Under a different scaling, a mean-field limit evolves a distribution of neurons and can learn features. Neither says every finite network has a convex objective, or that a billion parameters guarantee a better function. ReLU spline partitions describe local geometry, not generalization by themselves. Jacot et al., NTK · Mei et al., mean field · video
1:15:05–1:23:53Barron’s rate is dimension-free only for a function classConditional
For functions with bounded Fourier moment / Barron norm C, a width-m two-layer approximation has squared-error order C²/m without an explicit exponential in input dimension. C can itself grow with dimension or oscillation. Constant-in-most-coordinates is easy; “all high-dimensional learning escapes the curse” is not a corollary. Double descent is an observed risk curve, not a proof that every wide net generalizes. A useful theory must also survive tests on its assumptions. Barron (1993) · Belkin et al. · video
1:23:53–1:29:09A manifold plot is not a manifold proofDoes not follow
A 2D projection of activations can show a loop even when the original distribution has noise in every dimension. ReLU networks partition inputs into piecewise-affine regions; this neither proves nor refutes a low-dimensional data manifold. A modular-arithmetic circuit does not establish a circuit for every concept. And “in infinite dimensions every problem becomes convex” is simply false: f(x) = −‖x‖² is nonconvex on a Hilbert space of any dimension. Grokking circuits · video
Reward and constrained RL
1:29:09–1:40:01“Reward is enough” does not say data is freeConditional
You can define r(s,a) = 1 for the desired action and 0 otherwise. This proves a representation exists if the desired action is already known. With N indistinguishable actions and only one rewarding choice, even finding it may require O(N) trials; sparse feedback is costly. The original is a hypothesis about reward maximization, not a sample-complexity theorem. Control may identify an unknown plant; model-based RL may learn dynamics. The opposition “control has a model, RL has none” is too neat. Silver et al., Reward is Enough · video
1:35:00–1:42:15Expected rewards can hide failures and correlated shaping termsConditional
Expected cost E[C] ≤ B permits a 1% chance of cost 100B if the other 99% cost 0. Reward weights can double-count correlated outcomes and need retuning; an explicit constraint expresses a threshold instead. Neither is a deployment guarantee. The episode’s OpenAI Five appendix anecdote and “wrong lane” example are not treated as verified measurements here. A cited worst-case ReLU training difficulty also requires its missing assumptions; expected error alone cannot guarantee uniform error. OpenAI Five · CPO · video
1:42:16–1:47:56Constrained policies can need randomnessFollows
The two-action lab above proves it in one state: a budget B ∈ (0,10) is met optimally by playing risky with probability B/10. In finite discounted CMDPs, occupancy-measure constraints can be linear; a neural network’s parameter optimization need not be convex. A chance constraint limits violation probability; CVaR averages losses in a tail. CPO claims near-constraint satisfaction at each iteration
under its assumptions, not a blanket 99%-safe-per-update guarantee. Altman, CMDPs · Achiam et al., CPO · video
1:45–1:55Projection, dual penalties, expert constraints and maskingConditional
A penalty L(π,λ) = Jr(π) − λ(Jc(π)−B) can enforce an expected budget at a suitable saddle point; approximate updates can oscillate. Projection is only as good as the learned feasible set. Action masking can enforce known local limits; uncertain future state constraints require dynamics or a safety monitor. Generic constraint satisfaction can be hard; that does not forbid every hard-coded safe architecture. CPO · Altman · video
1:55:45–2:00:32Abstraction needs a contract, not a universal-approximation wishConditional
A lower-level system that guarantees a bounded error for a specified disturbance set offers a usable interface; a network’s capacity to approximate a function does not guarantee training finds it or that it stays safe. The conjectures about creativity, hierarchical RL’s popularity and fuzzy logic’s usefulness are not mathematical consequences of this contract. Hornik et al., approximation · video
World models and planning
2:00:33–2:04:37One name, several prediction targetsConditional
Sensor dynamics, latent transitions and video generation can all be called “world models”; the term alone says nothing about action conditioning or controllability. Ha and Schmidhuber trained a generative model and a separate agent in its own dream environment
. Dreamer learns a latent dynamics model with a policy. Whether to reserve “model-based RL” for models updated during interaction is the guest’s proposed vocabulary, not a theorem. World Models · DreamerV3 · video
2:04:38–2:08:13A perfect predictor does not select a good actionFollows
A complete transition rule T(s,a) gives each next state, not the best policy. With b actions and horizon H, exhaustive search considers bᴴ paths. Pruning, structure or a learned value function can save work; the rule alone does not. This refutes “prediction removes the need for decision-making,” but does not prove that RL specifically is necessary. World Models · MuZero (model + search) · video
2:08:14–2:14:14Rollouts, robot demonstrations, reliabilityConditional
A cheap simulator allows many trials only while its errors remain tolerable for the policy. The host’s unnamed ball-handling clip cannot establish generalization. For independent tasks with per-task success 0.9, all 20 succeed with probability 0.9²⁰ ≈ 0.12; correlated failures can be worse. Repeating a 51%-correct measurement helps only with sufficiently independent evidence—not as a generic Kalman-filter safety guarantee. Predictions about manufacturing adoption and “demo-first” methods remain opinions. DreamerV3 · video
Source: full interview, checked against the captioned 2:14:14 version. One caption gap near 1:54 is not interpreted; examples and derivations here are toy constructions, not reported measurements. A mathematical counterexample rules out a universal claim, not a particular experiment.