🕵️ SicherheitslückenHak5: Hackers Just Poisoned the Rust Supply Chain | Threat Wire(01.09.2026 um 14:00 Uhr)
🕵️ SicherheitslückenHak5: Hackers Found a Way Into Humanoid Robots | Threat Wire(04.09.2026 um 15:04 Uhr)
🔧 AI Nachrichten Bits und so #1021 (Passwort für Laufwerk)(31.08.2026 um 22:15 Uhr)
🔧 AI Nachrichten Bits und so #1022 (Wie Weißbier)(06.09.2026 um 20:39 Uhr)
🍏 iOS / Mac OSHue-App 6.0 ist da: das sind die Neuerungen(07.09.2026 um 17:21 Uhr)
🕵️ SicherheitslückenHak5: Hackers Just Poisoned the Rust Supply Chain | Threat Wire(01.09.2026 um 14:00 Uhr)
🕵️ SicherheitslückenHak5: Hackers Found a Way Into Humanoid Robots | Threat Wire(04.09.2026 um 15:04 Uhr)
🔧 AI Nachrichten Bits und so #1021 (Passwort für Laufwerk)(31.08.2026 um 22:15 Uhr)
🔧 AI Nachrichten Bits und so #1022 (Wie Weißbier)(06.09.2026 um 20:39 Uhr)
🍏 iOS / Mac OSHue-App 6.0 ist da: das sind die Neuerungen(07.09.2026 um 17:21 Uhr)

🔧 Programmierung 🕛 kürzlich 4 Min Lesezeit
0

Bayesian Neural Networks

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

Adapted from an appendix of my MS thesis.







Bayesian Neural Network



Deep neural networks (DNNs) are usually trained using a regularized maximum likelihood objective to find a single setting of parameters. However, large flexible models like neural networks can represent many functions, corresponding to different parameter settings, which fit the training data well, yet generalize in different ways. This phenomenon is known as underspecification [1].



Considering all of these different models together can lead to improved accuracy and uncertainty representation. This can be done by computing the posterior predictive distribution using Bayesian model averaging. The main challenges in applying Bayesian inference to DNNs are specifying suitable priors, and efficiently computing the posterior, which is challenging due to the large number of parameters and the large datasets [1].







p(yx,D)=p(yx,θ)p(θD)dθwherep(θD)p(θ)p(Dθ).



Consider the generalized MLP with

L1

hidden layers and a linear output as described in the companion MLP post [1].





f(x;θ)=WL(φ(W1x+b1))+bL.



The most common choice of priors is to use a factored Gaussian prior [1].





WN(0,α2I),bN(0,β2I).



Initializing this model’s parameters at a particular random value is like sampling a point from this prior over functions. In the limit of infinitely wide neural networks, a neural network defines a Gaussian process (see the companion Gaussian process series) with a fixed kernel. Indeed, an MLP with one hidden layer, whose width goes to infinity, and which has a Gaussian prior on all the parameters, converges to a Gaussian process with a well-defined kernel. These kernels can be used in modeling non-stationary covariance structure [1].



Monte Carlo dropout (MCD) is a very simple and widely used method for approximating the Bayesian predictive distribution. Usually stochastic dropout layers are added as a form of regularization during training, and are turned off at test time. However, the idea of MCD is to also perform random sampling during testing [1].



More precisely, we drop out each hidden unit according to a

Bernoulli(p)

distribution, and repeat this procedure

K

times to create

K

distinct models. We then create an equally weighted average of the predictive distribution for each of these models as shown below. One drawback of MCD is that it is slow at test time. However, this can be overcome by distilling the model’s predictions into a deterministic student network [1].





p(yx,D)K1k=1Kp(yx,θk).



Another very simply approximation to the posterior is to only be Bayesian about the weights in the final layer. This is called the “Bayesian last layer” approximation. In more detail, let

z=f(x;θ)

be the predicted outputs of the model before any optional final nonlinearity. We assume this has the form

z=wLϕ(x;θ)

, where

ϕ(x)

are the features extracted by the first

L1

layers. Then, we can use standard techniques, such as Markov chain Monte Carlo (described in the companion MCMC series), to compute

p(wLD)=N(μL,ΣL)

, given

ϕ()

[1].



MCMC methods like Hamiltonian Monte Carlo (HMC) are generally considered to be the gold standard for posterior approximation, since they do not make strong assumptions about the form of the posterior. However, a significant limitation of standard MCMC procedures, including HMC, is that they require access to the full training set at each step. Stochastic gradient MCMC methods, such as Stochastic gradient Langevin dynamics (SGLD), operate instead using mini-batches of data, offering a scalable alternative [1].



Many conventional approximate inference methods, such as variational inference, focus on approximating the posterior

p(θD)

in a local neighborhood around one of the posterior modes. While this is often not a major limitation in classical machine learning, modern deep neural networks have highly multi-modal posteriors, with parameters in different modes giving rise to very different functions [1].



On the other hand, the functions in a neighborhood of a single mode may make fairly similar predictions. So using a local approximation to compute the posterior predictive will underestimate uncertainty and generalize more poorly. A simple alternative method is to train multiple models, and then to approximate the posterior using an equally weighted mixture of delta functions as follows where

M

is the number of models. This approach is called deep ensembles [1].





p(θD)M1m=1Mδ(θθ^m).



Once we have an approximation

q(θD)

of the parameter posterior

p(θD)

, we can use it to approximate the posterior predictive distribution [1].





p(yx,D)=p(yx,θ)p(θD)dθ.



We often approximate this integral using Monte Carlo where

θsq(θD)

[1].





p(yx,D)S1s=1Sp(yx,θs).






References




  1. Kevin P. Murphy (2023) Probabilistic Machine Learning: Advanced Topics. MIT Press.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Hackers Just Poisoned the Rust Supply Chain | Threat Wire
1 Quelle
Hackers Found a Way Into Humanoid Robots | Threat Wire
1 Quelle
Bits und so #1021 (Passwort für Laufwerk)
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Bayesian Neural Networks

Thematisch verwandte Begriffe: Bayesian, Neural, Networks · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...