🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 3 Min Lesezeit
0

I Finally Understood Why Neural Networks Need Activation Functions

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




Today I Finally Understood Why We Plot the Derivative of Activation Functions



When I started learning Deep Learning, I thought activation functions were just mathematical equations we had to memorize.



Today, while implementing Sigmoid, Tanh, ReLU, Leaky ReLU, and Softmax from scratch in Python, I realized something much more interesting.



The activation function itself is only half of the story.



The derivative is what actually teaches the neural network how to learn.









My Experiment



I generated 100 equally spaced values between -10 and 10:




CODE
z = np.linspace(-10, 10, 100)






Then I implemented:




  • Sigmoid

  • Tanh

  • ReLU

  • Leaky ReLU

  • Softmax



For the first four, I plotted both the activation function and its derivative.



Initially, I wondered:




Why am I plotting the derivative? Isn't the activation function enough?




After learning about backpropagation, the answer became clear.









What the Activation Function Tells Us



The activation function answers:




"What output should this neuron produce?"




For example, Sigmoid compresses any input into a value between 0 and 1.



ReLU outputs 0 for negative inputs and passes positive inputs unchanged.



This is what happens during the forward pass.




CODE
Input

Weighted Sum

Activation Function

Prediction












What the Derivative Tells Us



The derivative answers a completely different question:




"How much should this neuron change to reduce the error?"




During the backward pass, the neural network computes gradients.



Those gradients come directly from the derivatives of the activation functions.



Without derivatives, the model wouldn't know how to update its weights.




CODE
Prediction

Calculate Error

Compute Derivatives

Update Weights






That was the moment when the plots finally made sense to me.









My Biggest Observation






Sigmoid



Beautiful S-shaped curve.



But its derivative almost becomes zero for very large positive or negative inputs.



That means learning slows down because the gradients almost disappear.



This is called the vanishing gradient problem.









Tanh



At first glance, it looked almost identical to Sigmoid.



But I noticed something important:



Its output is centered around zero.



That small difference helps optimization because the gradients are more balanced around the origin.









ReLU



The simplest graph was also the most surprising.




CODE
f(x)=max(0,x)






Its derivative is




CODE
0   (negative inputs)

1 (positive inputs)






This explains why ReLU trains deep neural networks much faster than Sigmoid.









Leaky ReLU



ReLU has one weakness.



Negative neurons completely stop learning because the derivative becomes zero.



Leaky ReLU fixes that with a tiny negative slope.



Instead of




CODE
Gradient = 0






it becomes




CODE
Gradient = 0.01






A small change mathematically—but an important one during training.









Softmax



Softmax was different from the others.



It doesn't transform a single value.



It transforms an entire vector into probabilities.



Example:




CODE
Input

[2,4,1]



Output

[0.114
0.844
0.042]






Every probability depends on every other input.



That explains why we usually don't visualize Softmax as a simple curve.



Today I stopped thinking of activation functions as just formulas.



Now I see them as two complementary ideas:




  • The activation function decides what a neuron outputs.

  • Its derivative decides how that neuron learns.



The graph explains the prediction.



The derivative explains the learning.



My next goal is to implement:




  • Forward propagation from scratch

  • Backpropagation from scratch

  • Gradient Descent

  • A simple neural network without TensorFlow or PyTorch



I want to understand the mathematics before relying on deep learning frameworks.



When you first learned neural networks, what concept changed your understanding the most—activation functions, backpropagation, or gradient descent?






DeepLearning #MachineLearning #Python #ArtificialIntelligence #NeuralNetworks #LearningInPublic #Bioinformatics #100DaysOfCode #DevCommunity

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten I Finally Understood Why Neural Networks Need Activation Functions

Thematisch verwandte Begriffe: Finally, Understood, Neural, Networks · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...