🔧 Programmierung 🕛 vor 2 Jahren 4 Min Lesezeit
0

Are we all prompting wrong? Balancing Creativity and Consistency in RAG.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

For a Boston native like myself, there are few things more heartwarming than Artificial Intelligence understanding the brilliance of Good Will Hunting. A few cursory prompts reveal that it views it as a "must-watch tale of redemption and self discovery".







Randomness and RAG 🎰



When building RAG based applications, we are often not as concerned with creativity as we are with facts. When dealing with facts, we want as little probability involved as possible. In other words, instead of sampling a probability distribution, its beneficial to just take the token with the maximum likelihood every time.



LLMWARE allows you to explore how random your generated results are, as well as augment how random you want them to be. Heres a quick demonstration:





Demo 🙌



Load the model




CODE
model = ModelCatalog().load_model("bling-stablelm-3b-tool",
sample=True,
temperature=0.3,
get_logits=True,
max_output=123)






In the load_model method, we make a few important selections. The bling 3B is one of our newest and highest performing models.



Setting the sample attribute to True or False will allow you to change between a stochastic approach and a top-token model.



The temperature can be an important tool to control the randomness of the output, with lower values making responses more focused and higher values increasing diversity in the generated text.



These key settings will allow you to see what kind of approach you want to take when it comes to the probabilistic nature of your model.



Run a simple inference model on some sample text




CODE
response = model.inference("What is a list of the key points?", sample)






This step is where your model is doing the heavy lifting, analyzing and summarizing the loaded-in documents.



Run a sampling analysis




CODE
sampling_analysis = ModelCatalog().analyze_sampling(response)
print("sampling analysis: ", sampling_analysis)






Now you get to see the analytics - giving you a better idea of how heavily your model samples from the lower-probability side of the distribution.



This analysis will include what percentage of the tokens selected by the model were also the highest probability output, and will note cases where the not-top-token was selected.



In cases where the top token was not selected, the below code will print out the exact entries of the outputs, including their token rank.




CODE
for i, entries in enumerate(sampling_analysis["not_top_tokens"]):
print("sampled choices: ", i, entries)






All these tools can help you make an informed decision on whether you want your model to think a little outside the box, or stick to the most likely answer. To see this process in action, check out our youtube video on consistent LLM output generation.







The full code for this example can be found in our to join. See you there!🚀🚀🚀

Vollständiger Original-Artikel
Den kompletten Beitrag mit allen Details direkt auf dev.to lesen.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Staatliche Hacker treiben laut Chainalysis einen Anstieg von 420 % bei Onchain-Malware voran
1 Quelle
Chinesische KI-Modelle verursachen Anstieg der auf der Blockchain platzierten Malware ...
1 Quelle
Millionen-Lösegeld nach Hacker-Angriff auf Revolut: Was Kunden jetzt wissen müssen
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Are we all prompting wrong? Balancing Creativity and Consistency in RAG.

Thematisch verwandte Begriffe: prompting, wrong, Balancing, Creativity · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...