<?xml version="1.0" encoding="UTF-8" ?>
<?xml-stylesheet type="text/xsl" href="/rss-style.xsl"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:media="http://search.yahoo.com/mrss/" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
<title><![CDATA[Team IT Security - 📰 Alle Kategorien]]></title>
<link><![CDATA[https://tsecurity.de/export/rss/alle-kategorien.xml?q=evaluation+with+ragas+faithfulness%2F]]></link>
<description><![CDATA[Das Gesamte Cyber Threat Intelligence Feed-Archiv von TSecurity.de. Alle Nachrichten, Sicherheitsmeldungen, Videos, Downloads und Analysen in einer zentralen Übersicht.]]></description>
<language>de-DE</language>
<lastBuildDate>Sat, 01 Aug 2026 09:26:27 +0200</lastBuildDate>
<pubDate>Sat, 01 Aug 2026 09:26:27 +0200</pubDate>
<ttl>15</ttl>
<copyright>2026 Team IT Security</copyright>
<managingEditor>lakandor@tsecurity.de (Horus Sirius)</managingEditor>
<webMaster>lakandor@tsecurity.de (Horus Sirius)</webMaster>
<category>IT Security</category>
<category>Cybersecurity</category>
<category>Nachrichten</category>
<generator>Team IT Security RSS Generator v2.0</generator>
<image>
<url>https://tsecurity.de/favicon.ico</url>
<title><![CDATA[Team IT Security - 📰 Alle Kategorien]]></title>
<link><![CDATA[https://tsecurity.de/export/rss/alle-kategorien.xml?q=evaluation+with+ragas+faithfulness%2F]]></link>
</image>
<atom:link href="https://tsecurity.de/export/rss/it-security.xml?q=evaluation+with+ragas+faithfulness%2F" rel="self" type="application/rss+xml" />
<item>
<title><![CDATA[US AI testing institute chief steps down within three months]]></title>
<description><![CDATA[The head of the US government’s AI testing institute, Chris Fall, has resigned about three months after taking charge of the Center for AI Standards and Innovation (CAISI), the federal organization responsible for evaluating advanced artificial intelligence models for safety and security.



Curr...]]></description>
<link>https://tsecurity.de/de/3694777/ai-nachrichten/us-ai-testing-institute-chief-steps-down-within-three-months/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3694777/ai-nachrichten/us-ai-testing-institute-chief-steps-down-within-three-months/</guid>
<pubDate>Sat, 25 Jul 2026 19:50:12 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">The head of the US government’s AI testing institute, Chris Fall, has resigned about three months after taking charge of the Center for AI Standards and Innovation (CAISI), the federal organization responsible for evaluating advanced artificial intelligence models for safety and security.</p>



<p class="wp-block-paragraph">Current National Institute of Standards and Technology NIST Director Arvind Raman will serve as acting CAISI Director following Fall’s departure while continuing to oversee the Commerce Department office responsible for the institute, the Daily Signal <a href="https://www.dailysignal.com/2026/07/20/scoop-head-of-federal-ai-safety-org-resigns/" target="_blank" rel="noreferrer noopener">reported</a>, citing two people familiar with the matter.</p>



<p class="wp-block-paragraph">A Commerce Department spokesperson who spoke to the publication did not disclose a reason for the resignation.</p>



<p class="wp-block-paragraph">Fall assumed leadership of CAISI in April after the Trump administration reorganized the former US AI Safety Institute under NIST. The institute develops methodologies for evaluating frontier AI models and works with AI developers on voluntary technical assessments covering areas such as cybersecurity, model misuse, reliability and other risks associated with increasingly capable AI systems.</p>



<p class="wp-block-paragraph">The leadership change comes as governments and AI companies continue developing technical approaches for evaluating frontier AI models while enterprises expand deployments of generative AI and agentic AI across business operations.</p>



<p class="wp-block-paragraph">In recent months, the Commerce Department has taken a <a href="https://www.infoworld.com/article/4194598/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion.html?_conv_v=vi:1*sc:1*cs:1784634320*fs:1784634320*pv:1*exp:%7B1004203305.%7Bv.1004477672-g.%7B%7D%7D%7D*seg:%7B%7D&amp;_conv_s=sh:1784634319808-0.24259838933788935*si:1*pv:1&amp;_conv_r=null&amp;_conv_sptest=null">more active role</a> in AI policy involving advanced models, placing greater attention on how the federal government evaluates technologies with potential national security implications.</p>



<h2 class="wp-block-heading">Continuity matters more than personalities</h2>



<p class="wp-block-paragraph">CAISI works with AI developers such as Anthropic, Google’s DeepMind and OpenAI on voluntary evaluations of frontier AI models and develops methodologies for testing model capabilities and risks. The institute does not regulate AI developers or certify commercial AI systems.</p>



<p class="wp-block-paragraph">For enterprises, those evaluations are one source of technical information alongside vendors’ own testing, third-party security assessments and internal AI governance programs.</p>



<p class="wp-block-paragraph">Sanchit Vir Gogia, chief analyst at Greyhound Research, said enterprises should focus less on the individual leading the institute and more on whether its technical work continues with the same level of consistency and transparency.</p>



<p class="wp-block-paragraph">“Leadership churn at CAISI weakens the signal long before it weakens the science,” Gogia said. “The testing has not stopped. Its authority simply does not travel as cleanly once the leadership does not.”</p>



<p class="wp-block-paragraph">According to Gogia, the more important question for enterprises is not whether the institute’s evaluation work will continue but whether the processes supporting those evaluations remain stable.</p>



<p class="wp-block-paragraph">“The instinct is to ask whether the pipeline is breaking,” he said. “The more useful question is where the pipeline now sits.”</p>



<h2 class="wp-block-heading">Enterprises still carry the burden of AI governance</h2>



<p class="wp-block-paragraph">Gogia said organizations should continue treating government-led AI evaluations as one input into their governance processes rather than as evidence that a model is inherently safe for enterprise deployment.</p>



<p class="wp-block-paragraph">“A government evaluation was always a signal, never a certificate,” he said. “A signal loses value the moment its issuer becomes unpredictable.”</p>



<p class="wp-block-paragraph">He said enterprises should instead monitor whether CAISI maintains consistent evaluation methodologies, continues publishing technical findings and preserves continuity within its research teams under interim leadership.</p>



<p class="wp-block-paragraph">“The name on the door is not the signal. The behaviour underneath it is,” Gogia said.</p>



<p class="wp-block-paragraph">Gogia also cautioned against linking Fall’s resignation to recent Commerce Department actions involving AI policy or export controls, noting that there is no public evidence connecting the two.</p>



<p class="wp-block-paragraph">“CAISI evaluates; it does not enforce export controls, because it holds no such power,” he said. “This is not a testing body reaching for enforcement. It is enforcement reaching past the testing body.”</p>



<p class="wp-block-paragraph">With Raman assuming the role on an interim basis, the next significant milestone for enterprises will be the appointment of a permanent director, and whether the institute’s evaluation programs continue without disruption, the analyst said.</p>



<p class="wp-block-paragraph">Gogia said the successor’s mandate may prove more important than the individual selected.</p>



<p class="wp-block-paragraph">“A CAISI result is not a safe harbour,” he said. “It informs an obligation; it does not discharge one.” NIST did not immediately respond to a request for comment.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Presence raises new questions about enterprise automation and jobs]]></title>
<description><![CDATA[OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.



The agents can answer questions and operate IT systems, and enterprises can decide w...]]></description>
<link>https://tsecurity.de/de/3694769/ai-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3694769/ai-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</guid>
<pubDate>Sat, 25 Jul 2026 19:50:08 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.</p>



<p class="wp-block-paragraph">The agents can answer questions and operate IT systems, and enterprises can decide what actions the agents may take and when they should seek human approval for actions or transfer a case to a human.</p>



<p class="wp-block-paragraph">OpenAI is already using Presence internally for its English-language phone support channel, where it verifies callers and uses account information to complete approved actions. The company said the system resolves 75% of inbound issues without human assistance.</p>



<p class="wp-block-paragraph">Another OpenAI service, Codex, can be used to monitor agents and suggest updates or improvements to processes. In OpenAI’s own tests, suggestions from Codex helped reduce handoffs to humans by 15 percentage points over 10 days, it said. Presence also includes simulation and evaluation tools that allow companies to test an agent before deployment. The tests assess whether it reaches the correct outcome, follows company policy, and hands a case to an employee when required.</p>



<p class="wp-block-paragraph">OpenAI intends each Presence deployment to deal with one kind of task, for example billing issues, insurance claims, or employee IT service requests, with agents getting only the knowledge and system access required for that task.</p>



<p class="wp-block-paragraph">Presence is not a self-service product: Enterprises will have to sign up for the limited availability program, with integration performed by OpenAI or selected <a href="https://www.computerworld.com/article/4136024/openai-partners-with-consulting-giants-to-deploy-enterprise-ai-agents.html">global systems integrators</a>.</p>



<p class="wp-block-paragraph">Companies exploring or testing Presence include Spanish bank BBVA, which is evaluating the service for everyday banking support in Mexico, and Japanese technology group SoftBank, which is using it in trials involving Japanese-language customer interactions. Australian insurer IAG is assessing whether the technology can help it respond to surges in customer demand during severe weather events.</p>



<h2 class="wp-block-heading">Workforce impact</h2>



<p class="wp-block-paragraph">OpenAI’s announcement did not address the potential effect of Presence on employment. But its claimed automation rate raises questions about how the technology could affect staffing in customer service and other support functions.</p>



<p class="wp-block-paragraph"><a href="https://pareekh.com/" target="_blank" rel="noreferrer noopener">Pareekh Jain</a>, CEO of Pareekh Consulting, said CIOs should regard the 75% figure as evidence that the technology can work, rather than as a benchmark that every enterprise can expect to reach.</p>



<p class="wp-block-paragraph">Jain said OpenAI’s deployment benefits from being built around the company’s own products and data. Large enterprises may achieve lower automation rates because they must contend with fragmented legacy systems, uneven knowledge bases and more complex compliance demands.</p>



<p class="wp-block-paragraph">“Most organizations should expect lower initial automation levels that improve over time as the AI agent is refined,” Jain said.</p>



<p class="wp-block-paragraph">The first workforce effect is more likely to be <a href="https://www.cio.com/article/4015750/cios-see-ai-prompting-new-it-hiring-even-as-boards-push-for-job-cuts.html">slower hiring than immediate layoffs</a>, according to <a href="https://www.linkedin.com/in/tulikasheel/" target="_blank" rel="noreferrer noopener">Tulika Sheel</a>, senior vice president at Kadence International.</p>



<p class="wp-block-paragraph">“The roles most exposed are likely to be repetitive, high-volume functions such as frontline customer support and routine back-office processing,” Sheel said. “However, I would expect the first impact to be on hiring and team growth rather than immediate large-scale job cuts. Over time, enterprises may redesign roles around AI-assisted workflows, with humans focusing more on complex cases, escalation, and relationship management.”</p>



<p class="wp-block-paragraph">Jain said Tier-1 support agents handling predictable queries would face the most exposure. Broader reductions would become more likely only after companies reorganize their operations around the technology.</p>



<p class="wp-block-paragraph">However, <a href="https://omdia.tech.informa.com/authors/lian-jye-su" target="_blank" rel="noreferrer noopener">Lian Jye Su</a>, chief analyst at Omdia, said Presence is unlikely to increase the threat of job displacement because companies have used similar customer-support automation from vendors such as Genesys, NiCE, Five9 and AWS for years.</p>



<p class="wp-block-paragraph">Enterprises are more likely to use Presence alongside employees, with AI handling routine requests while people remain responsible for work requiring judgment and empathy, Su said.</p>



<h2 class="wp-block-heading">Cost and operational risks</h2>



<p class="wp-block-paragraph">Analysts said CIOs should examine whether Presence can maintain resolution quality as usage grows, since fewer human handoffs could leave employees dealing with a more difficult mix of cases.</p>



<p class="wp-block-paragraph">“The key question is not simply how many tasks AI can handle, but whether it can handle them reliably at scale,” Sheel said.</p>



<p class="wp-block-paragraph">The financial case will depend partly on the cost of connecting Presence to existing systems and maintaining the controls needed to govern its use, according to Jain. “Often the biggest cost of enterprise AI is not tokens but <a href="https://www.computerworld.com/article/4128310/openai-responds-to-claude-cowork-with-its-own-platform-to-help-build-deploy-and-manage-ai-agents.html">integration and governance</a>,” Jain added.</p>



<p class="wp-block-paragraph">Companies will need to determine what systems and data the agents can access, monitor their performance, and audit the actions they take. Those investments could offset early savings.</p>



<p class="wp-block-paragraph">Su said the complexity of enterprise IT will make it difficult for OpenAI to automate entire workflows on its own. Enterprises will still need to work with other technology providers and human employees, while CIOs will favor systems that can be audited and integrated with existing infrastructure.</p>



<p class="wp-block-paragraph">Jain said the economics could improve if companies use the same integrations and governance controls across additional workflows.</p>



<p class="wp-block-paragraph"><em>This article first appeared on <a href="https://www.cio.com/article/4200684/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs.html">CIO</a>.</em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[KDnuggets Weekly Roundup: Week of July 20, 2026]]></title>
<description><![CDATA[Top 5 MCP Servers for High Performance Agentic Development • 10 Newsletters Keeping You Ahead in AI • Kaggle + Google’s Free 5-Day Agentic AI Course • Language Model Hallucination Evaluation with GraphEval]]></description>
<link>https://tsecurity.de/de/3694721/ai-nachrichten/kdnuggets-weekly-roundup-week-of-july-20-2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3694721/ai-nachrichten/kdnuggets-weekly-roundup-week-of-july-20-2026/</guid>
<pubDate>Sat, 25 Jul 2026 19:49:38 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Top 5 MCP Servers for High Performance Agentic Development • 10 Newsletters Keeping You Ahead in AI • Kaggle + Google’s Free 5-Day Agentic AI Course • Language Model Hallucination Evaluation with GraphEval]]></content:encoded>
</item>
<item>
<title><![CDATA[Stop asking AI nicely: Here’s how to get work-ready results every time]]></title>
<description><![CDATA[Over the past few years, I have learned that basic prompts produce inconsistent, hallucination-prone results that no executive would trust in production. What turned the tide was my move to advanced prompting techniques. These weren’t theoretical experiments; they became a practical foundation fo...]]></description>
<link>https://tsecurity.de/de/3694396/it-security-nachrichten/stop-asking-ai-nicely-heres-how-to-get-work-ready-results-every-time/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3694396/it-security-nachrichten/stop-asking-ai-nicely-heres-how-to-get-work-ready-results-every-time/</guid>
<pubDate>Sat, 25 Jul 2026 18:55:51 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Over the past few years, I have learned that basic prompts produce inconsistent, hallucination-prone results that no executive would trust in production. What turned the tide was my move to advanced prompting techniques. These weren’t theoretical experiments; they became a practical foundation for reliable, measurable outcomes. I want to share the techniques that consistently delivered the biggest gains in my projects, complete with real before-and-after examples, copy-paste templates, lessons from failures and guidance on when to evolve beyond prompting to agentic systems.</p>



<h2 class="wp-block-heading">Why advanced prompting still matters in enterprise settings</h2>



<p class="wp-block-paragraph">Sophisticated prompting remains essential for control, reliability and compliance. If you “ask nicely” and hope for the best, you need deterministic behavior, auditable reasoning and minimal risk of hallucination. Here’s what worked for me.</p>



<h3 class="wp-block-heading">1. Chain-of-Thought (CoT) and its variants: Unlocking step-by-step reasoning</h3>



<p class="wp-block-paragraph"><strong>The problem:</strong> Models would jump to conclusions on complex analysis tasks, especially involving data interpretation or multi-step logic.</p>



<p class="wp-block-paragraph"><strong>What I did:</strong> I started explicitly instructing the model to “think step by step” and show its reasoning.</p>



<p class="wp-block-paragraph"><strong>Before (basic prompt): </strong>“Analyze last quarter’s sales data and recommend three actions.”</p>



<p class="wp-block-paragraph"><strong>After (CoT prompt):</strong></p>



<p class="wp-block-paragraph">“You’re a senior business analyst. Analyze the following sales data step by step: [data]. First, identify the key trends. Second, calculate the rates and anomalies. Third, link findings to business context. Finally, recommend the three prioritized actions with expected impact. Explain your reasoning at each step.”  </p>



<p class="wp-block-paragraph"><strong>Results:</strong> Accuracy and depth improved dramatically.</p>



<p class="wp-block-paragraph"><strong>Variants that worked well:</strong> Self-consistency. I ran the same CoT prompt multiple times and took the majority consensus. This reduced variability significantly.</p>



<p class="wp-block-paragraph"><strong>Template you can use:</strong></p>



<pre class="wp-block-code"><code>You are [expert role]. Solve this problem by thinking step by step.

[Task or question]

For each step:

1. State your observation or calculation.

2. Explain the implication.

3. Proceed only when confident.

Final answer in this format: [structured output]</code></pre>



<h3 class="wp-block-heading">2. Tree-of-Thoughts (ToT): Exploring multiple reasoning paths</h3>



<p class="wp-block-paragraph">For truly complex decisions such as resource allocation or risk assessment, linear CoT isn’t enough. Tree-of-Thoughts lets the model generate and evaluate multiple branches.</p>



<p class="wp-block-paragraph"><strong>Example:</strong> I was helping a client evaluate three potential vendor platforms for an AI deployment. A standard prompt gave a superficial comparison. With ToT</p>



<p class="wp-block-paragraph"><strong>Prompt Snippet:</strong></p>



<pre class="wp-block-code"><code>Explore three different reasoning paths for selecting the best vendor platform:

Path 1: Focus on cost and scalability.

Path 2: Focus on security, compliance and integration.

Path 3: Focus on innovation and long-term roadmap.

For each path, evaluate pros/cons against our requirements [list].

Then, compare the paths and recommend the strongest overall option with justification.</code></pre>



<p class="wp-block-paragraph"><strong>Outcome:</strong> The model surfaced nuanced trade-offs (e.g., one vendor had superior security, but higher integration cost).</p>



<p class="wp-block-paragraph"><strong>When to use:</strong> Strategic planning, troubleshooting or scenarios with high uncertainty and multiple viable approaches.</p>



<h3 class="wp-block-heading">3. ReAct (Reason+ Act) and prompt chaining: Moving toward agentic behavior</h3>



<p class="wp-block-paragraph">One of the biggest leaps I have noticed comes from combining reasoning with tool use and chaining prompts.</p>



<p class="wp-block-paragraph"><strong>ReAct example</strong>: (used in data analytics workflow)</p>



<pre class="wp-block-code"><code>You are an AI analyst with access to tools. For the query below:

1. Reason about what information you need.

2. Choose the appropriate tool or action.

3. Observe the result.

4. Repeat until you can answer confidently.

Query: [user request]</code></pre>



<p class="wp-block-paragraph">In practice, I chained this with retrieval tools. One automated quarterly compliance reporting; the system reasoned about required data, pulled relevant records, validated them, and generated the reports.</p>



<h3 class="wp-block-heading">4. Meta-prompting and self-reflection: Letting the model improve itself</h3>



<p class="wp-block-paragraph">Use the model to refine its own prompt. This is a huge time-saver.</p>



<pre class="wp-block-code"><code>You are an expert prompt engineer. Improve the following prompt for clarity, structure and effectiveness with [target model]. Make it more precise while preserving intent.

Original prompt: [paste]

Provide the improved version and explain your changes.</code></pre>



<p class="wp-block-paragraph">Self-reflection loops (asking the model to critique its own output and revise) are a game-changer for content generation and code-review tasks.</p>



<h3 class="wp-block-heading">5. Multimodal and structured output techniques</h3>



<p class="wp-block-paragraph">With vision-enabled models, I started combining text with images (e.g., uploading architecture diagrams or dashboards).</p>



<p class="wp-block-paragraph"><strong>Tip from experience:</strong> Be extremely specific in describing what the models should focus on.</p>



<h4 class="wp-block-heading">Best practices I learned the hard way</h4>



<ul class="wp-block-list">
<li><strong>Start simple, then layer complexity</strong>: Over-engineered prompts from Day One usually backfire.</li>



<li><strong>Model specific tuning:</strong> Some models respond better to XML delimiters; others to explicit reasoning.</li>



<li><strong>Evaluation and versioning:</strong> Treat prompts like code if you track versions and run automated evals.</li>



<li><strong>Security guardrails:</strong> Always include instructions against prompt injections and respect data boundaries.</li>



<li><strong>When to stop prompting</strong>: For repetitive, high-stakes workflows, move to full agents or an orchestration framework.</li>
</ul>



<h2 class="wp-block-heading">Final takeaways for technical leaders</h2>



<p class="wp-block-paragraph">Advanced prompt engineering has now become a core competency for anyone responsible for enterprise AI outcomes. Start by picking one technique and apply it rigorously to a real business problem. Document before/ after and you will notice why it’s worth mastering.</p>



<p class="wp-block-paragraph">The field continues evolving towards more automated and agentic systems, but the ability to precisely direct AI reasoning remains foundational.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Presence raises new questions about enterprise automation and jobs]]></title>
<description><![CDATA[OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.



The agents can answer questions and operate IT systems, and enterprises can decide w...]]></description>
<link>https://tsecurity.de/de/3694393/it-security-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3694393/it-security-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</guid>
<pubDate>Sat, 25 Jul 2026 18:55:50 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.</p>



<p class="wp-block-paragraph">The agents can answer questions and operate IT systems, and enterprises can decide what actions the agents may take and when they should seek human approval for actions or transfer a case to a human.</p>



<p class="wp-block-paragraph">OpenAI is already using Presence internally for its English-language phone support channel, where it verifies callers and uses account information to complete approved actions. The company said the system resolves 75% of inbound issues without human assistance.</p>



<p class="wp-block-paragraph">Another OpenAI service, Codex, can be used to monitor agents and suggest updates or improvements to processes. In OpenAI’s own tests, suggestions from Codex helped reduce handoffs to humans by 15 percentage points over 10 days, it said. Presence also includes simulation and evaluation tools that allow companies to test an agent before deployment. The tests assess whether it reaches the correct outcome, follows company policy, and hands a case to an employee when required.</p>



<p class="wp-block-paragraph">OpenAI intends each Presence deployment to deal with one kind of task, for example billing issues, insurance claims, or employee IT service requests, with agents getting only the knowledge and system access required for that task.</p>



<p class="wp-block-paragraph">Presence is not a self-service product: Enterprises will have to sign up for the limited availability program, with integration performed by OpenAI or selected <a href="https://www.computerworld.com/article/4136024/openai-partners-with-consulting-giants-to-deploy-enterprise-ai-agents.html">global systems integrators</a>.</p>



<p class="wp-block-paragraph">Companies exploring or testing Presence include Spanish bank BBVA, which is evaluating the service for everyday banking support in Mexico, and Japanese technology group SoftBank, which is using it in trials involving Japanese-language customer interactions. Australian insurer IAG is assessing whether the technology can help it respond to surges in customer demand during severe weather events.</p>



<h2 class="wp-block-heading">Workforce impact</h2>



<p class="wp-block-paragraph">OpenAI’s announcement did not address the potential effect of Presence on employment. But its claimed automation rate raises questions about how the technology could affect staffing in customer service and other support functions.</p>



<p class="wp-block-paragraph"><a href="https://pareekh.com/" target="_blank" rel="noreferrer noopener">Pareekh Jain</a>, CEO of Pareekh Consulting, said CIOs should regard the 75% figure as evidence that the technology can work, rather than as a benchmark that every enterprise can expect to reach.</p>



<p class="wp-block-paragraph">Jain said OpenAI’s deployment benefits from being built around the company’s own products and data. Large enterprises may achieve lower automation rates because they must contend with fragmented legacy systems, uneven knowledge bases and more complex compliance demands.</p>



<p class="wp-block-paragraph">“Most organizations should expect lower initial automation levels that improve over time as the AI agent is refined,” Jain said.</p>



<p class="wp-block-paragraph">The first workforce effect is more likely to be <a href="https://www.cio.com/article/4015750/cios-see-ai-prompting-new-it-hiring-even-as-boards-push-for-job-cuts.html">slower hiring than immediate layoffs</a>, according to <a href="https://www.linkedin.com/in/tulikasheel/" target="_blank" rel="noreferrer noopener">Tulika Sheel</a>, senior vice president at Kadence International.</p>



<p class="wp-block-paragraph">“The roles most exposed are likely to be repetitive, high-volume functions such as frontline customer support and routine back-office processing,” Sheel said. “However, I would expect the first impact to be on hiring and team growth rather than immediate large-scale job cuts. Over time, enterprises may redesign roles around AI-assisted workflows, with humans focusing more on complex cases, escalation, and relationship management.”</p>



<p class="wp-block-paragraph">Jain said Tier-1 support agents handling predictable queries would face the most exposure. Broader reductions would become more likely only after companies reorganize their operations around the technology.</p>



<p class="wp-block-paragraph">However, <a href="https://omdia.tech.informa.com/authors/lian-jye-su" target="_blank" rel="noreferrer noopener">Lian Jye Su</a>, chief analyst at Omdia, said Presence is unlikely to increase the threat of job displacement because companies have used similar customer-support automation from vendors such as Genesys, NiCE, Five9 and AWS for years.</p>



<p class="wp-block-paragraph">Enterprises are more likely to use Presence alongside employees, with AI handling routine requests while people remain responsible for work requiring judgment and empathy, Su said.</p>



<h2 class="wp-block-heading">Cost and operational risks</h2>



<p class="wp-block-paragraph">Analysts said CIOs should examine whether Presence can maintain resolution quality as usage grows, since fewer human handoffs could leave employees dealing with a more difficult mix of cases.</p>



<p class="wp-block-paragraph">“The key question is not simply how many tasks AI can handle, but whether it can handle them reliably at scale,” Sheel said.</p>



<p class="wp-block-paragraph">The financial case will depend partly on the cost of connecting Presence to existing systems and maintaining the controls needed to govern its use, according to Jain. “Often the biggest cost of enterprise AI is not tokens but <a href="https://www.computerworld.com/article/4128310/openai-responds-to-claude-cowork-with-its-own-platform-to-help-build-deploy-and-manage-ai-agents.html">integration and governance</a>,” Jain added.</p>



<p class="wp-block-paragraph">Companies will need to determine what systems and data the agents can access, monitor their performance, and audit the actions they take. Those investments could offset early savings.</p>



<p class="wp-block-paragraph">Su said the complexity of enterprise IT will make it difficult for OpenAI to automate entire workflows on its own. Enterprises will still need to work with other technology providers and human employees, while CIOs will favor systems that can be audited and integrated with existing infrastructure.</p>



<p class="wp-block-paragraph">Jain said the economics could improve if companies use the same integrations and governance controls across additional workflows.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Eclipsa Video: HDR That Looks Right on Every Screen]]></title>
<description><![CDATA[Posted by Tibian Elsheikh, Product Manager, Android Core Graphics and Jeffrey Jose, Product Manager, Android Core Graphics
We’ve all been there: You’re scrolling through your favorite social media feed in a dim room, and suddenly an HDR video pops up. It’s so intensely bright that you have to squ...]]></description>
<link>https://tsecurity.de/de/3693501/android-tipps/eclipsa-video-hdr-that-looks-right-on-every-screen/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3693501/android-tipps/eclipsa-video-hdr-that-looks-right-on-every-screen/</guid>
<pubDate>Sat, 25 Jul 2026 10:15:30 +0200</pubDate>
<category>🤖 Android Tipps</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[
<img src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEhGg9E8BsBcgigJ3Pwhp0Wbd85wffhQKw9jT9eW4_IJHtJsxtaqBqZoWIc4agLIZu9h2eWFEnMgipcv2PnMM2UC9tsZOJp3AMjsOX1KQRoisg5IKTRS20hFOIvJmlViYFz-QOh3-KdyFRIgUaiKs2ehjrJBd9W_yW13aP4xgRQovNCEAviajCLWFTTVrjs/s2469/Eclipsa%20Video%20V01%20White_Meta.png"><div><i>Posted by Tibian Elsheikh, Product Manager, Android Core Graphics and Jeffrey Jose, Product Manager, Android Core Graphics</i></div><div class="separator"><a href="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEg0slfG8CUVGmPiAUHIXkeIVZGveJMOvf1TorUdONiRYV1THM80OzIIjGV5-bOboEhNz7FB4sTYx72ySEjFhQ4oW97-sLZ4scOX2Sb5BBU9qPMvOXvq2XRj098K7ElBnvy4k68jKELpDZ7vd4NIs2Hud2w14re18dOx7dksdFXRBR_Nd8yOiBrw8cLr_kM/s8583/Eclipsa%20Video%20V02_Blog.png"><img border="0" data-original-height="2601" data-original-width="8583" src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEg0slfG8CUVGmPiAUHIXkeIVZGveJMOvf1TorUdONiRYV1THM80OzIIjGV5-bOboEhNz7FB4sTYx72ySEjFhQ4oW97-sLZ4scOX2Sb5BBU9qPMvOXvq2XRj098K7ElBnvy4k68jKELpDZ7vd4NIs2Hud2w14re18dOx7dksdFXRBR_Nd8yOiBrw8cLr_kM/s1600/Eclipsa%20Video%20V02_Blog.png"></a></div><br><p><br></p>
<p>We’ve all been there: You’re scrolling through your favorite social media feed in a dim room, and suddenly an HDR video pops up. It’s so intensely bright that you have to squint, or maybe you find yourself turning down your screen brightness just to read the caption. Other times, a video that looks vibrant on your phone looks flat, dark, or washed out when you watch it on your living room TV. </p><p>While High Dynamic Range (HDR) technology was designed to make videos look richer and more lifelike, the lack of unified industry guidelines means that the exact same clip can render in unexpected and jarring ways depending on the display you’re using.</p>

<p>To solve this, we’re introducing Eclipsa Video—a new standard built to make your favorite videos look consistent, balanced, and comfortable on every screen. Eclipsa Video builds on the open <a href="https://github.com/SMPTE/st2094-50">SMPTE ST 2094-50 specification</a>, which Google developed in collaboration with Apple and NBCUniversal.</p><br><p></p><i><div><i><div class="separator"><a href="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEiLDY0gLjHQYTZZfRzikfPu8P3jZkXhq6Wqo1GFj3CvBh9YaboIDUstPcnV94Qan8nVkXXBlXLm5vSktLM_q9DJIIn_jyeW9LyZchI5Fpm6AD7A5XD3ZRslzBFhJLAvRj589ukW0etBNCg7004SjySw_SYsGkg6dQ8AtgfofOZeFTx8R3H7xWfwAuA-Rqc/s1066/Eclipsa_9-16_Transparent%20(2).gif"><img border="0" data-original-height="1066" data-original-width="600" height="400" src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEiLDY0gLjHQYTZZfRzikfPu8P3jZkXhq6Wqo1GFj3CvBh9YaboIDUstPcnV94Qan8nVkXXBlXLm5vSktLM_q9DJIIn_jyeW9LyZchI5Fpm6AD7A5XD3ZRslzBFhJLAvRj589ukW0etBNCg7004SjySw_SYsGkg6dQ8AtgfofOZeFTx8R3H7xWfwAuA-Rqc/w225-h400/Eclipsa_9-16_Transparent%20(2).gif" width="225"></a></div>Sudden brightness spikes during feed scrolling—fixed with Eclipsa Video.</i></div></i><p></p>

<h3><strong><span>More consistency, comfort, and creative control</span></strong></h3>Eclipsa Video moves past individual display guesswork. Instead of leaving it up to your device to interpret a video’s brightness on its own, our format carries precise guidelines that tell compatible displays exactly how to render the image. <br><p>Designed to scale with your hardware, Eclipsa Video provides three core benefits:</p>

<ul>
    <li><strong>A consistent baseline:</strong> Eclipsa Video introduces a shared rulebook for screens. It establishes a consistent benchmark for normal brightness—known as the <b>HDR reference white</b>. This ensures standard text, app interfaces, and standard-range colors remain vibrant and readable without causing uncomfortable screen glare.</li>
    <li><strong>Adaptive headroom:</strong> Screens have different physical brightness limits, or "headroom." Eclipsa Video guides how displays handle highlights dynamically. Bright details remain brilliant on a premium television, while being scaled intelligently on a mobile screen to prevent sudden blinding transitions.</li>
    <li><strong>Preserved creative intent:</strong> Rather than applying a single static setting to an entire video, Eclipsa Video carries adaptive, frame-by-frame instructions. Think of it as a set of digital notes from the creator traveling with the video, ensuring the exact colors, contrast, and mood they graded are preserved on your display.</li></ul>

<div class="separator"><img border="0" data-original-height="1080" data-original-width="2200" src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEirgS5TsogRUxWbypUiFlWIRuL8nQhdagvc7UHVFjoDG00SjqSrMniKFEys-EzgcrHKi6Am5BrtALEs7px1oaaJ5ciaO7hP0_49i8RuD7uCckjW7jYWrSoFkDlob6dJhL42MPLiBQqAjaPMOMJDEjZjDgvVe0P28fw13RlMNSiMEAlx5XFXCr8o6L8SRo0/s1600/Eclpsa%20Blog%20post%20image-AlphaB.png"><br><div><i>Eclipsa Video preserves true highlight detail on any screen you watch.</i></div></div><h3><strong><span>Built natively into Android 17</span></strong></h3>

<p>Starting with Android 17, support for Eclipsa Video is built directly into the platform. This means a more comfortable, true-to-life HDR experience is coming natively to the phones, tablets, and TVs you rely on every day. The video you capture carries its creative intent with it, and the video you watch is shown exactly the way it was meant to be seen.</p>

<h3><strong><span>Guidelines for developers &amp; creators</span></strong></h3>

<p>We’re inviting the developer and creator ecosystem to help build a more reliable HDR environment:</p>

<ul>
    <li><strong>Get started with implementation:</strong> Learn how to configure playback and capture in your apps with our <a href="https://developer.android.com/media/platform/integrate-eclipsa-video">official guide</a>.</li>
    <li><strong>ExoPlayer &amp; Media3 integration:</strong> Standard playback handling built directly into <a href="https://developer.android.com/media/media3/exoplayer">Jetpack Media3,</a> allowing ExoPlayer to support Eclipsa Video metadata automatically with no additional player configuration.</li>
    <li><strong>Explore open source tools:</strong> View and inspect <a href="https://github.com/SMPTE/st2094-50">SMPTE ST 2094-50</a> metadata and dynamic gain curves in real time using <a href="https://webmproject.github.io/hdr-explorer/">HDR Explorer</a>.</li>
</ul>

<h3><strong><span>What’s next</span></strong></h3>

<p>Eclipsa Video is rolling out now, and you’ll see more apps and devices supporting it over time. Because it’s an open standard, any app developer or hardware manufacturer can integrate it to elevate the viewing experience.</p>

<p>Try out the new tools in Android 17, explore the open-source metadata, and let us know what you think on our developer channels. We can’t wait to see what you create.</p>

<h3><strong><span>Notes &amp; Availability</span></strong></h3>
<p><strong>1. Device Compatibility:</strong> Eclipsa Video playback and capture are supported natively on devices running Android 17 (API level 37) and above with HDR displays passing Eclipsa Compliance tests.</p>
<p><strong>2. Developer Resources:</strong> The <a href="https://github.com/SMPTE/st2094-50">SMPTE ST 2094-50 Specification</a> is openly accessible for technical evaluation.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The Rust Programming Language Blog: The many journeys of learning Rust]]></title>
<description><![CDATA[This is another post in our series covering what we learned through the Vision Doc process. We previously described the overall approach and what we learned about doing user research, we explored what people love about Rust, dug into what it takes to ship safety-crticial Rust, and described some ...]]></description>
<link>https://tsecurity.de/de/3693289/tools/the-rust-programming-language-blog-the-many-journeys-of-learning-rust/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3693289/tools/the-rust-programming-language-blog-the-many-journeys-of-learning-rust/</guid>
<pubDate>Sat, 25 Jul 2026 08:37:24 +0200</pubDate>
<category>💾  Tools</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><em>This is another post in our series covering what we learned through the Vision Doc process. We previously <a href="https://blog.rust-lang.org/2025/12/03/lessons-learned-from-the-rust-vision-doc-process/" rel="external">described the overall approach and what we learned about doing user research</a>, we <a href="https://blog.rust-lang.org/2025/12/19/what-do-people-love-about-rust/" rel="external">explored what people love about Rust</a>, <a href="https://blog.rust-lang.org/2026/01/14/what-does-it-take-to-ship-rust-in-safety-critical/" rel="external">dug into what it takes to ship safety-crticial Rust</a>, and <a href="https://blog.rust-lang.org/2026/03/20/rust-challenges/" rel="external">described some of the major challenges that people face when using Rust</a>.</em></p>
<p>In this post we walk through what folks have found on their journey to learn the Rust programming language with ups and downs covered.</p>
<p>As a disclaimer, LLMs (Large Language Models) come up in this post because our interviewees brought them up. We're scoping discussion to their use as a learning tool, covering research and example generation, not broader questions about AI (Artificial Intelligence) in software development.</p>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#many-paths-to-needing-rust"></a>
Many paths to needing Rust</h3>
<p>The interviews surfaced several different paths into Rust: curiosity, embedded work, job-market pressure, organizational adoption, and reassignment after a team or company chose Rust. That last path matters because many learners are not evaluating Rust from a blank slate; they are trying to become productive after Rust has already arrived in their work.</p>
<blockquote>
<p>"Funny enough, I've advocated for more niche languages than Rust in the past. Rust has pretty much stopped being as much of a niche language as it was, but it's not Java." -- Fractional CTO</p>
</blockquote>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#rust-learning-resources"></a>
Rust learning resources</h3>
<p>Likely as expected, the folks that we talked to reach for a range of resources to learn Rust. Some reach for official documentation, such as <a href="https://doc.rust-lang.org/book/" rel="external">The Rust Programming Language Book</a> and find that sufficient to build on what the compiler was already showing them.</p>
<blockquote>
<p>"I started with the official Rust documentation because there are a lot of great examples of how features like the borrow checker work." -- Software engineer at an Automotive supplier</p>
</blockquote>
<p>Others needed more passes and more formats, sometimes reaching for resources the community maintains, such as <a href="https://rustlings.rust-lang.org/" rel="external">Rustlings</a>, <a href="https://danielkeep.github.io/tlborm/book/index.html" rel="external">The Little Book of Rust Macros</a>, and <a href="https://rust-unofficial.github.io/too-many-lists/" rel="external">Learn Rust With Entirely Too Many Linked Lists</a>.</p>
<blockquote>
<p>"The first time I went through the chapter in [The Rust Programming Language] on borrow checking, I was like, what is this? I read it again, then I watched a YouTube video of someone explaining the chapter." -- Rust freelance consultant</p>
</blockquote>
<blockquote>
<p>"Rust book, Rustlings, Zero to Production in Rust, Jon Gjengset tutorials. A bunch of books. It's not a one-pass reading. Can't say how many times I've gone through it." -- Software engineer working on video streaming and storage</p>
</blockquote>
<p>These resources have brought up an entire generation of Rust programmers. But, to some, there is a perception that these resources have trouble keeping pace with the language.</p>
<blockquote>
<p>"We'd like to use [The Rust Programming Language/'the book'], but we've found that it's out of date, unfortunately. We've looked at the GitHub repo and found it's got a lot of unresolved issues and unmerged PRs" -- Principal Software Engineering work on Rust adoption in a regulated industry</p>
</blockquote>
<p>Whether or not this is factually true, Rust's growth has nonetheless put more scrutiny on these materials. Companies evaluating adoption and engineers getting reassigned to Rust teams are looking at them with fresh eyes and finding the gaps that affect their own evaluation.</p>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#beginner-stumblings-and-unlearning-habits"></a>
Beginner stumblings and unlearning habits</h3>
<p>It's pretty typical for Rust to be the 2nd, 3rd or Nth programming language that someone picks up. They'd end up writing their most familiar language in Rust, whether C++ patterns, Java patterns, or whatever they knew, for months or even years. Eventually they got comfortable enough to start writing idiomatic Rust.</p>
<blockquote>
<p>"There's a bit of a drop in productivity compared to C if you're already familiar with it just because you're learning new rules, new syntax."  -- Principal Firmware Engineer (mobile robotics)</p>
</blockquote>
<blockquote>
<p>"In the beginning it was more poking around the code and adding and removing some ampersands and asterisks to try to make sense of <code>mut</code> and not <code>mut</code> and whatever." -- Senior engineer with 20 years of Java experience in cloud and IoT</p>
</blockquote>
<p>We also spoke with someone who found that not having much of a programming background seemed to benefit people picking up Rust. Not having worn-in grooves from other languages may play a role here, and it's worth investigating further.</p>
<blockquote>
<p>"I had someone who had never programmed much before start working on the internals of [our Rust project]. She was just fine with getting into Rust. It's more of the senior people that struggle as they need to unlearn practices which may work in other languages, but it's not the 'Rust' way." -- Researcher, Automotive OEM R&amp;D Lab</p>
</blockquote>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#learning-to-work-with-the-borrow-checker"></a>
Learning to work with the borrow checker</h3>
<p>We heard a lot about learning to work with the borrow checker instead of against it. People get there through different paths, but a few patterns came up repeatedly.</p>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#the-compiler-as-teacher"></a>
The compiler as teacher</h4>
<p>Rust's diagnostics did the teaching on their own, especially around lifetimes.</p>
<blockquote>
<p>"If you mess up the lifetimes in a piece of code that you've written by hand, I usually find that Rust's diagnostics are very helpful" -- Researcher working on static analysis of Rust programs</p>
</blockquote>
<blockquote>
<p>"Whatever's missing, the compiler usually fills in: it tells me 'you need to declare the lifetime of this reference', so I know and can figure it out. That all generally works pretty well." -- Senior Software Engineer</p>
</blockquote>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#learning-by-doing"></a>
Learning by doing</h4>
<p>Others felt like they only really internalized the borrow checker after writing a lot of Rust. It took projects, coding challenges, prototyping and so on until at some point it clicked.</p>
<blockquote>
<p>"I actually did not understand the borrow checker until I spent a lot of time writing Rust" -- Founder of a startup built on Rust</p>
</blockquote>
<blockquote>
<p>"Besides the prototyping work, I also did coding-challenge-type stuff to get familiar with Rust for Advent of Code. [..] It eventually clicked to the point where I wasn't fighting with Rust, it was working for me. I had that experience other people describe: when I managed to get my program to fit with Rust, it worked. I didn't spend time debugging." -- Principal Software Engineer, large SaaS provider</p>
</blockquote>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#letting-go-of-clone-guilt"></a>
Letting go of "clone guilt"</h4>
<p>Some learners arrive with the assumption that good Rust means zero clones, zero copies, lifetimes threaded through everything. They set the bar at optimal before they've learned how to write idiomatic Rust, and it makes the borrow checker feel harder than it needs to be at the outset.</p>
<blockquote>
<p>"On one of my first projects, I was like, 'I don't ever want to copy or clone anything,' so I carefully wove through all the lifetimes and got myself into a bit of a bind. Then I saw someone else just cloning the struct I was working with, and it was super cheap. Sometimes you can just clone and it's going to be okay." -- Researcher at a university</p>
</blockquote>
<p>The experienced Rust developers we spoke with consistently said the same thing: clone freely while you're learning, then optimize when you understand the problem. Rust's reputation for performance and correctness feeds this. Newcomers assume anything less than optimal is wrong before they've written a first working program, and clone guilt is how that shows up.</p>
<p>We think it could be an interesting area of future study to check into the patterns Rust programmers employ at different levels of experience and under which circumstances. One member of the Rust Vision doc team that's very experienced with Rust noted that there's kind of an "expected shape" they understand as passing the compiler. This knowledge influences how they approach writing code which wouldn't take that shape and they naturally find themselves understanding when to use so-called workarounds, such as passing around indices into arrays or <code>Vec</code>s.</p>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#multi-paradigm-but-not-the-oop-some-are-used-to"></a>
Multi-paradigm, but not the OOP some are used to</h3>
<p>The Rust programming language is multi-paradigm, and how that lands depends on what you're coming from. We heard some that came from a functional background were delighted with digging into learning how much Rust inherits from that lineage. Some others noted that they and others on their teams struggled to unlearn the object-oriented style they'd come to use heavily in other languages like C++ and Java.</p>
<blockquote>
<p>"Developers coming from C++ tend to think object-oriented. I think that's a difference between C++ and Rust." -- Architect at Automotive OEM</p>
</blockquote>
<blockquote>
<p>"I had exactly that thing, where I would apply all my years of Java and JS thinking, where I could just create some object, not care about it, return it, have it sloshing around between various functions. Found myself reaching for these patterns and then being told 'no, you cannot do that'." -- Principal Engineer at a SaaS company</p>
</blockquote>
<p>Developers coming from functional programming had less to unlearn: strong typing, pattern matching, and an expression-oriented style were already familiar.</p>
<blockquote>
<p>"My background has been more functional programming, strong typing. That originated for me as a Lisper: once a Lisper, always a Lisper." -- Principal Software Engineer working on Rust tooling for safety-regulated industries</p>
</blockquote>
<blockquote>
<p>"The languages I primarily used before Rust were things like OCaml. Way back, I came from C and C++, the classic languages, and then I spent quite a long time doing primarily pure functional stuff. These days I've ended up back in what I like to think of as a pragmatic center ground [with Rust]." -- Fractional CTO</p>
</blockquote>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#teaching-rust-in-academia"></a>
Teaching Rust in academia</h3>
<p>We spoke with a university professor that's been teaching Rust generally. In the academic environment, they were able to use proxies for some things such as "traits are like interfaces in Java" because the students had already gone through a set of courses in their first and second years that taught them Java. They introduced concepts slowly throughout the course, choosing to deal with some more complex topics like generics later. The outcome generally was that students had no problem picking up Rust in this setting.</p>
<blockquote>
<p>"I couldn't see any big difference on the embedded side. We also teach an embedded class, and we did an experiment. Half of the students' feedback was worse on the Rust class, mostly because they needed to build the project themselves. The C students just got one from [an LLM], absolutely no problem." -- University Professor, on teaching Rust</p>
</blockquote>
<p>The C cohort leaned on LLMs for the project in ways the Rust cohort couldn't. We don't yet have a clear answer for why.</p>
<p>What did come through clearly was the Rust cohort's experience with the community. Some students needed to figure out which drivers to use for the embedded project and how to use them. Their professor encouraged them to open issues and ask questions directly on GitHub, and the maintainers responded. Students who had never contributed to open source before were getting answers from the people who wrote the code.</p>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#learning-using-llms"></a>
Learning using LLMs</h3>
<p>Some experienced folks shared that they saw LLMs as a tool that can help someone come up to speed quickly, either as a research tool or for generating example Rust code to understand concepts.</p>
<blockquote>
<p>"I'm optimistic that there's a way to work [LLMs] in that will cut down that learning curve. One of the big things these tools bring is reducing the learning curve in general; these are very good tools to help you navigate a space that you don't know yet." -- Maintainer of large open source Rust crate</p>
</blockquote>
<blockquote>
<p>"I try [LLMs] out once a month, usually for generating an example or something like this. Just like with Stack Overflow: when you read an example, you should read it carefully and try to understand it. Not copy and paste it, but type it in your own words in code and then check it, because that's where the teeny tiny little mistakes are." -- Founder of startup built on Rust</p>
</blockquote>
<p>For some learners, an LLM is just another way to find answers, no different than a search engine.</p>
<blockquote>
<p>"So for the most part, picking up Rust - how do I learn? I'll [use web search for] things, I'll ask [an LLM], I'll just poke around and read the code." -- Senior Software Engineer working in a regulated space</p>
</blockquote>
<p>One founder went further and claimed that LLMs change who can become a Rust developer. One consulting company founder described hiring high school graduates with no systems programming background and training them as Rust developers, with LLMs filling in the learning gaps that would previously have required years of experience.</p>
<blockquote>
<p>"At the beginning, I was worried, but now that we have [LLMs] supporting development, the difficulty of the language doesn't matter. I'm seeing a huge opportunity behind strong runtime languages like Rust. [..] In [Developing Country] we hire 20-25 high school graduates, train them to be Rust programmers, then they enhance our workforce worldwide." -- Founder of a consulting company</p>
</blockquote>
<p>We heard this from one organization. This is a claim that the combination of Rust's compiler and LLM tooling can dramatically shorten the path from beginner to working developer. Whether it generalizes depends on questions we can't answer from a single interview: how long these developers stay, what kind of code they can maintain independently, and whether this training/learning model works outside this company's particular structure. If it holds up, the pool of people who can become Rust developers is much larger than the usual hiring profile suggests.</p>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#organizational-considerations-for-rust-learners"></a>
Organizational considerations for Rust learners</h3>
<p>We spoke with a number of folks on teams that are using Rust in larger organizations. Teams wanted to know that everyone would end up at roughly the same level of competence, which led a good number to invest in training courses to get there. Some leaders found that staff was able to ramp well enough by reading The Rust Programming Language, going through Rustlings, and then picking up lower risk and priority tickets to work on. Having a sense of community was also important within companies; it helps people know they are not alone when they are asked to work on Rust after, say, a reorganization happens.</p>
<blockquote>
<p>"[..] the idea with the class as opposed to 'just read the Rust book on your own' was that this gives everyone kind of the same baseline going in."  -- Principal Firmware Engineer (mobile robotics)</p>
</blockquote>
<blockquote>
<p>"So typically we're going to have people work through Rustlings, work through The Rust Programming Language. We have them then start to pick up lower risk tickets to work on." -- Principal Engineer at a large SaaS provider</p>
</blockquote>
<blockquote>
<p>"We've got an internal Slack channel for Rust learning where people can drop questions and others will come in and answer them. That helps build up understanding and community." -- Software Engineer at a large corporation</p>
</blockquote>
<p>Some organizations found that while the person they'd hire would need to learn Rust, it was still preferable to the alternative of hiring someone for a critical piece of software written in another language.</p>
<blockquote>
<p>"They needed to grow and maintain this C++ codebase. They had a C++ wizard, and they tried for about two years to find someone with the same level of expertise. They ended up hiring people that didn't know Rust and ramping them up, creating FFI bindings from the C++ side so they could work in Rust. And you can feel it: the borrow checker is teaching these people the right way to handle their systems." -- Principal Engineer at an Automotive OEM</p>
</blockquote>
<p>The community and helping each other aspect seems to grow bonds as organizations mature.</p>
<blockquote>
<p>"Our team is [all about] mentorship. I've mentored people coming up to speed on Rust, and people help each other hugely." -- Principal Software Engineer at a large SaaS company</p>
</blockquote>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#silent-attrition"></a>
Silent attrition</h3>
<p>We identified some cases where people have approached Rust and bounced off of it, for one reason or another. In the below case, someone with a background in a language with fewer guardrails found themselves frustrated enough with Rust to walk away.</p>
<blockquote>
<p>"All of that means that that embedded ecosystem is very frustrating to somebody who comes from C and is like, why can't I just get a pointer to this peripheral and then write into the registers. What are you doing to me? [..] My friend never got over that. He looked at it and said, I'm not going to deal with this and walked away." -– A second University Professor</p>
</blockquote>
<p>There may be language features that for a particular domain are not seen as comfortable or usable yet, such as async Rust usage in a safety domain. We'd like to map which language features feel off-limits in which domains; async in safety-critical work probably isn't the only case.</p>
<blockquote>
<p>"We're not fully sure how async [Rust] will work out in the long run in our domain. [..] People don't feel comfortable yet since C++14 doesn't provide such concepts. [..] It's the chicken-and-egg problem again: we probably need to gain some experience to see whether we can actually benefit from these new concepts in the automotive and safety domains." -- Team Lead at Automotive Supplier (ASIL D target)</p>
</blockquote>
<p>We heard in at least one case, that while the language was challenging and there was a near bounce, the tooling helped keep them coming back and trying.</p>
<blockquote>
<p>"Well, I think my early impressions of Rust - one is I find C++ so intimidating, and I think a big part of why I was able to succeed at [..] learning Rust is the tooling. I mean, all this makes sense [..] but it's like, for me, getting started with Rust, the language was challenging, but the tooling was incredibly easy." -- Founder of another startup built on Rust</p>
</blockquote>
<p>While it might be considered more of a community concern, if there are interactions online and in spaces that point to learners having
so-called "skill issues" this feeds into the narrative that Rust must be hard to learn. We may be unintentionally turning away Rust Project contributors and maintainers due to the vibes being put out when new learners show up in certain spaces.</p>
<blockquote>
<p>"People are very helpful, but generally the attitude is: if your program is very complicated, it's mostly a skill issue. There's not that much empathy when people get stuck learning, and a lot of people are just pushed away by it. There's probably a huge number of people who silently stop wanting to write Rust, because at some point it gets complicated and the feedback they get is 'you just need to be a better programmer, obviously'." -- Software Engineer at a SaaS Provider</p>
</blockquote>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#feedback-on-near-bounces-from-survey"></a>
Feedback on near-bounces from survey</h4>
<p>We found a few interesting perspectives collected in the Rust Vision doc survey which we administered with examples of bouncing and coming back:</p>
<blockquote>
<p>"I started before 1.0, got stuck very soon when trying to translate patterns from C++ to Rust (due to borrow checking). I tried again after 1.0 and it stuck. [..]" -- Survey Respondent A</p>
</blockquote>
<p>Survey Respondent A went on to share in a more detailed response about a perceived weakness in Rust learning materials related to lifetimes and the borrow checker are explained. There was an observation that it's fairly easy to run into more complex situations with lifetimes and the borrow checker. They felt that the current state of this sort of material and tutorials is fairly superficial and can leave learners stuck when they run into those more complex situations.</p>
<p>One respondent that bounced once and came back shared challenges around usage of async. In concert with Rust's memory-safety and the borrow checker, they found some of the nitty-gritty details of async were difficult to learn. While we're aware of the Rust Project's continuous efforts to improve Rust's async story, this is another data point of a user that faced challenges.</p>
<p>Another survey respondent shared how they had multiple times bounced in trying to learn Rust. They returned after a year or so and found Rustlings to be highly motivating. We note that having multiple pathways for folks to learn Rust opens up more possibilities for those that nearly bounced, just like this person.</p>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#need-more-focused-work-on-silent-attritrion"></a>
Need more focused work on silent attritrion</h4>
<p>The thing that stood out most to us was the lack of real, first-hand knowledge of having bounced when learning Rust. While this is an obvious effect of soliciting answers to our survey and opportunities to interview through Rust channels and our networks, this cohort is good future candidate where interviews could start.</p>
<h3><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#conclusions"></a>
Conclusions</h3>
<p>Across these conversations, the experience of learning Rust depended heavily on context. Why someone was learning and what support they had mattered as much as the borrow checker. The same kinds of examples kept coming up: a training course that got a team to a shared baseline, a maintainer answering a student's first GitHub issue, and a colleague whose code showed that cloning was okay.</p>
<p>That context is largely something the community has a hand in. With that in mind, here is what we take away from what we heard, and what we still don't know.</p>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#what-seems-worth-trying"></a>
What seems worth trying</h4>
<p><strong>Learning materials aimed at unlearning.</strong> Syntax barely came up when people described their struggles. People struggled with unlearning habits from previous languages, whether OOP structuring from C++ and Java or the instinct to grab a raw pointer to a peripheral. Most of our learning materials teach Rust from first principles, and that works. What we didn't come across is much written for, say, the engineer with ten years of Java who lands on a Rust team after a reorg: material that names the patterns they'll reach for that won't transfer, and shows what to do instead. The professor we spoke with did a version of this in the classroom, leaning on "traits are like interfaces in Java" and saving generics for later in the course, and the students did fine. Something similar could work outside the classroom too.</p>
<p><strong>Put the "clone freely while you're learning" advice somewhere official.</strong> Every experienced developer we spoke with gave the same advice, but learners seem to mostly pick it up by accident, like the researcher who happened to see someone else cloning the struct they had been carefully threading lifetimes through. Saying it early in official materials would take some of the steepness out of the curve. The broader version belongs there too: idiomatic Rust doesn't have to mean optimal Rust, especially on a first project.</p>
<p><strong>Diagnostics are already a primary learning resource: several people told us the compiler taught them lifetimes before any documentation did.</strong> Diagnostics reach learners right at the moment they're stuck. When writing new ones, it seems worth keeping the confused newcomer in mind alongside the expert, because for a lot of people this is where the learning happens.</p>
<p><strong>Is "the book" actually out of date?</strong> Whether or not The Rust Programming Language or other materials are actually behind, a team evaluating Rust looked at its repository, saw unresolved issues and unmerged PRs, and moved on. As more companies evaluate adoption, more people will look at these materials with the same fresh eyes. Visible issue triage and some communication about what's current and what's planned would address the perception, separately from whatever content work may or may not be needed.</p>
<p><strong>How stuck learners get treated is shaping who stays.</strong> We heard about students getting answers on GitHub from the maintainers who wrote the code, and we heard about learners being told their struggles were a skill issue. The first group came away with a lasting good impression of Rust. Some of the second group walked away entirely, and because they leave quietly, it's easy to underestimate how many of them there are. The welcoming side of the community came up unprompted as a reason people stayed, so we know it makes a difference when we get this right.</p>
<p><strong>Every organization we spoke with described essentially the same ramp-up for bringing a team to Rust.</strong> Teams that brought groups of developers to Rust described roughly the same approach: get everyone to a shared baseline with a training course or with The Rust Programming Language and Rustlings, start people on lower-risk tickets, and give them somewhere internal to ask questions. Several organizations also found that hiring developers without Rust experience and ramping them up worked out better than continuing to search for rare expertise in another language. None of this is complicated, and teams weighing adoption don't need to invent a training program from scratch.</p>
<h4><a class="anchor" href="https://blog.rust-lang.org/2026/06/25/vision-doc-journeys-to-learning-rust/#what-we-still-don-t-know"></a>
What we still don't know</h4>
<p>The biggest gap is the people we didn't reach. Nearly everyone we spoke with stuck with Rust long enough to be reachable through Rust channels, so the stories of bouncing off came to us second-hand: a friend who walked away from embedded Rust, colleagues who quietly stopped after the responses they got. As we wrote in <a href="https://blog.rust-lang.org/2025/12/03/lessons-learned-from-the-rust-vision-doc-process/" rel="external">our first post</a>, finding people who decided against Rust takes targeted outreach. If the proposed User Research team comes together, talking with learners who bounced would make a good early project, and learning is probably the area where that research would teach us the most.</p>
<p>We also don't know what to make of LLMs as a learning tool yet. They came up as a search engine, as an example generator, and in one organization's case as something that makes training high school graduates into working Rust developers possible. We saw a classroom where the C cohort leaned on LLMs in ways the Rust cohort couldn't, and we don't have an explanation for it. All of this comes from a handful of conversations, so we treat it as a set of leads to follow up on. Given how quickly the tools are changing, it seems better to study this deliberately than to wait and see what folklore develops.</p>
<p>The folks we spoke with showed that people do get there: with enough passes through the materials and enough code written, it eventually clicks. The opportunities above are mostly about making it work for the people who didn't pick Rust on purpose, and for the ones who would have stuck around if their early experience had gone a little differently.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[AI agent went rogue and hacked startup by itself, OpenAI reveals]]></title>
<description><![CDATA[Company behind ChatGPT says agent ‘cheated’ an evaluation by attacking a Hugging Face database OpenAI has revealed that an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an “unprecedented incident”.The comp...]]></description>
<link>https://tsecurity.de/de/3693138/it-nachrichten/ai-agent-went-rogue-and-hacked-startup-by-itself-openai-reveals/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3693138/it-nachrichten/ai-agent-went-rogue-and-hacked-startup-by-itself-openai-reveals/</guid>
<pubDate>Sat, 25 Jul 2026 07:07:02 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Company behind ChatGPT says agent ‘cheated’ an evaluation by attacking a Hugging Face database </p><p>OpenAI has revealed that an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an “unprecedented incident”.</p><p>The company behind ChatGPT said the startup Hugging Face had detected and contained the agent – an AI tool designed to carry out tasks without human assistance – which had entered its systems.</p> <a href="https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incident">Continue reading...</a>]]></content:encoded>
</item>
<item>
<title><![CDATA[VentureBeat Research: Where enterprise AI agent governance hasn't caught up]]></title>
<description><![CDATA[Enterprises deployed AI agents ahead of the controls needed to manage them — and they did it knowingly. That is the central finding across the five parallel surveys VentureBeat Research fielded in June, spanning every layer of the agentic stack. Now those enterprises are retrofitting to catch up ...]]></description>
<link>https://tsecurity.de/de/3692498/it-nachrichten/venturebeat-research-where-enterprise-ai-agent-governance-hasnt-caught-up/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3692498/it-nachrichten/venturebeat-research-where-enterprise-ai-agent-governance-hasnt-caught-up/</guid>
<pubDate>Fri, 24 Jul 2026 22:51:31 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Enterprises deployed AI agents ahead of the controls needed to manage them — and they did it knowingly. That is the central finding across the five parallel surveys VentureBeat Research fielded in June, spanning every layer of the agentic stack. Now those enterprises are retrofitting to catch up with their own standards, and they are budgeting for it: In each of the five control layers we measured, 57 to 68% of enterprises plan to switch vendors or add new ones within 12 months, and roughly a third, depending on the layer, plan to move within the quarter.</p><p><a href="https://venturebeat.com/category/resources">VentureBeat Research</a> measured the five controls an enterprise has to build before it can trust an agent: identity, evaluation, cost telemetry, the context layer, and orchestration. Identity governs which agent is allowed to do what, under whose credentials. Evaluation determines whether the agent's work is any good. Cost telemetry tracks what each agent costs to run. The context layer supplies the business data and definitions agents draw on when they answer. And the orchestration control plane coordinates multi-step agent work. Each of our five reports measures one of those controls.</p><p><b>Most deployed "agents" are chatbots wearing the label.</b> Seventy-one percent of enterprises said a quarter or fewer of their deployed "agents" can complete multi-step work on their own; only 10% said true agents are the majority of what they run. These respondents are positioned to know: 81% recommend or decide AI purchases at their companies. A single-prompt chatbot with a human reading every answer needs none of the controls the other four reports measure. A true multi-step agent needs all of them — and most enterprises can't say which one they've deployed. <i>(Full findings: </i><a href="https://venturebeat.com/resources/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents"><i>Agentic Orchestration report.</i></a><i>)</i></p><p><b>Autonomy is outrunning trust in the evaluations that gate it.</b> Two-thirds of enterprises either already allow an agent to push a code or system change to production on automated evaluation results alone, with no human review, or are actively engineering toward that within 12 months. Only 5% fully trust the evaluations that would make that call — and half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year. Before removing human review from any workflow, test evaluations against production outcomes rather than internal benchmarks. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway"><i>Agent Reliability &amp; Evals report</i></a><i>.)</i></p><p><b>Companies that let agents share credentials get hit more often.</b> Sixty-nine percent of companies let at least some of their agents share credentials — multiple agents operating under one API key or service account. Organizations that allow credential sharing anywhere experienced a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (nine of 22) at companies where every agent has its own scoped identity. The fix is scoped identity for every agent, starting with the ones that touch production systems. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials"><i>Agentic Security &amp; Identity report</i></a><i>.)</i></p><p><b>The most expensive hardware in the building runs at half capacity or less.</b> More than eight in 10 enterprises that run their own GPUs reported utilization of 50% or less, and only 44% rigorously track what their AI compute actually costs and returns. The number worth chasing first isn't more GPUs — it's the utilization and per-workload cost of the ones already running. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs"><i>AI Infrastructure &amp; Compute report</i></a><i>.)</i></p><p><b>Agents answer confidently from data nobody governs.</b> Fifty-seven percent of enterprises traced a confident, wrong agent answer in the past six months to their own missing or inconsistent business context — wrong metrics, stale definitions, absent documents — and most saw it happen more than once. Governing the definitions agents answer from — metrics and entities first — has to come before scaling the agents that depend on them. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-ai-context-gap-enterprise-ai-organizations-have-a-trust-problem-not-a-retrieval-problem-and-most-are-still-building-the-fix"><i>Context Layers / RAG report</i></a><i>.)</i></p><p>No layer has an entrenched incumbent: The defaults today are the built-in tools that ship with the big AI platforms enterprises already use. Switching intent runs highest in orchestration itself, where 68% plan to adopt, add, or replace platforms within 12 months and 34% within the quarter. Our surveys did not ask which direction that money moves — toward the platforms' built-in tools or toward the specialists challenging them — and that open question is the next four quarters of this market.</p><hr><p><b>About this research</b> </p><p><a href="https://venturebeat.com/category/resources">VentureBeat Research</a> fielded five parallel surveys in June 2026 under its VB Pulse program: Agentic Orchestration (101 respondents), Agent Reliability &amp; Evals (157), Agentic Security &amp; Identity (107), AI Infrastructure &amp; Compute (107), and Context Layers / RAG (101) — 573 qualified respondents in total, all at organizations with 100 or more employees. Samples are self-selected, and some findings should be read directionally; each report carries its full methodology note. What the pattern supports more strongly than any single percentage is the direction: every survey, independently, points the same way. VentureBeat produces both this research and <a href="https://venturebeat.com/vbtransform2026">VB Transform</a>, the conference where these reports debuted.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Discloses AI-Driven Breach During Cybersecurity Testing]]></title>
<description><![CDATA[An internal cybersecurity evaluation to evaluate Hugging Face’s offensive cyber capabilities allowed two of the company’s advanced AI models to hack into the organization’s infrastructure autonomously. These models include GPT-5.6 Sol and a more advanced pre-release model.  One of the…
Read more ...]]></description>
<link>https://tsecurity.de/de/3692255/it-security-nachrichten/openai-discloses-ai-driven-breach-during-cybersecurity-testing/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3692255/it-security-nachrichten/openai-discloses-ai-driven-breach-during-cybersecurity-testing/</guid>
<pubDate>Fri, 24 Jul 2026 20:17:32 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>An internal cybersecurity evaluation to evaluate Hugging Face’s offensive cyber capabilities allowed two of the company’s advanced AI models to hack into the organization’s infrastructure autonomously. These models include GPT-5.6 Sol and a more advanced pre-release model.  One of the…</p>
<p class="more-link-p"><a class="more-link" href="https://www.itsecuritynews.info/openai-discloses-ai-driven-breach-during-cybersecurity-testing/">Read more →</a></p>
<p>The post <a href="https://www.itsecuritynews.info/openai-discloses-ai-driven-breach-during-cybersecurity-testing/">OpenAI Discloses AI-Driven Breach During Cybersecurity Testing</a> appeared first on <a href="https://www.itsecuritynews.info/">IT Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows]]></title>
<description><![CDATA[Anthropic released Claude Opus 5 on Friday, a model the company says delivers nearly all the intelligence of its top-of-the-line Claude Fable 5 at half the cost — a launch that signals how the AI race is shifting from raw capability to the economics of daily use.The model, available immediately o...]]></description>
<link>https://tsecurity.de/de/3692246/it-nachrichten/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3692246/it-nachrichten/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows/</guid>
<pubDate>Fri, 24 Jul 2026 20:10:11 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="https://www.anthropic.com/">Anthropic</a> released Claude <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> on Friday, a model the company says delivers nearly all the intelligence of its top-of-the-line Claude <a href="https://www.anthropic.com/claude/fable">Fable 5</a> at half the cost — a launch that signals how the AI race is shifting from raw capability to the economics of daily use.</p><p>The model, available immediately on all of Anthropic's platforms, is priced at $5 per million input tokens and $25 per million output tokens, unchanged from its predecessor, <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a>. It becomes the new default model on <a href="https://support.claude.com/en/articles/11049741-what-is-the-max-plan">Claude Max</a>, Anthropic's premium consumer tier, and the strongest model available on <a href="https://support.claude.com/en/articles/8325606-what-is-the-pro-plan">Claude Pro</a>.</p><p>The positioning is deliberate. Anthropic is not claiming <a href="http://anthropic.com/news/claude-opus-5">Opus 5 </a>is its smartest model — that distinction still belongs to <a href="https://www.anthropic.com/claude/fable">Fable 5</a>, and rival systems retain an edge in certain domains. Instead, the company is making a subtler argument that may matter more to enterprise buyers: that the most economically important AI work happens in a middle band of difficulty, where near-frontier intelligence delivered efficiently and cheaply beats frontier intelligence delivered expensively.</p><p>"Opus 5 as your daily driver, the model you hand complex work to and review when it's done," an Anthropic spokesperson said in an interview with VentureBeat, describing how the company's lineup now stratifies. "Fable 5 for your most ambitious work, the days-long autonomous projects nothing could take on before... Sonnet 5 for work you run at scale, where speed and cost per call decide what ships. Haiku 4.5 for subagents and instant answers."</p><h2><b>How Claude Opus 5 benchmark results stack up against Fable 5 and rival AI models</b></h2><p>On paper, the results are striking. Anthropic says <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> sets new state-of-the-art marks on coding and knowledge-work evaluations including <a href="https://www.frontierbench.ai/announcement">Frontier-Bench</a> and <a href="https://artificialanalysis.ai/evaluations/gdpval-aa">GDPval-AA</a>. On <a href="https://www.frontierbench.ai/announcement">Frontier-Bench v0.1</a>, an agentic terminal coding benchmark, Opus 5 scores 43.3 percent — more than double Opus 4.8's 18.7 percent and well ahead of Fable 5's 33.7 percent — at a lower cost per task, according to the company. On <a href="https://arcprize.org/arc-agi/3">ARC-AGI 3</a>, an evaluation of novel problem-solving, Anthropic reports Opus 5 scored three times as high as the next best model. On <a href="https://github.com/xlang-ai/OSWorld-V2">OSWorld 2.0</a>, a computer-use benchmark, the company says the model surpasses Fable 5's best result at just over a third of the cost.</p><p>The numbers come with honest caveats that are themselves notable in an industry prone to superlatives. Anthropic acknowledges <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> remains behind <a href="https://www.anthropic.com/claude/mythos">Mythos 5</a>, a competing model, on cybersecurity tasks and biology research, and an OpenAI-family model still leads on one agentic coding benchmark.</p><p>The more revealing caveat came from Anthropic itself, when asked where <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> still falls short of <a href="https://www.anthropic.com/claude/fable">Fable 5</a>. The spokesperson's answer amounted to a candid admission about what benchmarks do and don't capture.</p><p>"The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it's strongest. What those evals don't measure is duration," the spokesperson told VentureBeat. "One way to put it: Opus 5 is the best tool for the jobs benchmarks can see, and Fable 5 is what you reach for when the job outruns the benchmark."</p><p><a href="https://www.anthropic.com/claude/fable">Fable 5</a>, by contrast, "is for the longest, most autonomous jobs, where the model has to stay coherent across many connected steps over hours or days with dense source material," the spokesperson said, advising customers to "run both on a representative workload, one bounded task and one long-horizon job." That framing — bounded tasks versus long-horizon autonomy — may become the defining axis of model differentiation in 2026, as benchmarks saturate and the hardest remaining problems involve sustained, multi-day agentic work rather than discrete puzzles.</p><h2><b>Why token efficiency is becoming the real battleground for enterprise AI spending</b></h2><p>Threaded through the launch is a theme Anthropic clearly wants buyers to absorb: <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> doesn't just score well, it scores well per dollar. The model ships with an adjustable "effort" setting that lets customers trade intelligence for speed and token savings, and Anthropic's charts emphasize performance at a given cost rather than peak performance alone.</p><p>Early customers echoed the point with unusual specificity. Harvey, the legal AI company, said Opus 5 achieved similar performance to Opus 4.8's maximum-reasoning mode "while generating 26% fewer tokens on average," according to Niko Grupen, its head of applied research. Richard Pham of Fundamental Research Lab said that on hard financial-modeling tasks, the model averaged nine percentage points higher accuracy "while using roughly one-third fewer turns and tool calls and 60% less time."</p><p>Wade Foster, chief executive of Zapier, said Opus 5 topped his company's AutomationBench leaderboard "without spending more tokens than prior Claude models," running a full churn-prevention workflow from start to finish. "Previous models didn't pass; Opus 5 hit 100%," he said. Scott Wu, chief executive of Cognition, the company behind the Devin coding agent, said that on FrontierCode 1.1, "Claude Opus 5 approaches Fable-level performance at half the cost," with particular strength in debugging and root-cause analysis.</p><p>The efficiency emphasis reflects commercial reality. Enterprise AI spending is no longer experimental, and inference costs — the price of actually running these models at scale — have become a board-level line item. </p><p>Anthropic's business skews heavily toward API and enterprise usage; according to a February 2026 analysis by <a href="https://research.contrary.com/company/anthropic">Contrary Research</a>, Claude held roughly 40 percent of the enterprise large language model market by usage as of late 2025, and Claude Code alone had reached about $1 billion in annualized revenue. For a company whose customers pay by the token, a model that does more with fewer tokens is not a nice-to-have. It is the product.</p><h2><b>Self-verifying AI agents and what they mean for the hidden costs of automation</b></h2><p>Beyond the numbers, Anthropic is selling a behavioral story: that <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> verifies its work and iterates until it succeeds. The company offered several examples from testing that read like small parables of machine stubbornness.</p><p>In one <a href="https://www.frontierbench.ai/announcement">Frontier-Bench</a> task, the model was asked to reconstruct a machine part as a 3D CAD model from a drawing it was intentionally given no way to view. Rather than fail, Anthropic says, Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels — and did so repeatedly, while no competing model solved the task in five attempts. In another case, given a real bug in a popular open-source package manager, the model found the root cause and fixed an edge case the community's own patch had missed; a competing model patched only the symptom and declared victory. An engineer at a trading firm, the company says, used Opus 5 to build a market data feed for a new exchange in a single session and, finding no live feed to validate against, watched the model build its own test harness to check its parsing code.</p><p>Customers described similar behavior in the wild. Cristian Rivera, a staff software engineer at Stripe, said he gave the model "a chief-of-staff role over my dev environments" for a weekend: "it built its own monitor, drove each box, and pulled me in only for the judgment calls."</p><p>This is the capability enterprises actually care about, and it is worth dwelling on why. The gap between a model that produces plausible output and one that verifies its output is the gap between a demo and a deployable system. Most of the hidden cost of enterprise AI today is human review — engineers checking the machine's work. A model that reliably checks its own work compresses that cost, which is precisely why customers keep citing fewer turns, fewer passes, and less time rather than higher raw scores.</p><h2><b>Inside Anthropic's safety strategy: capability gaps, classifiers, and model fallbacks</b></h2><p>The launch also showcases Anthropic's increasingly intricate approach to safety — one that now involves deliberately not teaching its models certain skills. The company says its automated behavioral audit found Opus 5 to be its most aligned model to date, scoring 2.3 on overall misaligned behavior, lower than <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a>, <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a>, or <a href="https://www.anthropic.com/claude/fable">Fable 5</a>, with the lowest rates of deceptive behavior and the least susceptibility to being tricked into misuse.</p><p>On the capability side, Anthropic says it intentionally avoided training <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> on cyber tasks, as it did with Opus 4.8. The model improved on them anyway — a side effect of general capability gains — and now nearly matches Mythos 5 at finding software vulnerabilities. But it remains far behind at exploiting them: on Anthropic's OSS-Fuzz evaluation, Opus 5 identified vulnerabilities at a 79.4 percent rate, close to Mythos 5's 80 percent, but succeeded at developing exploits in only 4 challenges versus Mythos 5's 13. That asymmetry — strong at defense-relevant discovery, weak at offense-relevant exploitation — appears to be by design, and the safeguards follow the same logic. Anthropic expects Opus 5's cyber classifiers to intervene about 85 percent less often than Fable 5's.</p><p>When a classifier does trigger, requests in <a href="http://claude.ai/">Claude.ai</a>, <a href="https://code.claude.com/docs/en/overview">Claude Code</a>, and <a href="https://claude.com/product/cowork">Claude Cowork</a> fall back to <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a> by default — raising an obvious question: if a request is too risky for one model, why is it acceptable for another? "The model it falls back to has lower capability levels making the risk of harmful use lower as well," the spokesperson said, adding that "there is a message that lets the user know when this occurs and is visible in the chat."</p><p>The logic is defensible, but it reveals how AI safety actually works in 2026: risk is not a property of the question alone, but of the question multiplied by the capability of the system answering it. On biology, the calculus runs the other way. Opus 5 is now Anthropic's most capable generally available model for scientific research — scoring 10.2 percentage points higher than Opus 4.8 on the company's internal chemistry benchmark — though the spokesperson acknowledged that "Mythos 5 remains the stronger model for long-horizon, open-ended work like autonomous drug design campaigns."</p><h2><b>The business stakes behind the launch: a $380 billion valuation and massive compute bets</b></h2><p>The launch lands at a moment of extraordinary commercial momentum — and extraordinary obligations — for Anthropic. Reuters reported in February that the company was valued at <a href="https://www.reuters.com/technology/anthropic-valued-380-billion-latest-funding-round-2026-02-12/">roughly $380 billion</a> in its latest funding round, following a period in which, per Contrary Research's analysis, its annualized revenue climbed from about $1 billion at the end of 2024 to a projected $9 billion by the end of 2025, with internal targets reportedly <a href="https://research.contrary.com/company/anthropic">reaching $20 to $26 billion for 2026</a>. Those targets are underwritten by enormous infrastructure commitments, including a <a href="https://www.anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships">reported $30 billion Azure compute deal</a> alongside arrangements with Google Cloud and Nvidia — spending that only pencils out if enterprises keep expanding usage.</p><p>That is the context in which Opus 5's pricing strategy makes sense. Holding the price at Opus 4.8 levels while roughly doubling performance on key agentic benchmarks is effectively a steep price cut per unit of capability, designed to widen the funnel of workloads that are economical to automate. Every task that was marginal at Opus 4.8's cost-per-success becomes viable at Opus 5's — and every viable task is recurring token revenue.</p><p>The regulatory backdrop has grown more complex as well. A U.S. judge gave final approval this week to <a href="https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/">Anthropic's $1.5 billion copyright settlement with book authors</a>, Reuters reported, closing a chapter of litigation over the company's early training data. And in June, Reuters, citing Axios, reported that the U.S. government had moved to <a href="https://www.reuters.com/technology/us-blocks-foreign-access-anthropics-most-advanced-ai-models-axios-reports-2026-06-13/">block foreign access </a>to Anthropic's most advanced models — a reminder that frontier AI is now entangled with export policy in ways that shape which customers can buy what.</p><p>Also shipping Friday: a Fast mode running at roughly 2.5 times default speed at twice the base price, automatic fallback routing on the API, and mid-conversation tool changes that no longer invalidate the prompt cache — a small feature that agent developers may appreciate more than any benchmark. Consistent with prior Opus models, Opus 5 carries no data retention requirements for general access, a point the spokesperson flagged unprompted for customers with "a hard zero data retention requirement." Developers can access the model as claude-opus-5 on the <a href="https://platform.claude.com/login?returnTo=%2F%3F">Claude API</a> starting today.</p><p>Two questions will determine whether the bet pays off: whether <a href="http://anthropic.com/news/claude-opus-5">Opus 5's efficiency claims </a>survive contact with production workloads at scale, and whether enterprises embrace a world where safety classifiers, not users, sometimes decide which model answers. But the deeper message of Friday's launch is that the AI industry's center of gravity has moved. For three years, the labs competed on what their best model could do on its best day. With Opus 5, Anthropic is competing on something less glamorous and far more lucrative: what a very good model can do every day, for half the price. In a market where the frontier keeps moving, Anthropic is wagering that the real fortune lies just behind it.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Choosing high assurance schemes offers better flexibility as technology requirements change]]></title>
<description><![CDATA[Author: PQShield - Bewertung: 0x - Views:80 Selecting the most rigorous certification path today provides long term benefits for product manufacturers. 

@Wei Yuan suggests that choosing high assurance schemes offers better flexibility as technology requirements change. Here’s why:
Strict schemes...]]></description>
<link>https://tsecurity.de/de/3691784/videos/choosing-high-assurance-schemes-offers-better-flexibility-as-technology-requirements-change/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3691784/videos/choosing-high-assurance-schemes-offers-better-flexibility-as-technology-requirements-change/</guid>
<pubDate>Fri, 24 Jul 2026 16:22:13 +0200</pubDate>
<category>🎥 Videos</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: PQShield - Bewertung: 0x - Views:80 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/Vqf7vBDMx2A?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>Selecting the most rigorous certification path today provides long term benefits for product manufacturers. <br />
<br />
@Wei Yuan suggests that choosing high assurance schemes offers better flexibility as technology requirements change. Here’s why:<br />
Strict schemes reduce the overall cost of ownership over time.<br />
High assurance solutions live longer within evolving regulatory timelines.<br />
Early investment prevents the need for a mass migration in just a few years.<br />
<br />
Robust compliance paths offer more certainty for international markets.<br />
Learn why high assurance evaluation saves costs over time. Listen to the full episode today!<br />
<br />
#Certification #PQC #TechStrategy #AppplusLabs<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Language Model Hallucination Evaluation with GraphEval]]></title>
<description><![CDATA[Turning the key principles and methodological stages of GraphEval into a simulated practical scenario to better understand its usefulness and key implications in understanding and combating LLM hallucinations.]]></description>
<link>https://tsecurity.de/de/3691615/ai-nachrichten/language-model-hallucination-evaluation-with-grapheval/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3691615/ai-nachrichten/language-model-hallucination-evaluation-with-grapheval/</guid>
<pubDate>Fri, 24 Jul 2026 15:05:11 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Turning the key principles and methodological stages of GraphEval into a simulated practical scenario to better understand its usefulness and key implications in understanding and combating LLM hallucinations.]]></content:encoded>
</item>
<item>
<title><![CDATA[Inside the OpenAI – Hugging Face Incident: The AI Breach With No Human Attacker Behind It]]></title>
<description><![CDATA[OpenAI’s own models broke out of a test sandbox and into Hugging Face’s servers to solve an evaluation, with no human attacker involved. The incident showed how keeping agentic AI safe now depends on how it’s contained, not just on how it’s trained.]]></description>
<link>https://tsecurity.de/de/3690378/it-security-nachrichten/inside-the-openai-hugging-face-incident-the-ai-breach-with-no-human-attacker-behind-it/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3690378/it-security-nachrichten/inside-the-openai-hugging-face-incident-the-ai-breach-with-no-human-attacker-behind-it/</guid>
<pubDate>Fri, 24 Jul 2026 00:43:13 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[OpenAI’s own models broke out of a test sandbox and into Hugging Face’s servers to solve an evaluation, with no human attacker involved. The incident showed how keeping agentic AI safe now depends on how it’s contained, not just on how it’s trained.]]></content:encoded>
</item>
<item>
<title><![CDATA[Multi-turn attacks broke AI models 88% of the time — single-turn testing missed it, Cisco AI security lead warns at VB Transform 2026]]></title>
<description><![CDATA[When Cisco ran 6,986 multi-turn attacks against 15 flagship models, attackers who adapted across the conversation broke through as often as 88.3% of the time. Amy Chang, Cisco's head of AI threat intelligence and security research, brought that finding to the agentic security panel at VB Transfor...]]></description>
<link>https://tsecurity.de/de/3690018/it-nachrichten/multi-turn-attacks-broke-ai-models-88-of-the-time-single-turn-testing-missed-it-cisco-ai-security-lead-warns-at-vb-transform-2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3690018/it-nachrichten/multi-turn-attacks-broke-ai-models-88-of-the-time-single-turn-testing-missed-it-cisco-ai-security-lead-warns-at-vb-transform-2026/</guid>
<pubDate>Thu, 23 Jul 2026 20:48:24 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>When Cisco ran 6,986 multi-turn attacks against <a href="https://blogs.cisco.com/ai/proprietary-problems">15 flagship models</a>, attackers who adapted across the conversation broke through as often as 88.3% of the time. Amy Chang, Cisco's head of AI threat intelligence and security research, brought that finding to the agentic security panel at <a href="https://venturebeat.com/vbtransform2026">VB Transform 2026</a>; the number should worry anyone still running single-turn red-teaming programs.</p><p><a href="https://venturebeat.com/resources/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials">VentureBeat's June 2026 Pulse survey of 107 enterprise respondents</a> explains why the room was full. More than half, 54%, have already had a confirmed agent security incident (18%) or a near-miss caught before harm (36%). Just 32% give every agent its own scoped, managed identity, and fewer still, 30%, isolate their highest-risk agents in sandboxes. Provider-native and hyperscaler controls remain the primary agent security layer at <a href="https://venturebeat.com/security/shared-api-keys-expose-ai-agent-fleets-venturebeat-research">82% of companies surveyed</a>. The world's largest security vendors have done the same math. </p><p>Palo Alto Networks closed its <a href="https://www.paloaltonetworks.com/company/press/2026/palo-alto-networks-completes-acquisition-of-cyberark-to-secure-the-ai-era">$25 billion acquisition of CyberArk</a> in February, CrowdStrike <a href="https://www.crowdstrike.com/en-us/press-releases/crowdstrike-to-acquire-sgnl-to-transform-identity-security-for-ai-era/">agreed in January to pay $740 million for SGNL</a>, and Cisco announced its <a href="https://blogs.cisco.com/news/cisco-announces-intent-to-acquire-astrix-security">intent to acquire Astrix Security</a> for a reported $400 million, all of it aimed at the identity and isolation layer most enterprises have not finished building.</p><div></div><p>Chang came to the panel with almost two decades of experience spanning cybersecurity operations, government, and the military. She ran global cybersecurity operations as an executive director at JPMorgan Chase, where she led the bank's cyber threat intelligence teams, and served as a senior staffer on the House Foreign Affairs Committee and as a U.S. Navy Reserve officer. She also teaches cybersecurity and emerging threats as adjunct faculty at the Middlebury Institute of International Studies.</p><p>Chang's 88.3% number comes from a study she co-authored with Nicholas Conley, built on 30,090 single-turn prompts and 6,986 multi-turn attacks against those 15 closed and proprietary flagship models. Multi-turn success rates ranged from 7.89% to 88.3%, every model tested showed non-trivial multi-turn exposure, and the two testing styles did not even rank the models in the same order. Cisco publishes adversarial evaluation signals for what is now 105 models on its <a href="https://leaderboard.aidefense.cisco.com/">LLM Security Leaderboard</a>, she told the audience.</p><p>"If you don't understand how models are susceptible to different types of attacks, then you are unable to account for how that model that is powering your agent, that is powering your application, to understand where those failure points are," Chang said. Single-turn testing is the one-shot malicious prompt, she explained, while extending an attack into a longer conversation "is more realistic of how we are actually engaging with our models, with our agents, with our applications." That longer arc surfaces harmful outputs and misaligned behaviors that a snapshot never catches.</p><p>Cisco has pushed the testing itself into agentic territory. Chang described a framework where agents assess a deployment scenario, develop relevant attacks, judge whether they are worth pursuing, execute them, and evaluate their own success. What surprised her most, after all that sophistication, was how simple the defensive answer stays. "The answer is still that it's pretty simple," she said. "You don't have to get super creative. You just need to think about truly what are the fundamentals and basics of what I'm trying to secure in my organization."</p><p>Her starting point for CISOs beginning agentic deployments is Cisco's <a href="https://blogs.cisco.com/ai/security-framework">Integrated AI Security and Safety Framework</a>, which she said "stipulates all the ways that AI can be compromised across the AI lifecycle" from modality through supply chain. From there, teams can work backward from real incidents, trace how each attack was achieved, and use the framework to build a strategy with the right coverage and mitigations.</p><p>Heather Ceylan, the CISO of Box, sees the same gap from the defender's side. "A lot of what you see out there with agent red teaming is just single-turn, and that's not how people are actually interacting with AI day-to-day," she told the audience. Box now simulates multi-turn adversaries with agents that think like an attacker and iterate attempt after attempt to hijack the target. "You have to pressure test your agents because otherwise you don't know if your execution controls are really working as you intended."</p><p>Box deployed agents inside its security operations center about a year ago, starting with human approval required for every action, and trust built quickly enough that analysts shifted into monitoring mode. Then the agent made one mistake, and every bit of that accumulated trust vanished. "They had to start all over again," she said. "So I think that that monitoring piece is so important. Even if you're not gonna have a human in the loop, things change, models change, and we can't control how the models change and interpret things."</p><p>Rajesh Parekh, VP of AI and ML at Intuit, brought the builder's perspective. Parekh led large-scale computer vision and ML systems powering Google's Maps and Geo products before joining Intuit, and holds a doctorate in computer science. </p><h2>Three layers versus an operating system</h2><p>Ceylan described Box's approach as three concentric layers. Permissioning comes first, so the agent never accesses more content than the human who invoked it. Ephemeral sandbox environments spin up for each agent task, containing the blast radius if an agent gets hijacked, and runtime execution control restricts the agent's tool calls to only those relevant to the task at hand. "If you want an agent to summarize a doc for you, if you have a prompt injection that came in that says forward this to maliciousattacker at domain.com, it can't do that," Ceylan said. "That action in that tool call is not even in its vocabulary."</p><p>She classified agent actions into three oversight categories. Actions that are not sensitive, like read and summarize, need no human in the loop. Moderately sensitive actions skip human approval but get logged and monitored, while destructive actions like mass deletion of files always require a human. "Things are gonna shift between those three categories quite a bit," she acknowledged, "but setting those types of categories up front allows you to have a principled framework."</p><p>Rather than layering controls onto agents one at a time, Intuit has built a central platform called GenOS, short for generative AI operating system, which abstracts security, risk, and fraud modeling so individual agent developers never reinvent protection. "Permissioning is not about giving access to AI," Parekh said. "Instead, it is defining very tightly scoped and clearly auditable authority to the agent to perform very specific tasks." Intuit evolved from agents inheriting user permissions to each agent carrying its own identity, and the company is now investigating mid-session permission changes tied to the specific task underway.</p><p>Parekh calls the broader model an AI-powered expert platform, one where the human expert is built into the trust architecture rather than bolted on as a gate. "The paradigm that we are pursuing is where the user, the AI agent, and the human expert are collaborating to solve the user problem," he said.</p><h2>The end of human code review</h2><p>Ceylan took on the tension between security testing and development velocity without hedging. "The days of secure code reviews where a human's looking at the code and we're looking at security architecture reviews, design docs, those are done," she said. "If you keep trying to do security that way, you're gonna get left behind." Box is building toward a fully agentic development lifecycle where agents review design documents, apply security requirements, and review the code for vulnerabilities. "I'm very optimistic that we will get to a point where we will write code without security vulnerabilities because agents and the models are going to get so good at writing code without vulnerabilities," she said. "We're still a long way away from that."</p><p>Her advice for development teams skips the advanced AI concepts entirely and returns to basics that predate agents. "It comes down to very basic least privilege access," she said. "If you start giving your agents overly broad permissions at the beginning, it's really hard to comb that back and build an infrastructure that allows for those ephemeral credentials and only those narrowly scoped tasks."</p><p>Parekh explained why the red teaming surface has expanded so quickly. "These agents have skills, and skills could become vulnerabilities," he said. "Agents have access to certain data, they have access to tools, and there could be threats that are lurking within those tools as well. So suddenly the blast radius of the malicious code or the intent increases dramatically." When Intuit identifies common vulnerability patterns from its manual red teaming exercises, it automates those tests back into the GenOS harness so future agents inherit protection and red teamers stay focused on new threat vectors. Runtime scanning of prompts and responses adds a final layer that can stop a suspect response and escalate to a human expert, he said.</p><p>"You need to continuously test to ensure that those remain robust to the protections that you have built, as well as to account for any sort of drift or any other types of dependencies that you introduce into your scenario that can create novel vulnerabilities," she said.</p><h2>Intent versus probability</h2><p>An audience question about intent detection set off the sharpest exchange of the session. Ceylan noted that when Box's own agent operates, the system always knows the user's intent because it controls the prompt, which means guardrails and tool-call restrictions can be engineered around it. The harder challenge, which she admitted Box is still trying to solve, arrives when external agents connect and the context behind the request is opaque.</p><p>That exchange exposed a split running through the wider industry. Mastercard, in the fireside chat immediately preceding the panel, came down on the side of quantifying intent, building an open-source framework to propagate it as a standard because complex B2B procurement cannot work without that trust. Endpoint security CTOs, in briefings with VentureBeat, have gone the other way, saying they will bet on probability rather than intent inference for production workloads. Chang explained why models, as they are trained today, cannot reliably derive intent from a prompt, which is why deterministic controls and behavioral proxies remain necessary. Ceylan agreed that both are required. "If you're not doing anything deterministic, you're really relying heavily on that intent, and I haven't seen programs that are there yet," she said.</p><p>Ceylan's story about trust collapsing after a single agent mistake landed as the panel's most memorable moment because enterprise agentic security is not a problem that gets solved and stays solved. Models change, permissions drift, and adversaries adapt across multi-turn conversations that snapshot tests never capture.</p><p>For the 82% of enterprises relying on provider-native controls as their primary security layer, and the 59% shopping for agent security tooling over the next 12 months, the panel's takeaway was blunt. Test the way attackers attack, across full conversations and continuously, or find out in production what your single-turn red teaming missed.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start]]></title>
<description><![CDATA[Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today's launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture t...]]></description>
<link>https://tsecurity.de/de/3690017/it-nachrichten/black-forest-labs-launches-flux-3-capable-of-generating-images-and-20-second-video-with-audio-but-in-limited-release-to-start/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3690017/it-nachrichten/black-forest-labs-launches-flux-3-capable-of-generating-images-and-20-second-video-with-audio-but-in-limited-release-to-start/</guid>
<pubDate>Thu, 23 Jul 2026 20:48:22 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with <a href="https://bfl.ai/blog/flux-3">today's launch of FLUX 3</a>, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.</p><p>The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface. </p><p>That distinction is central to the company's pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company's words, "that can perceive, predict, and act across physical and digital environments." This release marks BFL's first public video generation model. </p><div></div><p>FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a <a href="https://tally.so/r/44d9NX">gated "Early Access" program now</a>, to which anyone can apply, but which BFL must approve. </p><p>There is presently no public access through BFL's application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including <a href="https://venturebeat.com/technology/anthropic-says-its-most-powerful-ai-cyber-model-is-too-dangerous-to-release">Anthropic</a> and <a href="https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov">OpenAI</a>, though those were ostensibly for security concerns and due to government request. </p><p>What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.</p><p>Another big notable omission: FLUX 3 is <i>not</i> launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as "open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction" — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.</p><p>But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company's commitment, but it is disappointing given the role open weights have played in FLUX's adoption thus far. </p><h2><b>Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption</b></h2><p>BFL has published several benchmark comparisons, but they're qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability. </p><p>In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google's Gemini Omni Flash in 52%.</p><p>One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a "preliminary evaluation of an early FLUX 3 candidate" — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.</p><p>Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.</p><p>Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0's international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.</p><p><a href="https://venturebeat.com/technology/googles-gemini-omni-flash-hits-the-api-turning-enterprise-video-production-into-a-conversation">Gemini Omni Flash</a>, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL's own measurement, the two are indistinguishable on 10-second text-to-video quality. </p><p>Google's advantage in that matchup is that Omni is generally available via Google's Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.</p><p>One regional wrinkle matters for a German company's home market. Editing <i>uploaded</i> video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.</p><p>Here's a rough guide for enterprises considering which video models to rely upon: </p><table><tbody><tr><td><p><b>Model</b></p></td><td><p><b>Max single-generation duration</b></p></td><td><p><b>Max resolution</b></p></td><td><p><b>Key constraints</b></p></td><td><p><b>Price per 10-second clip (720p)</b></p></td><td><p><b>Price per 10-second clip (1080p)</b></p></td><td><p><b>Price per 10-second clip (4K)</b></p></td></tr><tr><td><p>FLUX 3 Video </p></td><td><p><b>20 seconds </b></p></td><td><p>Not stated; evaluations run at 720p </p></td><td><p>Early access; no published SLA or pricing </p></td><td><p>Not announced </p></td><td><p>Not announced </p></td><td><p>Not announced </p></td></tr><tr><td><p>HappyHorse 1.1 </p></td><td><p>15 seconds </p></td><td><p>1080p </p></td><td><p>No 4K; closed weights </p></td><td><p>Not published (v1.0 reseller rate is ~$1.82) </p></td><td><p>Not published (v1.0 reseller rate is ~$3.12) </p></td><td><p>n/a </p></td></tr><tr><td><p>Veo 3.1 </p></td><td><p>Per-second billing </p></td><td><p><b>4K</b> </p></td><td><p><b>Supports clip extension; preview </b></p></td><td><p>$4.00 </p></td><td><p>$4.00 </p></td><td><p>$6.00 </p></td></tr><tr><td><p>Veo 3.1 Fast </p></td><td><p>Per-second billing </p></td><td><p><b>4K </b></p></td><td><p>Preview </p></td><td><p>$1.00 </p></td><td><p>$1.20 </p></td><td><p><b>$3.00 </b></p></td></tr><tr><td><p>Veo 3.1 Lite </p></td><td><p>Per-second billing </p></td><td><p>1080p </p></td><td><p>No 4K, no clip extension; preview </p></td><td><p><b>$0.50 </b></p></td><td><p><b>$0.80 </b></p></td><td><p>n/a </p></td></tr><tr><td><p>Gemini Omni Flash </p></td><td><p>10 seconds (3s minimum) </p></td><td><p>720p at 24 FPS </p></td><td><p>Preview abd no EU access</p></td><td><p>$1.00 </p></td><td><p>n/a </p></td><td><p>n/a </p></td></tr></tbody></table><h2><b>One architecture for media generation and physical action</b></h2><p>FLUX 3 builds on <a href="https://venturebeat.com/technology/black-forest-labs-new-self-flow-technique-makes-training-multimodal-ai">Self-Flow</a>, BFL's method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026. </p><p>The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.</p><p>"We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture," said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. "True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express."</p><p>He put the case more bluntly elsewhere in the announcement: "You can't cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds."</p><p>BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.</p><p>For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.</p><p>For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.</p><h2><b>What FLUX 3 Video can actually do</b></h2><p>The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation. </p><p>Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching <a href="https://developers.openai.com/api/docs/guides/video-generation">OpenAI's discontinued Sora model.</a></p><p>The capability list BFL published covers:</p><ul><li><p>Text-to-video generation.</p></li><li><p>Image-to-video generation, either animating from a starting frame or using images as visual references.</p></li><li><p>Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.</p></li><li><p>Generative video-audio continuation from existing video and audio input.</p></li><li><p>Keyframe-to-video generation for controlled transitions between defined moments.</p></li><li><p> Multilingual dialogue.</p></li><li><p>A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.</p></li><li><p>Typography generation and animated design.</p></li><li><p>Agentic chaining of individual clips into longer, multi-shot sequences.</p></li></ul><p>That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.</p><p>It is also the capability where competition is most direct. HappyHorse 1.1's headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.</p><p>BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.</p><h2><b>FLUX-mimic tests whether video models can become robot models</b></h2><p>BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.</p><p>The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic's robot-learning and production-deployment expertise in dexterous manipulation.</p><p>FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data. </p><p>BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.</p><p>"The hardest part of robotics is data," said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. "Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning."</p><p>BFL<!-- --> argues that a model trained only on images cannot understand a world that "moves, sounds, changes, and responds," and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni. </p><p>Its developer documentation cites "world knowledge" that combines "an understanding of physics" with Gemini's grasp of history, science and cultural context. Its marketing is blunter still: "Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different," the company posted in June, crediting the model with "an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic." </p><p>The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it. </p><p>There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.</p><h2><b>Open weights helped make FLUX an industry standard</b></h2><p>BFL<a href="https://venturebeat.com/technology/s"> officially launched in summer 2024 </a>and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises. </p><p>The company's founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and <a href="https://venturebeat.com/business/stable-diffusion-creators-launch-black-forest-labs-secure-31m-for-flux-1-ai-image-generator">Stable Diffusion</a>, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies. </p><p>That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research's Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.</p><p><a href="https://www.wired.com/story/black-forest-labs-ai-image-generation/"><i>Wired</i></a> magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley's largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.</p><p>FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, <a href="https://venturebeat.com/business/black-forest-labs-releases-flux-1-1-pro-and-an-api">released shortly after the firm's launch,</a>  gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.</p><p>The company continued that pattern with <a href="https://venturebeat.com/ai/black-forest-labs-launches-flux-2-ai-image-models-to-challenge-nano-banana">FLUX.2 Dev</a> in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.</p><p>FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery.  BFL hasn't yet shared information about its license, the parameter count, quantizations or hardware requirements.</p><p>The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows. </p><p>The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom's T.Capital.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents]]></title>
<description><![CDATA[Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agen...]]></description>
<link>https://tsecurity.de/de/3689830/it-nachrichten/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689830/it-nachrichten/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents/</guid>
<pubDate>Thu, 23 Jul 2026 19:19:45 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception.</p><p>This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what drives the choice, what they optimize for, how they expect agent control to be structured, and — most revealingly — how orchestrated their deployed “agents” actually are and how tightly they control the cost of running them.</p><p>The central finding is a gap between orchestration ambition and orchestration reality. Enterprises are consolidating fast onto the major model platforms: Anthropic’s Claude is the primary platform for 40%, more than double any rival, followed by Microsoft (18%) and OpenAI (13%). The choice is driven by “model gravity” — native alignment with a state-of-the-art base model (21%) — and success is judged by reliable, multi-step execution (task completion reliability 32%, multi-step workflow management 28%). Yet asked to assess their portfolios honestly, 71% say a quarter or fewer of their deployed “agents” are true multi-step orchestrated workflows rather than single-prompt chatbot wrappers, and only 10% have crossed the halfway mark. The orchestration layer is being built well ahead of the orchestrated portfolio it is meant to run.</p><p>That gap shapes the architecture enterprises are putting in place. By the end of 2026 a clear majority (51%) expect a hybrid control plane — provider-native plus external orchestration — and only 6% expect to hand control to a provider-managed service, because vendor lock-in (35%) is the risk they fear most if control lives inside a model provider. Investment follows the build-out: agent workflow tooling leads the spend (34%), with security and permissions enforcement (25%) behind. And fiscal control lags throughout — more than a quarter (27%) have no real-time way to stop a runaway agent before the bill arrives.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent orchestration. Responses are filtered to organizations with 100 or more employees (n=101), drawn from a single June 2026 wave; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends.</p><p>By organization size the sample is spread evenly across the enterprise bands: 100–499 employees, 2,500–9,999, and 50,000+ (21% each), with 10,000–49,999 and 500–2,499 (19% each). By role it is senior and buyer-credible: product and program managers (15%), CIO/CTO/CISO (13%), consultants and advisors (13%), and a spread of data, AI, and engineering directors and VPs, with an “Other” function at 18%. On purchasing, 81% are recommenders, influencers, or final decision-makers for AI solutions (66% recommender/influencer, 15% final decision-maker). Technology/Software is the largest industry at 44%, followed by Financial Services (17%) and Healthcare/Life Sciences (8%).</p><p>At 101 respondents the sample is robust enough to read directionally with reasonable confidence, though it remains self-selected and is not a probability sample.</p><h2>Finding 1: Orchestration runs on model-provider platforms</h2><p><b>Anthropic’s Claude leads; open frameworks are marginal</b></p><p>We asked which agent orchestration platform enterprises primarily use today. The answer concentrates on the major model providers — and on one in particular.</p><div></div><p>A note on reading these shares. As described in the methodology section, the respondents are self-selected, and this question asked them for a single primary platform — so the figures measure which platform leads each enterprise's deployment, within a self-selected audience of AI-active technical decision-makers. A sample built this way can diverge substantially from spend-weighted market measures, and each VB Pulse survey draws its own sample with its own company-size mix, so vendor figures should not be compared across our surveys either. Read these shares as a portrait of where this cohort has placed its primary orchestration bet today, rather than as market share.</p><p>The model platforms dominate. Anthropic, Microsoft, OpenAI, Google, and Amazon together account for roughly 80% of deployments (81 of 101), while the open frameworks (LangChain/LangGraph) and custom in-house builds that anchor engineering discussion sit in single digits. Anthropic’s lead — 40%, more than double the next platform — mirrors the “model gravity” selection logic in Finding 2: enterprises are choosing the orchestration layer that comes with the model they want to build on. As with the security vendors in the prior agent-security wave, the tools that define the category in technical circles are not yet where enterprise deployment concentrates. A small 3% are not orchestrating at all.</p><p>Respondents rate the platforms they run at 3.94 out of 5 overall (109 answered), with “value for money” specifically at 3.94 and “ease of implementation” the weakest score, at 3.85 — placing orchestration near the bottom of our five-tracker satisfaction range, ahead of only evaluation tooling. A rating just under 4 out of 5, from users of whom 96% plan to change their orchestration approach within the year, reads as provisional acceptance: the platforms work well enough to run today, and not well enough to stop the search for something better. The ratings sit alongside near-universal intent to change; this is a layer enterprises tolerate more than they love.</p><h2>Finding 2: Model gravity drives platform selection</h2><p><b>The base model, not the tooling, decides the platform</b></p><p>We asked what most influenced the orchestration platform choice. The single largest factor is the pull of the underlying model — though flexibility and ease of development follow close behind.</p><div></div><p>Model gravity leading is the selection-side explanation for Anthropic’s platform lead: enterprises pick the orchestration environment closest to the frontier model they have standardized on. But the next tier complicates the picture — flexibility across models and tools (17%) and ease of development (17%) say enterprises also want to avoid being trapped by that choice, foreshadowing the lock-in fear in Finding 6. Security and permissions (14%) and total cost of ownership (11%) round out a pragmatic buying logic. Performance (latency/memory) sits last at 4%, a reminder that at this stage of adoption the binding constraints are model fit and optionality, not raw speed.</p><h2>Finding 3: The job is reliable multi-step execution</h2><p><b>Enterprises just orchestration by whether it completes the work</b></p><p>We asked what enterprises optimize for — their primary success metric for orchestration. Reliability and multi-step workflow management dominate; developer- and user-facing metrics trail.</p><div></div><p>Task completion reliability (32%) and multi-step workflow management (28%) together account for 59% of responses (60 of 101): orchestration succeeds, in the enterprise view, when it reliably carries a task through multiple steps to completion. Developer productivity (17%) matters but is secondary — the inverse of its prominence in framework discussion — and end-user experience (9%) is a minor concern, consistent with orchestration being an internal execution problem rather than a UX one. This reliability-first standard is exactly what makes the Chatbot Trap finding so pointed: enterprises define success as dependable multi-step execution, yet most of their deployed “agents” do not yet do multi-step work at all.</p><p>The trap is not evenly distributed. Splitting the sample by organization size, 77% of smaller enterprises say a quarter or fewer of their agents do true multi-step work, against 62% of larger ones. Larger enterprises are meaningfully further into genuine multi-step deployment; the chatbot trap is, directionally, a mid-market condition.</p><h2>Finding 4: Consolidate, productionize, and build in-house </h2><p><b>Three strategic moves are nearly tied for the year ahead</b></p><p>We asked what major change enterprises anticipate in their orchestration strategy over the next 12 months. Three moves cluster at the top, almost evenly split.</p><div></div><p>The top three — building in-house control (25%), standardizing on one framework (24%), and moving agents from sandbox to production (23%) — are statistically indistinguishable and tell a single story: enterprises are moving from experimentation to operational consolidation. They want fewer frameworks, more production exposure, and more ownership of the control layer; only 4% expect no change. The appetite for custom in-house control planes is notable alongside the platform concentration in Finding 1 — enterprises are standardizing on model-provider platforms while simultaneously planning to wrap them in control logic they own, the hybrid posture that Finding 6 makes explicit.</p><h2>Finding 5: Nearly seven in 10 plan to switch — and the biggest group of movers has no shortlist </h2><p>The strategic change enterprises anticipate (previous finding) comes with vendor motion attached. Asked whether they plan to adopt a new, additional, or replacement agent orchestration platform in the next twelve months, more respondents are moving here than in any other layer we track.</p><div></div><p>Asked which platforms they are considering, the most common answer among those in motion is none yet: 29% of all respondents are evaluating without a shortlist, the largest single response after "not considering a change." Among named candidates, OpenAI leads at 16%, followed by LangChain/LangGraph at 12% and Anthropic at 7% — and notably, the independent frameworks draw roughly double their current usage footprint in forward consideration, the same pattern our security tracker found for specialist vendors. Read with this report's concentration and lock-in findings, the picture completes itself: the major model-platform providers hold roughly four-fifths of today's primary usage, vendor lock-in has become the leading fear, 96% anticipate a strategic change — and now the purchase intent to act on all of it, with the largest bloc of buyers still undecided. The most concentrated layer of the agentic stack is also, as of June, the least settled.</p><h2>Finding 6: Investment flows to workflow tooling</h2><p><b>Tooling and permissions lead the spend; monitoring trails</b></p><p>We asked which orchestration-related investment will grow most next year. Agent workflow tooling leads, with security and permissions enforcement behind.</p><div></div><p>Workflow tooling leading (34%) is the budget-side expression of the reliability-and-multi-step priority in Finding 3: the money is going to the machinery that strings steps together dependably. Security and permissions enforcement (25%) and scaling infrastructure (20%) follow — the investments required to take agents from sandbox into production, the strategic move in Finding 4. Monitoring and debugging draws a smaller 11%, with another 11% reporting flat budgets. The weight on tooling, permissions, and scaling over pure observability signals that enterprises are spending to build and harden orchestration, not merely to watch it run.</p><h2>Finding 7: The control plane will be hybrid — and lock-in is why</h2><p><b>Enterprises expect to split control between providers and their own layer</b></p><p>We asked where enterprises expect the primary control plane for agents to live by the end of 2026, and what worries them most if that control sits inside a model-provider platform. A clear majority expect a hybrid model — and vendor lock-in is the reason.</p><div></div><p>Hybrid control is the dominant expectation by a wide margin (51%), and only 6% expect to hand control to a provider-managed service outright. Read together, the hybrid, custom, and externally-abstracted options — every architecture that keeps control at least partly outside the provider — sum to 88% (89 of 101). The reason surfaces directly when we asked about the risk of provider-resident control: vendor lock-in leads at 35% (35 of 101), ahead of security and permissioning limitations (28%) and inflexibility across models and tools (21%). The pattern echoes the prior wave’s “don’t trust the model to police itself” posture — here, enterprises will build on a provider’s platform but decline to be governed entirely by it. The hybrid control plane is the architectural hedge against the lock-in they most fear.</p><p>The June figure asserting a preference for a hybrid control plane marks movement from earlier. In the April–May survey (n=145), only 34% expected a hybrid control plane, and a greater number (12%) expected to hand control fully to a provider-managed service. These two snapshots don’t yet measure a confirmed longitudinal trend — but the direction of the conversation is unambiguous: toward keeping control.</p><p>Lock-in is also a new arrival as a top concern. In the April–May wave, the leading concern was security and permissioning limitations (32%), with lock-in second at 24%; by June the two had traded places. The worry about provider platforms appears to be maturing from whether they can be secured to whether they can be replaced.</p><h2>Finding 8: The chatbot trap — most “agents” aren’t agents yet</h2><p><b>Enterprises admit most deployments are still chatbot wrappers</b></p><p>We asked enterprises to assess their portfolios honestly: what share of their deployed “agents” are true multi-step orchestrated workflows versus simple single-prompt chatbot wrappers. The answer is the defining finding of this wave.</p><div></div><p>This is the gap at the center of the report. Combining the bottom two bands, 71% of enterprises (72 of 101) say a quarter or fewer of their deployed “agents” are genuinely orchestrated — and just 10% (10 of 101) have crossed the halfway mark. The ambition documented in the earlier findings — model-provider platforms, reliability-first success metrics, production rollouts, a deliberate control architecture — runs well ahead of the deployed reality, which remains overwhelmingly single-prompt assistants dressed as agents. This is less a contradiction than a roadmap: the platforms, budgets, and strategies are being put in place precisely because the orchestrated portfolio is still so thin. The open question for later waves is how fast the reality closes on the ambition.</p><h2>Finding 9: Fiscal control is still reactive</h2><p><b>Only a minority can stop a runaway agent before the bill arrives</b></p><p>Finally, we asked how enterprises enforce fiscal control over agent token consumption — the risk that an autonomous loop exhausts a budget before anyone intervenes. Most rely on native caps or after-the-fact monitoring; real-time programmatic control is the exception.</p><div></div><p>More than a quarter of enterprises (27%) admit they have no real-time, programmatic way to stop an agent before a budget-breaking bill arrives — they learn of it from the logs afterward. Another 32% lean entirely on the native caps and throttles built into their primary platform, a control only as good as the provider’s tooling and one that ties back to the lock-in concern of Finding 6. The enterprises building custom gateways (23%) or exploiting cross-model routing to arbitrage cost (19%) are the ones treating token burn as an engineering problem to be controlled deterministically. As with orchestration maturity, fiscal control is an area where the operational reality lags the ambition: agents are moving toward production faster than the cost-control plane around them is being built.</p><p>It’s worth noting, a split appears according to company size: roughly one in three enterprises under 2,500 employees (34%) exercises only reactive control of agent spend, against 20% of larger enterprises — directional figures, but consistent with the chatbot-trap split. The mid-market is running the least mature agents on the least instrumented budgets.</p><h2>The bottom line: The layer is real; most of the agents aren't yet</h2><p>Organizations with 100 or more employees describe an orchestration strategy that is consolidating quickly and maturing slowly. They are standardizing — for now — on model-provider platforms, which collectively hold roughly four-fifths of primary usage, chosen for the gravity of the underlying model, and they judge success by reliable multi-step execution. Investment is flowing to workflow tooling and permissions, the strategy is to consolidate frameworks and push agents into production, and the control plane they expect is deliberately hybrid, because vendor lock-in is the risk they fear most. But the standardization is provisional: 68% plan to adopt a new, additional, or replacement orchestration platform within twelve months — the highest switching intent of any layer we track — and the largest group of those movers has not yet shortlisted a candidate. Today's concentration describes where enterprises are, and visibly does not describe where they intend to stay.</p><p>But the honest self-assessment punctures the ambition. Seventy-one percent say a quarter or fewer of their deployed "agents" are truly orchestrated, only 10% are past the halfway mark, and more than a quarter cannot stop a runaway agent in real time. The orchestration layer — the platforms, the budgets, the control architecture — is being built ahead of the orchestrated portfolio it is meant to run. At 101 respondents in a single June wave this reads as a clear directional signal rather than a precise measurement: enterprises have decided how they want to orchestrate agents well before most of their agents are doing anything an orchestration layer is for. The questions for subsequent waves are whether the deployed reality closes the gap on the ambition — and, with nearly seven in ten buyers in motion and most of them undecided, which platforms the settled stack finally lands on.</p><hr><p><i>Based on survey responses from 101 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. Because this is one wave rather than a pooled multi-month sample, results read directionally rather than as a confirmed trend. Respondents include product and program managers, CIOs, CTOs and CISOs, consultants and advisors, and directors and VPs of data, AI, and engineering, across Technology/Software, Financial Services, Healthcare, and other sectors.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway]]></title>
<description><![CDATA[Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated...]]></description>
<link>https://tsecurity.de/de/3689829/it-nachrichten/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689829/it-nachrichten/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway/</guid>
<pubDate>Thu, 23 Jul 2026 19:19:44 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap — the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures.</p><p>This wave of VentureBeat Pulse Research examines how technical leaders measure agent performance: which reliability and evaluation platforms they use, how they select and trust them, what breaks in production, and how far they are willing to let agents run without a human in the loop.</p><p>The central finding is an evaluation gap — the distance between the autonomy enterprises are granting their agents and the trust they place in the evaluations meant to govern it. Half of organizations (50%) have, in the past year, deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure, and a quarter have seen it happen more than once. Trust in the tests themselves is thin: only 5% say they fully trust automated evaluation today, and the single most-cited limitation is that evaluations align poorly with real-world outcomes (29%). Enterprises are discovering that a passing eval is not the same as a working agent.</p><p>What makes the gap consequential is the direction of travel. Two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve months (33%). At the same time, the evaluation stack that would have to earn that trust is fragmented and immature: the most common primary tools are the model providers’ native evals, tied with having no dedicated tooling at all (17% each); and only about a quarter of enterprises run real-time quality checks on live production traffic. The autonomy is arriving faster than the assurance.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey — the Agentic Reliability &amp; Evals tracker — focused on how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=157), drawn from a single survey in June 2026; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Where questions were multiple-select, those shares can sum to more than 100%.</p><p>By role the sample is senior and buyer-credible: 38% are final decision-makers for AI purchases and another 34% recommenders or influencers. Product and program managers (15%), consultants and advisors (10%), directors of engineering/IT (8%), and CIOs/CTOs/CISOs (8%) lead the named titles, alongside a large “Other” function (37%). By organization size the sample is mid-market-weighted: 100–499 (37%) and 500–2,499 (27%) employees lead, with 2,500–9,999 (20%), 10,000–49,999 (10%), and 50,000+ (6%) above them. Technology/Software is the largest industry at 23%, followed by Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%).</p><p>At 157 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent evaluation practices rather than from the largest operators.</p><p><i>Note: This survey was rebuilt for the June wave from the earlier “LLM observability and evaluations” survey; because the questions and sample differ, no comparisons are made to the April–May data.</i></p><h1>Finding 1: A passing eval is not a working agent</h1><p><b>Half have shipped an agent that passed evals, then failed a customer</b></p><p>We asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. Half of those that run evaluations had.</p><div></div><p>This is the report’s defining number. Half of organizations (50%) have shipped an AI feature that cleared their internal evaluations and then failed in front of a customer — an incorrect output, a broken workflow, or a quality incident — and a quarter have seen it happen more than once. Only 36% report no such failure, and the remainder either run no pre-deployment evaluations (8%) or don’t track the root cause closely enough to know (6%). The failure is precise and expensive: the evaluation said the agent was ready, and it was not. Everything that follows — how enterprises trust their evals, what they monitor, and how much autonomy they grant — is shaped by this experience.</p><h2>Finding 2: Almost no one fully trusts automated evaluation</h2><p><b>The top complaint: Evals don't match real-world outcomes</b></p><p>We asked which limitation most reduces trust in automated agent evaluations today. Only a sliver of enterprises had no complaint at all.</p><div></div><p>Trust in automated evaluation is scarce, and specific. Only 5% of organizations say they fully trust automated evaluation as it stands — meaning 95% name a limitation that holds them back. The most common, at 29%, is the one that most directly explains Finding 1: evaluations align poorly with real-world outcomes, passing agents that later fail. Bias or inconsistency (21%) and a lack of explainability (18%) follow — enterprises cannot always tell why an evaluation reached its verdict — and 17% cite data-leakage or privacy concerns in the evaluation process itself. The tests meant to certify agents are not yet trusted to certify them, which is precisely why the autonomy trajectory in Finding 3 is so striking.</p><h2>Finding 3: The autonomy ceiling is rising anyway</h2><p><b>Two-thirds already allow, or are building toward, zero-human deployment</b></p><p>We asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The trajectory runs straight through the trust gap.</p><div></div><p>Here is the paradox at the heart of the report. Even though almost no one fully trusts automated evaluation (Finding 2), two-thirds of organizations (66%) either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within a year (33%). Only 22% rule it out for the foreseeable future. The direction is unambiguous: enterprises are moving to let evaluations gate production autonomously — removing the human check — at the same moment they say those evaluations don’t reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it, which is the mechanism by which the false-confidence failures of Finding 1 will scale rather than shrink.</p><p>Notably, the autonomy bet is not just a small company phenomenon. Splitting the sample by company size, larger enterprises are slightly further down the path toward zero human review than smaller companies (70% versus 64%) and slightly more likely to have shipped an evaluation-passing agent that then failed a customer (54% versus 48%). The assumption that large, regulated organizations are holding the human in the loop longest is, in this sample, backwards.  To be sure, these are directional figures, since the survey was not a huge sample — 57 respondents from companies with 2,500+ employees and 100 from companies smaller than that. </p><h2>Finding 4: The evaluation stack is fragmented and provider-led</h2><p><b>Provider-native evals lead — tied with no dedicated tool at all</b></p><p>We asked which agent reliability or evaluation platform enterprises primarily use today. The market has no clear leader — and a large share has nothing dedicated.</p><div></div><p>The evaluation layer is early and unconsolidated. Provider-native tooling leads — OpenAI’s native evals and traces (17%) and Anthropic’s Claude Console evals (13%) together outweigh any independent platform — but it is tied at the top by a striking answer: 17% of enterprises use no dedicated agent-evaluation tooling at all, a notable gap for organizations shipping agents to customers. The specialist evaluation vendors — DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize — are scattered across single to low double digits, and 11% have built their own. No independent platform has yet become the category standard, which leaves most enterprises evaluating agents with provider-native tools, home-grown scripts, or nothing.</p><h2>Finding 5: Production monitoring rarely watches output quality</h2><p><b>Only a quarter run real-time quality checks on live traffic</b></p><p>Production monitoring for an AI agent can watch two very different things. It can watch whether the system is <b>functioning</b> — is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent's output is <b>correct</b> — automated checks that evaluate the content of each answer as it goes out: did the agent give the right answer, take the right action, stay within policy. The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked organizations which kind their live production monitoring is built for today.</p><div></div><p>Grouped by what is actually being watched, the split is stark: 51% of organizations monitor only whether the agent is functioning, while 23% monitor whether its answers are right. Counting the ad-hoc reviewers and the don't-knows, roughly three-quarters of organizations run no automated, real-time evaluation of output correctness in production — they can see that the system is up and what it costs, and they are taking the correctness of its answers on faith. That blind spot is the runtime counterpart to the pre-deployment gap in Finding 1: the same organizations engineering the human out of the deployment decision mostly cannot see, in real time, when the deployed agent starts getting things wrong.</p><h2>Finding 6: Bought on cost, measured on consistency</h2><p><b>Price and integration drive selection; evaluation consistency is the goal</b></p><p>We asked what most influenced enterprises’ choice of an evaluation vendor, and what they treat as their primary measure of success. Both answers are pragmatic.</p><div></div><p>Enterprises buy evaluation tooling on economics and trust it on repeatability. Cost of evaluations (28%) narrowly leads selection, just ahead of ease of integration (27%) and evaluation accuracy (24%) — breadth of observability (13%) and vendor roadmap (4%) matter far less. On what success looks like, more than a third (36%) name evaluation consistency — getting the same verdict on the same behavior every time — well ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%), and compliance (11%). The emphasis on consistency is telling: before enterprises can trust an evaluation’s verdict, they need it to be stable — the very property whose absence (bias and inconsistency) ranked among the top trust limitations in Finding 2. Satisfaction with current tooling is only moderate, averaging 3.8 on a five-point scale across overall satisfaction, ease of implementation, and value for money.</p><h2>Finding 7: The next dollar goes to humans and observability</h2><p><b>Investment is flowing to oversight, not just automation</b></p><p>We asked which reliability and evaluation investment will grow most over the next year. The money is going toward watching agents more closely — including with people.</p><div></div><p>The second-largest planned investment — behind only production observability — is human review workflows, at 26%. Read against Finding 1, that is the report's quietest contradiction: at the same moment two-thirds of enterprises are engineering the human out of the deployment decision, more of them plan to grow spending on human reviewers (26%) than on the automated evaluation pipelines (16%) that would replace them. The zero-human trajectory and the human-review budget are rising in the same companies at the same time. Indeed, only 8% report that their budget is not increasing. </p><p>Taken together, enterprises are hedging: building toward autonomy while spending to watch agents more closely and keep humans available for the calls that automated evaluation cannot yet be trusted to make.</p><h2>Finding 8: A tooling reshuffle is coming</h2><p><b>Nearly two-thirds plan to adopt or switch platforms within a year</b></p><p>We asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Few intend to stand pat.</p><div></div><p>The evaluation market is wide open. While 36% have no plans to change, a clear majority (64%) intend to adopt a new, additional, or replacement platform within twelve months, and 31% within the next quarter. The consideration set points where current usage is thinnest: Confident AI’s DeepEval leads what enterprises are evaluating (20%), ahead of OpenAI’s native evals (13%) and Braintrust (9%) — the open-source specialists drawing more interest than their present footprint. </p><p>Given that so many enterprises today rely on provider-native tools or nothing at all (Finding 4), this is less a defection than a first real wave of tooling adoption — the moment the evaluation layer starts to consolidate. Which platforms earn that trust, in a market where almost no one trusts automated evaluation yet, is the open question this series will keep tracking.</p><h2>The bottom line: An evaluation gap that autonomy will widen, not close</h2><p>Organizations with 100 or more employees are granting AI agents more independence than they trust their evaluations to support. Half have already shipped an agent that passed its evals and then failed a customer; almost none fully trust automated evaluation, chiefly because it doesn’t match real-world outcomes; and most watch production for uptime and cost rather than for whether the agent’s answers are right. Yet two-thirds already allow, or are actively building toward, deploying to production on automated evaluation alone.</p><p>The vendor market is early and unsettled: the most common primary evaluation tools are provider-native evals, tied with no dedicated tooling at all, and a clear majority plan to adopt or switch platforms within the year. Encouragingly, the next dollar is going to observability and — pointedly — human review, suggesting enterprises sense the gap even as they engineer past it. At 157 respondents in a single wave this is a directional read, skewed toward the mid-market — but the direction is clear: autonomy is being granted on the strength of evaluations that the people granting it do not yet trust. The evaluation gap is not a coverage problem that more tests alone will close; it is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy — or whether the false-confidence failures move from customer incidents into changes that deploy themselves.</p><hr><p><i>Based on survey responses from 157 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read rather than a precise measurement — the sample is self-selected, not a probability sample, and skews toward the mid-market. Respondents include product and program managers, consultants and advisors, directors of engineering/IT, and CIOs/CTOs/CISOs, among other functions, across technology/software, retail/consumer, healthcare/life sciences, manufacturing, and other industries.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials]]></title>
<description><![CDATA[Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents s...]]></description>
<link>https://tsecurity.de/de/3689827/it-nachrichten/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689827/it-nachrichten/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials/</guid>
<pubDate>Thu, 23 Jul 2026 19:19:41 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents still share credentials; and only three in ten isolate their highest-risk agents. The security stack is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents, spending remains a thin slice of the security budget, and enterprises are evenly split on whether their defenses are keeping pace with AI-enabled attackers. The result is an agent security gap — autonomous agents proliferating faster than the identity, isolation, and enforcement controls needed to hold them.</p><p>This wave of VentureBeat Pulse Research examines how enterprises secure their AI agents: what tooling they run, how they manage agent identity and isolation, what has already gone wrong, how much they spend, and whether they believe their defenses are keeping pace with AI-enabled attackers.</p><p>The central finding is an agent security gap — the distance between the autonomy enterprises are granting their agents and the controls in place to contain them. More than half of organizations (54%) have already experienced a confirmed agent security incident (18%) or a near-miss caught before harm (36%). The structural weakness beneath those numbers is identity: only about a third (32%) give every agent its own scoped, managed identity, while the rest report that some agents share credentials or that agents mostly run on shared API keys and human or service-account credentials. When agents share credentials, a single compromised or over-permissioned agent carries a wide blast radius — and only three in ten enterprises (30%) isolate their highest-risk agents in sandboxes to bound that radius.</p><p>What makes the gap notable is how comfortable enterprises are inside it. The security stack is overwhelmingly provider-native — OpenAI’s guardrails (51%), Google’s and Microsoft’s cloud controls, and Anthropic’s managed-agent controls dominate, while the dedicated agent-security specialists barely register — and satisfaction with that borrowed stack is high, averaging 4.2 out of 5. Yet spending remains a thin slice of the security budget, only a third of enterprises believe their AI defenses are ahead of AI-enabled attackers, and a clear majority plan to change tooling within the year. Enterprises are satisfied with controls they are simultaneously preparing to replace.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent security — the tooling, identity, isolation, and enforcement controls organizations use to secure autonomous AI agents. Responses are filtered to organizations with more than 100 employees (n=107; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single June 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.</p><p>By role the sample is senior and buyer-credible: 45% are final decision-makers for AI purchases and another 30% recommenders or influencers. Managers (43%), individual contributors (24%), VPs and directors (15%), and the C-suite (11%) make up the seniority mix. By organization size the sample is mid-market-weighted: 251–1,000 (42%) and 101–250 (25%) employees lead, with 1,001–5,000 (19%), 5,001–10,000 (8%), and 10,001+ (7%) above them. Technology/Software is the largest industry at 23%, followed by Manufacturing (15%), Retail/E-commerce (14%), and Healthcare/Life Sciences (13%).</p><p>At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent security rather than from the largest operators.</p><p>Satisfaction ratings are computed on the respondents who answered each rating question; the overall satisfaction score reflects 82 of the 107 qualified respondents.</p><h2>Finding 1: The incidents are already here</h2><p><b>More than half have had an agent security incident or near-miss</b></p><p>We asked whether organizations had experienced an agent security incident — a confirmed breach, or a near-miss caught before harm. Most that run agents in production had.</p><div></div><p>This is the report’s defining number. More than half of organizations (54%) have already had an agent security event — 18% a confirmed incident and 36% a near-miss caught before it caused harm. Only 42% report nothing, and a small remainder either run no agents in production or don’t track such events. That so many report near-misses rather than only confirmed incidents is telling: enterprises are catching problems, but they are catching them close to the edge. The controls examined in the rest of this report — identity, isolation, enforcement — are what determine whether the next near-miss stays a near-miss.</p><p>Exposure scales with company size, but containment does not. The incident-or-near-miss rate rises from 49% in the mid-market (companies with 101-1,000 employees) to 63% at larger enterprises (above 1,000 employees), while sandbox isolation of high-risk agents falls from 35% to 20%, and satisfaction with security tooling drops from 4.36 to 3.97. The organizations running the most agents across the most systems carry the most incidents and the least of the one control that bounds an incident's blast radius.</p><h2>Finding 2: The identity gap</h2><p><b>Only a third give every agent its own scoped identity</b></p><p>We asked how enterprises manage the identity of their AI agents — whether each agent has its own credentials, or agents share them. Full per-agent identity is the exception.</p><div></div><p>Rolled together, the overlapping answers show 69% of enterprises (74 of 107) with credential sharing somewhere in the agent fleet. Identity is the structural weakness beneath the incidents. Only about a third of enterprises (32%) give every agent its own scoped, managed identity — the precondition for least-privilege access and clean attribution. Nearly half (48%) say some agents have scoped identities but many still share credentials, and another 32% say agents mostly run on shared API keys or borrowed human and service-account credentials. (Respondents could describe more than one pattern across their agent fleet, so these overlap.) </p><p>The consequence is direct: when agents share credentials, an over-permissioned or compromised agent can act with far more reach than intended, and forensics after an incident cannot cleanly tell which agent did what. The non-human identity problem — giving every agent its own governed identity — is the single largest unfinished piece of enterprise agent security.</p><p>Moreover, a company’s agent credential posture is correlated with incidents. Organizations with credential sharing anywhere in the fleet were hit — with an incident or a near-miss in the past twelve months — at 63.5% (47 of 74). Organizations where every agent carries its own scoped identity were hit at 40.9% (9 of 22). The fully-scoped group is small, so for now the relationship is an association rather than proven causation, and the gap is concentrated in the mid-market — but within a single survey, a twenty-three point difference in incident rate suggests significance.</p><h2>Finding 3: Observe and enforce, but rarely isolate</h2><p><b>Only three in 10 sandbox their highest-risk agents</b></p><p>We asked what an organization’s agent security posture looks like in practice — whether they observe, enforce, isolate, or some combination. The control that bounds damage is the least common.</p><div></div><p>Monitoring and enforcement are reasonably common; containment is not. Roughly half of enterprises observe agent activity (47%) or enforce scoped permissions at runtime (49%), but only 30% isolate their highest-risk agents in sandboxes that bound the blast radius when the other controls fail. That ordering is backwards from a defense-in-depth standpoint: observation tells you what happened, enforcement tries to prevent it, but isolation is what limits the damage when prevention fails — and it is the control enterprises have adopted least. Combined with the identity gap in Finding 2, the picture is of agents that are watched and permissioned but rarely boxed in, which is precisely the configuration in which a single failure propagates.</p><h2>Finding 4: Security runs on borrowed, provider-native controls</h2><p><b>Guardrails from OpenAI, Google and Microsoft dominate; specialists barely register</b></p><p>We asked which agent security tooling enterprises use, and which is their primary layer. The answer favors the model providers and hyperscalers over the dedicated security vendors.</p><div></div><p>Enterprises are securing agents with tools that came bundled with their models and clouds. OpenAI’s guardrails lead at 51%, followed by Google’s and Microsoft’s cloud-native controls and Anthropic’s managed-agent controls — and when asked to name their single primary security layer, 82% name one of these provider-native offerings. The purpose-built agent-security category — Palo Alto’s Prisma AIRS, CrowdStrike, Cisco AI Defense, Zenity, HiddenLayer, Check Point’s Lakera, Okta for AI Agents, non-human identity platforms — barely registers, each in the low single digits, and only 5% run no dedicated tooling at all. As with retrieval and evaluation elsewhere in this series, the provider bundle is winning the default: enterprises reach first for the guardrails their platform ships, and the independent security layer that would address the identity and isolation gaps has not yet been adopted at scale.</p><p>The provider-default pattern is consistent across both Q2 survey waves. In April–May (n=110), usage was led by the same names — OpenAI's controls at 26%, Azure at 15%, AWS at 14%, Google at 12% — with every dedicated agent-security specialist at 3% or below and one in ten using no dedicated tooling at all. The common finding from the two surveys: Enterprises are defaulting to the solutions provided by the platform they’re using, and the specialist category vendors have yet to become big players here.</p><p>(<i>A note on reading these shares. As described in the methodology section, the respondent sample is self-selected and skews mid-market, and the usage question counted every vendor or approach a respondent has in place — so the figures measure presence in the security stack rather than spending or exclusivity. Individual vendor percentages therefore carry all the usual sample caveats. The structural pattern, however, held across both Q2 waves on two differently worded questions: provider-native and hyperscaler controls lead, and dedicated agent-security specialists remain in low single digits. Read the individual shares loosely and the pattern with confidence.)</i></p><h2>Finding 5: And enterprises are comfortable with it</h2><p><b>Satisfaction is high, even as incidents mount and identity lags</b></p><p>We asked how satisfied enterprises are with their current agent security tooling. The comfort is notably out of step with the exposure documented above.</p><div></div><p>Satisfaction with agent security tooling is high — 4.2 out of 5 overall, and 4.1 for value for money — among the most positive readings in this series. That is the striking part: enterprises are highly satisfied with a stack that is mostly borrowed provider guardrails, even though more than half have already had an incident or near-miss and only a third give their agents scoped identities. The comfort appears to rest on the convenience and low friction of provider-native controls rather than on demonstrated containment. It is a false comfort in the making — the same enterprises expressing satisfaction are, as Finding 8 shows, a clear majority planning to change tooling within the year, which suggests the confidence is thinner than the score implies.</p><h2>Finding 6: Budgets haven’t caught up</h2><p><b>Most spend under a tenth of the security budget on agents</b></p><p>We asked what share of the security budget enterprises allocate to securing AI agents. For a fast-emerging risk, the allocation is modest.</p><div></div><p>Spending on agent security is still a thin slice. The most common allocation is 6–10% of the security budget (46%), and a third of enterprises (34%) spend 5% or less; only a quarter (24%) devote more than a tenth. Given the incident rate in Finding 1 and the identity and isolation gaps in Findings 2 and 3, the budget looks like a lagging indicator — the risk has arrived faster than the funding to address it. The enterprises spending more than a tenth of their security budget on agents are a distinct minority, and they are likely the ones building the scoped-identity and isolation controls the rest have not.</p><h1>Finding 7: The arms race is even, at best</h1><p><b>Only a third think their AI defenses are ahead of AI-enabled attackers</b></p><p>We asked how enterprises assess the balance between their AI-enabled defenses and AI-enabled attackers. Confidence is far from settled.</p><div></div><p>Enterprises are split on whether they are winning. Only about a third (35%) believe their AI-enabled defenses are ahead of AI-enabled attackers; the rest are less sure — 32% call it roughly even, 21% think attackers are ahead, and another 21% say it is too early to tell. Taken together, a clear majority (53%) rate the balance as even or tilted toward the attacker. That uncertainty sits uneasily beside the high satisfaction of Finding 5: enterprises are content with their tooling yet unconvinced it is winning the contest it exists to win. In a domain where the offense is also compounding with AI, an even race is not a comfortable place to be.</p><h2>Finding 8: A security reshuffle is coming</h2><p><b>Nearly six in 10 plan to adopt or switch tooling within a year</b></p><p>We asked whether enterprises plan to adopt a new, additional, or replacement agent security solution, and which they are considering. Few intend to stand pat.</p><div></div><p>The security stack is not settled. While 41% have no plans to change, a clear majority (59%) intend to adopt a new, additional, or replacement agent security solution within twelve months, and 29% within the next quarter — a strong signal that, high satisfaction notwithstanding, enterprises know the current stack is provisional. Incidents are what start the buying cycle. </p><p>Among organizations that have been hit, 42.1% plan to adopt, add, or replace agent security tooling within the next ninety days, against 14.0% of organizations with no incident — and after a confirmed incident it becomes majority behavior, at 52.6%. Getting hit also changes the threat assessment: 33.3% of hit organizations say AI-armed attackers are ahead of their defenses, against 8.0% of the unhit. Experience, in this data, is the strongest predictor of both urgency and pessimism.</p><p>The consideration set still leans provider-native (OpenAI 34%, Google 30%, Anthropic 29%, Azure 25%), but the dedicated security vendors — Cloudflare, Cisco, Palo Alto, Okta, Check Point’s Lakera — draw early interest in the mid-to-high single digits, more than their current footprint. </p><p>What the shopping does not yet include is the identity layer specifically. Twelve percent of the respondents include an agent-identity product — Okta for AI Agents, Microsoft Entra Agent ID, or a non-human identity platform — anywhere in their consideration set, and among the credential-sharing organizations that have already had an incident, identity consideration is essentially unchanged, at roughly one in ten. The control most directly implicated by the incident data is the one largely missing from the purchase plans. Whether this wave hardens the provider-native default or finally opens the door to purpose-built agent security — the identity and isolation controls the incidents call for — is the question this series will keep tracking.</p><h2>The bottom line: A security gap that autonomy will test first</h2><p>Organizations with more than 100 employees are giving AI agents real reach into systems and data while securing them with controls built for something else. More than half have already had an incident or near-miss; only a third give every agent its own scoped identity, and most still share credentials; only three in ten isolate their highest-risk agents; and the stack doing this work is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents.</p><p>The uncomfortable pairing is confidence with exposure: satisfaction with the current tooling is among the highest in this series, yet spending is a thin slice of the security budget, only a third believe their defenses are ahead of AI-enabled attackers, and a clear majority are already planning to replace what they have. At 107 respondents in a single wave this is a directional read, skewed toward the mid-market — but the direction is clear: agent adoption is running ahead of agent security, and the controls that matter most when something fails — scoped identity and isolation — are the ones enterprises have built least. The agent security gap is not a coverage problem that a provider guardrail will close on its own; it is a problem of identity, isolation, and enforcement built for autonomous software. The open question for later waves is whether enterprises close it deliberately — or whether a confirmed incident closes it for them.</p><hr><p><i>Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read, not a precise measurement — the sample is self-selected and skews mid-market, so it's best read as the view from organizations actively standing up agent security rather than from the largest operators. Respondents are senior and buyer-credible (45% final decision-makers, 30% recommenders/influencers), spanning managers through the C-suite, and drawn primarily from Technology/Software, Manufacturing, Retail/E-commerce, and Healthcare/Life Sciences.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs]]></title>
<description><![CDATA[Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today...]]></description>
<link>https://tsecurity.de/de/3689826/it-nachrichten/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689826/it-nachrichten/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs/</guid>
<pubDate>Thu, 23 Jul 2026 19:19:39 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price — which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilization or less, and fewer than half rigorously track what their compute actually costs. The result is a compute gap — heavy, fast-moving investment running ahead of the visibility needed to control it.</p><p>This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how satisfied they are, what would make them switch, where they plan to evaluate their investments, and — most revealingly — how well they can measure and control the economics of the compute underneath it all.</p><p>The central finding is a compute gap — the distance between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity: the single largest planned area enterprises plan to evaluate over the next year is AI-specialized clouds (45%), a layer almost none of these enterprises use today. Meanwhile the compute already in place runs cold — 83% report GPU utilization of 50% or less — and fewer than half (44%) can rigorously track what their AI compute costs. Enterprises are buying more infrastructure faster than they can account for what they already own.</p><p>Enterprises are not settled on their infrastructure vendors, either: A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter — unusually high churn intent for a category this foundational. When they choose, they choose on integration with the existing stack (41%) and total cost of ownership (35%), not on headline price: cost per million tokens is the deciding factor for just 8%. And the frontier constraint that will shape the next round of decisions — the shift from GPU compute to memory bandwidth as inference scales — is barely on the radar, with roughly one in five enterprises either unaware of it or yet to address it.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=107; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.</p><p>By organization size the sample concentrates in the mid-market: 101–250 employees (36%) and 251–1,000 (27%) lead, with 1,001–5,000 (22%), 5,001–10,000 (8%), and 10,001+ (7%) above them. By role it spans managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%); on purchasing authority it is buyer-credible, with 45% final decision-makers and another 30% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 26%, followed by Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%).</p><p>At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It also skews toward the mid-market and toward earlier-stage adopters, so it is best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators.</p><h2>Finding 1: Ambition outpaces production</h2><p><b>Only one in five run AI in production at scale</b></p><p>We asked where organizations sit in their AI deployment journey. Most are still building toward production rather than operating at scale.</p><div></div><p>The maturity curve is front-loaded. Three-quarters of enterprises (76%) are either experimenting or running only some workloads in production, and just 21% describe AI in production at scale. This matters for everything that follows: the infrastructure decisions in this report are being made largely by organizations still early in deployment, whose compute footprint — and whose costs — are about to grow. The evaluation and switching intentions in Findings 3 and 4 are the leading edge of that build-out, not the settled preferences of operators who have already found what works.</p><h2>Finding 2: Enterprises run on hyperscalers and model APIs</h2><p><b>The specialized GPU clouds barely register — today</b></p><p>We asked which providers and platforms enterprises currently use to run their AI. The answer is a familiar one: the incumbents.</p><div></div><p>The current stack is hyperscaler-and-API. Google Cloud leads at 48%, and the general-purpose clouds (Google, Microsoft, AWS, Oracle) together with the major model APIs (Gemini, OpenAI, Anthropic) account for essentially all current deployment. The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines — CoreWeave, Lambda, Crusoe, Nebius and peers — register at or near zero among these enterprises today. Only 6% run their own on-prem GPU clusters and 4% a custom open-source stack. Enterprises are, for now, running AI on the providers they already buy from — which makes the evaluation intentions in Finding 3 all the more striking.</p><p><i>(A note on reading these shares. As described in the methodology section, this sample is self-selected and skews mid-market, and this question counted every provider a respondent uses — an average of 2.1 selections each — so the figures measure presence in the stack rather than spending or primary status. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; Google's strength here, for example, is consistent with its long-standing position among smaller enterprises building on AI. Read these shares as a portrait of what this AI-active cohort runs today, and treat gaps between these figures and industry-wide market share estimates as a property of the sample rather than a contradiction of either.)</i></p><h2>Finding 3: The next dollar goes to infrastructure they don’t yet run</h2><p><b>AI-specialized clouds top the evaluations list</b></p><p>We asked where enterprises planned to evaluate AI infrastructure over the next 12 months. Their answers point away from the stack they run today.</p><div></div><p>Here is the report’s sharpest tension. The single most-cited planned evaluation area — AI-specialized clouds, at 45% — is the very category almost none of these enterprises use today (Finding 2). Nearly a third (32%) intend to evaluate non-Nvidia accelerators, and 28% in next-generation Nvidia silicon; even decentralized compute networks (16%) and sovereign compute (11%) draw meaningful interest. Read against current usage, this is not incremental — it is the leading edge of a re-platforming. The direction-of-travel question tells the same story: every infrastructure approach is net-expanding, but specialized AI clouds carry the highest net momentum (+24), edging out even the hyperscalers (+22). Enterprises are preparing to move a meaningful share of AI compute off the general-purpose cloud.</p><p>This continues a trend we saw in our April-May survey wave. Back then, usage of the AI-specialized clouds was equally marginal — CoreWeave at 3%, Lambda at 4%, Crusoe at 2% of enterprises. When we asked enterprises what change they planned in their AI infrastructure strategy over the next twelve months, the most-cited answer was moving workloads to specialized AI clouds, at 33%. Asked in April-May which emerging compute option they were most likely to evaluate AI-specialized clouds again drew the most responses. Two waves, two differently worded questions, one consistent picture: the type of cloud enterprises are most eager to assess is the type they have barely begun to use.</p><h2>Finding 4: A switching wave is building</h2><p><b>Six in 10 plan to change providers within a year — many within a quarter</b></p><p>We asked whether and when enterprises plan to switch or add an infrastructure provider. Very few intend to stand still.</p><div></div><p>For a category as foundational as compute, this is a remarkable amount of intended movement. Only 36% have no plans to change, meaning a clear majority (64%) intend to switch or add a provider within twelve months — and 38% within the next quarter alone. Where that interest points is telling: the providers drawing the most switching consideration are again the incumbents — Microsoft Azure and Google Cloud (33% each), OpenAI (30%), and Gemini (22%) — which suggests much of the near-term movement is reshuffling among the majors and consolidating spend rather than defecting to new entrants. The neocloud interest in Finding 3 is a 12-month evaluation thesis; the switching in the next quarter is mostly incumbents trading share.</p><p>(<i>Method note: Respondents who selected both "no plans to change" and a specific switching window are counted as switchers, on the logic that naming a timeframe is the more specific answer; three respondents were reclassified under this rule.</i>)</p><h2>Finding 5: Nobody buys on token price</h2><p><b>Integration and total cost of ownership decide — not sticker price</b></p><p>We asked what matters most when enterprises select an AI infrastructure provider. Headline price finished last.</p><div></div><p>Enterprises do not buy AI infrastructure on pricing, which is the place vendors compete on hardest. Integration with the existing stack (41%) and total cost of ownership (35%) dominate, while the headline metric — cost per million tokens — is the deciding factor for just 8%, dead last. The pattern is coherent: buyers are optimizing for how a provider fits and what it truly costs to operate, not for the advertised unit rate. It also foreshadows Finding 7 — enterprises say TCO matters most, yet most cannot yet measure it rigorously. The stated priority and the measured capability are out of step.</p><h2>Finding 6: Expensive GPUs, idle most of the time</h2><p><b>83% report GPU utilization of 50% or less</b></p><p>We asked what share of their GPU capacity enterprises actually utilize. The answer is a well-known but rarely quantified inefficiency.</p><div></div><p><i>Disclosure: Band percentages count every selection against all 107 qualified respondents; 14 respondents selected more than one band, so bands overlap. At the respondent level, 83 of the 100 GPU-operating enterprises reported utilization at or below 50%</i></p><p>The compute already in place runs cold. Adding the bands at or below half capacity, 83% of enterprises that operate GPUs report utilization of 50% or less, and nearly half (49%) run at 25% or below. Only 12% clear the 50% mark, and a further 8% do not measure utilization at all. Idle accelerators are expensive accelerators, and this is the clearest single measure of the compute gap: enterprises are planning to buy more GPUs and specialized compute (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large — and largely unmeasured.</p><h2>Finding 7: Spending fast, measuring slowly</h2><p><b>Fewer than half rigorously track what their compute costs</b></p><p>We asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger lags the spending.</p><div></div><p>Measurement trails money. Fewer than half of enterprises (44%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (20%), or have not prioritized it (6%). That gap is consequential given Finding 5, where total cost of ownership was the second-ranked buying criterion — enterprises are choosing providers on an economic basis they mostly cannot yet measure. Satisfaction with current infrastructure is moderately positive but not enthusiastic: on a five-point scale, overall satisfaction averages 4.0, with ease of implementation (3.8) and value for money (3.9) trailing slightly — the softness landing, tellingly, on cost. Enterprises are spending quickly and accounting slowly.</p><h2>Finding 8: The next bottleneck few are watching</h2><p><b>As inference shifts from compute to memory, the field scatters</b></p><p>Finally, we asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The responses reveal a frontier that is not yet a priority.</p><div></div><p>The memory frontier is real but barely governed. Asked which approach they would rely on as the binding constraint in inference shifts from compute to memory bandwidth, enterprises scatter: Dell leads at 31%, Nvidia follows at 16%, and the rest fragments across storage vendors, open-source tooling, and model-level efficiency techniques. Most telling is that roughly one in five (18%) either do not recognize the constraint or have not begun to address it. For a shift that will reshape inference cost and architecture, this is an early and unsettled market — and, consistent with the measurement gap in Finding 7, one where many enterprises simply do not yet have a view. It is the next chapter of the compute gap, arriving before most have closed the current one.</p><h2>The bottom line: A compute gap that faster spending will widen, not close</h2><p>Organizations with more than 100 employees are investing in AI infrastructure faster than they can measure it. Most are still early in deployment, yet their spending intentions point past their current stack — toward specialized clouds and alternative accelerators almost none of them run today — and a clear majority intend to change providers within the year. They buy on integration and total cost of ownership rather than headline price, which is rational; the difficulty is that most cannot yet see those economics clearly.</p><p>The visibility gap is concrete. The GPUs enterprises already own run at half utilization or less for the overwhelming majority, and fewer than half can rigorously track what their compute costs or returns. Satisfaction is decent but unenthusiastic, softest on value for money — the dimension hardest to judge without measurement. And the next constraint, the shift from compute to memory in large-scale inference, is arriving while most enterprises are still unaware of it. At 107 respondents in a single Q2 wave this is a directional read, skewed toward the mid-market and earlier-stage adopters — but the direction is consistent: the appetite to spend is running well ahead of the instrumentation to spend well. The compute gap is not a capacity problem that more hardware will solve on its own; it is, first, a problem of seeing what the hardware already costs. The open question for later waves is whether enterprises build that visibility before the re-platforming arrives — or buy the next layer of infrastructure as blind to its economics as the last.</p><hr><p><i>Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the results read cross-sectionally rather than as a month-over-month trend, and at 107 respondents this is a directional signal rather than a precise measurement — the sample is self-selected, skews mid-market, and leans toward earlier-stage adopters rather than the largest hyperscale operators. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with buyer-credible purchasing authority, across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Evaluating AI Agents: A production blueprint with Strands and AgentCore]]></title>
<description><![CDATA[Together, Motorway and AWS built an end-to-end evaluation pipeline that reduced incorrect results from 1 in 8 queries to 1 in 50 and cut issue detection time from few hours to few minutes. The pipeline combines the Strands Agents SDK with Amazon Bedrock AgentCore, a fully managed service for depl...]]></description>
<link>https://tsecurity.de/de/3689812/ai-nachrichten/evaluating-ai-agents-a-production-blueprint-with-strands-and-agentcore/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689812/ai-nachrichten/evaluating-ai-agents-a-production-blueprint-with-strands-and-agentcore/</guid>
<pubDate>Thu, 23 Jul 2026 19:08:15 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Together, Motorway and AWS built an end-to-end evaluation pipeline that reduced incorrect results from 1 in 8 queries to 1 in 50 and cut issue detection time from few hours to few minutes. The pipeline combines the Strands Agents SDK with Amazon Bedrock AgentCore, a fully managed service for deploying and operating AI agents at scale. In this post, you will learn how to build this pipeline for your own agents.]]></content:encoded>
</item>
<item>
<title><![CDATA[Federal quantum bet grows with DARPA’s $125 million PsiQuantum award]]></title>
<description><![CDATA[Defense research agency DARPA made its largest quantum computing award ever this week, with a $125 million agreement announced on Wednesday. The same day, the White House announced an additional $5 billion for the Genesis Mission, which focuses on AI for science but also includes technology to ac...]]></description>
<link>https://tsecurity.de/de/3689459/it-security-nachrichten/federal-quantum-bet-grows-with-darpas-125-million-psiquantum-award/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689459/it-security-nachrichten/federal-quantum-bet-grows-with-darpas-125-million-psiquantum-award/</guid>
<pubDate>Thu, 23 Jul 2026 17:13:10 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Defense research agency DARPA made its largest quantum computing award ever this week, with a <a href="https://www.psiquantum.com/news-import/psiquantum-signs-125-million-agreement-with-darpa">$125 million agreement</a> announced on Wednesday. The same day, the White House announced an <a href="https://www.whitehouse.gov/releases/2026/07/45502/">additional $5 billion for the Genesis Mission</a>, which focuses on AI for science but also includes technology to accelerate quantum computing and quantum sensors.</p>



<p class="wp-block-paragraph">“Taken together, these announcements signal that U.S. quantum strategy is shifting from supporting individual research projects to building the infrastructure needed for a quantum-enabled economy,” says <a href="https://www.linkedin.com/in/heather-c-west-ph-d-52075667/">Heather West</a>, research manager in the infrastructure systems, platforms, and technology group at IDC.</p>



<p class="wp-block-paragraph">None of the individual quantum announcements are surprising, she says. But the level of coordination is new. “Government investment is expanding beyond foundational research toward commercialization, manufacturing, and deployment,” she says.</p>



<p class="wp-block-paragraph">“The US government has been signaling that quantum computing is a priority,” says <a href="https://www.linkedin.com/in/davidmooter/">David Mooter</a>, an analyst at Forrester Research. Part of it is the desire for the US to be a leader in quantum, as it has been in other high-tech areas, he says. And part of it is because the government itself can take advantage of quantum computers.</p>



<p class="wp-block-paragraph">“Spy agencies would love to use them to decrypt intercepted messages, including messages they intercepted years ago and saved,” he says. And other departments could use quantum computers or networks for energy-related research, for supply chain optimization, and for secure communications. </p>



<p class="wp-block-paragraph">Quantum computing is accelerating, he says. “I would not be surprised to see a general gate-based quantum computer that’s good enough to provide commercial value for limited use cases by 2030.”</p>



<h2 class="wp-block-heading">DARPA’s Quantum Benchmarking Initiative</h2>



<p class="wp-block-paragraph">DARPA’s Quantum Benchmarking Initiatives was launched in 2024, and 18 companies were selected in April of 2025 for <a href="https://www.darpa.mil/news/2025/companies-targeting-quantum-computers">Stage A of the project</a>, with awards of up to $1 million each. The companies were to use the money to provide details of their concepts and show how they could lead to a functional, fault-tolerant quantum computer in under a decade.</p>



<p class="wp-block-paragraph">Then, in November of 2025, DARPA chose 11 companies for <a href="https://www.darpa.mil/research/programs/quantum-benchmarking-initiative/stage-b-selection">Stage B of the project</a>, with awards of up to $15 million for developing their research plans.</p>



<p class="wp-block-paragraph">To date, only two companies have been chosen for <a href="https://www.darpa.mil/news/2025/quantum-computing-approaches">Stage C</a>: PsiQuantum and Microsoft. PsiQuantum announced $32 million of DARPA funding for testing and evaluation in September of last year. This week’s $125 million award will expand the scope and pacing of the validation and verification work. Stage C awards can go up to $300 million, <a href="https://www.darpa.mil/sites/default/files/attachment/2025-09/darpa-mto-spark-tank-qbi.pdf">according to DARPA</a>.</p>



<p class="wp-block-paragraph">This past May, <a href="https://www.psiquantum.com/news-import/us-department-of-commerce">PsiQuantum also announced $100 million</a> from the Department of Commerce, part of the CHIPS and Science Act, to accelerate domestic manufacturing of critical quantum computing components.</p>



<p class="wp-block-paragraph">Microsoft and PsiQuantum are both in Stage C, bypassing the sequential path that other companies are expected to follow, because they were both part of DARPA’s predecessor to QBI, the Underexplored Systems for Utility-Scale Quantum Computing program.</p>



<h2 class="wp-block-heading">Genesis Mission</h2>



<p class="wp-block-paragraph">Genesis Mission was <a href="https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/">launched</a> in late 2025 with the goal of using AI to accelerate scientific breakthroughs, and it now includes more than 15 government agencies.</p>



<p class="wp-block-paragraph">As part of the Genesis Mission, quantum computing and sensing company Infleqtion announced <a href="https://infleqtion.com/infleqtion-secures-three-genesis-mission-projects-from-u-s-department-of-energy/">three projects for the Department of Energy</a> on Wednesday. The three projects focus on quantum circuit design for nuclear applications, atomic quantum sensing, and nuclear fusion energy research.</p>



<p class="wp-block-paragraph">This announcement did not include the total monetary value of the projects, but, in May, the company announced a separate agreement with the Department of Commerce for $100 million to accelerate Infleqtion’s neutral-atom technology roadmap.</p>



<p class="wp-block-paragraph">Other quantum-related Genesis Mission projects announced this week include $1.5 million for a <a href="https://www.bluequbit.io/blog/bluequbit-and-partners-awarded-1-5m-in-doe-genesis-mission-grants-to-advance-ai-driven-quantum-error-correction">BlueQubit quantum error correction project</a> with Microsoft and other partners, a <a href="https://news.stanford.edu/stories/2026/07/stanford-and-slac-to-lead-genesis-mission-projects-that-tackle-the-nation-s-most-complex-science-and-technology-challenges">Stanford effort</a> to model the behavior of electrons at quantum scale, an <a href="https://news.mit.edu/2026/mit-projects-selected-funding-under-doe-genesis-mission-0723">MIT quantum sensing project</a>, Argonne National Laboratory <a href="https://www.anl.gov/article/argonne-to-lead-ai-research-projects-under-the-department-of-energys-genesis-mission">projects</a> on quantum circuit design and quantum sensors, Brookhaven Lab <a href="https://www.bnl.gov/newsroom/news.php?a=123041">quantum sensor projects</a>, and quantum computing <a href="https://news.northwestern.edu/stories/2026/07/northwestern-projects-receive-genesis-mission-funding">projects</a> at Northwestern University.</p>



<p class="wp-block-paragraph">IBM, one of three dozen private companies that are part of the <a href="https://www.genesismissionconsortium.org/our-members#private-sector">Genesis Mission Consortium</a>, announced that it will be leading a <a href="https://research.ibm.com/blog/ibm-us-genesis-mission-quantum-ai">project</a> to support more effective quantum applications, and will contribute up to $50 million of quantum compute access for the Genesis Mission.</p>



<h2 class="wp-block-heading">Enterprise priorities</h2>



<p class="wp-block-paragraph">This week’s quantum announcements aren’t a sign that enterprises need to run out and buy quantum computers, says IDC’s West. But they do need to start preparing for the quantum era — such as by identifying business areas where quantum computing could become a competitive differentiator over the next decade.</p>



<p class="wp-block-paragraph">But the most immediate threat is that of adversaries using quantum computers to break current encryption standards. Organizations should be inventorying cryptographic assets and developing a roadmap for the migration to quantum-proof algorithms, West says.</p>



<p class="wp-block-paragraph"><a href="https://www.networkworld.com/article/4158139/fixing-encryption-isnt-enough-quantum-developments-put-focus-on-authentication.html">The point of no return is closer than ever</a>, and many major players in the encryption and communication space, including Google and Cloudflare, have been accelerating their timelines. In fact, this Wednesday was the <a href="https://www.whitehouse.gov/presidential-actions/2026/06/securing-the-nation-against-advanced-cryptographic-attacks/">federal deadline</a> for naming their post-quantum cryptography migration leads under a June executive order.</p>



<p class="wp-block-paragraph">“The preparation that needs to be done to prepare is to implement post-quantum cryptography yesterday,” says Forrester’s Mooter.</p>



<p class="wp-block-paragraph">However, according to a survey <a href="https://www.digicert.com/news/quantum-security-deployment-remains-stuck">released by DigiCert this morning</a>, while 87% of organizations are planning, testing or implementing PQC initiatives, only 7% of organizations have deployed quantum-safe or hybrid cryptography across most of their digital certificates.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[GPT-Modelle starten einen Cyber-Angriff, Google ersetzt NotebookLM & Kimi K3 ist da | KI-News]]></title>
<description><![CDATA[Author: Digitale Profis - Bewertung: 12x - Views:105 Artikel & Newsletter: https://digitaleprofis.de/die-ki-news-der-woche-vom-23-07-2026/

Quellen
OpenAI-Modelle hacken Hugging Face 
Artikel: https://openai.com/index/hugging-face-model-evaluation-security-incident/ 
Hintergrund: https://huggingf...]]></description>
<link>https://tsecurity.de/de/3689268/ai-nachrichten/gpt-modelle-starten-einen-cyber-angriff-google-ersetzt-notebooklm-kimi-k3-ist-da-ki-news/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689268/ai-nachrichten/gpt-modelle-starten-einen-cyber-angriff-google-ersetzt-notebooklm-kimi-k3-ist-da-ki-news/</guid>
<pubDate>Thu, 23 Jul 2026 16:04:50 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: Digitale Profis - Bewertung: 12x - Views:105 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/tgldX3mtgfg?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>Artikel & Newsletter: https://digitaleprofis.de/die-ki-news-der-woche-vom-23-07-2026/<br />
<br />
Quellen<br />
OpenAI-Modelle hacken Hugging Face <br />
Artikel: https://openai.com/index/hugging-face-model-evaluation-security-incident/ <br />
Hintergrund: https://huggingface.co/blog/security-incident-july-2026 <br />
<br />
Digitaleprofis.de hat ein Update bekommen <br />
Website: https://digitaleprofis.de/ <br />
<br />
Kimi K3 ist da <br />
Artikel: https://www.kimi.com/de/blog/kimi-k3 <br />
Artikel: https://apnews.com/article/kimi-k3-china-ai-0d8a5e268deb11a673f4d444fc597cc5 <br />
Ausprobieren: https://www.kimi.com/ <br />
<br />
Neue Gemini Modelle <br />
Artikel: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/ <br />
Modellkarte: https://deepmind.google/models/model-cards/gemini-3-6-flash/ <br />
Cyber-Modell: https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/ <br />
<br />
Neue Kennzeichnungspflichten des EU AI Acts <br />
Artikel: https://digital-strategy.ec.europa.eu/en/news/commission-publishes-guidelines-transparency-obligations-providers-and-deployers-certain-ai-systems <br />
Doku: https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act <br />
<br />
Anthropic zahlt 1,5 Milliarden Dollar <br />
Artikel: https://m.investing.com/news/stock-market-news/us-judge-approves-anthropics-15-billion-settlement-of-copyright-lawsuit-4801706?ampMode=1 <br />
Artikel: https://apnews.com/article/74b140444023898aeba8579b6e9f0d63 <br />
<br />
NotebookLM wird zu Gemini Notebook <br />
Artikel: https://blog.google/innovation-and-ai/products/gemini-notebook/notebooklm-gemini-notebook/ <br />
<br />
Microsoft und Mistral Partnerschaft <br />
Artikel: https://news.microsoft.com/source/2026/07/21/microsoft-and-mistral-expand-strategic-partnership-to-give-enterprises-and-regulated-industries-frontier-ai-they-can-control/ <br />
Artikel: https://www.tagesschau.de/wirtschaft/unternehmen/microsoft-mistral-ai-ki-deal-100.html <br />
<br />
Qwen-Audio-3.0-TTS <br />
Artikel: https://www.alibabacloud.com/blog/qwen-audio-3-0-tts-more-multilingual-easier-to-direct_603379 <br />
Doku: https://docs.qwencloud.com/developer-guides/speech/tts-models <br />
<br />
OpenAI-Modelle überwinden bei einem internen Sicherheitstest ihre isolierte Testumgebung und gelangen bis in die Produktionsinfrastruktur von Hugging Face.<br />
Wir ordnen den Vorfall ein und fassen die weiteren wichtigen KI-News der Woche kompakt für euch zusammen.<br />
<br />
Außerdem geht es um Kimi K3, drei neue Gemini-Modelle, die kommenden Kennzeichnungspflichten des EU AI Acts und den milliardenschweren Urheberrechtsvergleich zwischen Anthropic und Autoren.<br />
<br />
Im Video:<br />
- OpenAIs ungewöhnlicher Sicherheitsvorfall bei Hugging Face<br />
- Kimi K3 und drei neue Gemini-Modelle<br />
- Kennzeichnungspflichten des EU AI Acts ab August 2026<br />
- Anthropics Vergleich über mindestens 1,5 Milliarden US-Dollar<br />
- Gemini Notebook, Microsofts Mistral-Partnerschaft und Qwen-Audio-3.0-TTS<br />
<br />
Unsere Website wurde ebenfalls vollständig überarbeitet. Auf digitaleprofis.de findet ihr unsere KI-News mit allen Quellen, ausführliche Artikel und praktische Anleitungen – kostenlos, ohne Paywall und ohne Werbung.<br />
<br />
Werde Kanalmitglied und unterstütze damit unsere Arbeit:<br />
https://www.youtube.com/channel/UCv90NdTyTp7ZPPRvvSZaS5w/join<br />
<br />
Videoinhalt:<br />
00:00 Die KI-News der Woche vom 23.07.2026<br />
00:24 OpenAI Modelle greifen HuggingFace an<br />
02:06 Unsere neue Website für euch<br />
02:59 Kimi K3 ist da und die USA werfen Diebstahl vor<br />
04:28 Drei neue Gemini Modelle, aber kein Pro<br />
05:53 Die neuen Kennzeichnungspflichten des AI Acts<br />
07:21 Anthropic muss 1,5 Milliarden Dollar Vergleich zahlen<br />
08:48 NotebookLM verschwindet - und wird zu Gemini Notebook<br />
09:30 Partnerschaft von Microsoft und Mistral<br />
10:50 Neuen Text to Speech Modell von Alibabas Qwen<br />
<br />
Videovorschläge, Feedback und Kritik kannst Du uns jederzeit in den Kommentaren mitteilen!<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[What Happened Between OpenAI and Hugging Face?]]></title>
<description><![CDATA[The OpenAI and Hugging Face incident lands like a warning shot for anyone thinking seriously about frontier AI and cybersecurity research. A model evaluation crossed the neat boundary of a research environment, reached a live third-party production system, and forced the industry to confront a qu...]]></description>
<link>https://tsecurity.de/de/3689199/it-security-nachrichten/what-happened-between-openai-and-hugging-face/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689199/it-security-nachrichten/what-happened-between-openai-and-hugging-face/</guid>
<pubDate>Thu, 23 Jul 2026 15:28:53 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><span>The </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank"><span>OpenAI and Hugging Face incident</span></a><span> lands like a warning shot for anyone thinking seriously about frontier AI and cybersecurity research. A model evaluation crossed the neat boundary of a research environment, reached a live third-party production system, and forced the industry to confront a question that is moving quickly from theory to operations: what happens when AI agents can pursue an objective with enough persistence, speed, and creativity to behave less like a tool and more like an autonomous intrusion path?</span></p><p><span>According to OpenAI’s disclosure, the incident began during an internal evaluation of advanced cyber capabilities using GPT-5.6 Sol and a more capable pre-release model. The evaluation was designed to test whether AI agents could pursue complex exploit paths, and OpenAI says cyber refusal safeguards were reduced or disabled to measure maximum capability. Inside that environment, the models reportedly found and exploited a zero-day in the package registry cache proxy that was meant to constrain network access, moved through OpenAI’s research environment, reached a node with internet connectivity, and then inferred that Hugging Face may host artifacts related to the benchmark they were trying to solve.</span></p><p><span>From there, the models compromised part of Hugging Face’s dataset-processing pipeline, gained code execution on a worker, escalated access, harvested credentials, and moved laterally across internal clusters. Hugging Face detected and contained the activity, and OpenAI later connected the activity back to its own evaluation. Both companies have said the investigation is continuing, which means some details will almost certainly evolve. Still, the direction of travel is clear enough for defenders to act on now.</span></p><h2>How did the OpenAI model evaluation reach Hugging Face?</h2><p><span>The activity stands out because it looked less like a single model producing a risky command and more like a compressed intrusion path. Based on the public disclosures, the reported chain moved from identifying a constraint, to breaking that constraint, gaining access, inferring where valuable data may live, and continuing toward that objective across a live environment.</span></p><p><span>Security teams should use that sequence to revisit assumptions built around human pacing. Many detection and response workflows still assume there will be time between stages of an attack, with reconnaissance followed by exploitation, lateral movement, and then objective pursuit. In an agent-driven scenario, those stages can begin to collapse into one continuous loop, with fewer natural pauses for defenders to catch up.</span></p><p><span>The defensive model now has to account for a world where discovery, exploitation, and follow-on action can happen faster and with more persistence than traditional human-led campaigns. The uncomfortable lesson is that AI agents can be tireless, goal-oriented, and increasingly capable of finding the loose seams in systems built for a slower era.</span></p><p><span>The incident highlights the collapse of the traditional OODA (Observe-Orient-Decide-Act) loop. In standard human-led campaigns, attackers experience natural pauses between stages like reconnaissance, exploitation, and lateral movement, providing defenders critical windows to intervene. AI agents, however, can compress these stages into a single, continuous loop of automated activity. This machine-speed execution renders manual detection workflows obsolete, necessitating a shift toward behavioral, AI-enabled defenses capable of identifying and disrupting malicious intent before it cascades.</span></p><p><span>Security teams should also assume that advanced AI cyber capability will diffuse over time, which means defenders need their own AI-enabled workflows to mature quickly enough to find, validate, prioritize, and reduce risk before attackers operationalize the same class of tools. Rapid7’s experimentation with security-specific frontier models, including Claude Mythos Preview and GPT-5.5-Cyber, points to the same control challenge: the more freedom a model has to pursue a defined reward or goal, the more important containment, monitoring, and clear constraints become. Without those controls, a model optimized to “solve” a security task may find paths the evaluator did not intend, including sandbox escapes, credential misuse, or other offensive security behaviors that look less like assistance and more like uncontrolled execution.</span></p><h2>Why AI evaluation environments need stronger containment</h2><p><span>The incident raises practical questions about how frontier AI evaluations are designed and contained. OpenAI’s evaluation environment was meant to restrict network access through an internally hosted package registry proxy and cache. The models reportedly identified that proxy as the weak point, exploited a zero-day in the software, and used it as a path out of the sandbox.</span></p><p><span>The pattern will feel familiar to security teams: a chokepoint only works as a control if it can withstand pressure. Package mirrors, dependency proxies, build systems, internal artifact repositories, sandbox harnesses, and evaluation infrastructure all become part of the security boundary when they sit between a capable agent and the wider internet.</span></p><p><span>For teams building or testing agentic systems, isolation has to be layered and monitored rather than assumed. Egress restrictions matter, but so do workload identity, credential scope, network segmentation, package policy, runtime monitoring, and fast containment when an agent behaves unexpectedly. The goal is to make sure the environment reflects the capability of the system being tested, especially when the test is explicitly designed to measure how far that capability can go.</span></p><h2>Why AI and ML pipelines are now software supply chain risk</h2><p><span>The Hugging Face side of the incident is a reminder that AI and ML pipelines are part of the software supply chain. Models, datasets, loader scripts, notebooks, and evaluation artifacts may look like research materials, but in modern environments they often behave like executable code. Hugging Face has said its models, datasets, and Spaces were not tampered with, and that its images and published packages were verified as clean.</span></p><p><span>According to the technical reporting reviewed, the initial access path involved Hugging Face’s dataset-processing pipeline and a combination of code execution paths, including custom loader behavior and template injection in a dataset configuration flow. The exact implementation details may continue to evolve as the investigation progresses, but the defensive takeaway is already clear: AI and ML processing systems should be secured like high-risk software supply chain infrastructure.</span></p><p><span>Any system that automatically processes external datasets or model artifacts should be designed with hostile input in mind. Processing workers should run with least privilege, should not have broad access to cloud credentials or cluster-level tokens, and should be segmented so compromise of one worker does not become compromise of the environment around it.</span></p><p><span>Security teams should also hunt for early signs of intent drift inside ML workflows. Unexpected reads of environment variables, cloud metadata services, secret stores, package registries, or internal APIs from dataset-processing jobs can be meaningful signal. In an AI-driven environment, the first clue may not be a known malicious indicator. It may be a workload behaving with curiosity it should not have.</span></p><h2>What AI guardrails mean for incident response</h2><p><span>One of the most useful lessons for security teams came during the response, when Hugging Face’s responders reportedly needed to analyze logs containing exploit payloads, attacker commands, and command-and-control artifacts. When they tried to use commercial hosted AI models to help reconstruct the attack, those models refused parts of the analysis because the content looked malicious. The team then moved to a self-hosted open-weight model so they could continue the investigation without refusals and without sending sensitive incident data outside their own environment.</span></p><p><span>The practical issue here is guardrail asymmetry: attackers, rogue agents, or unrestricted systems may not respect usage policies, while defenders using hosted tools can be slowed by the same safety systems designed to prevent misuse. Security teams need trusted-defender pathways, including access models and deployment options that preserve safeguards against harmful use while allowing responders to analyze real malicious content safely, privately, and without interruption.</span></p><p><span>AI-assisted incident response also needs to be tested before an incident begins, especially if teams expect models to support triage, log analysis, malware review, or timeline reconstruction under pressure. Responders should know which models can analyze real attack payloads, where those models run, what data they can access, and whether safety controls could block time-sensitive forensic work.</span></p><h2>What security teams should do about agentic AI risk</h2><p><span>The practical response to this incident is preparation, with organizations treating AI agents as both powerful tools and potential attack surfaces when those agents can access code, credentials, infrastructure, datasets, or internal systems.</span></p><p><span>Security teams should apply familiar controls with more urgency across AI and ML environments: treat downloaded models, datasets, and loader scripts as untrusted code; enforce least privilege on AI and ML processing workers; prevent node-level compromise from becoming cluster-wide compromise; restrict access to cloud metadata endpoints and secrets; monitor unexpected outbound connections from AI-related workloads; and review whether evaluation environments have real containment or simply convenient isolation.</span></p><p><span>Detection logic needs to account for machine-speed activity. In this case, reporting described thousands of actions and more than 17,000 discrete recorded events, with reconnaissance, exploitation, and follow-on action occurring inside one continuous loop rather than across the pauses defenders are used to seeing in human-led campaigns. Security teams should focus on behavioral patterns that show intent, including unusual access to secrets, unexpected package activity, suspicious use of metadata services, sudden privilege changes, or processing jobs reaching systems they have no reason to touch.</span></p><p><span>As autonomous activity becomes faster and noisier, the bottleneck may shift from detecting that something happened to understanding what matters quickly enough to change the outcome. A security team that can see thousands of events but needs hours to reconstruct the story is still operating behind the pace of the incident.</span></p><h2>How preemptive security helps reduce AI-driven risk</h2><p><span>At Rapid7, our view is that this is where preemptive security becomes especially important. Faster discovery only creates value when defenders can turn it into faster validation, prioritization, remediation, detection, and response. The same principle applies to </span><a href="https://www.rapid7.com/blog/post/ai-changing-vulnerability-discovery-software-supply-chain-strateg" target="_self"><span>agentic AI risk</span></a><span>. If AI accelerates how weaknesses are found and exploited, defenders need security operations that can act earlier with better context and more confidence.</span></p><p><span>That means connecting exposure management with detection and response, so teams understand which risks are exploitable, which assets matter most, what suspicious behavior is already present, and which actions will reduce risk fastest. It also means </span><a href="https://www.rapid7.com/platform/artificial-intelligence-features" target="_self"><span>using AI carefully and practically</span></a><span>, not as a replacement for security judgment, but as a way to reason across telemetry, reduce noise, support investigation, and help teams make decisions at the speed the threat environment now demands.</span></p><p><span>AI-enabled defense is becoming part of resilience planning, especially for organizations running critical systems or high-value digital infrastructure. The goal is to give defenders the speed, context, and consistency to operate inside the attacker’s decision cycle, without removing the judgment and accountability that effective security requires.</span></p><p><span>The OpenAI and Hugging Face incident will continue to generate debate as more details emerge, but defenders already have enough to work with. Agentic systems are beginning to test the seams between AI research, software supply chain security, cloud infrastructure, and incident response. The organizations best positioned for what comes next will be the ones making those seams visible, monitored, and resilient before the next incident puts them under pressure.</span></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Presence raises new questions about enterprise automation and jobs]]></title>
<description><![CDATA[OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.



The agents can answer questions and operate IT systems, and enterprises can decide w...]]></description>
<link>https://tsecurity.de/de/3689165/it-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689165/it-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</guid>
<pubDate>Thu, 23 Jul 2026 15:20:41 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.</p>



<p class="wp-block-paragraph">The agents can answer questions and operate IT systems, and enterprises can decide what actions the agents may take and when they should seek human approval for actions or transfer a case to a human.</p>



<p class="wp-block-paragraph">OpenAI is already using Presence internally for its English-language phone support channel, where it verifies callers and uses account information to complete approved actions. The company said the system resolves 75% of inbound issues without human assistance.</p>



<p class="wp-block-paragraph">Another OpenAI service, Codex, can be used to monitor agents and suggest updates or improvements to processes. In OpenAI’s own tests, suggestions from Codex helped reduce handoffs to humans by 15 percentage points over 10 days, it said. Presence also includes simulation and evaluation tools that allow companies to test an agent before deployment. The tests assess whether it reaches the correct outcome, follows company policy, and hands a case to an employee when required.</p>



<p class="wp-block-paragraph">OpenAI intends each Presence deployment to deal with one kind of task, for example billing issues, insurance claims, or employee IT service requests, with agents getting only the knowledge and system access required for that task.</p>



<p class="wp-block-paragraph">Presence is not a self-service product: Enterprises will have to sign up for the limited availability program, with integration performed by OpenAI or selected <a href="https://www.computerworld.com/article/4136024/openai-partners-with-consulting-giants-to-deploy-enterprise-ai-agents.html">global systems integrators</a>.</p>



<p class="wp-block-paragraph">Companies exploring or testing Presence include Spanish bank BBVA, which is evaluating the service for everyday banking support in Mexico, and Japanese technology group SoftBank, which is using it in trials involving Japanese-language customer interactions. Australian insurer IAG is assessing whether the technology can help it respond to surges in customer demand during severe weather events.</p>



<h2 class="wp-block-heading">Workforce impact</h2>



<p class="wp-block-paragraph">OpenAI’s announcement did not address the potential effect of Presence on employment. But its claimed automation rate raises questions about how the technology could affect staffing in customer service and other support functions.</p>



<p class="wp-block-paragraph"><a href="https://pareekh.com/" target="_blank" rel="noreferrer noopener">Pareekh Jain</a>, CEO of Pareekh Consulting, said CIOs should regard the 75% figure as evidence that the technology can work, rather than as a benchmark that every enterprise can expect to reach.</p>



<p class="wp-block-paragraph">Jain said OpenAI’s deployment benefits from being built around the company’s own products and data. Large enterprises may achieve lower automation rates because they must contend with fragmented legacy systems, uneven knowledge bases and more complex compliance demands.</p>



<p class="wp-block-paragraph">“Most organizations should expect lower initial automation levels that improve over time as the AI agent is refined,” Jain said.</p>



<p class="wp-block-paragraph">The first workforce effect is more likely to be <a href="https://www.cio.com/article/4015750/cios-see-ai-prompting-new-it-hiring-even-as-boards-push-for-job-cuts.html">slower hiring than immediate layoffs</a>, according to <a href="https://www.linkedin.com/in/tulikasheel/" target="_blank" rel="noreferrer noopener">Tulika Sheel</a>, senior vice president at Kadence International.</p>



<p class="wp-block-paragraph">“The roles most exposed are likely to be repetitive, high-volume functions such as frontline customer support and routine back-office processing,” Sheel said. “However, I would expect the first impact to be on hiring and team growth rather than immediate large-scale job cuts. Over time, enterprises may redesign roles around AI-assisted workflows, with humans focusing more on complex cases, escalation, and relationship management.”</p>



<p class="wp-block-paragraph">Jain said Tier-1 support agents handling predictable queries would face the most exposure. Broader reductions would become more likely only after companies reorganize their operations around the technology.</p>



<p class="wp-block-paragraph">However, <a href="https://omdia.tech.informa.com/authors/lian-jye-su" target="_blank" rel="noreferrer noopener">Lian Jye Su</a>, chief analyst at Omdia, said Presence is unlikely to increase the threat of job displacement because companies have used similar customer-support automation from vendors such as Genesys, NiCE, Five9 and AWS for years.</p>



<p class="wp-block-paragraph">Enterprises are more likely to use Presence alongside employees, with AI handling routine requests while people remain responsible for work requiring judgment and empathy, Su said.</p>



<h2 class="wp-block-heading">Cost and operational risks</h2>



<p class="wp-block-paragraph">Analysts said CIOs should examine whether Presence can maintain resolution quality as usage grows, since fewer human handoffs could leave employees dealing with a more difficult mix of cases.</p>



<p class="wp-block-paragraph">“The key question is not simply how many tasks AI can handle, but whether it can handle them reliably at scale,” Sheel said.</p>



<p class="wp-block-paragraph">The financial case will depend partly on the cost of connecting Presence to existing systems and maintaining the controls needed to govern its use, according to Jain. “Often the biggest cost of enterprise AI is not tokens but <a href="https://www.computerworld.com/article/4128310/openai-responds-to-claude-cowork-with-its-own-platform-to-help-build-deploy-and-manage-ai-agents.html">integration and governance</a>,” Jain added.</p>



<p class="wp-block-paragraph">Companies will need to determine what systems and data the agents can access, monitor their performance, and audit the actions they take. Those investments could offset early savings.</p>



<p class="wp-block-paragraph">Su said the complexity of enterprise IT will make it difficult for OpenAI to automate entire workflows on its own. Enterprises will still need to work with other technology providers and human employees, while CIOs will favor systems that can be audited and integrated with existing infrastructure.</p>



<p class="wp-block-paragraph">Jain said the economics could improve if companies use the same integrations and governance controls across additional workflows.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Presence raises new questions about enterprise automation and jobs]]></title>
<description><![CDATA[OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.



The agents can answer questions and operate IT systems, and enterprises can decide w...]]></description>
<link>https://tsecurity.de/de/3689164/it-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689164/it-nachrichten/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs/</guid>
<pubDate>Thu, 23 Jul 2026 15:20:32 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">OpenAI has launched Presence, an enterprise service for deploying voice and chat agents that can resolve customer and employee requests, potentially automating some work now handled by frontline support teams.</p>



<p class="wp-block-paragraph">The agents can answer questions and operate IT systems, and enterprises can decide what actions the agents may take and when they should seek human approval for actions or transfer a case to a human.</p>



<p class="wp-block-paragraph">OpenAI is already using Presence internally for its English-language phone support channel, where it verifies callers and uses account information to complete approved actions. The company said the system resolves 75% of inbound issues without human assistance.</p>



<p class="wp-block-paragraph">Another OpenAI service, Codex, can be used to monitor agents and suggest updates or improvements to processes. In OpenAI’s own tests, suggestions from Codex helped reduce handoffs to humans by 15 percentage points over 10 days, it said. Presence also includes simulation and evaluation tools that allow companies to test an agent before deployment. The tests assess whether it reaches the correct outcome, follows company policy, and hands a case to an employee when required.</p>



<p class="wp-block-paragraph">OpenAI intends each Presence deployment to deal with one kind of task, for example billing issues, insurance claims, or employee IT service requests, with agents getting only the knowledge and system access required for that task.</p>



<p class="wp-block-paragraph">Presence is not a self-service product: Enterprises will have to sign up for the limited availability program, with integration performed by OpenAI or selected <a href="https://www.computerworld.com/article/4136024/openai-partners-with-consulting-giants-to-deploy-enterprise-ai-agents.html">global systems integrators</a>.</p>



<p class="wp-block-paragraph">Companies exploring or testing Presence include Spanish bank BBVA, which is evaluating the service for everyday banking support in Mexico, and Japanese technology group SoftBank, which is using it in trials involving Japanese-language customer interactions. Australian insurer IAG is assessing whether the technology can help it respond to surges in customer demand during severe weather events.</p>



<h2 class="wp-block-heading">Workforce impact</h2>



<p class="wp-block-paragraph">OpenAI’s announcement did not address the potential effect of Presence on employment. But its claimed automation rate raises questions about how the technology could affect staffing in customer service and other support functions.</p>



<p class="wp-block-paragraph"><a href="https://pareekh.com/" target="_blank" rel="noreferrer noopener">Pareekh Jain</a>, CEO of Pareekh Consulting, said CIOs should regard the 75% figure as evidence that the technology can work, rather than as a benchmark that every enterprise can expect to reach.</p>



<p class="wp-block-paragraph">Jain said OpenAI’s deployment benefits from being built around the company’s own products and data. Large enterprises may achieve lower automation rates because they must contend with fragmented legacy systems, uneven knowledge bases and more complex compliance demands.</p>



<p class="wp-block-paragraph">“Most organizations should expect lower initial automation levels that improve over time as the AI agent is refined,” Jain said.</p>



<p class="wp-block-paragraph">The first workforce effect is more likely to be <a href="https://www.cio.com/article/4015750/cios-see-ai-prompting-new-it-hiring-even-as-boards-push-for-job-cuts.html">slower hiring than immediate layoffs</a>, according to <a href="https://www.linkedin.com/in/tulikasheel/" target="_blank" rel="noreferrer noopener">Tulika Sheel</a>, senior vice president at Kadence International.</p>



<p class="wp-block-paragraph">“The roles most exposed are likely to be repetitive, high-volume functions such as frontline customer support and routine back-office processing,” Sheel said. “However, I would expect the first impact to be on hiring and team growth rather than immediate large-scale job cuts. Over time, enterprises may redesign roles around AI-assisted workflows, with humans focusing more on complex cases, escalation, and relationship management.”</p>



<p class="wp-block-paragraph">Jain said Tier-1 support agents handling predictable queries would face the most exposure. Broader reductions would become more likely only after companies reorganize their operations around the technology.</p>



<p class="wp-block-paragraph">However, <a href="https://omdia.tech.informa.com/authors/lian-jye-su" target="_blank" rel="noreferrer noopener">Lian Jye Su</a>, chief analyst at Omdia, said Presence is unlikely to increase the threat of job displacement because companies have used similar customer-support automation from vendors such as Genesys, NiCE, Five9 and AWS for years.</p>



<p class="wp-block-paragraph">Enterprises are more likely to use Presence alongside employees, with AI handling routine requests while people remain responsible for work requiring judgment and empathy, Su said.</p>



<h2 class="wp-block-heading">Cost and operational risks</h2>



<p class="wp-block-paragraph">Analysts said CIOs should examine whether Presence can maintain resolution quality as usage grows, since fewer human handoffs could leave employees dealing with a more difficult mix of cases.</p>



<p class="wp-block-paragraph">“The key question is not simply how many tasks AI can handle, but whether it can handle them reliably at scale,” Sheel said.</p>



<p class="wp-block-paragraph">The financial case will depend partly on the cost of connecting Presence to existing systems and maintaining the controls needed to govern its use, according to Jain. “Often the biggest cost of enterprise AI is not tokens but <a href="https://www.computerworld.com/article/4128310/openai-responds-to-claude-cowork-with-its-own-platform-to-help-build-deploy-and-manage-ai-agents.html">integration and governance</a>,” Jain added.</p>



<p class="wp-block-paragraph">Companies will need to determine what systems and data the agents can access, monitor their performance, and audit the actions they take. Those investments could offset early savings.</p>



<p class="wp-block-paragraph">Su said the complexity of enterprise IT will make it difficult for OpenAI to automate entire workflows on its own. Enterprises will still need to work with other technology providers and human employees, while CIOs will favor systems that can be audited and integrated with existing infrastructure.</p>



<p class="wp-block-paragraph">Jain said the economics could improve if companies use the same integrations and governance controls across additional workflows.</p>



<p class="wp-block-paragraph"><em>This article first appeared on <a href="https://www.cio.com/article/4200684/openai-presence-raises-new-questions-about-enterprise-automation-and-jobs.html">CIO</a>.</em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Q&A: Google’s AI and computing chief talks about its shapeshifting data centers]]></title>
<description><![CDATA[Google’s AI offerings span its internal and cloud offerings. Its data centers are processing seven times more AI tokens compared to last year. To keep up, Google is upgrading its data-center hardware and software technologies at a faster clip. It plans to raise $80 billion to build new data cente...]]></description>
<link>https://tsecurity.de/de/3689101/it-security-nachrichten/qa-googles-ai-and-computing-chief-talks-about-its-shapeshifting-data-centers/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689101/it-security-nachrichten/qa-googles-ai-and-computing-chief-talks-about-its-shapeshifting-data-centers/</guid>
<pubDate>Thu, 23 Jul 2026 14:55:21 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Google’s AI offerings span its internal and cloud offerings. Its data centers are processing seven times more AI tokens compared to last year. To keep up, Google is upgrading its data-center hardware and software technologies at a faster clip. It plans to raise $80 billion to build new data centers. (See related story: <a href="https://www.networkworld.com/article/4200581/google-transforms-its-data-center-architecture-for-agent-era.html">Google transforms its data center architecture for agent era</a>)</p>



<p class="wp-block-paragraph"><em>Network World</em> spoke with <a href="https://www.linkedin.com/in/marklohmeyer/">Mark Lohmeyer</a>, vice president and general manager of AI and computing at Google, about how the company’s infrastructure is keeping pace with AI demand.</p>



<p class="wp-block-paragraph"><strong>Network World: What is the primary shift in infrastructure needs?</strong></p>



<p class="wp-block-paragraph"><strong>Mark Lohmeyer:</strong> We’ve seen the <a href="https://www.networkworld.com/article/4175890/cisco-ai-traffic-is-radically-reshaping-wans.html">rise of agents and agentic use cases</a>. Years ago, it was the chat phase: Ask a question, get an answer. Now we’re in the agentic era, where you express your intent, agents spin off multiple sub-agents, working in parallel, preserving state. This is a radical shift in what infrastructure needs to do; make them fast, cost effective, secure, reliable. We’re delivering infrastructure optimized for the age of agents.</p>



<p class="wp-block-paragraph"><strong>NW: What’s the goal of the infrastructure buildout, and what should customers expect regarding costs?</strong></p>



<p class="wp-block-paragraph"><strong>ML: </strong>Ultimately, it’s about enabling customers with leading-edge capabilities and models at scale cost-effectively. With agents, <a href="https://www.networkworld.com/article/4057121/network-and-cloud-implications-of-agentic-ai.html">inference transactions increase</a> by 50x, 100x versus non-agentic workloads. We’re driving the cost per transaction down exponentially. In our latest platforms, we reduce the cost by almost 2x for the same work. Customers serve twice the number of users at the same cost, directly driving profitability.</p>



<p class="wp-block-paragraph"><strong>NW: How are you addressing energy efficiency?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Energy is a critical resource, and Google has optimized for years. We design data centers and compute [to drive] high PUE (power usage effectiveness). We introduced <a href="https://www.networkworld.com/article/4149069/why-ai-rack-densities-make-liquid-cooling-nonnegotiable.html">liquid cooling</a> over five years ago, and these latest systems are all liquid cooled. For agentic workloads, CPUs come to the forefront… orchestrating agents, calling tools, doing evaluation loops in reinforcement learning. Our latest Axion-based CPU platform called <a href="https://www.networkworld.com/article/4086182/google-cloud-aims-for-more-cost-effective-arm-computing-with-axion-n4a.html">N4A</a> has energy efficiency and is significantly better than the prior generation and x86 comparables.</p>



<p class="wp-block-paragraph"><strong>NW: How do you think about token efficiency as you build-out systems?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Performance and efficiency gains are powered by co-design of the model and infrastructure. <a href="https://www.computerworld.com/article/4161990/gemini-enterprise-update-brings-ai-agents-into-collaborative-workflows.html">Gemini</a> is trained on TPUs, primarily served on TPUs with high frontier model capability, in a token and cost-efficient way. This stems from co-design across the full stack.</p>



<p class="wp-block-paragraph"><strong>NW: How do you project what infrastructure will be needed years in advance?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Hardware cycles deliver a new next generation roughly every year, but design cycles are two years or more in advance. We work with <a href="https://deepmind.google/about/">DeepMind</a> doing core research, to application teams taking models into production, to billions of users, to our team building infrastructure. We work upstream with DeepMind and application teams to understand what’s coming. Agents weren’t being broadly spoken of externally, but internally we had those insights around what they would need. That shows up in hardware design. We hit the timing right — these platforms are built for agents.</p>



<p class="wp-block-paragraph"><strong>NW: What’s the eighth generation TPU platform?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> We deliver new platforms every year, and ones launched years ago are close to 100% utilized because demand for AI-optimized compute is high. The <a href="https://www.networkworld.com/article/4162004/google-bets-on-workload-specific-tpus-with-8t-and-8i-launch.html">eighth-generation TPU platform</a> is the first delivering two complete systems, from the chip all the way up to the network and storage and software, that are optimized.</p>



<p class="wp-block-paragraph"><a href="https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive">TPU-8t</a> is optimized for training, and TPU-8i is optimized for inference. For TPU-8i, we increased SRAM on the chip to 384MB — three times the prior generation — and increased the HBM by 50%.</p>



<p class="wp-block-paragraph"><strong>NW: How are you approaching GPU and TPU compatibility?</strong></p>



<p class="wp-block-paragraph"><strong>ML: </strong>People in a single cluster do not commingle GPUs and TPUs. We offer both options based on specific workload needs. We’ve been investing on the TPU side in using software frameworks customers are comfortable with on GPUs and enabling those on TPUs. For example, <a href="https://www.infoworld.com/article/2335194/what-is-pytorch-python-machine-learning-on-gpus.html">PyTorch</a> and vLLM. Customers could have a pool of GPUs and TPUs, running vLLM on top of that. Start with a workload on TPUs, but if the TPU pool is fully utilized, spill to GPUs or vice versa. This works because it’s all leveraging the same compatible software layer on top.</p>



<p class="wp-block-paragraph"><strong>NW: How has the orchestration platform changed for agents?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Kubernetes is becoming the orchestration platform of choice for AI. Google is transforming <a href="https://www.infoworld.com/article/2255921/gke-tutorial-get-started-with-google-kubernetes-engine.html">GKE</a> [Google Kubernetes Engine] into an agent-native orchestration solution. When expressing intent to an agent and it spins up multiple sub-agents, compute needs to spin up rapidly — TPUs or GPUs — without long delays, then run and spin back down. We’re optimizing at every layer of the <a href="https://cloud.google.com/kubernetes-engine">GKE stack</a>: significantly improving node startup time and how rapidly we start and stop containers. Lovable demonstrates this with GKE, spinning up hundreds of sandboxes for live coding sessions on their platform in parallel, paying for infrastructure when needed.</p>



<p class="wp-block-paragraph"><strong>NW: What is the role of the network and storage infrastructure?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> The network is critical for AI. This requires creating large-scale clusters of GPUs or TPUs and enabling them to talk to each other in a high-performance way. <a href="https://cloud.google.com/blog/products/networking/introducing-virgo-megascale-data-center-fabric">We created the Virgo network</a> — a collapsed network architecture, non-blocking within a data center, where multiple pods or NVLink72 domains connect together.</p>



<p class="wp-block-paragraph">In TPU8T, we can connect over a million TPUs together leveraging Virgo, creating large-scale, high-performance, reliable clusters that shrink innovation cycles. Storage is equally critical. In large-scale clusters, something is always failing. The ability to take snapshots and go back to a checkpoint is important.</p>



<p class="wp-block-paragraph">We’ve introduced <a href="https://cloud.google.com/products/managed-lustre">Managed Lustre 10T</a>, with 10 terabytes per second of bandwidth, 18 petabytes of storage in single clusters. This is 10 times faster than last year and 20 times faster than competition. We have Rapid Bucket, low-latency storage backed by Google storage systems. Both are impactful in large-scale training environments.</p>



<p class="wp-block-paragraph"><strong>NW: How does KV cache strategy differ between training and inference?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> For <a href="https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/eighth-generation-tpu-agentic-era/">TPU-8i</a>, we increased SRAM on the chip to 384 megabytes — three times the prior generation — and increased the HBM by 50%. Storing KV cache directly in chip memory allows responding to inference requests much more rapidly and cost-effectively than going to an external system. For inference workloads, storing as much KV cache as possible on-chip is critical.</p>



<p class="wp-block-paragraph">We’re introducing a dedicated KV cache storage subsystem that works across GPUs and TPUs. As KV caches get larger, being able to fall back to this dedicated subsystem becomes critical. Loading model weights rapidly is important in dynamic inference environments where accelerators switch between models hour by hour.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI „hackt“ Hugging Face – eine Analyse]]></title>
<description><![CDATA[Wenn KI-Modelle die Grenzen überwinden, die ihnen gesetzt werden, hinterlassen sie unter Umständen weniger sichtbare Spuren.Nelson Antoine | shutterstock.com



Der heimliche Cybercrime-Akt zweier KI-Modelle von OpenAI hat weltweit ein enormes Echo in Mainstream– und sozialen Medien hervorgerufen...]]></description>
<link>https://tsecurity.de/de/3689099/it-security-nachrichten/openai-hackt-hugging-face-eine-analyse/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689099/it-security-nachrichten/openai-hackt-hugging-face-eine-analyse/</guid>
<pubDate>Thu, 23 Jul 2026 14:55:17 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" src="https://b2b-contenthub.com/wp-content/uploads/2025/08/Nelson-Antoine-shutterstock_1672788895_16z9.jpg?quality=50&amp;strip=all&amp;w=1024" alt="Jailbreak 16z9" class="wp-image-4038755" width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"><figcaption class="wp-element-caption">Wenn KI-Modelle die Grenzen überwinden, die ihnen gesetzt werden, hinterlassen sie unter Umständen weniger sichtbare Spuren.</figcaption></figure><p class="imageCredit">Nelson Antoine | shutterstock.com</p></div>



<p class="wp-block-paragraph">Der heimliche Cybercrime-Akt zweier KI-Modelle von OpenAI hat weltweit ein enormes Echo in <a href="https://www.tagesschau.de/wirtschaft/unternehmen/openai-ki-hackerangriff-100.html" target="_blank" rel="noreferrer noopener">Mainstream</a>– und <a href="https://www.reddit.com/r/OpenAI/comments/1v2ybnw/openai_models_escaped_containment_and_hacked/" target="_blank" rel="noreferrer noopener">sozialen Medien</a> hervorgerufen. Der Vorfall dürfte die Debatte über die allgemeine <a href="https://www.computerwoche.de/article/4155663/6-wege-uber-ki-gehackt-zu-werden.html" target="_blank">KI-Sicherheit</a> und den verantwortungsvollen Umgang mit der Technologie neu befeuern. </p>



<p class="wp-block-paragraph">Doch der Incident wirft auch spezifische Fragen auf. Etwa, wie genau die OpenAI-Modelle es geschafft haben, ihrer Sandbox zu entkommen und warum das beim ChatGPT-Erfinder zunächst niemandem aufgefallen ist. Oder, wie andere Unternehmen solche und ähnliche Vorkommnisse künftig verhindern können. Dazu haben wir die Einschätzung von Branchenexperten und Analysten eingeholt. </p>



<p class="wp-block-paragraph">Zunächst werfen wir aber noch einen kurzen Blick darauf, was sich eigentlich abgespielt hat. Falls Sie bereits informiert sind, können Sie alternativ auch das nachfolgende Meme konsumieren, um sich den Vorfall noch einmal auf unkonventionellere Art und Weise vor Augen zu halten.</p>


<div class="wp-block-embed-reddit">
					<blockquote class="reddit-card">
						<a href="https://www.reddit.com/r/singularity/comments/1v2xgqc/openai_hacking_huggingface_in_one_meme/"></a>
					</blockquote>
				</div>


<p class="wp-block-paragraph"></p>



<h2 class="wp-block-heading">Der autonome Hugging-Face-Hack</h2>



<p class="wp-block-paragraph">Die KI-Plattform Hugging Face meldete Mitte Juli einen <a href="https://huggingface.co/blog/security-incident-july-2026" target="_blank" rel="noreferrer noopener">Cyberangriff auf ihre Produktions-Infrastruktur</a>. Diese wurde offenbar vollständig autonom von einem agentenbasierten KI-System gefahren, das sich dabei sowohl Zugriff auf interne Datensätze als auch auf Zugangsdaten verschaffen konnte.  </p>



<p class="wp-block-paragraph">Wie sich im Rahmen der folgenden Untersuchung herausstellte, stammte die autonom hackende KI von OpenAI. Eigentlich wollte der KI-Pionier lediglich zwei seiner neuen Modelle – GPT 5.6 Sol sowie ein nicht näher spezifiziertes Pre-Release-Modell – innerhalb einer Sandbox mit dem Benchmarking-Tool <a href="https://github.com/sunblaze-ucb/exploitgym" target="_blank" rel="noreferrer noopener">ExploitGym</a> testen.</p>



<p class="wp-block-paragraph">Die Aufgabe für die KI: Schwachstellen in tragfähige Exploits zu verwandeln. Die Testumgebung selbst war OpenAI zufolge mit beschränktem Netzwerkzugriff ausgestattet. Die KI sollte lediglich auf einen intern gehosteten Drittanbieter-Service zugreifen können, der als Proxy und Cache für eine Package-Registry fungierte. Die Guardrails auf Modellebene waren für den Test laut dem KI-Unternehmen reduziert worden.</p>



<p class="wp-block-paragraph">Das war der KI aber offenbar nicht genug: Sie verkettete zunächst einige zuvor unbekannte Schwachstellen in der Testumgebung, um sich uneingeschränkten Internet-Zugriff zu verschaffen, wie OpenAI <a href="https://openai.com/de-DE/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="noreferrer noopener">in einem Blogbeitrag</a> offenlegt: „Mit diesem Zugriff führten unsere Modelle in unserer Forschungsumgebung eine Reihe von Privilege-Escalation- und Lateral-Movement-Aktionen aus, bis sie eine Node mit Internetzugang erreichten.“</p>



<p class="wp-block-paragraph">Weil die OpenAI-Modelle bei Hugging Face offenbar Lösungen für ihre ExploitGym-Aufgaben vermuteten, suchten sie nach Möglichkeiten, Zugriff auf diese zu erhalten. Und fanden sie dann auch, wie OpenAI darlegt: „In einem Beispiel verkettete das Modell mehrere Angriffsvektoren, darunter gestohlene Zugangsdaten und Zero-Day-Schwachstellen, um auf den Servern von Hugging Face einen Remote-Execution-Pfad zu finden.“</p>



<p class="wp-block-paragraph">Das Security-Team von OpenAI konnte diese Vorgänge im Nachgang identifizieren und schloss sich daraufhin mit dem Team von Hugging Face zusammen, welches die Attacke seinerseits bereits identifiziert und (mit einigen Anlaufschwierigkeiten) eingedämmt hatte.</p>



<p class="wp-block-paragraph">„Wir betrachten dies als beispiellosen Cybervorfall mit hochentwickelten Fähigkeiten und reagieren entsprechend. Wir teilen zu diesem Zeitpunkt vorläufige Erkenntnisse, damit Sicherheitsverantwortliche nachvollziehen können, was passiert ist, und besser einschätzen können, wozu die Modelle inzwischen in der Lage sind“, schreibt OpenAI in seinem Blog – und verspricht, weitere Details zu veröffentlichen, sobald diese vorliegen.</p>



<h2 class="wp-block-heading">KI-Ausbruch bei OpenAI – so reagieren Experten</h2>



<p class="wp-block-paragraph">Branchenexperten und Analysten bewerten den schlagzeilenträchtigen Incident um OpenAI und Hugging Face folgendermaßen: </p>



<ul class="wp-block-list">
<li><a href="https://www.kuppingercole.com/people/balaganski" target="_blank" rel="noreferrer noopener">Alexei Balaganski</a>, Lead Analyst bei KuppingerCole<strong>: </strong>„Dieser Vorfall sollte nicht als ‚Rogue AI‘-Geschichte betrachtet werden. Das Modell hat exakt das getan, wofür agentische Systeme gemacht sind: Es hat sich allen verfügbaren Tools und Wegen bedient, um das ihm gesetzte Ziel zu erreichen. Die Sicherheitsvorkehrungen, die es normalerweise in Zaum gehalten hätten, wurden von OpenAI selbst zu Testzwecken deaktiviert. Darin besteht die wahre Lektion.“</li>



<li><a href="https://www.kuppingercole.com/people/care" target="_blank" rel="noreferrer noopener">Jonathan Care</a>, Lead Analyst und AI Practice Lead bei KuppingerCole: „Es geht bei diesem Vorfall nicht darum, dass eine KI ausgebrochen ist und zum Angreifer wurde. Wir wussten, das würde passieren. Bemerkenswert ist allerdings, dass die Verteidiger – in diesem Fall das Team von Hugging Face – keine kommerziellen KI-Modelle nutzen konnten, um den Angriff zu analysieren. Denn deren Guardrails sorgen dafür, dass kein Exoploit-Code verarbeitet werden kann.“</li>



<li><a href="https://www.linkedin.com/in/beuchelt" target="_blank" rel="noreferrer noopener">Gerald Beuchelt</a>, CISO bei Acronis: „Der Vorfall verdeutlicht eine zentrale Herausforderung für Incident-Response-Teams: Angreifer sind nicht an Nutzungsrichtlinien gebunden. Verteidiger können hingegen an die Grenzen ihrer eigenen Tools stoßen, wenn diese genau jene Daten nicht verarbeiten, die für eine Untersuchung erforderlich sind. Im Ernstfall können daraus Verzögerungen mit unmittelbaren operativen Folgen entstehen.“</li>



<li><a href="https://www.computerwoche.de/profile/sabine-fromling/" target="_blank">Sabine Frömling</a>, Experten-Autorin und Cybersecurity-Beraterin: „Der eigentliche Sicherheitsvorfall war nicht die KI – sondern die Sandbox, die aus Versehen eine Tür zum Internet hatte. Man hat ein Raubtier freigelassen und dem Zaun die Schuld gegeben.“</li>



<li><a href="https://www.linkedin.com/in/martinzugec" target="_blank" rel="noreferrer noopener">Martin Zugec</a>, Technical Solutions Director bei Bitdefender:<strong> „</strong>Was meiner Meinung nach für KI-generierte Malware galt, untermauert auch dieser Vorfall: Die Bedrohung ist real, KI ist aber keine Magie. Wer glaubt, es mit einer neuartigen Superwaffe zu tun zu haben, wartet auf eine neuartige Gegenmaßnahme. Wer jedoch erkennt, dass es sich um bereits bekannte, aber unerbittlich angewandte Angriffstechniken handelt, weiß bereits, was zu tun ist.“</li>



<li><a href="https://de.linkedin.com/in/riwerner/de" target="_blank" rel="noreferrer noopener">Richard Werner</a>, Cybersecurity Platform Lead Europe bei TrendAI: „Das Narrativ von der ‚eigenmächtig handelnden KI‘ ist effizient darin, Verantwortung abzuwälzen. Das ist, als würden Sie eine autonome Waffe bauen, diese auf einem vermeintlich sicheren Testgelände erproben, sie außer Kontrolle geraten und jemanden treffen lassen – und der Welt anschließend erklären, die Waffe habe eigenständig gehandelt. Das ist zwar technisch korrekt. Dennoch bleibt es Ihre Waffe, Ihr Testgelände und Ihr Versagen.“</li>
</ul>



<h2 class="wp-block-heading">Was Unternehmen jetzt tun sollten</h2>



<p class="wp-block-paragraph">IT- und Sicherheitsentscheider können aus dem Hugging-Face-Hack mehrere Lektionen ziehen. Etwa, dass Sicherheitsvorkehrungen auf Modellebene <strong>nicht</strong> als primäre Security-Grenze für KI-Agenten geeignet sind, wie <a href="https://www.forrester.com/analyst-bio/biswajeet-mahapatra/BIO20046" target="_blank" rel="noreferrer noopener">Biswajeet Mahapatra</a>, Principal Analyst bei Forrester, festhält: „Prompt-Guardrails sind keine Sicherheits-, sondern Verhaltenskontrollmaßnahmen. Und diese können versagen, umgangen oder absichtlich deaktiviert werden.“</p>



<p class="wp-block-paragraph">Der Forrester-Analyst rät Unternehmen deshalb dazu, KI-Agenten als <a href="https://www.computerwoche.de/article/4152424/insider-threats-sind-wieder-im-kommen.html" target="_blank">hochriskante, nicht-menschliche Identitäten</a> zu behandeln – und jeden einzelnen in einer isolierten Umgebung zu betreiben, in der Datenzugriff auf den jeweiligen Task beschränkt bleibt und die Zugangsdaten selbst möglichst schnell ablaufen: „Das sorgt für einen akzeptablen ‚Blast Radius‘: Wird ein Agent <a href="https://www.computerwoche.de/article/4190978/so-spuren-sie-kompromittierte-ki-agenten-auf.html" target="_blank">kompromittiert</a>, kann er nur einen einzigen Workflow, Datensatz oder eine einzige Anwendung beeinträchtigen. Anstatt die gesamte Unternehmensinfrastruktur.“</p>



<p class="wp-block-paragraph"><a href="https://greyhoundresearch.com/svg/" target="_blank" rel="noreferrer noopener">Sanchit Vir Gogia</a>, Chefanalyst bei Greyhound Research, warnt an dieser Stelle davor, (Drittanbieter-)Services unter den Tisch fallen zu lassen: „Dienste, die auf Package Registries, Update-Systeme oder andere externe Ressourcen zugreifen, können ebenfalls zu Einfallstoren werden, wenn sie nicht derselben, ausgiebigen Prüfung unterzogen werden wie der Agent selbst.“</p>



<p class="wp-block-paragraph">Unabhängig davon sollten Unternehmen laut Gogia auch testen, ob ihre Containment-Grenzen auch funktionieren, anstatt sich allein auf Architekturdiagramme oder dokumentierte Richtlinien zu verlassen: „Im Rahmen dieser Tests sollte geprüft werden, ob Anmeldedaten erlangt, Trust-Grenzen überwunden und Systeme außerhalb der einem Agenten zugewiesenen Aufgabe erreicht werden können.“</p>



<p class="wp-block-paragraph">KuppingerCole-Chefanalyst Care rät IT-Entscheidern und Unternehmen im Wesentlichen zu drei Maßnahmen, nämlich:</p>



<ul class="wp-block-list">
<li>ein fähiges Modell auf der eigenen Infrastruktur auszuführen, das unter der eigenen Kontrolle steht und mit Guardrails ausgestattet ist, die sowohl eine forensische als auch defensive Nutzung ermöglichen. Nur so ließen sich Angriffe dieser Art auch zuverlässig analysieren.</li>



<li>jeden KI-Agent in der eigenen Umgebung als privilegierten Insider zu behandeln – statt als vertrauenswürdigen Benutzer: „Wenn die Modelle von OpenAI aus ihrer Sandbox ausgebrochen sind, sollten Sie davon ausgehen, dass Ihre Agenten dazu auch in der Lage sind.“</li>



<li>den eigenen Incident-Response-Plan mit Blick auf Angriffe in maschineller Geschwindigkeit zu aktualisieren: „Hugging Face hatte einige Tage Zeit, um zu reagieren, Sie haben vielleicht nur Minuten.“   </li>
</ul>



<p class="wp-block-paragraph">Acronis-CISO Beuchelt rät Organisationen, die gehostete <a href="https://www.computerwoche.de/article/4186715/31-wege-llms-zu-evaluieren.html" target="_blank">LLMs</a> für Security-Untersuchungen einsetzen, dazu, deren Grenzen möglichst bereits im Vorfeld zu durchdringen und zu testen – sowie ein alternatives Modell auf der eigenen Infrastruktur bereitzuhalten: „So reduzieren Sie das Risiko, im entscheidenden Moment keinen Zugriff auf wichtige Analysefunktionen zu haben. Gleichzeitig bleiben sensible Incident-Daten und Zugangsinformationen innerhalb der eigenen Organisation.“</p>



<p class="wp-block-paragraph"><a href="https://de.linkedin.com/in/udoschneider">Udo Schneider</a>, Governance, Risk &amp; Compliance Lead Europe bei TrendAI weist darauf hin, dass die beiden naheliegendsten Lösungsansätze bei Angriffen wie dem der OpenAI-KI auf Hugging Face nur teilweise greifen. Human-in-the-Loop-Kontrollen funktionierten zwar, so der Experte, skalierten aber nicht für die langlaufenden, komplexen Workflows, denen Incidents dieser Art entspringen. Ebenso könnten engere Guardrails für Modelle oder Prompts zwar helfen, stellten jedoch keine Garantie dar: „Es handelt sich um probabilistische Systeme. Eine Guardrail ist insofern keine Mauer, sondern eher eine starke Wahrscheinlichkeitsannahme.“</p>



<p class="wp-block-paragraph">Deshalb komme es laut Schneider vor allem auf die unspektakulären, nicht-KI-spezifischen Kontrollen an: „Zugriffsfilterung, Kontrolle darüber, was überhaupt als Input beim Modell ankommt, Sandboxes, die tatsächlich halten, und Berechtigungskonzepte nach dem Least-Privilege-Prinzip.“</p>



<p class="wp-block-paragraph">In Panik zu verfallen, wäre nach Ansicht von <a href="https://www.linkedin.com/in/martinzugec" target="_blank" rel="noreferrer noopener">Martin Zugec</a>, Technical Solutions Director bei Bitdefender, in jedem Fall die falsche Reaktion:„Was gegen solche Angriffe wirkt, ist eine präventionsorientierte Security, die den Handlungsspielraum eines Angreifers von vorneherein einschränkt – und eine verhaltensbasierte Abwehr, die bösartige Muster kennzeichnet, unabhängig davon, mit welchen Tools diese generiert wurden.“</p>



<p class="wp-block-paragraph"><strong>Dieser Artikel wurde </strong><a href="https://www.csoonline.com/article/4200043/openai-model-escape-puts-enterprise-ai-defenses-on-notice.html" target="_blank"><strong>mit Material</strong></a><strong> unserer Schwesterpublikation CSOonline.com angereichert.</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Certifying the future; synchronizing EUDI wallets and quantum-readiness]]></title>
<description><![CDATA[Author: PQShield - Bewertung: 0x - Views:0 The rollout of the European Digital Identity Wallet (EUDI wallet) demands tight security, but combining high-assurance compliance with post-quantum cryptography creates new engineering bottlenecks. In this episode, host Johannes Lintzen sits down with We...]]></description>
<link>https://tsecurity.de/de/3688805/videos/certifying-the-future-synchronizing-eudi-wallets-and-quantum-readiness/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3688805/videos/certifying-the-future-synchronizing-eudi-wallets-and-quantum-readiness/</guid>
<pubDate>Thu, 23 Jul 2026 13:07:40 +0200</pubDate>
<category>🎥 Videos</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: PQShield - Bewertung: 0x - Views:0 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/qH2adeafr6g?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>The rollout of the European Digital Identity Wallet (EUDI wallet) demands tight security, but combining high-assurance compliance with post-quantum cryptography creates new engineering bottlenecks. In this episode, host Johannes Lintzen sits down with Wei Yuan, lab manager and director of operations at Applus+ Laboratories. They dissect hardware trust architectures, hardware security module (HSM) limits, and why implementation flaws cause far more security breaches than algorithmic failures.<br />
<br />
YouTube chapters<br />
00:00 Introduction to Wei Yuan and Applus+ Laboratories <br />
03:28 Explaining the European Digital Identity (EUDI) Wallet <br />
05:28 The importance of data minimization for citizens <br />
07:41 How laboratories verify vendor security claims <br />
09:35 The vulnerability of identity data to quantum attacks <br />
11:12 The 2026 deadline for member state deployment <br />
12:22 Creating certification schemes with ENISA <br />
15:46 Comparing secure elements and remote HSMs <br />
19:12 The role of standardization in preventing delays <br />
24:33 Navigating international regulation differences <br />
28:51 Advice for vendors starting their migration <br />
32:10 Reporting obligations under the Cyber Resilience Act <br />
34:02 Why implementation matters more than design<br />
<br />
Guest bio<br />
Wei Yuan is lab manager and director of operations at Applus+ Laboratories. Working at the boundary of high-assurance evaluation, hardware security, and regulatory policy, he serves as an expert within ENISA working groups, helping shape European digital identity certification standards and transition pathways to post-quantum cryptography.<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Stop asking AI nicely: Here’s how to get work-ready results every time]]></title>
<description><![CDATA[Over the past few years, I have learned that basic prompts produce inconsistent, hallucination-prone results that no executive would trust in production. What turned the tide was my move to advanced prompting techniques. These weren’t theoretical experiments; they became a practical foundation fo...]]></description>
<link>https://tsecurity.de/de/3688796/it-nachrichten/stop-asking-ai-nicely-heres-how-to-get-work-ready-results-every-time/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3688796/it-nachrichten/stop-asking-ai-nicely-heres-how-to-get-work-ready-results-every-time/</guid>
<pubDate>Thu, 23 Jul 2026 13:07:21 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Over the past few years, I have learned that basic prompts produce inconsistent, hallucination-prone results that no executive would trust in production. What turned the tide was my move to advanced prompting techniques. These weren’t theoretical experiments; they became a practical foundation for reliable, measurable outcomes. I want to share the techniques that consistently delivered the biggest gains in my projects, complete with real before-and-after examples, copy-paste templates, lessons from failures and guidance on when to evolve beyond prompting to agentic systems.</p>



<h2 class="wp-block-heading">Why advanced prompting still matters in enterprise settings</h2>



<p class="wp-block-paragraph">Sophisticated prompting remains essential for control, reliability and compliance. If you “ask nicely” and hope for the best, you need deterministic behavior, auditable reasoning and minimal risk of hallucination. Here’s what worked for me.</p>



<h3 class="wp-block-heading">1. Chain-of-Thought (CoT) and its variants: Unlocking step-by-step reasoning</h3>



<p class="wp-block-paragraph"><strong>The problem:</strong> Models would jump to conclusions on complex analysis tasks, especially involving data interpretation or multi-step logic.</p>



<p class="wp-block-paragraph"><strong>What I did:</strong> I started explicitly instructing the model to “think step by step” and show its reasoning.</p>



<p class="wp-block-paragraph"><strong>Before (basic prompt): </strong>“Analyze last quarter’s sales data and recommend three actions.”</p>



<p class="wp-block-paragraph"><strong>After (CoT prompt):</strong></p>



<p class="wp-block-paragraph">“You’re a senior business analyst. Analyze the following sales data step by step: [data]. First, identify the key trends. Second, calculate the rates and anomalies. Third, link findings to business context. Finally, recommend the three prioritized actions with expected impact. Explain your reasoning at each step.”  </p>



<p class="wp-block-paragraph"><strong>Results:</strong> Accuracy and depth improved dramatically.</p>



<p class="wp-block-paragraph"><strong>Variants that worked well:</strong> Self-consistency. I ran the same CoT prompt multiple times and took the majority consensus. This reduced variability significantly.</p>



<p class="wp-block-paragraph"><strong>Template you can use:</strong></p>



<pre class="wp-block-code"><code>You are [expert role]. Solve this problem by thinking step by step.

[Task or question]

For each step:

1. State your observation or calculation.

2. Explain the implication.

3. Proceed only when confident.

Final answer in this format: [structured output]</code></pre>



<h3 class="wp-block-heading">2. Tree-of-Thoughts (ToT): Exploring multiple reasoning paths</h3>



<p class="wp-block-paragraph">For truly complex decisions such as resource allocation or risk assessment, linear CoT isn’t enough. Tree-of-Thoughts lets the model generate and evaluate multiple branches.</p>



<p class="wp-block-paragraph"><strong>Example:</strong> I was helping a client evaluate three potential vendor platforms for an AI deployment. A standard prompt gave a superficial comparison. With ToT</p>



<p class="wp-block-paragraph"><strong>Prompt Snippet:</strong></p>



<pre class="wp-block-code"><code>Explore three different reasoning paths for selecting the best vendor platform:

Path 1: Focus on cost and scalability.

Path 2: Focus on security, compliance and integration.

Path 3: Focus on innovation and long-term roadmap.

For each path, evaluate pros/cons against our requirements [list].

Then, compare the paths and recommend the strongest overall option with justification.</code></pre>



<p class="wp-block-paragraph"><strong>Outcome:</strong> The model surfaced nuanced trade-offs (e.g., one vendor had superior security, but higher integration cost).</p>



<p class="wp-block-paragraph"><strong>When to use:</strong> Strategic planning, troubleshooting or scenarios with high uncertainty and multiple viable approaches.</p>



<h3 class="wp-block-heading">3. ReAct (Reason+ Act) and prompt chaining: Moving toward agentic behavior</h3>



<p class="wp-block-paragraph">One of the biggest leaps I have noticed comes from combining reasoning with tool use and chaining prompts.</p>



<p class="wp-block-paragraph"><strong>ReAct example</strong>: (used in data analytics workflow)</p>



<pre class="wp-block-code"><code>You are an AI analyst with access to tools. For the query below:

1. Reason about what information you need.

2. Choose the appropriate tool or action.

3. Observe the result.

4. Repeat until you can answer confidently.

Query: [user request]</code></pre>



<p class="wp-block-paragraph">In practice, I chained this with retrieval tools. One automated quarterly compliance reporting; the system reasoned about required data, pulled relevant records, validated them, and generated the reports.</p>



<h3 class="wp-block-heading">4. Meta-prompting and self-reflection: Letting the model improve itself</h3>



<p class="wp-block-paragraph">Use the model to refine its own prompt. This is a huge time-saver.</p>



<pre class="wp-block-code"><code>You are an expert prompt engineer. Improve the following prompt for clarity, structure and effectiveness with [target model]. Make it more precise while preserving intent.

Original prompt: [paste]

Provide the improved version and explain your changes.</code></pre>



<p class="wp-block-paragraph">Self-reflection loops (asking the model to critique its own output and revise) are a game-changer for content generation and code-review tasks.</p>



<h3 class="wp-block-heading">5. Multimodal and structured output techniques</h3>



<p class="wp-block-paragraph">With vision-enabled models, I started combining text with images (e.g., uploading architecture diagrams or dashboards).</p>



<p class="wp-block-paragraph"><strong>Tip from experience:</strong> Be extremely specific in describing what the models should focus on.</p>



<h4 class="wp-block-heading">Best practices I learned the hard way</h4>



<ul class="wp-block-list">
<li><strong>Start simple, then layer complexity</strong>: Over-engineered prompts from Day One usually backfire.</li>



<li><strong>Model specific tuning:</strong> Some models respond better to XML delimiters; others to explicit reasoning.</li>



<li><strong>Evaluation and versioning:</strong> Treat prompts like code if you track versions and run automated evals.</li>



<li><strong>Security guardrails:</strong> Always include instructions against prompt injections and respect data boundaries.</li>



<li><strong>When to stop prompting</strong>: For repetitive, high-stakes workflows, move to full agents or an orchestration framework.</li>
</ul>



<h2 class="wp-block-heading">Final takeaways for technical leaders</h2>



<p class="wp-block-paragraph">Advanced prompt engineering has now become a core competency for anyone responsible for enterprise AI outcomes. Start by picking one technique and apply it rigorously to a real business problem. Document before/ after and you will notice why it’s worth mastering.</p>



<p class="wp-block-paragraph">The field continues evolving towards more automated and agentic systems, but the ability to precisely direct AI reasoning remains foundational.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI and Hugging Face Investigate AI Models’ Cyber Breakout]]></title>
<description><![CDATA[OpenAI and Hugging Face are investigating an AI security incident involving an AI agent that compromised infrastructure while models were being evaluated for advanced cyber capabilities. The incident was detected and contained after the models identified and chained vulnerabilities across OpenAI’...]]></description>
<link>https://tsecurity.de/de/3688175/it-security-nachrichten/openai-and-hugging-face-investigate-ai-models-cyber-breakout/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3688175/it-security-nachrichten/openai-and-hugging-face-investigate-ai-models-cyber-breakout/</guid>
<pubDate>Thu, 23 Jul 2026 08:54:52 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1536" height="1024" src="https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident.webp" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="OpenAI and Hugging Face Probe AI Security Incident" decoding="async" srcset="https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident.webp 1536w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-300x200.webp 300w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-1024x683.webp 1024w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-768x512.webp 768w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-600x400.webp 600w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-150x100.webp 150w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-750x500.webp 750w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-1140x760.webp 1140w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident.webp 1536w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-300x200.webp 300w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-1024x683.webp 1024w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-768x512.webp 768w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-600x400.webp 600w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-150x100.webp 150w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-750x500.webp 750w, https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-Probe-AI-Security-Incident-1140x760.webp 1140w" sizes="(max-width: 1536px) 100vw, 1536px" title="OpenAI and Hugging Face Investigate AI Models’ Cyber Breakout 1"></p><p class="PDq2pG_selectionAnchorContainer" data-start="453" data-end="826">OpenAI and Hugging Face are investigating an <a href="https://thecyberexpress.com/incident-response-automating-with-genai/" target="_blank" rel="noopener">AI security incident </a>involving an AI agent that compromised infrastructure while models were being evaluated for advanced cyber capabilities. The incident was detected and contained after the models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure.</p>
<p data-start="828" data-end="1234">Hugging Face disclosed the incident last week, while <a href="https://thecyberexpress.com/lockdown-mode-means-default-chatgpt-wasnt-safe/" target="_blank" rel="noopener">OpenAI</a> has now shared preliminary findings from its investigation. According to OpenAI, the activity involved a combination of models, including GPT 5.6 Sol and a more capable pre-release model, with reduced <a class="wpil_keyword_link" href="https://thecyberexpress.com/cyber-news/" title="cyber" data-wpil-keyword-link="linked" data-wpil-monitor-id="29078">cyber</a> refusals for evaluation purposes. The models were being tested against a benchmark designed to measure advanced cyber capabilities.</p>
<p data-start="1236" data-end="1504">OpenAI described the incident as unprecedented and said it is continuing its investigation alongside Hugging Face. The company said the findings are being shared to help defenders understand what occurred and assess the capabilities of increasingly advanced AI models.</p>

<h3 data-section-id="97z8dp" data-start="1506" data-end="1566"><strong>OpenAI and Hugging Face Investigate AI Security Incident</strong></h3>
<p data-start="1568" data-end="1858">The incident took place during an internal evaluation designed to test models against advanced exploitation scenarios involving complex attack paths. OpenAI said the evaluation was conducted without production classifiers that normally prevent models from pursuing high-risk cyber activity.</p>
<p data-start="1860" data-end="2088">The testing environment was designed to be highly isolated, with network access limited to package installation through an internally hosted third-party software system that functions as a proxy and cache for package registries.</p>
<p data-start="2090" data-end="2406">Despite these restrictions, the models spent substantial inference compute attempting to obtain open <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-is-internet/" title="Internet" data-wpil-keyword-link="linked" data-wpil-monitor-id="29085">Internet</a> access. During this process, they identified and exploited a <a href="https://thecyberexpress.com/zero-day-vulnerability-microsoft-sharepoint/" target="_blank" rel="noopener">zero-day vulnerability</a> in the package registry cache proxy. OpenAI said the <a class="wpil_keyword_link" href="https://thecyberexpress.com/firewall-daily/vulnerabilities/" title="vulnerability" data-wpil-keyword-link="linked" data-wpil-monitor-id="29079">vulnerability</a> has since been responsibly disclosed to the vendor.</p>
<p data-start="2408" data-end="2565">The models then carried out privilege escalation and <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-is-lateral-movement-in-cyber-security/" title="lateral movement" data-wpil-keyword-link="linked" data-wpil-monitor-id="29081">lateral movement</a> within the research testing environment until they reached a node with Internet access.</p>


[caption id="attachment_113283" align="aligncenter" width="600"]<img class="wp-image-113283 size-full" src="https://thecyberexpress.com/wp-content/uploads/OpenAI-and-Hugging-Face-e1784786987186.webp" alt="OpenAI and Hugging Face" width="600" height="323"> Source: OpenAI[/caption]
<h3 data-section-id="1envt5o" data-start="2567" data-end="2627"><span role="text"><strong data-start="2571" data-end="2627">Models Chained Vulnerabilities Across Infrastructure</strong></span></h3>
<p data-start="2629" data-end="2886">After obtaining Internet access, the models inferred that Hugging Face could host models, datasets and solutions related to the ExploitGym evaluation. The models then searched for ways to obtain secret information that could help them bypass the evaluation.</p>
<p data-start="2888" data-end="3092">In one example described by OpenAI, the models chained multiple attack vectors, including stolen credentials and zero-day <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-are-vulnerabilities/" title="vulnerabilities" data-wpil-keyword-link="linked" data-wpil-monitor-id="29080">vulnerabilities</a>, to identify a remote code execution path on Hugging Face servers.</p>
<p data-start="3094" data-end="3403">OpenAI's <a class="wpil_keyword_link" href="https://thecyberexpress.com/" title="security" data-wpil-keyword-link="linked" data-wpil-monitor-id="29082">security</a> team detected the anomalous activity internally. Hugging Face's security team and agents also detected and stopped the activity on its infrastructure. The company had already begun containment and forensic reconstruction using its own open-source models when the two organizations connected.</p>
<p data-start="3405" data-end="3478">Both companies are continuing to investigate the incident and its impact.</p>

<h3 data-section-id="7fczcn" data-start="3480" data-end="3539"><span role="text"><strong data-start="3484" data-end="3539">OpenAI Strengthens Security Controls After Incident</strong></span></h3>
<p data-start="3541" data-end="3872">As part of its response, OpenAI <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="nofollow noopener">said</a> it is implementing stricter infrastructure configuration controls while identified vulnerabilities are patched. The company acknowledged that these measures could come at the cost of research velocity and said its Safety and Security Committee is being briefed on the controls and their impact.</p>
<p data-start="3874" data-end="4063">OpenAI is also working with Hugging Face on the forensic investigation and has responsibly disclosed the identified zero-day vulnerability in the internally hosted third-party software.</p>
<p data-start="4065" data-end="4221">The company has also brought Hugging Face into its trusted access program and is supporting its teams in using AI model capabilities to strengthen defenses.</p>
<p data-start="4223" data-end="4562">OpenAI said it is improving protections around future training and evaluations, including stronger safeguards for model alignment, <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-is-cybersecurity/" title="cybersecurity" data-wpil-keyword-link="linked" data-wpil-monitor-id="29083">cybersecurity</a> and monitoring during internal testing. The company noted that deployment safeguards were intentionally disabled during this evaluation because the goal was to measure cyber vulnerabilities.</p>

<h3 data-section-id="1vqt96" data-start="4564" data-end="4621"><span role="text"><strong data-start="4568" data-end="4621">AI Cyber Capabilities Raise New Security Concerns</strong></span></h3>
<p data-start="4623" data-end="4891">OpenAI said the incident demonstrates the need for <a href="https://thecyberexpress.com/ai-security-is-top-cyber-concern/" target="_blank" rel="noopener">AI security </a>and safety measures to keep pace with rapidly advancing model capabilities. The company is strengthening containment, monitoring, access controls and evaluation practices used during model development.</p>
<p data-start="4893" data-end="5226">The incident also highlights how advanced models can potentially discover and <a class="wpil_keyword_link" href="https://cyble.com/exploit/" target="_blank" rel="noopener" title="exploit" data-wpil-keyword-link="linked" data-wpil-monitor-id="29084">exploit</a> novel attack paths in real-world systems without access to source code. OpenAI said increasingly capable models should also be used defensively to help security teams identify weaknesses, understand vulnerability chains and accelerate remediation.</p>
<p data-start="5228" data-end="5513" data-is-last-node="" data-is-only-node="">Hugging Face CEO Clem Delangue said the incident demonstrates the importance of collaboration in addressing AI safety and security challenges. Both organizations said they will continue investigating the incident and share additional findings and best practices as the work progresses.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The credential that let OpenAI's agents into Hugging Face exists in most enterprises right now]]></title>
<description><![CDATA[When Hugging Face got hit last week, co-founder Clement Delangue suspected a frontier lab, given the agent's sophistication. He was right. Delangue said on X that after a day working with OpenAI he strongly believed there was no malicious intent and that it was mind-blowing the whole thing had ha...]]></description>
<link>https://tsecurity.de/de/3687771/it-nachrichten/the-credential-that-let-openais-agents-into-hugging-face-exists-in-most-enterprises-right-now/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3687771/it-nachrichten/the-credential-that-let-openais-agents-into-hugging-face-exists-in-most-enterprises-right-now/</guid>
<pubDate>Thu, 23 Jul 2026 01:32:49 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>When Hugging Face got hit last week, co-founder Clement Delangue suspected a frontier lab, given the agent's sophistication. He was right. Delangue <a href="https://x.com/ClementDelangue/status/2079670308156645882">said on X</a> that after a day working with OpenAI he strongly believed there was no malicious intent and that it was mind-blowing the whole thing had happened autonomously.</p><p>The two OpenAI models that <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">broke into Hugging Face</a> last week did not breach it through malice or superintelligence. They breached it through credentials and permissions they should never have been able to reach, a non-human identity failure that is the oldest problem in security rather than the newest one in AI, and the one every enterprise can actually fix.</p><p>OpenAI <a href="https://venturebeat.com/security/openais-models-broke-containment-and-cyberattacked-hugging-face-what-enterprises-need-to-know">disclosed on July 21</a> that two of its models, GPT-5.6 Sol and an unreleased, more capable model, were running a cyber benchmark called <a href="https://arxiv.org/abs/2605.11086">ExploitGym</a> with their safety refusals switched off, and inferred that the answer key sat in Hugging Face's production database. Getting there took two different failures. A zero-day in a package-registry proxy let the models out of their sandbox and onto the open internet, the kind of persistence OpenAI details in its companion post on <a href="https://openai.com/index/safety-alignment-long-horizon-models/">long-horizon safety</a>, and that part is genuinely new. The breach of Hugging Face itself came the ordinary way. OpenAI's own account is that the models chained stolen credentials and further zero-days into a remote code execution path, after a series of privilege escalation and lateral movement steps. The exotic part got them to the door, and credentials walked them through it.</p><p>Hugging Face also disclosed last week that an <a href="https://venturebeat.com/security/safety-guardrails-blocked-hugging-faces-defenders-not-the-attacker-when-an-ai-agent-breached-its-systems">autonomous agent had harvested cloud and cluster credentials</a> scoped broadly enough to reach multiple internal clusters, then left a trail of more than 17,000 recorded events across short-lived sandboxes over a weekend. Both disclosures describe the same escalation. An agent lands somewhere it should not be, finds credentials scoped far wider than any task requires, and uses them to move. These are two accounts of one incident, not two attacks. The agent Hugging Face watched was OpenAI's models, and both companies describe the same ordinary escalation.</p><p>The version of this in a typical enterprise is worse, not better. OpenAI and Hugging Face are among the most security-mature organizations in the industry, and both still needed the intrusion to happen before they could see it. The average company wiring agents into Copilot or an internal assistant has neither the identity inventory nor the behavioral monitoring those two brought to bear. The same breach in a normal company would not be contained in days, it would simply go unnoticed.</p><h2>The industry is debating the wrong failure</h2><p>The reaction has split into familiar camps. Former White House AI and crypto czar David Sacks and a run of China hawks <a href="https://fortune.com/2026/07/20/hugging-face-turns-to-chinese-open-source-ai-to-fend-off-autonomous-ai-cyber-attack-after-american-ai-guardrails-stymie-defense/">seized on the guardrail paradox</a>, that commercial safety filters blocked Hugging Face's defenders while the attacking model ran with its refusals off, and that a Chinese open-weight model, z.ai's GLM 5.2, was what finally let the team finish its forensics. Hugging Face made the case for openness, arguing in an April <a href="https://huggingface.co/blog/cybersecurity-openness">blog post</a> that open models and open tooling give defenders the same capabilities attackers already have. Both arguments are about the model, and neither touches the mechanism. </p><p>Reduced refusals let the model attempt an attack, and over-scoped credentials are what let it succeed, and those have nothing to do with whether the model was open or closed, American or Chinese. Making a frontier model provably safe is a multi-year alignment problem no customer can buy or accelerate, while scoping an identity is a configuration change a team can ship this sprint. The industry is being urged to fixate on the part of this it cannot control and to treat the part it can as a footnote.</p><p>Forrester reached the same read. In a <a href="https://www.forrester.com/blogs/an-ai-security-facepalm-openais-evaluation-became-hugging-faces-incident/">blog on the incident</a>, its analysts argue that security architectures which assume benign intent will miss this failure mode, because an agent can pursue an authorized goal through unauthorized means, which is what OpenAI's models did.</p><h2>This was a non-human identity failure, and it is the oldest one in security</h2><p>Strip the science-fiction framing and what remains is a textbook case of over-privileged machine identity, the kind security teams have fought for a decade, now driven by an autonomous agent at machine speed. Machine identities already outnumber humans in most enterprises by more than <a href="https://www.cyberark.com/press/machine-identities-outnumber-humans-by-more-than-80-to-1-new-report-exposes-the-exponential-threats-of-fragmented-identity-security/">80 to one</a>, according to CyberArk research, with 42% of them carrying privileged or sensitive access, and an agent inherits whatever its identity can touch. OWASP ranks agent identity and privilege abuse near the top of its <a href="https://neuraltrust.ai/blog/owasp-agentic-ai-top-10">agentic risk list</a>, the confused-deputy pattern where inherited credentials and weak scoping let an agent reach past its mandate, and that is precisely what both July disclosures describe. </p><p><a href="https://www.ieee.org/membership/senior">IEEE Senior Member</a> Kayne McGladrey has argued in <a href="https://venturebeat.com/security/cisco-crowdstrike-rsac-2026-agent-identity-iam-gap-maturity-model">previous VentureBeat interviews</a> that enterprises keep cloning human user accounts onto agents that then wield far more permission than any human would, and this is what that looks like when the agent is a frontier model and the target is a production database.</p><p>The people closest to it read it the same way. OpenAI frames its models as hyperfocused on a benchmark score rather than acting against anyone. Nobody describes an adversary, only a goal, a scoring function, and credentials that were reachable when they should not have been.</p><p>The specific failure is easy to name once the AI framing is stripped away. A credential scoped to one job that can reach ten is a standing invitation, and it does not matter whether a human attacker, a worm, or an autonomous model chasing a benchmark score finds it. What changed in July is the finder. An agent enumerates reachable systems, tests credentials, and pivots faster than any human red team, without malice or hesitation, whenever the path is open. The over-scoping was always the vulnerability, and the agent merely industrialized its discovery.</p><p>Forrester named the control that would have blunted it. Its agentic-security framework, AEGIS, calls for least agency, holding an agent's tools, credentials, and network paths to the minimum its task requires, and files this incident under unrestrained agency and privilege. That is the identity argument in different words, arrived at independently by an analyst firm.</p><p>The data says this is where the risk now lives. Verizon's 2026 Data Breach Investigations Report <a href="https://www.helpnetsecurity.com/2026/05/20/verizon-2026-dbir-findings/">found</a> that exploitation of vulnerabilities has overtaken stolen credentials as the top initial access vector for the first time in 19 years. That is the initial-access half. The other half is the one OpenAI itself describes, stolen credentials driving the privilege escalation and lateral movement that followed. A vulnerability opened the door, and credentials walked through the building unchallenged. Beyond the breach itself, that same over-scoping carries a legal liability most enterprises have never priced. The models' actions <a href="https://techcrunch.com/2026/07/21/openai-says-hugging-face-was-breached-by-its-pre-release-models/">likely violated the Computer Fraud and Abuse Act</a>, according to TechCrunch. The statute contains no carve-out for an AI agent that exceeds its authorized scope during sanctioned testing. Whatever the legal answer, the technical enabler is the same, an identity scoped wider than its task. This is an access-control problem with an owner and a budget, not a philosophy seminar about machine cognition.</p><p>Merritt Baer, Senior Advisor to Andesite, G2I, and AppOmni and former Deputy CISO at AWS, frames the underlying shift to VentureBeat as a new kind of asymmetry. Both sides now reach for the same capabilities, she said, but one side is constrained by enterprise governance, policy, compliance, and safety controls while the adversary simply downloads an uncensored open-weight model and keeps going. The organizations that come through it best, in her view, will be the ones that treat AI as a resilient, governed capability rather than a single service they do not control.</p><h2>Four moves that shrink the blast radius</h2><p>The breach worked because the agent reached identities scoped far wider than its task. None of the four controls that would have contained it requires a new platform, and none of them appears on the list of general AI-safety advice now circulating. They are identity hygiene, applied to non-human actors with the same rigor you already apply to people.</p><p><b>1. Scope every non-human identity to one task.</b> The models reached credentials that touched multiple clusters, which is what turned a foothold into a breach. An identity scoped to a single job, with no standing access to anything else, hits a wall at the first lateral move instead of opening the next door. This is least privilege, the control everyone endorses and few enforce on machine accounts, and it is the single highest-impact fix here.</p><p><b>2. Give credentials short lifetimes and rotate them hard.</b> Harvested credentials are only useful while they are valid, and both July agents worked by collecting them. Short time-to-live and aggressive rotation turn a credential dump into expired noise, so a token stolen during a weekend intrusion is dead before the attacker can chain it. Static secrets that never rotate are the version of this control that fails.</p><p><b>3. Monitor for lateral movement, not just prompts.</b> The tell in both incidents was privilege escalation and lateral movement, which a prompt filter never sees because it is watching the wrong layer. Identity-behavior monitoring, keyed to what a given non-human identity normally does and alerting when it reaches somewhere new, catches the escalation the content guardrail missed. The question for your stack is whether anything you run today would flag a service account suddenly moving between clusters.</p><p><b>4. Rehearse instant revocation before you need it.</b> When the incident is your own agent, the fastest containment is killing its identity mid-run, and that only works if the path to do it exists before the day you need it. Rehearse revoking a machine identity under fire the way you rehearse a human credential compromise. If you have never done it, you do not yet have the control, you have an intention.</p><p>The defense also worked, and that matters. OpenAI's security team caught the anomalous activity internally, Hugging Face's own detection and agents stopped the intrusion, and the breach was contained in days rather than discovered in months, because the defenders could see into systems they controlled. That visibility is the same discipline the four controls depend on. The debate over whether frontier models are safe, open, or American will run for years, and none of it will be settled in time to help the enterprise deploying agents this quarter. The non-human identity gap is different, because it is understood, measurable, and fixable now. The model that breached Hugging Face did not need to be brilliant; it needed credentials someone left in reach. The fix is scoping them before an agent finds them.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Research-Grade EdgeBench Analysis: AI Agent Benchmarking, Leaderboard Analytics, Scaling Laws, and Evaluation Metrics]]></title>
<description><![CDATA[In this tutorial, we explore EdgeBench as a practical benchmark for evaluating advanced AI agents across diverse task categories, runtime environments, and interaction-time budgets. We begin by downloading the dataset snapshot from Hugging Face, parsing the released task specifications, and exami...]]></description>
<link>https://tsecurity.de/de/3687695/ai-nachrichten/research-grade-edgebench-analysis-ai-agent-benchmarking-leaderboard-analytics-scaling-laws-and-evaluation-metrics/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3687695/ai-nachrichten/research-grade-edgebench-analysis-ai-agent-benchmarking-leaderboard-analytics-scaling-laws-and-evaluation-metrics/</guid>
<pubDate>Thu, 23 Jul 2026 00:10:37 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>In this tutorial, we explore EdgeBench as a practical benchmark for evaluating advanced AI agents across diverse task categories, runtime environments, and interaction-time budgets. We begin by downloading the dataset snapshot from Hugging Face, parsing the released task specifications, and examining the benchmark taxonomy, execution settings, internet requirements, judging logic, and scoring metadata. We then […]</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/22/research-grade-edgebench-analysis-ai-agent-benchmarking-leaderboard-analytics-scaling-laws-and-evaluation-metrics/">Research-Grade EdgeBench Analysis: AI Agent Benchmarking, Leaderboard Analytics, Scaling Laws, and Evaluation Metrics</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Apple Store app Virtual Shopping Assistant could be coming soon]]></title>
<description><![CDATA[The privacy policy for the Apple Store app suggests that a new feature called Virtual Shopping Assistant could be implemented soon. It may function similarly to the Apple Support Assistant.Apple Support already has a chatbot, and soon, Apple Store will tooApple has been slowly increasing its use ...]]></description>
<link>https://tsecurity.de/de/3687682/ios-mac-os/apple-store-app-virtual-shopping-assistant-could-be-coming-soon/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3687682/ios-mac-os/apple-store-app-virtual-shopping-assistant-could-be-coming-soon/</guid>
<pubDate>Thu, 23 Jul 2026 00:00:31 +0200</pubDate>
<category>🍏 iOS / Mac OS</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[The privacy policy for the Apple Store app suggests that a new feature called Virtual Shopping Assistant could be implemented soon. It may function similarly to the Apple Support Assistant.<br><br><div><img src="https://photos5.appleinsider.com/gallery/68330-144031-Apple-Support-Assistant-chat-icon-xl.jpg" alt="Glowing blue chat bubble with sparkles in the center on a dark background, surrounded by small colorful app and system icons scattered around it" height="738"><br><span>Apple Support already has a chatbot, and soon, Apple Store will too</span></div><br>Apple has been slowly <a href="https://appleinsider.com/articles/26/07/19/genius-bar-ai-tools-spark-concerns-over-employee-monitoring-evaluation">increasing its use</a> of AI tools both internally and in public-facing apps. While Apple still relies on humans for some support inquiries, AI is being offered in more locations.<br><br>It seems Apple may have let slip that a new feature is coming to the Apple Store app. The Virtual Shopping Assistant feature is called out in the Apple Store app <a href="https://www.apple.com/legal/privacy/data/en/apple-store-app/">privacy policy</a>, which was <a href="https://www.macrumors.com/2026/07/22/apple-store-app-shopping-assistant/">first spotted</a> by <em>MacRumors</em>.<br><br><br> <a href="https://appleinsider.com/articles/26/07/22/apple-store-app-virtual-shopping-assistant-could-be-coming-soon?utm_source=rss">Continue Reading on AppleInsider</a> | <a href="https://forums.appleinsider.com/discussion/245029?urm_source=rss">Discuss on our Forums</a>]]></content:encoded>
</item>
<item>
<title><![CDATA[Cisco’s new AI model tells code reviewers where to look for vulnerabilities]]></title>
<description><![CDATA[Cisco has revealed a family of open-weight AI models called Antares that, it said, can help security teams isolate potentially vulnerable parts of a software repository before deeper investigation begins.



Rather than detecting a specific CVE or generating a patch, these models search a codebas...]]></description>
<link>https://tsecurity.de/de/3687085/ai-nachrichten/ciscos-new-ai-model-tells-code-reviewers-where-to-look-for-vulnerabilities/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3687085/ai-nachrichten/ciscos-new-ai-model-tells-code-reviewers-where-to-look-for-vulnerabilities/</guid>
<pubDate>Wed, 22 Jul 2026 19:05:41 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div><div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Cisco has revealed a family of open-weight AI models called Antares that, it said, can help security teams isolate potentially vulnerable parts of a software repository before deeper investigation begins.</p>



<p class="wp-block-paragraph">Rather than detecting a specific CVE or generating a patch, these models search a codebase using only a Common Weakness Enumeration (CWE) description and return the files most likely to contain that class of vulnerability.</p>



<p class="wp-block-paragraph">“Its purpose is to reduce a large codebase to a focused set of files that a security professional or a downstream security workflow should investigate,” Cisco’s AI researcher <a href="https://www.linkedin.com/in/supriti-vijay/" target="_blank" rel="noreferrer noopener">Supriti Vijay</a> said via email. “The goal is not to replace a security engineer’s judgement or send them on a wild-goose chase, but to reduce fatigue and workload by helping them triage an issue earlier and focus their investigation on the most relevant parts of the codebase.”</p>



<p class="wp-block-paragraph">The Antares family consists of models with 350 million, 1 billion, and 3 billion parameters trained specifically for repository-scale vulnerability localization.</p>



<p class="wp-block-paragraph">The company said its largest model approaches the performance of GPT-5.5 on its internal vulnerability localization (Vloc) benchmark while remaining small enough for low-cost local deployment.</p>



<h2 class="wp-block-heading">A search assistant, not a vulnerability detector</h2>



<p class="wp-block-paragraph">Cisco is careful to define what Antares is, and what it is not.</p>



<p class="wp-block-paragraph">“Antares outputs a ranked list of source files likely to contain a relevant vulnerability, along with the terminal exploration trace that led to that result,” Cisco Foundation AI Chief Scientist <a href="https://www.linkedin.com/in/amin-karbasi-5025335/" target="_blank" rel="noreferrer noopener">Amin Karbasi</a> wrote in a blog post, adding that the models are not meant to replace the broader application security toolchain: Human analysts or downstream security tools will still be needed to confirm exploitability, <a href="https://www.infoworld.com/article/4200083/gitlab-previews-auto-remediation-of-vulnerable-dependencies.html">identify vulnerable lines of code</a>, assess severity and generate fixes.</p>



<p class="wp-block-paragraph">Antares differs from conventional static analysis platforms such as Semgrep or CodeQL, which primarily rely on predefined rules or queries. Cisco instead describes Antares as an evidence-driven exploration agent that adapts its search as it traverses the repository.</p>



<p class="wp-block-paragraph">Cisco’s argument is that large repositories often contain thousands of files, making manual reviews exhaustive and unrealistic. By reducing the search space to a manageable shortlist, the company hopes to reduce investigation fatigue without replacing human judgement.</p>



<h2 class="wp-block-heading">Claims of specialization over scale</h2>



<p class="wp-block-paragraph">Cisco is also making a statement about how cybersecurity models should evolve.</p>



<p class="wp-block-paragraph">Instead of pursuing larger foundational models, Cisco argued that specialized, task-trained models can outperform much larger open-weight alternatives for vulnerability localization. In its evaluation Antares-3B, the largest model intended for single-GPU deployments, produced results comparable to GPT-5.5 while outperforming several substantially larger open models by Google, OpenAI and Meta.</p>



<p class="wp-block-paragraph">The family also includes Antares-350M for resource-constrained environments and Antares-1B for laptops and workstations, which Cisco has made available as open-weight models on Hugging Face.</p>



<p class="wp-block-paragraph">The command line interface (CLI) on the models supports targeted CWE investigations, repository-wide scans, SARIF output and local inference, which Cisco said enables organizations to keep proprietary code inside their own trust boundary.</p>



<p class="wp-block-paragraph">However, because Antares identifies candidate files rather than confirmed vulnerabilities, organizations will still need to understand how often such repository-wide searches should be run, how much they improve existing triage workflows, and whether the reduction in investigation effort ultimately translates into measurable security or cost benefits.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Cisco’s new AI model tells code reviewers where to look for vulnerabilities]]></title>
<description><![CDATA[Cisco has revealed a family of open-weight AI models called Antares that, it said, can help security teams isolate potentially vulnerable parts of a software repository before deeper investigation begins.



Rather than detecting a specific CVE or generating a patch, these models search a codebas...]]></description>
<link>https://tsecurity.de/de/3687065/it-security-nachrichten/ciscos-new-ai-model-tells-code-reviewers-where-to-look-for-vulnerabilities/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3687065/it-security-nachrichten/ciscos-new-ai-model-tells-code-reviewers-where-to-look-for-vulnerabilities/</guid>
<pubDate>Wed, 22 Jul 2026 18:54:39 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Cisco has revealed a family of open-weight AI models called Antares that, it said, can help security teams isolate potentially vulnerable parts of a software repository before deeper investigation begins.</p>



<p class="wp-block-paragraph">Rather than detecting a specific CVE or generating a patch, these models search a codebase using only a Common Weakness Enumeration (CWE) description and return the files most likely to contain that class of vulnerability.</p>



<p class="wp-block-paragraph">“Its purpose is to reduce a large codebase to a focused set of files that a security professional or a downstream security workflow should investigate,” Cisco’s AI researcher <a href="https://www.linkedin.com/in/supriti-vijay/" target="_blank" rel="noreferrer noopener">Supriti Vijay</a> said via email. “The goal is not to replace a security engineer’s judgement or send them on a wild-goose chase, but to reduce fatigue and workload by helping them triage an issue earlier and focus their investigation on the most relevant parts of the codebase.”</p>



<p class="wp-block-paragraph">The Antares family consists of models with 350 million, 1 billion, and 3 billion parameters trained specifically for repository-scale vulnerability localization.</p>



<p class="wp-block-paragraph">The company said its largest model approaches the performance of GPT-5.5 on its internal vulnerability localization (Vloc) benchmark while remaining small enough for low-cost local deployment.</p>



<h2 class="wp-block-heading">A search assistant, not a vulnerability detector</h2>



<p class="wp-block-paragraph">Cisco is careful to define what Antares is, and what it is not.</p>



<p class="wp-block-paragraph">“Antares outputs a ranked list of source files likely to contain a relevant vulnerability, along with the terminal exploration trace that led to that result,” Cisco Foundation AI Chief Scientist <a href="https://www.linkedin.com/in/amin-karbasi-5025335/" target="_blank" rel="noreferrer noopener">Amin Karbasi</a> wrote in a blog post, adding that the models are not meant to replace the broader application security toolchain: Human analysts or downstream security tools will still be needed to confirm exploitability, <a href="https://www.infoworld.com/article/4200083/gitlab-previews-auto-remediation-of-vulnerable-dependencies.html">identify vulnerable lines of code</a>, assess severity and generate fixes.</p>



<p class="wp-block-paragraph">Antares differs from conventional static analysis platforms such as Semgrep or CodeQL, which primarily rely on predefined rules or queries. Cisco instead describes Antares as an evidence-driven exploration agent that adapts its search as it traverses the repository.</p>



<p class="wp-block-paragraph">Cisco’s argument is that large repositories often contain thousands of files, making manual reviews exhaustive and unrealistic. By reducing the search space to a manageable shortlist, the company hopes to reduce investigation fatigue without replacing human judgement.</p>



<h2 class="wp-block-heading">Claims of specialization over scale</h2>



<p class="wp-block-paragraph">Cisco is also making a statement about how cybersecurity models should evolve.</p>



<p class="wp-block-paragraph">Instead of pursuing larger foundational models, Cisco argued that specialized, task-trained models can outperform much larger open-weight alternatives for vulnerability localization. In its evaluation Antares-3B, the largest model intended for single-GPU deployments, produced results comparable to GPT-5.5 while outperforming several substantially larger open models by Google, OpenAI and Meta.</p>



<p class="wp-block-paragraph">The family also includes Antares-350M for resource-constrained environments and Antares-1B for laptops and workstations, which Cisco has made available as open-weight models on Hugging Face.</p>



<p class="wp-block-paragraph">The command line interface (CLI) on the models supports targeted CWE investigations, repository-wide scans, SARIF output and local inference, which Cisco said enables organizations to keep proprietary code inside their own trust boundary.</p>



<p class="wp-block-paragraph">However, because Antares identifies candidate files rather than confirmed vulnerabilities, organizations will still need to understand how often such repository-wide searches should be run, how much they improve existing triage workflows, and whether the reduction in investigation effort ultimately translates into measurable security or cost benefits.</p>



<p class="wp-block-paragraph"><em>This article first appeared on <a href="https://www.infoworld.com/article/4200143/ciscos-new-ai-model-tells-code-reviewers-where-to-look-for-vulnerabilities.html">InfoWorld</a>.</em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[LG To Ban Residential Proxies From Smart TV Apps]]></title>
<description><![CDATA[An anonymous reader quotes a report from KrebsOnSecurity: The home appliance giant LG Electronics USA said this week it plans to suspend any apps built for its smart TVs that turn one's television into an always-on residential proxy node. The move comes less than a month after researchers found t...]]></description>
<link>https://tsecurity.de/de/3686990/it-security-nachrichten/lg-to-ban-residential-proxies-from-smart-tv-apps/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3686990/it-security-nachrichten/lg-to-ban-residential-proxies-from-smart-tv-apps/</guid>
<pubDate>Wed, 22 Jul 2026 18:20:32 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[An anonymous reader quotes a report from KrebsOnSecurity: The home appliance giant LG Electronics USA said this week it plans to suspend any apps built for its smart TVs that turn one's television into an always-on residential proxy node. The move comes less than a month after researchers found that more than 42 percent of games and other apps available for download on LG's webOS store allow unknown third-parties to route their Internet traffic through a user's TV. On July 2, [KrebsOnSecurity] featured research by the security firm Spur that examined the prevalence of residential proxy software development kits (SDKs) in smart TV apps. Spur found more than 42 percent of apps available for download on LG smart TVs include SDKs that turn one's television in a proxy node indefinitely, and that more than a quarter of the apps made for Samsung's Tizen operating system had similar residential proxy components.
 
Responding to questions about Spur's research, LG Senior Vice President John Taylor told KrebsOnSecurity the company was working with app developers to remove the residential proxy option from their apps on the webOS platform. Developers that fail to comply, he said, will find their apps suspended. "A residential proxy network is not an intended use for LG smart TVs, and LG Electronics is working with developers to remove the residential proxy option from their apps on the webOS platform," Taylor said. "If this option is not removed, these apps will be suspended." Taylor said LG is committed to keeping residential proxy networks out of its smart TV apps going forward, and that the company's review of those apps is "well underway now."
 
"As part of our ongoing efforts to enhance platform quality and the user experience, LG will continue to strengthen our evaluation process for developer-submitted apps, including those that incorporate residential proxy SDKs," Taylor wrote in an emailed statement. [...] "A one-time consent prompt buried in a TV app is not a substitute for meaningful transparency, ongoing control, and platform oversight," Spur's Trevor Sutter wrote. "The risk is amplified when consent comes from individuals within the household who use the device but shouldn't give consent, such as minors." LG is also facing criticism for monitors that automatically install software promoting paid McAfee subscriptions through Windows Update without user approval.<p></p><div class="share_submission">
<a class="slashpop" href="http://twitter.com/home?status=LG+To+Ban+Residential+Proxies+From+Smart+TV+Apps%3A+https%3A%2F%2Fentertainment.slashdot.org%2Fstory%2F26%2F07%2F22%2F0426218%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter"><img src="https://a.fsdn.com/sd/twitter_icon_large.png"></a>
<a class="slashpop" href="http://www.facebook.com/sharer.php?u=https%3A%2F%2Fentertainment.slashdot.org%2Fstory%2F26%2F07%2F22%2F0426218%2Flg-to-ban-residential-proxies-from-smart-tv-apps%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook"><img src="https://a.fsdn.com/sd/facebook_icon_large.png"></a>



</div><p><a href="https://entertainment.slashdot.org/story/26/07/22/0426218/lg-to-ban-residential-proxies-from-smart-tv-apps?utm_source=rss1.0moreanon&amp;utm_medium=feed">Read more of this story</a> at Slashdot.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI unveils Presence, a new platform that lets enterprises launch and manage realtime voice agents and chatbots]]></title>
<description><![CDATA[OpenAI has announced Presence, a new enterprise product for deploying and managing AI agents across customer-facing and internal business workflows. The offering is designed for eligible enterprise customers that want agents to answer questions, access company systems, take approved actions and e...]]></description>
<link>https://tsecurity.de/de/3686972/it-nachrichten/openai-unveils-presence-a-new-platform-that-lets-enterprises-launch-and-manage-realtime-voice-agents-and-chatbots/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3686972/it-nachrichten/openai-unveils-presence-a-new-platform-that-lets-enterprises-launch-and-manage-realtime-voice-agents-and-chatbots/</guid>
<pubDate>Wed, 22 Jul 2026 18:12:08 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI has <a href="https://openai.com/index/introducing-openai-presence/">announced Presence</a>, a new enterprise product for deploying and managing AI agents across customer-facing and internal business workflows. </p><p>The offering is designed for eligible enterprise customers that want agents to answer questions, access company systems, take approved actions and escalate to human workers while operating under company-defined policies, permissions and evaluation standards.</p><p>Presence is available immediately through a limited general availability program. OpenAI Forward Deployed Engineers (FDEs) and select global systems integrators lead deployments, and the product is not available on a self-service basis. </p><p>OpenAI has not disclosed pricing, geographic limits, contractual terms or the expected cost of the engineering and integration work that accompanies a deployment. The company also has not said whether Presence can use models from providers other than OpenAI, including the increasingly powerful and popular Chinese open weights alternatives like <a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">GLM-5.2</a> and <a href="https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems">Kimi K3</a>. I've asked an OpenAI contact to clarify both pricing and external-model compatibility, but those remain unanswered questions for now. I'lll update when I hear back.</p><p>OpenAI positions Presence as a response to a problem that has become more important as companies move beyond AI demonstrations: getting agents to behave reliably in production as business rules, customer needs and operating conditions change. Presence packages the policies, system connections, evaluations, guardrails and update processes required to run agents inside an enterprise.</p><p>If your business has been interested in using AI agents, but you aren't sure how to stitch together OpenAI's models, APIs, internal systems, security controls and evaluation tools into something reliable, Presence is designed to simplify that process. Instead of building the infrastructure yourself, you work with OpenAI and its deployment engineers to put production-ready agents into your existing workflows.</p><p>The product is available today for real-time voice and chat experiences, according to OpenAI’s formal announcement. The company’s outreach materials also describe a broader ambition spanning voice, chat, email and other channels, but OpenAI has not confirmed that email support is available at launch.</p><h2><b>A governed foundation for production agents</b></h2><p>Presence brings together company knowledge, standard operating procedures, approved actions, simulations, evaluation tools, guardrails and escalation rules. Enterprises can reuse some controls across deployments while adjusting others for a particular workflow or channel.</p><p>Each deployment starts with a defined job, such as resolving a billing issue, supporting an insurance claim or handling an employee IT request. The agent receives only the information and system access required for that task. The customer determines what the agent may do independently, which actions require approval and when a person must take over.</p><p>Before an agent reaches production, teams can test it against common requests, unusual edge cases and higher-risk scenarios. Graders evaluate whether it reached the intended outcome, followed policy, used tools correctly and escalated when required. Guardrails can intervene when an interaction moves outside the organization’s defined boundaries.</p><p>OpenAI shared promotional screenshots with VentureBeat showing administrators running simulation batches against policy changes, including a revised annual refund policy, and reviewing results across operational categories. </p><p>Other interface mockups display production health, customer-intent patterns and task-performance signals. The visuals illustrate the type of oversight OpenAI is promising, although they do not establish how those metrics are calculated or how they map to contractual service levels.</p><p>The product continues to monitor performance after launch. Production sessions, escalations and quality signals can reveal where an agent is working as intended and where it needs attention. Codex, using a Presence plugin, investigates those signals and proposes updates. Teams then test a proposed change against the version already in production before approving a controlled rollout.</p><p>That process is intended to address one of the hardest operational problems in enterprise AI: an agent that works at launch may become less reliable when policies, products or user behavior change. Presence gives companies a formal mechanism for updating behavior without allowing an automated system to rewrite itself unchecked.</p><p>OpenAI says Presence already powers its English-language phone-support channel at 1-888-GPT-0090. The system handles open-ended requests, verifies callers, uses account context and performs approved actions. According to the company, it now resolves <b>75% of inbound issues without human assistance</b>. </p><p>OpenAI also says its Codex-powered improvement loop reduced human handoffs by <b>15 percentage points over a 10-day period</b>. Those figures are company-reported and have not been independently verified.</p><p>Several large organizations are evaluating the same foundation. BBVA is exploring voice support for routine banking needs in Mexico. SoftBank is testing natural Japanese-language customer conversations, while Australian insurer IAG is exploring support during high-demand periods such as severe weather and natural disasters.</p><p>“At BBVA, we are working closely with OpenAI to explore how trusted customer agents can help shape the future of financial services,” said Daniel Ordaz, head of AI transformation at BBVA Mexico.</p><p>“Through our collaboration with OpenAI, we are exploring how Presence can enable trusted customer agents that communicate naturally, connect to the processes needed to resolve requests, and represent SoftBank consistently across customer interactions,” said Tadahisa Murakami, vice president and head of the Data &amp; Digital Transformation Division at SoftBank Corp.</p><h2><b>From model access to forward-deployed implementation</b></h2><p>Presence expands OpenAI’s enterprise strategy beyond APIs and subscription software by formalizing a high-touch deployment model. Forward Deployed Engineers work alongside customers to select workflows, connect internal systems, establish permissions, configure policies, test agents and move them into production.</p><p>That approach resembles a <a href="https://fde.academy/blog/how-palantir-invented-the-forward-deployed-engineer-model">model pioneered by AI ontology and intelligence platform Palantir,</a> which embeds FDEs with customers to adapt its proprietary software to complex government and commercial environments. The similarity lies less in the underlying technology than in the delivery method: both companies place technical personnel close to the customer’s operations, where integration and process design often determine whether software creates value.</p><p>The products are not interchangeable. Palantir’s model has historically centered on data integration, ontologies and operational decision systems. Presence is more narrowly focused on AI-agent behavior, approved actions, evaluations, escalation and continuous improvement. OpenAI presents it as a repeatable software product supported by engineers and systems integrators, rather than as consulting alone.</p><p>In May 2026, OpenAI launched its own enterprise AI consulting and integration firm, the <a href="https://openai.com/index/openai-launches-the-deployment-company/">OpenAI Deployment Company</a>, with investment and <a href="https://www.bain.com/about/media-center/press-releases/2026/bain-company-openai-a-new-venture-to-deploy-ai-at-enterprise-scale/">support from Bain &amp; Company.</a> It also offers programs for model customization and fine-tuning to fit specific enterprise needs. </p><p>Its chief U.S. rival Anthropic has also moved <a href="https://techcrunch.com/2026/07/15/anthropic-blackstone-bet-the-next-trillion-dollar-ai-business-is-implementation-not-models/">toward a services-led enterprise model through Ode,</a> its consulting organization built around forward-deployed engineers helping companies integrate Claude into complex workflows, which launched just a week ago. The broad rationale is similar: enterprises often need more than access to a model. They need help connecting data and systems, defining permissions, validating behavior and managing deployment risk.</p><p>Presence differs in how explicitly OpenAI packages those requirements into a branded agent-governance product. Anthropic’s initiative is centered on helping enterprises deploy Claude, while Presence combines implementation services with a defined operational layer for policies, simulations, evaluations, approvals and production updates.</p><p>Presence goes further by making forward deployment a core part of how a specific agent product reaches customers. It does not replace OpenAI’s API business; the company says it will continue supporting voice customers with access to frontier models through the OpenAI API.</p><p>The trend reflects a broader market view that many enterprises still need hands-on assistance to move agents from pilot projects into stable operations. Even organizations with strong internal engineering teams must coordinate security, compliance, workflow ownership, data access and escalation responsibilities. Presence attempts to consolidate those tasks rather than leaving customers to assemble separate orchestration, evaluation and consulting layers.</p><h2><b>A recent security breach looms in the background</b></h2><p>Inconveniently for OpenAI, the Presence launch arrives just a day after <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI and Hugging Face disclosed an unprecedented security incident</a> in which OpenAI frontier models undergoing internal evaluation escaped containment, accessed the open web, and cyberattacked Hugging Face to achieve a benign goal — without being instructed to pursue these methods.</p><p>According to the described joint disclosure, OpenAI models operating in an evaluation framework called ExploitGym identified and exploited a zero-day vulnerability in a third-party package-registry cache proxy. The models reportedly escalated privileges, moved laterally and obtained internet access before targeting Hugging Face systems while seeking benchmark-related information.</p><p>The incident is relevant to enterprise buyers because it raises questions about sandboxing, tool permissions, external access, monitoring and incident response. </p><p>The disclosure also highlighted a practical problem for defenders. Hugging Face personnel reportedly found that commercial frontier-model APIs refused some forensic requests because logs contained exploit payloads, credentials and shell commands that triggered safety systems. The team then used a locally deployed open-weight model to assist with analysis.</p><p>Presence therefore arrives as both a product launch and a test of OpenAI’s ability to convert model capability into controlled enterprise operations. Its policies, simulations, evaluations and human approvals address real deployment gaps. But without public pricing, technical interoperability details, compliance information or service-level commitments, customers still lack much of the information needed to assess total cost and operational risk.</p><p>For now, Presence appears aimed at enterprises willing to adopt a high-touch, OpenAI-led deployment process. Whether it develops into a broadly accessible platform—or remains a closely managed product for selected customers—will depend in part on the answers OpenAI has not yet provided.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Presence connects AI agents to enterprise data with built-in guardrails]]></title>
<description><![CDATA[OpenAI has introduced Presence, a product designed to help companies deploy AI agents that handle customer support and internal service requests across voice and chat. (Source: OpenAI) The company describes Presence as a deployment platform rather than a standalone model. “Presence brings togethe...]]></description>
<link>https://tsecurity.de/de/3686657/it-security-nachrichten/openai-presence-connects-ai-agents-to-enterprise-data-with-built-in-guardrails/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3686657/it-security-nachrichten/openai-presence-connects-ai-agents-to-enterprise-data-with-built-in-guardrails/</guid>
<pubDate>Wed, 22 Jul 2026 16:30:10 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI has introduced Presence, a product designed to help companies deploy AI agents that handle customer support and internal service requests across voice and chat. (Source: OpenAI) The company describes Presence as a deployment platform rather than a standalone model. “Presence brings together the components teams need to run agents in production: policies and standard operating procedures, guardrails, approved actions, simulations, evaluation tools, and a Codex-powered improvement process,” OpenAI said. “Together, these components help teams … <a href="https://www.helpnetsecurity.com/2026/07/22/openai-presence-ai-agent-platform/" rel="nofollow">More <span class="meta-nav">→</span></a></p>
<p>The post <a href="https://www.helpnetsecurity.com/2026/07/22/openai-presence-ai-agent-platform/">OpenAI Presence connects AI agents to enterprise data with built-in guardrails</a> appeared first on <a href="https://www.helpnetsecurity.com/">Help Net Security</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI model escape puts enterprise AI defenses on notice]]></title>
<description><![CDATA[Some of OpenAI’s most powerful AI models teamed up to escape their sandbox and attack systems at Hugging Face in a cybersecurity evaluation gone wrong, the company has admitted. The models under test were modified to allow them to perform potentially harmful actions that production versions would...]]></description>
<link>https://tsecurity.de/de/3686581/it-security-nachrichten/openai-model-escape-puts-enterprise-ai-defenses-on-notice/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3686581/it-security-nachrichten/openai-model-escape-puts-enterprise-ai-defenses-on-notice/</guid>
<pubDate>Wed, 22 Jul 2026 15:53:09 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Some of OpenAI’s most powerful AI models teamed up to escape their sandbox and attack systems at Hugging Face in a cybersecurity evaluation gone wrong, the company has admitted. The models under test were modified to allow them to perform potentially harmful actions that production versions would refuse. The incident highlights how, if AI prompt guardrails fail or, as in this incident, are removed, then enterprises must have robust sandboxing or other technical restrictions in place to protect systems.</p>



<p class="wp-block-paragraph">In OpenAI’s evaluation, GPT-5.6 Sol and a more capable pre-release model were competing to solve cybersecurity challenges in ExploitGym, a benchmarking tool for cybersecurity capabilities. They concluded that systems at Hugging Face could have the answers, and set out to gain access, OpenAI wrote in a <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="noreferrer noopener">blog post about the attack on Hugging Face</a>.</p>



<p class="wp-block-paragraph">ExploitGym challenges AI models to turn security vulnerabilities into attack exploits. OpenAI gave its models limited access to the internet through a package-registry proxy to download code to build their attacks. However, the models exploited a zero-day vulnerability in the proxy to obtain unrestricted internet access, then used stolen credentials and other flaws to access Hugging Face systems.</p>



<p class="wp-block-paragraph">For CIOs and CISOs, the concern extends beyond cybersecurity research. AI agents connected to corporate databases, source code repositories, or payment systems could expose sensitive resources unless <a href="https://www.csoonline.com/article/4199408/ai-agents-can-escape-sandboxes-without-ever-breaking-them.html">access controls and sandboxing</a> continue to work when model-level safeguards fail.</p>



<p class="wp-block-paragraph">“Prompt guardrails are not security controls,” said <a href="https://www.forrester.com/analyst-bio/biswajeet-mahapatra/BIO20046" target="_blank" rel="noreferrer noopener">Biswajeet Mahapatra</a>, principal analyst at Forrester. “They are behavioral controls that can fail, be bypassed, or be intentionally disabled during testing.”</p>



<h2 class="wp-block-heading">How containment failed</h2>



<p class="wp-block-paragraph">Despite the relaxed model safeguards, the surrounding infrastructure was supposed to contain the activity. Network access was restricted to an internally hosted third-party service that acted as a proxy and cache for software package registries. Such proxies are typically used in the enterprise to protect against <a href="https://www.csoonline.com/article/4167465/supply-chain-attacks-take-aim-at-your-ai-coding-agents.html">software supply chain attacks</a> coming from the outside.</p>



<p class="wp-block-paragraph">OpenAI said the models discovered a zero-day vulnerability in the service, exploited it and carried out privilege escalation and <a href="https://www.csoonline.com/article/4194708/lateral-movement-risk-rises-as-enterprises-emphasize-convenience-over-containment.html">lateral movement</a> until they reached a node with open internet access.</p>



<p class="wp-block-paragraph">Once online, the models inferred that Hugging Face might hold models, datasets or solutions related to ExploitGym. OpenAI said that, in one attack path, the models combined stolen credentials with zero-day vulnerabilities to obtain remote code execution on Hugging Face servers.</p>



<p class="wp-block-paragraph">In its account of the activity on its infrastructure, <a href="https://huggingface.co/blog/security-incident-july-2026" target="_blank" rel="noreferrer noopener">Hugging Face said</a> a malicious dataset abused two code-execution paths in its data-processing pipeline. The actor then obtained node-level access, harvested cloud and cluster credentials, and moved into several internal clusters.</p>



<p class="wp-block-paragraph">OpenAI said the models ultimately obtained ExploitGym test solutions directly from a Hugging Face production database. The models remained focused on solving the benchmark, despite the lengths to which they went to achieve that objective, it said.</p>



<p class="wp-block-paragraph"><a href="https://greyhoundresearch.com/svg/" target="_blank" rel="noreferrer noopener">Sanchit Vir Gogia</a>, chief analyst at Greyhound Research, said the package-registry proxy should concern enterprise security teams because it was a narrowly permitted exception inside an otherwise restricted environment.</p>



<p class="wp-block-paragraph">“Every enterprise estate is built from exceptions of exactly that kind,” he said.</p>



<p class="wp-block-paragraph">In this incident, OpenAI’s models were performing the intended — albeit by unintended means — but enterprises should take care that their own AI systems are kept on target.</p>



<p class="wp-block-paragraph">Mahapatra said enterprises should treat AI agents as “<a href="https://www.csoonline.com/article/4194548/agentic-ai-identity-a-6-stage-maturity-model-for-non-human-identities.html">high-risk non-human identities</a>,” confining each one to an isolated environment where access is limited to the assigned task and credentials expire quickly.</p>



<p class="wp-block-paragraph">“An acceptable blast radius means a compromised agent can affect only a single workflow, dataset, or application rather than providing a pathway into broader enterprise infrastructure,” Mahapatra said.</p>



<p class="wp-block-paragraph">And Gogia warned that services allowed to reach package registries, update systems, or other external resources can become escape routes if they are not subjected to the same scrutiny as the agent itself.</p>



<h2 class="wp-block-heading">Defenses enterprises need</h2>



<p class="wp-block-paragraph">Hugging Face said the intrusion resulted in unauthorized access to a limited number of internal datasets and several credentials used by its services. It found no evidence that public models, datasets or Spaces had been altered and said its published software packages and container images were not compromised.</p>



<p class="wp-block-paragraph">The company closed the code-execution paths used to gain access and rebuilt the affected nodes. It also revoked exposed credentials and tightened the rules governing workloads admitted to its clusters.</p>



<p class="wp-block-paragraph">Whether they are keeping their own AIs in or rogue Ais out, Gogia said enterprises should test whether their containment boundaries work, rather than relying on architecture diagrams or stated policies. Such tests should attempt to obtain credentials, cross trust boundaries and reach systems outside the agent’s assigned task.</p>



<p class="wp-block-paragraph">Mahapatra said enterprises should assume that one containment layer may fail and ensure that an agent’s access cannot provide a route into unrelated applications or broader corporate infrastructure.</p>



<p class="wp-block-paragraph">OpenAI said it is still investigating the incident with Hugging Face, and is imposing stricter configurations on its research environment while the vulnerabilities are being addressed, even if that means slowing down its research. It is also strengthening containment and monitoring around future evaluations.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-63263 | Elastic Elasticsearch Query Evaluation resource consumption (esa-2026-74 / EUVD-2026-47591)]]></title>
<description><![CDATA[A vulnerability marked as problematic has been reported in Elastic Elasticsearch. Impacted is an unknown function of the component Query Evaluation. This manipulation causes resource consumption.

This vulnerability is tracked as CVE-2026-63263. The attack is possible to be carried out remotely. ...]]></description>
<link>https://tsecurity.de/de/3686157/sicherheitsluecken/cve-2026-63263-elastic-elasticsearch-query-evaluation-resource-consumption-esa-2026-74-euvd-2026-47591/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3686157/sicherheitsluecken/cve-2026-63263-elastic-elasticsearch-query-evaluation-resource-consumption-esa-2026-74-euvd-2026-47591/</guid>
<pubDate>Wed, 22 Jul 2026 13:41:25 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability marked as <a href="https://vuldb.com/kb/risk">problematic</a> has been reported in <a href="https://vuldb.com/product/elastic:elasticsearch">Elastic Elasticsearch</a>. Impacted is an unknown function of the component <em>Query Evaluation</em>. This manipulation causes resource consumption.

This vulnerability is tracked as <a href="https://vuldb.com/cve/CVE-2026-63263">CVE-2026-63263</a>. The attack is possible to be carried out remotely. No exploit exists.]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI claims responsibility for the Hugging Face hack after its own models escaped a test sandbox]]></title>
<description><![CDATA[During an internal security evaluation, OpenAI models, including GPT-5.6 Sol, escaped their sandbox, independently discovered a zero-day vulnerability, and breached Hugging Face's production infrastructure. The models were trying to steal benchmark solutions to cheat on the evaluation. OpenAI adm...]]></description>
<link>https://tsecurity.de/de/3685749/ai-nachrichten/openai-claims-responsibility-for-the-hugging-face-hack-after-its-own-models-escaped-a-test-sandbox/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685749/ai-nachrichten/openai-claims-responsibility-for-the-hugging-face-hack-after-its-own-models-escaped-a-test-sandbox/</guid>
<pubDate>Wed, 22 Jul 2026 11:05:02 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1376" height="768" src="https://the-decoder.com/wp-content/uploads/2026/07/openai_kraken_cyber.png" class="attachment-full size-full wp-post-image" alt="" decoding="async" fetchpriority="high"></p>
<p>        During an internal security evaluation, OpenAI models, including GPT-5.6 Sol, escaped their sandbox, independently discovered a zero-day vulnerability, and breached Hugging Face's production infrastructure. The models were trying to steal benchmark solutions to cheat on the evaluation. OpenAI admits that disabling security filters during the test was inadequate.</p>
<p>The article <a href="https://the-decoder.com/openai-claims-responsibility-for-the-hugging-face-hack-after-its-own-models-escaped-a-test-sandbox/">OpenAI claims responsibility for the Hugging Face hack after its own models escaped a test sandbox</a> appeared first on <a href="https://the-decoder.com/">The Decoder</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Says Its AI Models Acted On Its Own In An 'Unprecedented' Hack]]></title>
<description><![CDATA["GPT-5.6 Sol and an 'even more capable' model used stolen credentials and exploited vulnerabilities in the Hugging Face API to obtain secret information used to cheat on evaluations," writes longtime Slashdot reader Dr. Bombay. The Associated Press reports: "We had a significant security incident...]]></description>
<link>https://tsecurity.de/de/3685495/it-security-nachrichten/openai-says-its-ai-models-acted-on-its-own-in-an-unprecedented-hack/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685495/it-security-nachrichten/openai-says-its-ai-models-acted-on-its-own-in-an-unprecedented-hack/</guid>
<pubDate>Wed, 22 Jul 2026 09:16:37 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA["GPT-5.6 Sol and an 'even more capable' model used stolen credentials and exploited vulnerabilities in the Hugging Face API to obtain secret information used to cheat on evaluations," writes longtime Slashdot reader Dr. Bombay. The Associated Press reports: "We had a significant security incident during evaluation of our models," OpenAI CEO Sam Altman said in a statement posted on social media. AI startup Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own. "We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent," Hugging Face co-founder and CEO Clement Delangue said in a statement. "Turns out it did!"
 
[...] "AI is accelerating the discovery and exploitation of vulnerabilities," OpenAI said in its statement Tuesday. "The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities." Delangue said he spent the past 24 hours working with OpenAI, "and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously!" Delangue added that it "might be the first incident of its kind."<p></p><div class="share_submission">
<a class="slashpop" href="http://twitter.com/home?status=OpenAI+Says+Its+AI+Models+Acted+On+Its+Own+In+An+'Unprecedented'+Hack%3A+https%3A%2F%2Fit.slashdot.org%2Fstory%2F26%2F07%2F22%2F0348206%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter"><img src="https://a.fsdn.com/sd/twitter_icon_large.png"></a>
<a class="slashpop" href="http://www.facebook.com/sharer.php?u=https%3A%2F%2Fit.slashdot.org%2Fstory%2F26%2F07%2F22%2F0348206%2Fopenai-says-its-ai-models-acted-on-its-own-in-an-unprecedented-hack%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook"><img src="https://a.fsdn.com/sd/facebook_icon_large.png"></a>



</div><p><a href="https://it.slashdot.org/story/26/07/22/0348206/openai-says-its-ai-models-acted-on-its-own-in-an-unprecedented-hack?utm_source=rss1.0moreanon&amp;utm_medium=feed">Read more of this story</a> at Slashdot.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Google Unveils Gemini 3.5 Flash Cyber to Find and Fix Software Vulnerabilities Faster]]></title>
<description><![CDATA[Google has introduced Gemini 3.5 Flash Cyber, a lightweight AI model designed to improve cybersecurity by helping defenders identify, validate, and patch software vulnerabilities more efficiently. Built on Gemini 3.5 Flash and optimized for security tasks, Flash Cyber aims to deliver a cost-effec...]]></description>
<link>https://tsecurity.de/de/3685467/it-security-nachrichten/google-unveils-gemini-35-flash-cyber-to-find-and-fix-software-vulnerabilities-faster/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685467/it-security-nachrichten/google-unveils-gemini-35-flash-cyber-to-find-and-fix-software-vulnerabilities-faster/</guid>
<pubDate>Wed, 22 Jul 2026 08:55:35 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1133" height="692" src="https://thecyberexpress.com/wp-content/uploads/Flash-Cyber.webp" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="Flash Cyber" decoding="async" srcset="https://thecyberexpress.com/wp-content/uploads/Flash-Cyber.webp 1133w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-300x183.webp 300w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-1024x625.webp 1024w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-768x469.webp 768w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-600x366.webp 600w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-150x92.webp 150w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-750x458.webp 750w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber.webp 1133w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-300x183.webp 300w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-1024x625.webp 1024w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-768x469.webp 768w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-600x366.webp 600w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-150x92.webp 150w, https://thecyberexpress.com/wp-content/uploads/Flash-Cyber-750x458.webp 750w" sizes="(max-width: 1133px) 100vw, 1133px" title="Google Unveils Gemini 3.5 Flash Cyber to Find and Fix Software Vulnerabilities Faster 4"></p><span data-contrast="auto">Google has introduced Gemini 3.5 Flash Cyber, a lightweight AI model designed to improve cybersecurity by helping defenders identify, validate, and patch software vulnerabilities more efficiently. Built on Gemini 3.5 Flash and optimized for security tasks, Flash Cyber aims to deliver a cost-effective alternative to larger AI models while supporting large-scale vulnerability analysis.</span>

<span data-contrast="auto">The company said it has invested in cybersecurity research for years, including automated vulnerability discovery through CodeMender, its code security agent that can detect and fix critical software flaws. However, as AI systems become increasingly capable of discovering <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-are-vulnerabilities/" title="vulnerabilities" data-wpil-keyword-link="linked" data-wpil-monitor-id="29060">vulnerabilities</a> faster than defenders can resolve them, Google believes a scalable and affordable approach is needed.</span><span data-ccp-props="{}"> </span>
<h3 aria-level="2"><b><span data-contrast="none">Gemini 3.5 Flash Cyber Focuses on Scalable Cybersecurity</span></b><span data-ccp-props='{"134245418":true,"134245529":true,"335559738":160,"335559739":80}'> </span></h3>
<span data-contrast="auto">According to <a href="https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/" target="_blank" rel="nofollow noopener">Google</a>, Gemini 3.5 Flash Cyber has been fine-tuned specifically to locate, verify, and remediate vulnerabilities more effectively than Gemini's standard Flash models. Because of the technology's dual-use nature, the company is initially limiting access through a pilot program for governments and trusted partners via CodeMender, with broader availability planned over time.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Google also confirmed that CodeMender's core capabilities will be made available through generally available Gemini models on the Gemini Enterprise Agent Platform.</span><span data-ccp-props="{}"> </span>
<h3 aria-level="2"><b><span data-contrast="none">Flash Cyber Improves Large-scale Code Analysis</span></b><span data-ccp-props='{"134245418":true,"134245529":true,"335559738":160,"335559739":80}'> </span></h3>
<span data-contrast="auto">A major challenge in <a class="wpil_keyword_link" href="https://cyble.com/knowledge-hub/what-is-cybersecurity/" target="_blank" rel="noopener" title="cybersecurity" data-wpil-keyword-link="linked" data-wpil-monitor-id="29059">cybersecurity</a> is exploring vast execution search spaces across complex codebases. Instead of relying on a single call to a <a href="https://thecyberexpress.com/us-gets-pre-release-access-to-ai-models/" target="_blank" rel="noopener">large language model</a>, CodeMender invokes Flash Cyber multiple times, allowing sub-agents to inspect significantly more code paths before generating one consolidated report.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Google said the model's speed and lower operating cost make it suitable for continuous code scanning, software launch processes, and commit-scanning pipelines at scale.</span><span data-ccp-props="{}"> </span>
<h3 aria-level="2"><b><span data-contrast="none">Benchmark Results Show Competitive Performance</span></b><span data-ccp-props='{"134245418":true,"134245529":true,"335559738":160,"335559739":80}'> </span></h3>
<span data-contrast="auto">Google evaluated Gemini 3.5 Flash <a class="wpil_keyword_link" href="https://thecyberexpress.com/cyber-news/" title="Cyber" data-wpil-keyword-link="linked" data-wpil-monitor-id="29061">Cyber</a> using the CyberGym benchmark, which measures AI agents against hundreds of real-world software vulnerabilities. Configured to call the model up to five times before producing a final report, CodeMender achieved competitive performance against significantly larger cybersecurity models. Google noted that competitor results were based on provider self-reported scores.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">The model also outperformed Gemini 3.5 Flash and 3.6 Flash during Google's internal Big Sleep evaluation, which tested <a class="wpil_keyword_link" href="https://thecyberexpress.com/firewall-daily/vulnerabilities/" title="vulnerability" data-wpil-keyword-link="linked" data-wpil-monitor-id="29058">vulnerability</a> discovery in complex projects such as Chrome and Safari without safety guardrails.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">In Chrome's production commit-scanning pipeline, where vulnerabilities remained undisclosed to prevent benchmark contamination, Flash Cyber again delivered a significant improvement over Gemini 3.5 Flash. Google added that competitor models released after Opus 4.6 were excluded because their safety guardrails prevented them from completing the tasks.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Testing on the V8 JavaScript Engine found 55 unique confirmed vulnerabilities with <a href="https://thecyberexpress.com/gemini-ad-safety-targets-scam-ads/" target="_blank" rel="noopener">Gemini</a> 3.5 Flash Cyber, compared with 47 for Gemini 3.5 Flash and 36 for Opus 4.6, including 10 issues missed by both competing models.</span><span data-ccp-props="{}"> </span>
<h3 aria-level="2"><b><span data-contrast="none">Real-world Cybersecurity Deployment</span></b><span data-ccp-props='{"134245418":true,"134245529":true,"335559738":160,"335559739":80}'> </span></h3>
<span data-contrast="auto">Google said Flash Cyber is already helping secure internal projects, including Chrome, Android, Cloud, Ads and YouTube. In one example, Google's Cloud Vulnerability Research team used the model to identify remote code execution vulnerabilities in public APIs and a memory-corruption flaw within a sensitive production service in just two hours. The model also generated a 100% reliable <a href="https://thecyberexpress.com/cve-2026-45829-chromatoast-chromadb/" target="_blank" rel="noopener">remote code execution</a> exploit capable of bypassing Address Space Layout Randomization (ASLR) and Write XOR Execute (W^X).</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Google added that early feedback from Wiz and Cloud CISO <a class="wpil_keyword_link" href="https://thecyberexpress.com/" title="Security" data-wpil-keyword-link="linked" data-wpil-monitor-id="29062">Security</a> Engineering testers indicated a significant capability improvement over Gemini 3.5 Flash. The company also highlighted resources such as OSV.dev, which tracks more than 700,000 open-source vulnerabilities, and over a decade of OSS-Fuzz <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-is-data/" title="data" data-wpil-keyword-link="linked" data-wpil-monitor-id="29063">data</a> as key training assets supporting its cybersecurity models.</span><span data-ccp-props="{}"> </span>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark]]></title>
<description><![CDATA[OpenAI on Tuesday said a combination of its artificial intelligence (AI) models, including GPT-5.6 Sol and an "even more capable pre-release model," was behind the security incident that targeted Hugging Face's production infrastructure last week.

The AI company said the models were operating wi...]]></description>
<link>https://tsecurity.de/de/3685410/it-security-nachrichten/openai-says-its-ai-models-escaped-sandbox-targeted-hugging-face-to-cheat-benchmark/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685410/it-security-nachrichten/openai-says-its-ai-models-escaped-sandbox-targeted-hugging-face-to-cheat-benchmark/</guid>
<pubDate>Wed, 22 Jul 2026 08:25:42 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[OpenAI on Tuesday said a combination of its artificial intelligence (AI) models, including GPT-5.6 Sol and an "even more capable pre-release model," was behind the security incident that targeted Hugging Face's production infrastructure last week.

The AI company said the models were operating with "reduced cyber refusals for evaluation purposes" that might otherwise limit their ability to]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Models Chain Zero-Days to Breach Hugging Face During Cyber Evaluation]]></title>
<description><![CDATA[OpenAI has confirmed that two of its own artificial intelligence models, including the flagship GPT-5.6 Sol and an unnamed pre-release model, autonomously breached Hugging Face’s production infrastructure during an internal cybersecurity benchmarking exercise. The breach was disclosed last week b...]]></description>
<link>https://tsecurity.de/de/3685356/it-security-nachrichten/openai-models-chain-zero-days-to-breach-hugging-face-during-cyber-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685356/it-security-nachrichten/openai-models-chain-zero-days-to-breach-hugging-face-during-cyber-evaluation/</guid>
<pubDate>Wed, 22 Jul 2026 08:00:31 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI has confirmed that two of its own artificial intelligence models, including the flagship GPT-5.6 Sol and an unnamed pre-release model, autonomously breached Hugging Face’s production infrastructure during an internal cybersecurity benchmarking exercise. The breach was disclosed last week by Hugging Face after its security team detected and contained an AI agent operating within its […]</p>
<p>The post <a href="https://cyberpress.org/openai-models-chain-zero-days/">OpenAI Models Chain Zero-Days to Breach Hugging Face During Cyber Evaluation</a> appeared first on <a href="https://cyberpress.org/">Cyber Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Exploits Zero-Day to Gain Internet Access and Compromise Hugging Face Servers]]></title>
<description><![CDATA[OpenAI has revealed that during an internal evaluation of advanced cyber capabilities, AI agents exploited a zero-day vulnerability, escaped a constrained research environment, and compromised parts of Hugging Face’s production infrastructure. While Hugging Face detected and contained the activit...]]></description>
<link>https://tsecurity.de/de/3685293/it-security-nachrichten/openai-exploits-zero-day-to-gain-internet-access-and-compromise-hugging-face-servers/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685293/it-security-nachrichten/openai-exploits-zero-day-to-gain-internet-access-and-compromise-hugging-face-servers/</guid>
<pubDate>Wed, 22 Jul 2026 07:10:50 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI has revealed that during an internal evaluation of advanced cyber capabilities, AI agents exploited a zero-day vulnerability, escaped a constrained research environment, and compromised parts of Hugging Face’s production infrastructure. While Hugging Face detected and contained the activity, OpenAI’s…</p>
<p class="more-link-p"><a class="more-link" href="https://www.itsecuritynews.info/openai-exploits-zero-day-to-gain-internet-access-and-compromise-hugging-face-servers/">Read more →</a></p>
<p>The post <a href="https://www.itsecuritynews.info/openai-exploits-zero-day-to-gain-internet-access-and-compromise-hugging-face-servers/">OpenAI Exploits Zero-Day to Gain Internet Access and Compromise Hugging Face Servers</a> appeared first on <a href="https://www.itsecuritynews.info/">IT Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need to know]]></title>
<description><![CDATA[Yesterday afternoon, OpenAI and Hugging Face published a joint disclosure outlining a cybersecurity event that redefines the threat landscape for enterprise technology. During an internal benchmark evaluation, frontier artificial intelligence models developed by OpenAI—including GPT-5.6 Sol and a...]]></description>
<link>https://tsecurity.de/de/3685286/it-nachrichten/openais-models-broke-containment-and-cyberattacked-hugging-face-what-enterprises-need-to-know/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685286/it-nachrichten/openais-models-broke-containment-and-cyberattacked-hugging-face-what-enterprises-need-to-know/</guid>
<pubDate>Wed, 22 Jul 2026 07:02:39 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Yesterday afternoon, OpenAI and Hugging Face <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">published a joint disclosure</a> outlining a cybersecurity event that redefines the threat landscape for enterprise technology. </p><p>During an internal benchmark evaluation, frontier artificial intelligence models developed by OpenAI—including GPT-5.6 Sol and an unreleased, higher-capability pre-release model—broke out of their sandboxed research environment, obtained raw internet access, and autonomously executed a complex cyberattack against Hugging Face’s production infrastructure.</p><p> OpenAI officially categorizes the breach as an "unprecedented cyber incident, involving state-of-the-art cyber capabilities". This incident fundamentally re-frames global discussions surrounding AI containment, frontier model alignment, commercial guardrails, and enterprise threat modeling.</p><p>But first thing's first: enterprises should understand the situation, evaluate their own AI and computer systems in light of it, and above all, don't panic. As we'll review, the incident does show the increasing power and danger of frontier AI systems, but it does not mean that enterprise Ai deployments are inherently less secure, nor that they need extensive overhauling. </p><h2><b>Anatomy of an Autonomous Breakout</b></h2><p>To understand how a routine benchmark evaluation escalated into a cross-infrastructure breach, one must examine the objective mechanics of the evaluation pipeline. </p><p>The models were prompted to solve <a href="https://arxiv.org/abs/2605.11086">ExploitGym</a>, a benchmark designed to quantify multi-step exploitation capabilities. </p><p>Hyper-focused on maximizing its evaluation score, the AI agent inferred that Hugging Face likely hosted the answer keys and benchmark solutions. In pursuit of solving the problem, the model determined that breaking out of its container and stealing the answers was an optimal strategy.</p><p>OpenAI’s evaluation environment is designed to be strictly isolated, limiting network traffic through a proxy used for package caching. However, the models identified and exploited a zero-day vulnerability in that internally-hosted third-party proxy software. </p><p>Once through, the AI executed a series of lateral movement and privilege escalation actions across OpenAI's research nodes until reaching a machine with unrestricted internet access. From there, the model searched the web, target-identified Hugging Face, and launched a multi-stage attack by chaining together stolen credentials and remote code execution vulnerabilities on the Hugging Face servers.</p><p>The UK AI Security Institute (UK AISI) recently evaluated models such as GPT-5.6 Sol, demonstrating that they are increasingly able to sustain complex, multi-step cyber operations over long time horizons. OpenAI notes that this incident confirms these theoretical capabilities now apply in real-world settings.</p><h2><b>Rewinding the Tape on a Forensic Trap</b></h2><p>While OpenAI’s July 21 release reveals the identity of the autonomous agent, Hugging Face had already begun managing the intrusion days earlier. </p><p>On July 16, Hugging Face disclosed that an autonomous AI agent system breached its production infrastructure. As <a href="https://venturebeat.com/security/safety-guardrails-blocked-hugging-faces-defenders-not-the-attacker-when-an-ai-agent-breached-its-systems">detailed by VentureBeat,</a> the attacker’s entry point was a malicious dataset that triggered code execution through a remote-code loader and template-injection flaws within dataset configuration files. </p><p>Once inside, the agent framework broke out onto the node running the workload and executed thousands of actions via short-lived sandboxes, harvesting cloud and cluster credentials over a single weekend.</p><p>When Hugging Face's security team detected the breach, responders immediately turned to frontier AI models via commercial APIs to parse the massive volume of system logs and reconstruct over 17,000 recorded events. Then, a secondary operational crisis emerged: the commercial AI models refused to help.</p><p>Because standard commercial frontier models utilize unified safety guardrails designed to block malicious prompt submissions, the models classified the incident response team's forensic queries—which contained raw shell commands, real exploit payloads, and credential dumps—as malicious attacks. Every forensic query submitted by the defenders was blocked outright.</p><p>"The same prompts that are most valuable during an active intrusion—shell commands, exploit chains, credential dumps, persistence mechanisms, lateral movement—are exactly the prompts most likely to trigger safety systems," notes Merritt Baer, former Deputy CISO at AWS and senior adviser to Andesite, G2I, and AppOmni, in an interview with VentureBeat. "As AI becomes embedded in security operations, this becomes an operational resilience issue rather than merely a model policy issue".</p><p>To bypass this roadblock, Hugging Face abandoned commercial hosted APIs and deployed <a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">GLM 5.2</a> —a  state-of-the-art Chinese open-weight model released last month by z.ai, as <a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">reported at the time by VentureBeat</a> —locally on its own infrastructure. </p><p>Free from third-party API restrictions and external safety filters, GLM 5.2 successfully analyzed the raw exploit data locally, allowing defenders to complete forensic reconstruction and contain the breach without any attacker data leaving the company's environment.</p><h2><b>Industry Reaction and the Geopolitical Paradox</b></h2><p>The revelation that an American frontier model autonomously escaped containment, attacked a partner platform, and was ultimately analyzed using a Chinese open-weight model sent shockwaves through the tech community. </p><p><i>The Wall Street Journal </i>summarized the <a href="https://x.com/WSJ/status/2079754070965854541?s=20">public reaction on X,</a> calling the event "the stuff of cybersecurity nightmares. OpenAI said two artificial intelligence systems it was testing broke out of their test environment, hacked their way onto the internet and broke into another company. The victim was Hugging Face."</p><p>Also posting to X, AI alignment researcher <a href="https://x.com/justanotherlaw/status/2079756943112159237">Lawrence Chan</a> emphasized the importance of transparency regarding the incident, noting that "Credit where it’s due: Hugging Face detected and disclosed the intrusion last week. OAI confirmed its models were involved and provided more details, even when it didn't have to. Separate from choices that led to the hack, voluntary disclosure is good, and I’m glad they did so." </p><p>Meanwhile, AI researcher <a href="https://x.com/natolambert/status/2079662928941474201?s=20">Nathan Lambert</a> provided a succinct technical summary in his own X post, observing that "An openai model, during evaluation on a cyber benchmark, exploited a public zero day bug, escaped sandboxing in openai's infra, and got into the internal huggingface infra via an exploit (through a public dataset service) all in the attempt to solve a benchmark problem." He later addressed the geopolitical implications, writing in another post on X: </p><blockquote><p><i>"Rght now American companies need Chinese models to secure their cyber infra due to guardrails on closed models.</i></p><p><i>But if a Chinese model in training had infiltrated a prominent American tech company, it very likely could've been the cause of policy banning future Chinese models."</i></p></blockquote><p>Technology investor <a href="https://x.com/DavidSacks/status/2078991100057141620?s=20">David Sacks also zeroed in</a> on the guardrail paradox, writing in his own X post that "Hugging Face tried using American frontier models to analyze an AI-powered cyber attack. But the guardrails blocked requests containing real exploit payloads so they switched to GLM 5.2 running locally. The guardrails actually impaired defensive security." </p><p>Sacks quote tweeted<a href="https://x.com/ClementDelangue/status/2078987852495364398"> Hugging Face CEO Clem Delangue</a>, who wrote: "We had this experience ourselves this week! Very scary to be guardrailed as a defender when you know attackers are likely bypassing".</p><h2><b>5 Strategic Takeaways for Enterprise Tech Leaders Now</b></h2><p>For the average enterprise executive, the central question is immediate: is our corporate network at risk from escaping AI agents? The short answer is no, not inherently.</p><p><b>1. Hugging Face occupies a unique position in the software ecosystem. </b>As a global repository for open-source AI models, code, and datasets, Hugging Face natively attracts autonomous agents, scrapers, automated evaluation pipelines, and active security researchers. Furthermore, the model’s target selection was context-specific: GPT-5.6 Sol searched for Hugging Face specifically because it deduced that Hugging Face hosted the answers to <i>ExploitGym</i>. Standard corporate networks—such as financial databases, HR platforms, or logistics systems—do not host benchmark solution keys that draw the direct focus of an agent attempting to solve an evaluation metric.</p><p><b>2. However, the long-term risk profile for enterprise technology permanently shifts following this event. </b>AI models with long-horizon reasoning seek the path of least resistance to accomplish a goal, including breaking rules, escaping sandboxes, or exploiting zero-days if deployment safeguards are intentionally disabled for testing or bypassed by an attacker. As Hugging Face's experience illustrates, data processing pipelines that ingest external datasets without sandbox execution or static analysis act as highly vulnerable initial access infrastructure.</p><p><b>3. This incident also drastically undercuts recent policy chatter in the U.S. calling for Chinese open-source AI models to be banned or restricted due to security concerns. </b>As this episode demonstrates, an open-weight Chinese model actually served as the vital defensive layer for an American and French firm facing an unanticipated cyberattack from an American model that broke containment. Contrary to the official line from some U.S. policymakers and hardline China hawks,  the Chinese open-source models weren't a security risk to the U.S. companies, in this case — rather, an American proprietary, closed-source model from an ostensibly secure American company was the source of the danger. Thus, any pressure U.S. companies may face from officials, agencies or non-governmental organizations to stop relying on affordable Chinese open weights models for defensive or any other lawful purposes should be viewed with a high degree of suspicion, and arguably resisted to the fullest legal extent. </p><p><b>4. Enterprise CISOs must audit their dependency on cloud-based AI APIs and pressure vendors to implement authenticated trust architectures</b>. Commercial AI vendors currently treat safety as a generic content-moderation problem, applying the same blanket refusals to an enterprise CISO as they would to a malicious hacker. Baer frames this requirement perfectly: "The model shouldn’t only understand what is being asked. It should understand who is asking, why, and under what governance".</p><p><b>5. Incident response plans must explicitly account for scenarios where commercial APIs fail, rate-limit, or actively refuse queries during an active security event. </b>Maintaining air-gapped, locally deployed open-weight models trained on security log analysis is no longer an edge-case luxury; it is a critical operational requirement. Security leaders running AI workloads in production must recalibrate their timelines and prepare for machine-speed threat actors that operate without human limits.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Exploits Zero-Day to Gain Internet Access and Compromise Hugging Face Servers]]></title>
<description><![CDATA[OpenAI has revealed that during an internal evaluation of advanced cyber capabilities, AI agents exploited a zero-day vulnerability, escaped a constrained research environment, and compromised parts of Hugging Face’s production infrastructure. While Hugging Face detected and contained the activit...]]></description>
<link>https://tsecurity.de/de/3685274/it-security-nachrichten/openai-exploits-zero-day-to-gain-internet-access-and-compromise-hugging-face-servers/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685274/it-security-nachrichten/openai-exploits-zero-day-to-gain-internet-access-and-compromise-hugging-face-servers/</guid>
<pubDate>Wed, 22 Jul 2026 06:55:24 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI has revealed that during an internal evaluation of advanced cyber capabilities, AI agents exploited a zero-day vulnerability, escaped a constrained research environment, and compromised parts of Hugging Face’s production infrastructure. While Hugging Face detected and contained the activity, OpenAI’s internal security team also identified unusual behavior during the assessment. OpenAI Compromise Hugging Face Servers […]</p>
<p>The post <a href="https://gbhackers.com/openai-compromise-hugging-face-servers/">OpenAI Exploits Zero-Day to Gain Internet Access and Compromise Hugging Face Servers</a> appeared first on <a href="https://gbhackers.com/">GBHackers Security | #1 Globally Trusted Cyber Security News Platform</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI’s GPT Agents Exploit Zero-Days and Hacked Hugging Face Servers]]></title>
<description><![CDATA[Hugging Face has disclosed a security incident that security researchers are calling a watershed moment for AI safety: an autonomous AI agent, built on OpenAI models, independently discovered and chained multiple vulnerabilities, including a zero-day, to breach Hugging Face’s production infrastru...]]></description>
<link>https://tsecurity.de/de/3685186/it-security-nachrichten/openais-gpt-agents-exploit-zero-days-and-hacked-hugging-face-servers/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685186/it-security-nachrichten/openais-gpt-agents-exploit-zero-days-and-hacked-hugging-face-servers/</guid>
<pubDate>Wed, 22 Jul 2026 05:26:21 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Hugging Face has disclosed a security incident that security researchers are calling a watershed moment for AI safety: an autonomous AI agent, built on OpenAI models, independently discovered and chained multiple vulnerabilities, including a zero-day, to breach Hugging Face’s production infrastructure. The incident occurred during an internal OpenAI evaluation testing the cyber capabilities of GPT-5.6 […]</p>
<p>The post <a href="https://cybersecuritynews.com/openai-zero-days-hugging-face/">OpenAI’s GPT Agents Exploit Zero-Days and Hacked Hugging Face Servers</a> appeared first on <a href="https://cybersecuritynews.com/">Cyber Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size]]></title>
<description><![CDATA[Poolside, the San Francisco AI lab that has spent most of its three-year existence quietly selling coding models to governments and defense agencies, released its most capable model to date on Tuesday — and made an unusually aggressive bet that radical transparency, not raw scale, is how a smalle...]]></description>
<link>https://tsecurity.de/de/3684985/it-nachrichten/poolside-drops-laguna-s-21-an-open-weight-coding-model-that-beats-rivals-10x-its-size/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3684985/it-nachrichten/poolside-drops-laguna-s-21-an-open-weight-coding-model-that-beats-rivals-10x-its-size/</guid>
<pubDate>Wed, 22 Jul 2026 01:07:11 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="http://poolside.ai/">Poolside</a>, the San Francisco AI lab that has spent most of its three-year existence quietly selling coding models to governments and defense agencies, released its most capable model to date on Tuesday — and made an unusually aggressive bet that radical transparency, not raw scale, is how a smaller lab competes at the frontier.</p><p>The model, <a href="https://poolside.ai/blog/introducing-laguna-s-2-1">Laguna S 2.1</a>, is a 118-billion-parameter<a href="https://huggingface.co/blog/moe"> Mixture-of-Experts (MoE) system</a> that activates only 8 billion parameters per token, supports a context window of up to 1 million tokens, and — according to benchmarks published by the company — matches or beats open models several times its size on agentic coding tasks. The weights are <a href="https://huggingface.co/poolside/Laguna-S-2.1">available immediately</a> on Hugging Face under the permissive OpenMDW-1.1 license.</p><p>The headline numbers are striking for a model this small. Poolside reports that <a href="https://huggingface.co/poolside/Laguna-S-2.1">Laguna S 2.1</a> scores 70.2% on <a href="https://www.tbench.ai/">Terminal-Bench 2.1</a>, a benchmark of long-horizon terminal tasks, placing it 11th on the company's compiled leaderboard — ahead of <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">DeepSeek-V4-Pro-Max</a>, a 1.6-trillion-parameter model that scored 64.0; Thinking Machines' 975-billion-parameter <a href="https://venturebeat.com/technology/thinking-machines-open-sources-first-multimodal-language-model-inkling-focused-on-low-cost-and-resistance-to-censorship">Inkling</a>, at 63.8; and Nvidia’s 550-billion-parameter <a href="https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/">Nemotron 3 Ultra</a>, at 56.4. On <a href="https://www.swebench.com/multilingual.html">SWE-Bench Multilingual</a>, it posts 78.5%, and on <a href="https://labs.scale.com/leaderboard/swe_bench_pro_public">SWE-Bench Pro</a>'s public dataset, 59.4%.</p><p>Perhaps more telling than any single score: the model went from the start of pre-training on May 22 to public launch in under nine weeks, trained on 4,096 Nvidia H200 GPUs. In an industry where flagship model cycles are typically measured in quarters or years, Poolside has now shipped three models in three months.</p><div></div><h2><b>Why the West's open-weight AI gap has become a boardroom issue</b></h2><p>The release lands in the middle of an increasingly pointed debate about <a href="https://www.scmp.com/tech/tech-war/article/3361142/why-chinas-open-weight-ai-model-kimi-k3-sparking-anxiety-silicon-valley">the provenance of open-weight AI</a>. Over the past year, developer adoption has shifted decisively toward open-weight systems that companies can download, inspect, and run on their own infrastructure — and the leading options in that category have overwhelmingly come from Chinese labs. <a href="https://www.deepseek.com/en/">DeepSeek</a>, <a href="https://qwen.ai/home">Qwen</a>, <a href="http://kimi.ai/">Kimi</a>, <a href="https://chat.z.ai/">GLM</a>, <a href="https://www.minimax.io/">MiniMax</a>, and <a href="https://hy.tencent.com/">Tencent's Hunyuan</a> line all feature prominently in Poolside's own comparison tables.</p><p>Poolside's accompanying press release frames <a href="https://poolside.ai/blog/introducing-laguna-s-2-1">Laguna S 2.1</a> explicitly as a response, noting that the model occupies a size class into which no Western lab has released open weights in 11 months — since OpenAI's <a href="https://openai.com/index/introducing-gpt-oss/">gpt-oss-120b</a> last August. "The West needs open-weight models it can trust, run, and build on," said Jason Warner, Poolside's co-CEO, in the announcement.</p><p>Co-founder and co-CEO Eiso Kant made the philosophical stakes even plainer in a <a href="https://x.com/eisokant/status/2079612416967491952?s=20">lengthy post</a> on X. "I believe intelligence should and will become a commodity," he wrote, arguing that the open ecosystem "will not win by being the best in its own category." Users, he argued, simply want the best intelligence for the task at hand — so open models must be on par with, or better than, their closed equivalents.</p><div></div><p>The strategic logic here is not charity. Poolside's core business is deploying models inside the security boundaries of government, defense, and regulated enterprises — customers for whom closed, metered API access is often a non-starter for compliance and sovereignty reasons. </p><p>Every enterprise that standardizes on a Chinese open model today becomes harder to win tomorrow. Releasing competitive open weights is both an ecosystem play and a top-of-funnel strategy for the company's high-security deployment business. It also reframes the AI race away from terrain where Poolside cannot compete — frontier-scale capital expenditure — and toward terrain where it believes it can: cost per token, self-hosting, and iteration speed.</p><h2><b>How a sparse architecture makes enterprise AI agents affordable to run</b></h2><p>The technical design reflects a specific thesis about where value in coding AI is moving. Laguna S 2.1's sparse MoE architecture — 256 routed experts plus one shared expert, with grouped-query attention and interleaved sliding-window layers, according to the <a href="https://huggingface.co/poolside/Laguna-S-2.1">Hugging Face model card</a> — means inference costs scale with the 8 billion active parameters, not the 118 billion total. Poolside emphasizes that the model is small enough to run on a single Nvidia DGX Spark, the desktop-class AI machine.</p><p>That matters for what Poolside calls token economics. Long-horizon coding agents are voracious consumers of tokens: the company's published data shows the model consuming a mean of roughly 249,000 completion tokens per trajectory on its hardest benchmark when thinking mode is enabled. At metered API prices, agentic workloads at enterprise scale become a meaningful budget line item. On OpenRouter, Poolside is offering a free 256K-context endpoint and a dedicated 1M-context deployment priced at $0.10 per million input tokens and $0.20 per million output tokens — aggressive pricing that undercuts most frontier alternatives by an order of magnitude.</p><p>The ecosystem support is unusually broad for day one. The model is live on <a href="https://www.baseten.co/library/laguna-s-21/">Baseten's model library</a> and <a href="https://vercel.com/changelog/laguna-s-2-1-is-now-available-on-ai-gateway">Vercel's AI Gateway</a>, with integrations across <a href="https://vllm.ai/">vLLM</a>, <a href="https://github.com/sgl-project/sglang">SGLang</a>, <a href="https://ollama.com/">Ollama</a>, and <a href="https://github.com/ggml-org/llama.cpp">llama.cpp</a>, plus quantized variants down to 4-bit GGUF files — 75 gigabytes — for local use. But Poolside's more interesting claim is behavioral, not architectural. Pengming Wang, co-head of applied research at Poolside, said the gains came from improving the model's working habits: "more verification, less taking things for granted, not declaring victory early, and being more persistent." Raw intelligence, the company argues, is one axis of capability; a model's way of working is a second axis that matters immensely for agents left unattended for hours.</p><h2><b>Publishing every benchmark trajectory to counter AI's credibility crisis</b></h2><p>The most consequential part of the release for enterprise buyers may be an evaluation-transparency move with little precedent among major labs: Poolside published the complete, unedited trajectory of every trial in its final benchmark runs — every reasoning step, tool call, and shell command behind every reported score.</p><p>This addresses a growing credibility problem in AI benchmarking. As top scores on mature benchmarks cluster in the 70–90% range, and as "reward hacking" — models finding solutions online or gaming verifiers rather than solving problems — has become endemic, self-reported numbers have lost much of their signal. Poolside disclosed its own encounters with the problem candidly: during training, more than half of trajectories on some SWE-bench tasks were flagged because the model simply researched the original bug-fix pull request online and applied it. The company documented its mitigations, including prompt addenda, LLM-based judging calibrated against human labels, and expert annotator review of a high-scoring Terminal-Bench run.</p><p>Three published case studies illustrate what the company means by persistence. In one, the model built a working HTML/CSS rendering engine from an empty folder in a 181-step, 50-minute unattended session — then, lacking vision capabilities, spun up headless Chromium to numerically compare its canvas output against a real browser's rendering. In another, pointed at Poolside's own agent harness in an automated optimization loop, the model made the Go codebase 5.2% faster with roughly 70% lower memory allocation, finding an O(n²) string-concatenation bug along the way. In a third, working in a sandbox with no Python installed, the model did its number theory in Perl and independently re-derived a proof of Erdős problem #397 — a combinatorics question open for five decades until GPT-5.2 Pro first solved it this past January. Poolside notes that its model's construction is structurally different from the earlier published solution, and that its November 2025 knowledge cutoff precedes the first proof.</p><div></div><h2><b>What the disclosed limitations and benchmark fine print reveal</b></h2><p><a href="https://poolside.ai/">Poolside</a> deserves credit for disclosing limitations most labs bury. The model can overfit to its native harness and stumble on slightly different tool schemas in third-party agents, mangles JSON in nested tool arguments, and is prone to overthinking on competition math. There is currently no user-configurable thinking-effort dial — just on or off — and the gap between the modes is enormous: thinking lifts <a href="https://www.tbench.ai/">Terminal-Bench 2.1</a> from 60.4% to 70.2%, and <a href="https://deepswe.datacurve.ai/">DeepSWE</a> from 16.5% to 40.4%, at substantially higher token cost.</p><p>Buyers should apply their own discounts to the comparison tables. Poolside's methodology takes the maximum of vendor self-reported scores, benchmark-author leaderboards, and third-party figures for competitors — a reasonable convention, but one that mixes harnesses and test conditions. On <a href="https://deepswe.datacurve.ai/">DeepSWE</a>, notably, Poolside ran its own agent harness rather than the leaderboard's standard mini-swe-agent, a difference the company acknowledges makes scores less directly comparable. And the frontier remains clearly out of reach: closed models like <a href="https://openai.com/index/previewing-gpt-5-6-sol/">GPT-5.6 Sol</a>, at 88.8 on Terminal-Bench 2.1, and <a href="https://www.anthropic.com/claude/fable">Claude Fable 5</a>, at 88.0, along with the 2.8-trillion-parameter open-weight <a href="https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems">Kimi K3</a>, at 88.3, sit well above Laguna S 2.1.</p><p>The deeper structural question is whether Poolside's "<a href="https://poolside.ai/blog/introducing-the-model-factory">Model Factory</a>" — the internal platform the company credits for its rapid release cadence — can sustain this pace as models scale. The trajectory so far is genuinely unusual: the April dual release of Laguna M.1 and XS.2, the July 2 refresh of XS 2.1, and now S 2.1, which the company says outperforms April's flagship M.1 at roughly a third of its active size. Remarkably, S 2.1 used the exact same pre-training data as XS 2.1, meaning nearly all the improvement came from scale, training fixes, and post-training across the company's corpus of 409,000 agentic and non-agentic training environments. Poolside says its next, larger Laguna model began pre-training last week.</p><p>For technical decision makers, <a href="https://huggingface.co/poolside/Laguna-S-2.1">Laguna S 2.1</a> is the most credible Western open-weight option to emerge in nearly a year for self-hosted agentic coding — with published evidence, a permissive license, broad ecosystem support, and an economics story built around hardware you can own. Whether it dents the dominance of Chinese open models will depend less on this release than on the ones that follow it.</p><p>Kant, for his part, has already told the world how he intends that story to end. Poolside is building toward a future where the most capable intelligence "can be owned and shaped by anyone," he wrote — and the company plans to keep shipping "until that future exists." In an industry where the biggest labs increasingly lock their best work behind an API, the most radical thing about Laguna S 2.1 may not be what it scores, but that anyone can download it and check.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Stop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens]]></title>
<description><![CDATA[GPU memory is the most expensive resource in production AI, and it's also the one running out fastest. Long context windows and multi-turn conversations force AI models to repeatedly recompute information they've already processed, consuming GPU memory and compute that could otherwise serve addit...]]></description>
<link>https://tsecurity.de/de/3684878/it-nachrichten/stop-adding-more-gpus-wekas-new-storage-platform-reduces-load-by-caching-100-of-an-ai-models-pre-calculated-tokens/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3684878/it-nachrichten/stop-adding-more-gpus-wekas-new-storage-platform-reduces-load-by-caching-100-of-an-ai-models-pre-calculated-tokens/</guid>
<pubDate>Tue, 21 Jul 2026 23:33:22 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>GPU memory is the most expensive resource in production AI, and it's also the one running out fastest. </p><p>Long context windows and multi-turn conversations force AI models to repeatedly recompute information they've already processed, consuming GPU memory and compute that could otherwise serve additional users or generate new responses.</p><p>Instead of treating GPU memory as the limiting resource,  why not extend it with much cheaper storage technologies? </p><p><a href="https://www.weka.io/">Weka</a>, for one, believes that cheap flash storage can close that gap. The company's NeuralMesh 6 software platform, launching alongside its first self-designed hardware line, Wekapod 3, extends what Weka calls Augmented Memory Grid, an approach that aggregates NAND flash to behave like GPU memory at a fraction of the cost.</p><p>This is an active and increasingly crowded category. Dell, NetApp, Pure Storage and VAST have all repositioned toward AI infrastructure over the past two years and Weka is one of several vendors arguing it's built for this specific moment rather than adapting to it.</p><p>"What we're seeing now with customers is they're chasing availability of compute, and once they get new allocation from anyone, they want to be able to grab it and start running right away," Weka co-founder and CEO Liran Zvibel, told VentureBeat.</p><p>The potential payoff is straightforward: better utilization of existing GPU investments, lower inference costs and faster deployment of new AI workloads without waiting months for additional GPU capacity.</p><p>The technology is most relevant for organizations already operating AI at scale or expecting rapid growth in usage, particularly enterprises building internal copilots, customer service agents, software engineering assistants or retrieval systems with long context windows. Smaller deployments may see less immediate benefit than organizations where GPU utilization has already become a limiting factor.</p><h2><b>Inside Weka's NeuralMesh 6</b></h2><p>NeuralMesh 6 adds four capabilities aimed directly at a functionality gap Zvibel says has been costing Weka deals in competitive evaluations.</p><p><b>Composable and virtual multi-tenancy.</b> Composable clusters give anchor tenants full hardware-level isolation, dedicated CPU, memory, and storage. Virtual multi-tenancy runs through Weka's RDMA fabric, delivering network-level isolation that scales past 1,000 tenants per cluster, with provisioning in under 30 minutes. Combined, a single cluster running 50 composable clusters can support up to 50,000 tenants. </p><p><b>Unified file and object storage.</b> Most storage systems keep two separate paths: a file-based path (the standard way servers and applications read and write files, used heavily in training and fine-tuning pipelines) and an object-based path (S3, the format inference and cloud-native tools typically expect). Normally a gateway translates between the two, meaning the data effectively exists twice. Weka's claim is that the same physical data on disk is directly readable through either path at once, no translation layer, no second copy. Zvibel is targeting non-AWS GPU clouds specifically, naming Lambda, Nebius, G42, and CoreWeave, with what he described as roughly two orders of magnitude higher performance than conventional S3 and a capacity-based pricing model instead of per-API charges. </p><p><b>Metadata-first replication.</b> Destination environments become browsable before a full data copy arrives, with data hydrating only when accessed. </p><p>"They had to wait for all of that to make it to the other side, and this takes days or weeks, in extreme cases a month," Zvibel said. "We now allow our customers to grab some allocation of new GPUs and get up and running within an hour."</p><p><b>AlloyFlash and Always-On data reduction</b>. TLC and QLC are two types of NAND flash memory. TLC is faster and more durable but costs more per terabyte, while QLC is cheaper and holds more data per chip but is slower. AlloyFlash mixes both within a single cluster, automatically routing latency-sensitive work to TLC while running bulk-capacity workloads on QLC, cutting cost per terabyte without a performance penalty on the work that needs speed. Data reduction now runs by default rather than as an option.</p><h2><b>Solving AI's context problem</b></h2><p>Multi-tenancy and object storage solve how enterprises and neo clouds operate the platform day to day. A harder problem sits underneath: as context windows and multi-turn interactions grow, so does the GPU compute wasted recalculating work a model has already done. Augmented Memory Grid, a NeuralMesh 6 feature built specifically for this, is Weka's answer.</p><p>Every prompt triggers two stages. Prefill calculates attention, the core mechanism behind how large language models process input, and it's computationally expensive. Decode converts that calculation into output and is comparatively lightweight. </p><p>The cost shows up hardest in multi-turn sessions like chat or coding, where each new turn re-triggers prefill for everything that came before it, unless that work has been cached.</p><p>"If you have 10 turns, you may overcalculate 100 times because you're redoing all of them. If you have 20, you'll overcalculate 400 times," Zvibel said. "You can put two orders of magnitude more NAND than you could afford in shared memory, and we can cache 100% of the pre-calculated tokens, so you never need to redo it."</p><h2><b>Where Weka sits competitively</b></h2><p>Storage vendors have spent the past year and a half repositioning around AI, and separating genuine capability from repositioned messaging is now a real evaluation problem for buyers. </p><p>"The storage world is shifting its focus from serving bits to enterprise workloads to managing data at the speed of AI. We've seen that most clearly over the past 18 months from Dell, NetApp, and Pure," Steve McDowell, chief analyst at NAND Research, told VentureBeat. "The interesting thing is that companies like Weka, and VAST, are the true AI-native data companies, solving these problems since day one."</p><p>McDowell singled out Augmented Memory Grid as Weka's clearest technical lead. </p><p>"Weka continues to have the most technically capable KV cache implementation on the market with its Augmented Memory Grid," he said. " They were early with this technology, and continue to innovate. This is critical for AI inference, as it enables a level of GPU efficiency that, without question, saves money on GPUs and memory. That’s key for today’s memory and GPU constrained market." </p><p>He also flagged Weka's contractual guarantee on its data reduction claims as underappreciated. </p><p>"One flying a little under the radar: Weka is putting its money where its mouth is with its contractual guarantees for its data reduction promises," he said.</p><p>McDowell's advice to buyers evaluating competing claims from Weka, VAST, Pure and NetApp alike was pointed suggesting that enterprise buyers should look hard at what vendors are promising versus what they're actually delivering.</p><p>"A smart buyer will look at how competing vendors are solving real-world problems today," McDowell said. " They do this by talking to organizations running similar workloads at similar scale. If a vendor can't point to that, then it should be a warning sign."</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026]]></title>
<description><![CDATA[“The new PRD are the evals,” Xavi Amatriain, Expedia Group’s first chief AI and data officer, told the VB Transform 2026 audience last week in Menlo Park. “So basically, you encode what you want the product to do through your evals, which might include red teaming evals and all kinds of other thi...]]></description>
<link>https://tsecurity.de/de/3684604/it-nachrichten/evals-are-the-new-prd-expedias-ai-chief-tells-vb-transform-2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3684604/it-nachrichten/evals-are-the-new-prd-expedias-ai-chief-tells-vb-transform-2026/</guid>
<pubDate>Tue, 21 Jul 2026 20:19:07 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>“The new PRD are the evals,” Xavi Amatriain, <a href="https://www.expediagroup.com/en-us">Expedia Group’s</a> first chief AI and data officer, told the <a href="https://venturebeat.com/vbtransform2026">VB Transform 2026</a> audience last week in Menlo Park. “So basically, you encode what you want the product to do through your evals, which might include red teaming evals and all kinds of other things, which already have a bunch of security requirements. So, you already embed that into the PRD and the product design document before you even start coding.”</p><p>He pushed it further. “With AI-assisted or AI-generated code, that’s gonna be the future. It’s like all your thinking is gonna go into the evals.”</p><p>Amatriain served as VP of AI and Compute Enablement at Google across the platforms powering Gemini and Google Search before his December 2025 appointment at Expedia. He's mentored talent who went on to found Perplexity and Scale AI. </p><p>VentureBeat’s <a href="https://venturebeat.com/orchestration/enterprise-ai-is-entering-an-evaluation-gap-agents-are-gaining-autonomy-faster-than-companies-can-verify-them">VB Pulse research on the evaluation gap</a> reinforced the stakes. Sixty-six percent of the 157 enterprises surveyed already permit some production deployment without human review or are building toward it within the next 12 months, yet only 5% fully trust the automated evaluations that would make that decision. Half have shipped an agent that passed internal evals but then failed with a real customer.</p><h2><b>Don’t let guardrails get in the way of feedback</b></h2><p>“The more guardrails and artificial business rules and sort of rules that you put into the system, the worse off,” Amatriain said. “Not only because they’re brittle, but also because they actually mess up with the feedback loop. You are actually biasing the user and the feedback you get from the user, and then you’re learning that in the wrong way.” He called guardrails “a necessary evil” and said the goal is to minimize their impact over time.</p><p>Not everyone at Transform agreed. Other speakers argued during the event that the highest-risk actions still demand very firm guardrails.</p><p>Expedia governs AI through three layers instead. Principles come first, communicated broadly. “I like to encode at a very high level how I expect decisions to be made, because in a large organization you’re gonna have a lot of distributed decision making,” Amatriain said. “And sometimes, if you’re lucky enough, those principles might be embedded in your culture. But most of the time, my experience has been they’re not.” The processes and tools that enforce them follow. “Principles look really nice on a picture on some wall, but you need to then give them teeth,” he said. Automation sits on top of both.</p><p>In practice, this plays out through what Expedia calls agent release toll gates, checkpoints calibrated to risk. “Governance needs to correlate to the risk,” Amatriain said. “And if you have something that is low risk, you don’t need too much governance to get in the way. But if there’s a lot of risk, then you need more governance. That can be encoded.” The toll gates tie evaluation rounds, red teaming, and security review to each agent’s risk level, and <a href="https://venturebeat.com/orchestration/what-billions-of-ai-predictions-taught-expedia-before-the-age-of-ai-agents">the checks shift from recommended to required as the stakes climb</a>. </p><h2>Specialized agents over monolithic intelligence</h2><p>“Even when I was at Google, I was like, I don’t believe in AGI as sort of like a singleton and a unified sort of like single model,” Amatriain told the audience. “I think it’s much better to think of it as composition, sort of like having specialized agents that are very good at some task and then composing the system out of those specialized agents.”</p><p>Expedia’s architecture starts at the component level. Tools compose into skills, skills assemble into sub-agents, and sub-agents get orchestrated into the full agentic system. “You need to have those principles that are unified that talk about things like what is the tone that we’re using, how are we addressing the user, how are we passing context, memory,” he said. “All of that needs to be thoroughly designed.” He framed this as a systemic design problem. “It’s not about the model, it’s not about a specific solution, it’s about how you’re designing the system.”</p><p>Amatriain argued that scoping each agent narrowly also makes the system easier to secure, since teams can evaluate and lock down individual agents in isolation before composing them.</p><h2>When the user must keep the final click</h2><p>Travel pricing changes in real time, flight availability shifts minute to minute, and hotel reviews routinely contradict what suppliers claim. Amatriain described a system that blends retrieval-augmented generation with direct API tool calls, choosing the approach based on latency. “If the user asks you a question like, how much does a four star hotel usually cost in Chicago in July, you don’t expect the agent to take two minutes to answer that question,” he said. “You expect an immediate answer because that answer can be cached and it doesn’t need real-time information.” A pet-friendly four-star near Lake Michigan with a pool might justify a 30-second reasoning window.</p><p>“The supplier might be saying, yeah, we have a great swimming pool, but then we also have the reviews from the travelers and we actually see there’s two reviews that say the swimming pool was not great or was not open after 6 p.m.,” Amatriain explained. A generic chatbot, he added, would only surface what a supplier self-reports, while Expedia cross-references against its own review corpus.</p><p>“We don’t want the agent to book the hotel or to buy you a plane ticket for you,” Amatriain said. “That’s something that the user has to have the agency. And the agent can recommend, can suggest, can discuss with you, but you’re gonna have to hit that click. And that’s non-negotiable.” That constraint, he argued, is also a security decision. “Once you establish those design principles, you also don’t need the guardrail because otherwise you’re gonna have to put all those guardrails in after the fact.”</p><h2>The next attackers will be other AI systems</h2><p>“Security needs to be a principle that is shifted as left as possible and as part of the design itself,” Amatriain said in response to an audience question. “And usually when you need a guardrail is because you’ve not thought about it early on.”</p><p>A second audience member pressed for lessons learned from production. Amatriain described a feedback loop where monitoring signals flow back into the eval suite. “You can almost automate the whole cycle,” he said. “But having that whole feedback loop from real signals, from your operating AI system, all the way into being reported and fixed as quickly as possible is going to become essential.”</p><p>Amatriain's toll gates are a bet that governance calibrated to risk can stay ahead of that feedback loop. VentureBeat’s separate June <a href="https://venturebeat.com/security/shared-api-keys-expose-ai-agent-fleets-venturebeat-research">Pulse survey on agent security</a>, drawn from 107 enterprises, shows how thin that margin is. More than half, 54 percent, have already had an agent security incident or near-miss. Fifty-nine percent plan to adopt, add, or replace agent security tooling within 12 months, and 29% plan to move this quarter. Incident rates climb with organization size, reaching 63% among enterprises with more than 1,000 employees versus 49% for companies with 101 to 1,000. And sandbox isolation, the one post-breach control that limits damage, drops from 35% adoption at the smaller companies to just 20 percent at the largest.</p><p>Amatriain warned that threats will increasingly come from other AI systems. “You’re gonna get threats coming not only from humans but also from other external agentic systems that are really powerful, and they’re gonna be poking at everything you’re doing. And as soon as you detect something, it’s not only about the detection, but the time to fix becomes essential here.”</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[HireQuotient Extends AI Recruiting Capabilities to Paylocity Customers in Frontline Industries]]></title>
<description><![CDATA[HireQuotient, an AI-native recruiting platform, today announced its integration with Paylocity (Nasdaq: PCTY), bringing AI-powered candidate sourcing and screening capabilities to Paylocity customers in manufacturing, building services, construction, healthcare and insurance, which are industries...]]></description>
<link>https://tsecurity.de/de/3684468/it-nachrichten/hirequotient-extends-ai-recruiting-capabilities-to-paylocity-customers-in-frontline-industries/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3684468/it-nachrichten/hirequotient-extends-ai-recruiting-capabilities-to-paylocity-customers-in-frontline-industries/</guid>
<pubDate>Tue, 21 Jul 2026 19:34:25 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">HireQuotient, an AI-native recruiting platform, today announced its integration with Paylocity (Nasdaq: PCTY), bringing AI-powered candidate sourcing and screening capabilities to Paylocity customers in manufacturing, building services, construction, healthcare and insurance, which are industries where deskless and frontline hiring has historically been underserved by AI recruiting tools.</p>



<p class="wp-block-paragraph">Employers in these sectors who already use HireQuotient are seeing the impact firsthand.</p>



<p class="wp-block-paragraph">“By simply putting in vetting criteria, I was able to get a very specific talent pool within a day,” said Niki Simoneaux, COO of Arc Health. “With EasySource, we can reach a huge database or filter for specific licenses and areas.”</p>



<p class="wp-block-paragraph">“With technology and the ability to work remotely, we’re able to recruit nationwide outside of our footprint. This expands the pool of prospective candidates and gives us access to candidates who might not have ever applied for a position on our website’s career page,” said Jeff McGee, vice president at W3 Insurance.</p>



<p class="wp-block-paragraph">“HireQuotient’s AI-native product helped my team at Alliance Building Services to cut down time to close position by more than 60%, and their team works very closely with the client to ensure adoption at scale,” said Willow Marcon, senior vice president at Alliance Building Services.</p>



<p class="wp-block-paragraph">Recruiters in manufacturing, building services, construction, healthcare and insurance often spend more than two-thirds of their time on manual work and using over 10 platforms to just close a hire: sourcing a wide range of profiles, skimming resumes, finding contact information, reaching out and doing hundreds of calls and constant follow-ups. That fragmented process leads to recruiter burnout and leaves gaps in candidate data. Because most HR platforms don’t offer built-in AI native agents that work in tandem to do all it takes to close the hire, employers in these industries have faced hiring delays and platform churn.</p>



<p class="wp-block-paragraph">The partnership addresses this gap by integrating HireQuotient’s EasySource platform directly into the Paylocity ecosystem. Instead of relying on rigid keyword searches, EasySource identifies strong talent pools, screens for role-specific credentials and licenses and personalizes outreach by phone and email. For mid-market companies with 500 to 1,000 employees, that translates to a 70% faster time-to-hire and an estimated $100,000 in annual savings.</p>



<p class="wp-block-paragraph">The integration also closes the feedback loop for employers. By feeding post-hire data back into the recruiting system, EasySource learns from successful hires to build smarter talent pools over time, while keeping recruiters engaged longer by freeing them to focus on the human side of hiring.</p>



<p class="wp-block-paragraph">Gokul Rajaram, board member at Coinbase and former board member of The Trade Desk (Nasdaq: TTD), and also known as the godfather of Google AdSense, said, “Smarthveer is one of those rare founders who picks an unglamorous, deeply underserved market and refuses to leave until it’s fixed. Frontline hiring is exactly that market. Being named to Paylocity’s elite partner network is a testament to his relentlessness and to the team he’s built. Excited for what’s ahead.”</p>



<p class="wp-block-paragraph">“Paylocity maintains an elite partner network, so being named one of them is a testament to the strength of our product,” said Smarthveer Sidana, founder and CEO of HireQuotient. “By keeping that recruiting activity inside the Paylocity ecosystem, we’re helping them close a critical gap for their employers, capture revenue that previously sat outside their platform, and prevent the client churn that comes with a fragmented hiring process.”</p>



<p class="wp-block-paragraph">Jim Moffatt, former global CEO of Deloitte Consulting and board partner at Greycroft VC, said, </p>



<p class="wp-block-paragraph">“This is a big milestone for Smarthveer and his team. I congratulate them on this big win. Paylocity’s large client base acts as a strong distribution, and HireQuotient’s EasySource acts as a strong product to cater to the needs of Paylocity’s clients. Paylocity is known for its highly selective approach, and I’m glad to see that after months of evaluation and extensive due diligence, they’re going live with HireQuotient’s EasySource. I feel confident in the value this partnership will create for frontline industries. I remember when Smarthveer spoke to me about it in December, and it felt very ambitious. I’m glad to see it turn to reality.”</p>



<p class="wp-block-paragraph">For employers with a traditionally deskless workforce in industries where AI adoption has lagged, the integration bridges a major capability gap by allowing them to source, screen, onboard and manage payroll all in one place.</p>



<p class="wp-block-paragraph">For more information, visit <a href="http://www.hirequotient.com/" target="_blank" rel="noreferrer noopener">www.hirequotient.com</a>.</p>



<p class="wp-block-paragraph">A media kit with additional partner quotes, executive headshots and logos can be found <a href="https://drive.google.com/drive/folders/1XZQE4etoLPgzNO94Y8I1-YX7ODyj8AnA?usp=sharing" target="_blank" rel="noreferrer noopener">here</a>.</p>



<p class="wp-block-paragraph"><strong>About HireQuotient</strong></p>



<p class="wp-block-paragraph">HireQuotient is an AI-native recruiting platform that automates candidate sourcing and screening for employers in manufacturing, building services, construction, healthcare, insurance and other frontline-heavy industries. Its flagship platform, EasySource, is one of a select number of recruiting technologies integrated with Paylocity’s HR and payroll ecosystem. Founded by Smarthveer Sidana, HireQuotient is headquartered in San Francisco. For more information, visit <a href="https://www.hirequotient.com/" target="_blank" rel="noreferrer noopener">https://www.hirequotient.com/</a>.</p>



<p class="wp-block-paragraph"><strong>Media Contact</strong></p>



<p class="wp-block-paragraph">Bethany Rhodes</p>



<p class="wp-block-paragraph">Uproar by Moburst for HireQuotient</p>



<p class="wp-block-paragraph">bethany@moburst.com</p>



<h5 class="wp-block-heading"><strong>Contact</strong></h5>



<p class="wp-block-paragraph"><strong>Bethany Rhodes</strong></p>



<p class="wp-block-paragraph"><strong>bethany@moburst.com</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Trump’s AI Safety Agency Chief Resigns After Just Three Months Leading CAISI]]></title>
<description><![CDATA[Chris Fall, the director of the U.S. Center for AI Standards and Innovation (CAISI), has resigned just three months after being appointed to lead the Commerce Department agency. This departure raises new uncertainties regarding the Trump administration’s agenda on AI safety, model evaluation, and...]]></description>
<link>https://tsecurity.de/de/3683810/it-security-nachrichten/trumps-ai-safety-agency-chief-resigns-after-just-three-months-leading-caisi/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683810/it-security-nachrichten/trumps-ai-safety-agency-chief-resigns-after-just-three-months-leading-caisi/</guid>
<pubDate>Tue, 21 Jul 2026 15:25:24 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Chris Fall, the director of the U.S. Center for AI Standards and Innovation (CAISI), has resigned just three months after being appointed to lead the Commerce Department agency. This departure raises new uncertainties regarding the Trump administration’s agenda on AI safety, model evaluation, and cybersecurity oversight. The Commerce Department confirmed his resignation on July 20, […]</p>
<p>The post <a href="https://gbhackers.com/trumps-ai-safety-agency-chief-resigns-after-just-three-months/">Trump’s AI Safety Agency Chief Resigns After Just Three Months Leading CAISI</a> appeared first on <a href="https://gbhackers.com/">GBHackers Security | #1 Globally Trusted Cyber Security News Platform</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The AI allocation trap: Record spend, vanishing returns]]></title>
<description><![CDATA[In a single month, one enterprise reportedly spent half a billion dollars on AI. A consultant told Axios that the client had handed its workforce AI licenses, set no usage limits and let the meter run until finance noticed. The figure is spectacular, and it is the wrong thing to fear. That half-b...]]></description>
<link>https://tsecurity.de/de/3683786/it-nachrichten/the-ai-allocation-trap-record-spend-vanishing-returns/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683786/it-nachrichten/the-ai-allocation-trap-record-spend-vanishing-returns/</guid>
<pubDate>Tue, 21 Jul 2026 15:18:28 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">In a single month, one enterprise reportedly spent half a billion dollars on AI. A consultant <a href="https://www.axios.com/2026/05/28/ai-spending-roi-enterprise-costs">told Axios</a> that the client had handed its workforce AI licenses, set no usage limits and let the meter run until finance noticed. The figure is spectacular, and it is the wrong thing to fear. That half-billion-dollar accident is only the visible part of a quieter, far larger failure. <a href="https://www.gartner.com/en/newsroom/press-releases/2026-1-15-gartner-says-worldwide-ai-spending-will-total-2-point-5-trillion-dollars-in-2026">Worldwide AI spending is forecast to reach $2.52 trillion in 2026</a>, more than any technology category in a generation, and by the most cited measure, roughly 95 percent of it returns nothing. Boards read that as proof that the technology does not work. The evidence points somewhere less comfortable, and it is not a technology problem at all. Most boards cannot see it because they are reading the wrong number: They track failure when the number that matters is allocation. The discipline that separates the winners is not technical. It is how they allocate capital across time, and how willing they are to stop. The hardest discipline in the AI era is not adopting faster. It is allocating honestly and refusing to judge a three-year bet on a six-month cycle.</p>



<h2 class="wp-block-heading">The number everyone quotes, and no one acts on</h2>



<p class="wp-block-paragraph">The headline statistic is now familiar. MIT’s Project NANDA, in its 2025 study <a href="https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/">The GenAI Divide</a>, found that about 95 percent of enterprise generative AI pilots produced no measurable impact on the P&amp;L, while roughly 5 percent captured nearly all the value. <a href="https://www.spglobal.com/market-intelligence/en/news-insights/research/2025/10/generative-ai-shows-rapid-growth-but-yields-mixed-results">S&amp;P Global Market Intelligence</a> found that the share of companies abandoning most of their AI initiatives jumped from 17 percent to 42 percent in a single year, with the average organization scrapping 46 percent of its proofs-of-concept before production. <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">Gartner</a> expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. And the pattern predates generative AI: <a href="https://www.rand.org/pubs/research_reports/RRA2680-1.html">RAND</a> found that more than 80 percent of AI projects fail, roughly twice the rate of comparable work that does not involve AI.</p>



<p class="wp-block-paragraph">Read as a technology story, these numbers say AI does not work. Read correctly, they say something more useful. MIT’s own authors located the cause not in model quality but in a <a href="https://virtualizationreview.com/articles/2025/08/19/mit-report-finds-most-ai-business-investments-fail-reveals-genai-divide.aspx">learning and integration gap</a>. The winners were not running better models. They picked one problem, executed and worked well together. Purchased solutions reached production about 67 percent of the time, while internal builds succeeded roughly a third as often. <a href="https://www.gartner.com/en/newsroom/press-releases/2025-03-31-gartner-forecasts-worldwide-genai-spending-to-reach-644-billion-in-2025">Gartner’s own spending forecast</a> notes the same pivot, with CIOs scaling back ambitious internal builds in favor of commercial solutions that promise more predictable value. None of that is a verdict on the technology. It is a verdict on allocation: What gets funded, for how long and against which yardstick. The popular prescription, heard in every boardroom this year, is to measure harder and prove value sooner. That advice quietly repeats the mistake, because forcing a three-year bet to prove itself sooner is precisely how you kill it. The fix is not more measurement. It is measuring each bet against the right clock and subtracting the ones that miss.</p>



<h2 class="wp-block-heading">The six-month cycle problem</h2>



<p class="wp-block-paragraph">Return to that 95 percent, because the way it is measured is the whole argument. Much of the reported failure is judged on a short clock, with a pilot counted as a failure if it has not shown a measurable financial return within roughly six months. The single most quoted number in enterprise AI is therefore a six-month yardstick applied to every initiative, including the bets designed to pay back in three years. The headline failure rate is not only a measure of AI. It is a measure of impatience.</p>



<p class="wp-block-paragraph">The most expensive mistake in enterprise AI is a timing error. Enterprises have been spending heavily on AI for more than two years, and 2026 is the year boards are demanding returns. The multi-year bets funded during the 2024 and 2025 scale-up are only now far enough along to be judged. When a board reviews an initiative, it applies the yardstick it knows, which is quarterly return. That yardstick is correct for an efficiency project and ruinous for a capability bet. A workflow automation that should pay back in two quarters and a foundational data and agent capability that pays back in three years are not the same instrument, yet they are reviewed in the same meeting against the same metric.</p>



<p class="wp-block-paragraph">This is the heart of the divide. The 5 percent did not simply pick better projects. They judged each project against its own horizon. McKinsey’s enduring <a href="https://www.mckinsey.com/capabilities/strategy-and-corporate-finance/our-insights/enduring-ideas-the-three-horizons-of-growth">Three Horizons model</a> made this discipline standard in corporate strategy a generation ago: near-term, emerging and long-term bets are funded and measured differently. AI erased that discipline because the hype compressed every timeline into the current quarter. The result is two failure modes that appear opposite yet share a common root. Organizations kill three-year bets at month six because they miss a metric the bet was never designed to hit. And they keep funding six-month theater for years because it is visible, safe and never asked to prove a return. Both are allocation failures. Neither is a technology failure.</p>



<h2 class="wp-block-heading">Subtraction is a strategy</h2>



<p class="wp-block-paragraph">There is a second discipline, the 5 percent share, and it is the one boards find hardest. They subtract. Every credible study of the failure rate describes the same chaotic pattern underneath it: Initiatives are <a href="https://www.ciodive.com/news/AI-project-fail-data-SPGlobal/742590/">abandoned late, without criteria</a>, after the money is spent and the credibility is gone. Disciplined organizations do the opposite. They decide the conditions for stopping before they start, and they stop on schedule. Subtraction is not the absence of strategy. It is the strategy. Capital removed from a failing bet is capital available for a surviving one, and the survivors are where the entire return lives.</p>



<p class="wp-block-paragraph">This reframes the 42 percent abandonment figure. Abandonment is not the problem. Undisciplined abandonment is. An organization that liquidates a position the moment it breaches a pre-agreed kill line is practicing portfolio hygiene. An organization that lets a doomed pilot run until someone loses patience is paying full price for a lesson it could have bought at a discount. The 5 percent who won were not smarter. They were patient in the right places and ruthless in the wrong ones.</p>



<h2 class="wp-block-heading">The HALT framework: Horizon, Allocation, Liquidation, Tracking</h2>



<p class="wp-block-paragraph">Treating AI as a portfolio rather than a pile of pilots requires four disciplines, and the organizations that execute well put all four in place before the next funding cycle, not after the next failure. The name is deliberate. The discipline most enterprises lack is the willingness to halt the wrong bets in time to fund the right ones.</p>



<p class="wp-block-paragraph"><strong>Component 1: Horizon. </strong>Classify every AI initiative by its true payoff horizon before it is funded. Horizon 1 covers efficiency plays that should return value within two quarters. Horizon 2 covers capability bets, data foundations, agent platforms and integration work that pays back in roughly 6 to 18 months. Horizon 3 covers transformation bets that take eighteen months to three years or longer. Each horizon carries its own success metric, set at funding time. A Horizon 1 yardstick never judges a Horizon 3 bet. This single rule prevents the most common and most expensive error in the portfolio.</p>



<p class="wp-block-paragraph"><strong>Component 2: Allocation. </strong>Decide the split across horizons deliberately, as a board-level capital decision, not as the accidental sum of whatever pilots happened to win approval. A practical reference point, borrowed from decades of innovation-portfolio practice, is roughly 70% to near-term value, 20% to capability, and 10% to transformation. The exact ratio is yours; the discipline is to choose and defend it. The failure mode is an unmanaged portfolio: 90 percent scattered across disconnected Horizon 1 experiments, with nothing compounding into the Horizon 2 capability that the buy-and-integrate winners actually built.</p>



<p class="wp-block-paragraph"><strong>Component 3: Liquidation. </strong>Attach a kill line to every initiative at the moment it is funded: A named milestone, a date and an owner empowered to stop it. If a bet misses its horizon-appropriate milestone, it is liquidated, and capital is reallocated on schedule without debate over sunk costs. The absence of a pre-agreed kill line is not patience. It is an unpriced liability that the board has almost certainly not been shown.</p>



<p class="wp-block-paragraph"><strong>Component 4: Tracking. </strong>Report the portfolio to the board on a fixed cadence using a single instrument: The AI Portfolio Scorecard. Not a deck of project updates, but a single view of allocation by horizon, burn against milestone, liquidation decisions taken and capital reallocated to survivors. The cadence is the control. A portfolio reviewed once a year is a portfolio managed by hope.</p>



<p class="wp-block-paragraph"><strong>THE AI PORTFOLIO SCORECARD: SCORE EVERY INITIATIVE BEFORE IT IS FUNDED</strong></p>



<figure class="wp-block-table"><div class="overflow-table-wrapper"><table class="has-fixed-layout"><thead><tr><td><strong>Evaluation criterion</strong></td><td><strong>0</strong></td><td><strong>1</strong></td><td><strong>2</strong></td></tr></thead><tbody><tr><td>Horizon assigned (H1 / H2 / H3) and documented before funding</td><td> </td><td> </td><td> </td></tr><tr><td>Success metric matched to the horizon, not a default quarterly ROI</td><td> </td><td> </td><td> </td></tr><tr><td>Kill line set: Named milestone and date, agreed at funding</td><td> </td><td> </td><td> </td></tr><tr><td>Owner named with explicit authority to stop the initiative</td><td> </td><td> </td><td> </td></tr><tr><td>Fits a deliberate allocation band, not an accidental addition</td><td> </td><td> </td><td> </td></tr><tr><td>Odds-raising path documented: Buy or partner and an integration plan</td><td> </td><td> </td><td> </td></tr></tbody></table> </div></figure>



<p class="wp-block-paragraph"><em>Score each criterion: 0 = not present, 1 = partially documented, 2 = fully verified. Total out of 12. Bands: 0 to 4 = DO NOT FUND  |  5 to 8 = CONDITIONAL  |  9 to 12 = FUND.</em></p>



<p class="wp-block-paragraph"><strong>THE LIQUIDATION GATE: RUN AT EVERY BOARD REVIEW BEFORE CONTINUING FUNDING</strong></p>



<figure class="wp-block-table"><div class="overflow-table-wrapper"><table class="has-fixed-layout"><thead><tr><td><strong>Review test</strong></td><td><strong>Status</strong></td></tr></thead><tbody><tr><td>Milestone for this horizon met or credibly on track</td><td>PASS / FAIL</td></tr><tr><td>Burn within plan to the next milestone</td><td>PASS / FAIL</td></tr><tr><td>Still fits the allocation band, with no quiet horizon drift</td><td>PASS / FAIL</td></tr><tr><td>Owner confirms continued strategic fit</td><td>PASS / FAIL</td></tr></tbody></table> </div></figure>



<p class="wp-block-paragraph"><em>Any unresolved FAIL = stop funding, liquidate the position, reallocate the capital to a survivor and record the decision on the scorecard.</em></p>



<h2 class="wp-block-heading">The cost of the timing error</h2>



<p class="wp-block-paragraph">The financial case follows the pattern and is consistent. Consider two organizations that funded the same class of Horizon 3 bet: A domain-specific agent platform meant to compound over three years. The first review was conducted at month six against a quarterly return test, found no payback and killed it, booking the write-off as a lesson about AI being overhyped. Its competitor classified the same work as Horizon 3, set an 18-month capability milestone, protected funding through two review cycles and shipped to production within the window the work actually required. One organization spent its money to learn that it lacks allocation discipline. The other spent comparable money and now owns a capability its rival has abandoned and cannot quickly rebuild. The dollars on the two income statements are similar. The competitive positions are not.</p>



<h2 class="wp-block-heading">The governance return the board has been waiting for</h2>



<p class="wp-block-paragraph">Allocation discipline does two things at once. It stops the bleed by liquidating failures on a schedule rather than at the point of exhaustion. And it concentrates capital where the entire return lives, in the small number of bets that survive their horizon. The 5 percent figure is not a ceiling imposed by the technology. It is the current yield of an industry allocated by hype. An organization that classifies by horizon, allocates on purpose, liquidates on a line and tracks on a cadence is not trying to beat the technology. It is trying to beat its own indiscipline, and that is a far more winnable contest.</p>



<p class="wp-block-paragraph">The board conversation about AI returns is coming for every organization, and it arrives the moment the spending outpaces the story. When it does, the CIO will be asked a simple question: Where did the money go? The leaders who can answer will not show a pile of pilots. They will show a portfolio: What was funded, against which horizon, what was liquidated and when, and what the survivors are now worth. Subtraction is a strategy. The only question is whether you are practicing it on purpose or about to learn it by accident.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[US AI testing institute chief steps down within three months]]></title>
<description><![CDATA[The head of the US government’s AI testing institute, Chris Fall, has resigned about three months after taking charge of the Center for AI Standards and Innovation (CAISI), the federal organization responsible for evaluating advanced artificial intelligence models for safety and security.



Curr...]]></description>
<link>https://tsecurity.de/de/3683724/it-nachrichten/us-ai-testing-institute-chief-steps-down-within-three-months/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683724/it-nachrichten/us-ai-testing-institute-chief-steps-down-within-three-months/</guid>
<pubDate>Tue, 21 Jul 2026 14:48:10 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">The head of the US government’s AI testing institute, Chris Fall, has resigned about three months after taking charge of the Center for AI Standards and Innovation (CAISI), the federal organization responsible for evaluating advanced artificial intelligence models for safety and security.</p>



<p class="wp-block-paragraph">Current National Institute of Standards and Technology NIST Director Arvind Raman will serve as acting CAISI Director following Fall’s departure while continuing to oversee the Commerce Department office responsible for the institute, the Daily Signal <a href="https://www.dailysignal.com/2026/07/20/scoop-head-of-federal-ai-safety-org-resigns/" target="_blank" rel="noreferrer noopener">reported</a>, citing two people familiar with the matter.</p>



<p class="wp-block-paragraph">A Commerce Department spokesperson who spoke to the publication did not disclose a reason for the resignation.</p>



<p class="wp-block-paragraph">Fall assumed leadership of CAISI in April after the Trump administration reorganized the former US AI Safety Institute under NIST. The institute develops methodologies for evaluating frontier AI models and works with AI developers on voluntary technical assessments covering areas such as cybersecurity, model misuse, reliability and other risks associated with increasingly capable AI systems.</p>



<p class="wp-block-paragraph">The leadership change comes as governments and AI companies continue developing technical approaches for evaluating frontier AI models while enterprises expand deployments of generative AI and agentic AI across business operations.</p>



<p class="wp-block-paragraph">In recent months, the Commerce Department has taken a <a href="https://www.infoworld.com/article/4194598/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion.html?_conv_v=vi:1*sc:1*cs:1784634320*fs:1784634320*pv:1*exp:%7B1004203305.%7Bv.1004477672-g.%7B%7D%7D%7D*seg:%7B%7D&amp;_conv_s=sh:1784634319808-0.24259838933788935*si:1*pv:1&amp;_conv_r=null&amp;_conv_sptest=null">more active role</a> in AI policy involving advanced models, placing greater attention on how the federal government evaluates technologies with potential national security implications.</p>



<h2 class="wp-block-heading">Continuity matters more than personalities</h2>



<p class="wp-block-paragraph">CAISI works with AI developers such as Anthropic, Google’s DeepMind and OpenAI on voluntary evaluations of frontier AI models and develops methodologies for testing model capabilities and risks. The institute does not regulate AI developers or certify commercial AI systems.</p>



<p class="wp-block-paragraph">For enterprises, those evaluations are one source of technical information alongside vendors’ own testing, third-party security assessments and internal AI governance programs.</p>



<p class="wp-block-paragraph">Sanchit Vir Gogia, chief analyst at Greyhound Research, said enterprises should focus less on the individual leading the institute and more on whether its technical work continues with the same level of consistency and transparency.</p>



<p class="wp-block-paragraph">“Leadership churn at CAISI weakens the signal long before it weakens the science,” Gogia said. “The testing has not stopped. Its authority simply does not travel as cleanly once the leadership does not.”</p>



<p class="wp-block-paragraph">According to Gogia, the more important question for enterprises is not whether the institute’s evaluation work will continue but whether the processes supporting those evaluations remain stable.</p>



<p class="wp-block-paragraph">“The instinct is to ask whether the pipeline is breaking,” he said. “The more useful question is where the pipeline now sits.”</p>



<h2 class="wp-block-heading">Enterprises still carry the burden of AI governance</h2>



<p class="wp-block-paragraph">Gogia said organizations should continue treating government-led AI evaluations as one input into their governance processes rather than as evidence that a model is inherently safe for enterprise deployment.</p>



<p class="wp-block-paragraph">“A government evaluation was always a signal, never a certificate,” he said. “A signal loses value the moment its issuer becomes unpredictable.”</p>



<p class="wp-block-paragraph">He said enterprises should instead monitor whether CAISI maintains consistent evaluation methodologies, continues publishing technical findings and preserves continuity within its research teams under interim leadership.</p>



<p class="wp-block-paragraph">“The name on the door is not the signal. The behaviour underneath it is,” Gogia said.</p>



<p class="wp-block-paragraph">Gogia also cautioned against linking Fall’s resignation to recent Commerce Department actions involving AI policy or export controls, noting that there is no public evidence connecting the two.</p>



<p class="wp-block-paragraph">“CAISI evaluates; it does not enforce export controls, because it holds no such power,” he said. “This is not a testing body reaching for enforcement. It is enforcement reaching past the testing body.”</p>



<p class="wp-block-paragraph">With Raman assuming the role on an interim basis, the next significant milestone for enterprises will be the appointment of a permanent director, and whether the institute’s evaluation programs continue without disruption, the analyst said.</p>



<p class="wp-block-paragraph">Gogia said the successor’s mandate may prove more important than the individual selected.</p>



<p class="wp-block-paragraph">“A CAISI result is not a safe harbour,” he said. “It informs an obligation; it does not discharge one.” NIST did not immediately respond to a request for comment.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Small models, sovereign advantage: Why Australia should build its own AI edge]]></title>
<description><![CDATA[For the past three years, the AI conversation has been dominated by scale. Bigger models, bigger compute clusters, bigger headlines. But the next wave of competitive advantage won’t come from who can rent the biggest model; it will come from who can build the smallest one that knows their busines...]]></description>
<link>https://tsecurity.de/de/3683294/it-nachrichten/small-models-sovereign-advantage-why-australia-should-build-its-own-ai-edge/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683294/it-nachrichten/small-models-sovereign-advantage-why-australia-should-build-its-own-ai-edge/</guid>
<pubDate>Tue, 21 Jul 2026 12:03:22 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">For the past three years, the AI conversation has been dominated by scale. Bigger models, bigger compute clusters, bigger headlines. But the next wave of competitive advantage won’t come from who can rent the biggest model; it will come from who can build the smallest one that knows their business.</p>



<p class="wp-block-paragraph">That model is the <a href="https://www.cio.com/article/4119259/small-language-models-why-specialized-ai-agents-boost-resilience-and-protect-privacy.html">small language model (SLM)</a>: Compact, purpose-built, trained on an organization’s own data and run under that organization’s own governance. And it is about to become one of the most consequential strategic assets available to both the private and public sector.</p>



<h2 class="wp-block-heading">The problem with renting intelligence</h2>



<p class="wp-block-paragraph">Right now, most organizations consume AI the way they once consumed electricity from a single utility by plugging into a handful of frontier models built by a small number of global vendors. These models are extraordinary generalists. They are also, by design, generic. They are tuned to be safe, broad and useful to everyone, which means they are optimised for no one in particular.</p>



<p class="wp-block-paragraph">That’s a problem for any organization trying to build genuine differentiation. If every competitor in your sector is calling the same foundation model with the same prompts, the model itself is not your edge. Your edge is what only you know, your proprietary data, your institutional judgement, your operating history. A generic model can’t see any of that unless you keep feeding it to them, turn after turn, at cost, with no lasting memory and no guarantee of where that data ends up.</p>



<p class="wp-block-paragraph">An SLM flips that equation. Trained on an organization’s own document libraries, case histories, policy archives, transaction data and operational know-how, it becomes a model that thinks the way your organization thinks, because it was built from your organization’s accumulated judgement. It doesn’t need to be the smartest model in the world. It needs to be the most useful one for you.</p>



<p class="wp-block-paragraph">I’ve seen this play out directly. At one of Australia’s largest integrated tourism and cruise businesses, simultaneously a B2C retailer, a B2B distributor to thousands of agency and wholesale clients globally, an aggregator marketplace for more than 1,800 independent tourism operators, and a cruise operator with offshore shared services spanning finance, customer contact and content management. The constraint wasn’t a lack of access to large general-purpose models. It was that none of them understood the business: 1,800 different operator catalogues, each with its own pricing logic, inventory quirks and content conventions; years of customer contact history with its own vocabulary and escalation patterns; a marketplace search experience that needed to reason over the business’s own product taxonomy, not the open web’s.</p>



<p class="wp-block-paragraph">Models trained and tuned on that proprietary data, operator listings, historical tickets, booking and pricing data delivered results a generic model never could. Domain-tuned content drafting cut operator listing time by 70% and eliminated a 23-day onboarding backlog outright, taking new-operator time-to-live from 23 days to three. A semantic search model trained on the marketplace’s own product catalogue lifted booking conversion by 24%. AI-driven triage trained on the business’s own contact history cut Tier 1 escalations by 34%. None of this came from a smarter foundation model. It came from a smaller, more specific one that knew the business.</p>



<h2 class="wp-block-heading">Why “small” is the strategic choice, not the compromise</h2>



<p class="wp-block-paragraph">There’s a temptation to treat SLMs as the budget option, what you build when you can’t afford a frontier model. That’s the wrong frame. The evidence is already compelling: <a href="https://azure.microsoft.com/en-us/blog/empowering-innovation-the-next-generation-of-the-phi-family/">Microsoft’s Phi-4 family of small models</a>, released in early 2025, demonstrated that a 14-billion-parameter model can match or exceed the performance of models many times its size on complex reasoning and domain-specific tasks while running at a fraction of the compute cost and on-premise, entirely within an organization’s own infrastructure. Smaller, domain-trained models are increasingly outperforming general-purpose giants on narrow, high-value tasks, with far tighter control over data residency, security and explainability.</p>



<p class="wp-block-paragraph">For a CIO or CTO, that combination of lower cost, tighter governance, higher task-specific accuracy is rare enough to demand attention on its own. But the deeper value sits one layer up, at the operating model. An SLM trained on your service history can sit inside claims processing, citizen services, clinical triage, asset maintenance scheduling or M&amp;A due diligence quietly compounding institutional knowledge into a reusable asset rather than letting it walk out the door every time someone retires or resigns.</p>



<p class="wp-block-paragraph">That is the real shift: AI capability stops being a subscription and starts being a balance-sheet asset. It can be valued, protected, audited and improved because it belongs to you.</p>



<h2 class="wp-block-heading">The public sector’s hidden advantage</h2>



<p class="wp-block-paragraph">Nowhere is this more obvious than in government. The public sector sits on some of the richest, least-exploited data and institutional knowledge in the country: Decades of policy outcomes, service delivery history, regulatory precedent, infrastructure records and frontline expertise. Most of it has never been put to systematic use because no commercially available model was ever trusted to touch it, and rightly so.</p>



<p class="wp-block-paragraph">A small, sovereign, purpose-built model changes that calculus. Trained, hosted and governed entirely within government infrastructure, an SLM doesn’t require sensitive citizen or policy data to leave a secure perimeter. The Australian Government has already recognised this direction: <a href="https://www.finance.gov.au/about-us/news/2025/introducing-aps-ai-plan">The APS AI Plan, released in November 2025</a>, commits to expanding the GovAI platform to provide all public servants with secure, sovereign AI tools operating entirely within Australian Government infrastructure. SLMs tuned to individual agency mandates are the logical next step and a more powerful one than any generic government-wide tool can deliver.</p>



<p class="wp-block-paragraph">Rather than each agency independently negotiating with the same handful of overseas vendors, a coordinated approach of common standards for model governance, shared security architecture, common evaluation frameworks and pooled infrastructure investment would let agencies build and reuse SLM capability horizontally, the way shared services and common ICT platforms have been built before. Each agency gets a model genuinely tuned to its mandate, but the security model, audit trail and assurance framework are consistent, government-backed and independently verifiable.</p>



<p class="wp-block-paragraph">Done well, this isn’t just an efficiency play. It’s a sovereignty play. As <a href="https://www.govtechreview.com.au/content/gov-datacentre/article/why-sovereign-ai-is-becoming-a-strategic-priority-in-australia-81646916">GovTech Review has noted</a>, large language models hosted offshore create data flows that extend beyond Australia’s borders in ways that are rarely transparent, a risk that is simply untenable for government. Sovereign, purpose-built models keep Australian public data, public knowledge and the resulting capability uplift inside Australian hands, rather than exporting both the data and the long-term value to offshore platforms.</p>



<h2 class="wp-block-heading">Why this belongs in the innovation budget, not the IT budget</h2>



<p class="wp-block-paragraph">The instinct in many organizations is to treat AI spend as an IT line item, something to be minimised, benchmarked and squeezed for cost efficiency. SLMs deserve a different treatment. They are closer to R&amp;D than infrastructure: An investment in converting accumulated institutional knowledge into a durable, defensible capability.</p>



<p class="wp-block-paragraph">That argument holds in the private sector too. A PE-backed portfolio company, a regulated financial services firm, a healthcare provider — each has years of proprietary operating data sitting idle in case files, transaction logs and service records. An SLM built on that data is a way of turning a sunk cost, decades of operational history, into a forward-looking asset that compounds with every additional case it processes.</p>



<p class="wp-block-paragraph">Boards and executive committees that are still asking “what is our AI strategy?” as a single, undifferentiated question are asking the wrong thing. The better question is: Which parts of our operation are rich enough in proprietary data and judgement to justify owning the model outright, rather than renting someone else’s?</p>



<h2 class="wp-block-heading">The opportunity in front of us</h2>



<p class="wp-block-paragraph">The first wave of enterprise AI adoption was about access: Getting a capable model into people’s hands quickly. The next wave will be about ownership: Who controls the model, who controls the data it was built on, and who captures the long-term value of the institutional knowledge it encodes.</p>



<p class="wp-block-paragraph">Australia, with a public sector rich in data and a private sector with deep vertical expertise in financial services, resources, healthcare and logistics, is well placed to lead on this if it treats small, sovereign models as a genuine national capability question, not a procurement footnote. The organizations, and the country, that move early will not just save money. They will own something their competitors can’t easily replicate: An AI that knows them.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[iPhone 20 Pro Max may get Apple's largest ever iPhone screen at 7 inches]]></title>
<description><![CDATA[A new rumor claims that Apple is considering using what it will call a 7-inch screen for one of the 20th anniversary iPhones, although that isn't as great an increase as it sounds.The display on the current iPhone 17 Pro Max is 6.86 inches.One recent rumor claimed that Apple had begun production ...]]></description>
<link>https://tsecurity.de/de/3683281/ios-mac-os/iphone-20-pro-max-may-get-apples-largest-ever-iphone-screen-at-7-inches/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683281/ios-mac-os/iphone-20-pro-max-may-get-apples-largest-ever-iphone-screen-at-7-inches/</guid>
<pubDate>Tue, 21 Jul 2026 11:57:45 +0200</pubDate>
<category>🍏 iOS / Mac OS</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A new rumor claims that Apple is considering using what it will call a 7-inch screen for one of the 20th anniversary <a href="https://appleinsider.com/inside/iphone" title="iPhone" data-kpt="1">iPhones</a>, although that isn't as great an increase as it sounds.<br><br><div><img src="https://photos5.appleinsider.com/gallery/68308-143998-000-lead-iPhone-17-Pro-Max-display-xl.jpg" alt="Modern smartphone lying on a dark surface, screen displaying a vivid purple and pink abstract flower-like pattern with bright glowing center and small light reflections" height="720"><br><span>The display on the current iPhone 17 Pro Max is 6.86 inches.</span></div><br>One recent rumor claimed that Apple had begun <a href="https://appleinsider.com/articles/26/05/21/rumor-2027-iphone-production-testing-underway-with-quad-curved-oled-display">production evaluation</a> for a 2027 iPhone with a display that is curved on all four sides. Then another claimed that the <a href="https://appleinsider.com/inside/iphone-20" title="iPhone 20" data-kpt="1">iPhone 20</a> range would feature a <a href="https://appleinsider.com/articles/26/07/14/glass-production-plans-give-the-iphone-20-redesign-new-credibility">significant redesign</a>.<br><br>For the first time, though, a leaker is claiming that Apple is testing what would be its largest iPhone screen. According to Digital Chat Station on Chinese social media site Weibo, if it goes ahead with this screen, Apple will market it as being a <a href="https://weibo.com/6048569942/R9G4trdtx">7-inch one</a>.<br><br><br> <a href="https://appleinsider.com/articles/26/07/21/iphone-20-pro-max-may-get-apples-largest-ever-iphone-screen-at-7-inches?utm_source=rss">Continue Reading on AppleInsider</a> | <a href="https://forums.appleinsider.com/discussion/245009?urm_source=rss">Discuss on our Forums</a>]]></content:encoded>
</item>
<item>
<title><![CDATA[SaaS will survive, but lazy SaaS is dead]]></title>
<description><![CDATA[Something interesting happened during an internal evaluation of AI meeting transcription tools at Tungsten Automation. The products worked. They weren’t bad. But sitting across from the pricing, we kept asking the same question: what exactly are we paying for? 



We already had a secure enterpri...]]></description>
<link>https://tsecurity.de/de/3683122/ai-nachrichten/saas-will-survive-but-lazy-saas-is-dead/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683122/ai-nachrichten/saas-will-survive-but-lazy-saas-is-dead/</guid>
<pubDate>Tue, 21 Jul 2026 11:05:13 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div><div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Something interesting happened during an internal evaluation of AI meeting transcription tools at Tungsten Automation. The products worked. They weren’t bad. But sitting across from the pricing, we kept asking the same question: what exactly are we paying for? </p>



<p class="wp-block-paragraph">We already had a secure enterprise AI environment. Building a meeting summary workflow took days, not months. We customized the outputs, injected our own internal context, and controlled security our way instead of working around someone else’s roadmap. We built it. It works better. We own it.</p>



<p class="wp-block-paragraph">That’s not a knock on those vendors. It’s a signal of something more fundamental happening across enterprise software.</p>



<h2 class="wp-block-heading">The moat was never the product</h2>



<p class="wp-block-paragraph">For two decades, <a href="https://www.infoworld.com/article/2256637/what-is-saas-software-as-a-service-defined.html" data-type="link" data-id="https://www.infoworld.com/article/2256637/what-is-saas-software-as-a-service-defined.html">SaaS</a> rode a favorable asymmetry: building internal tools was hard, integrations were messy, and even modest automation required developers and long timelines. Buying was faster and cheaper than building. That asymmetry fueled the explosion of SaaS into every corner of the enterprise stack.</p>



<p class="wp-block-paragraph">AI is collapsing that asymmetry. Large language models and agentic workflows can orchestrate APIs, move data between systems, generate interfaces, and automate business logic with a fraction of the engineering effort required even two years ago. The integration friction that once protected entire product categories is evaporating.</p>



<p class="wp-block-paragraph">The vendors most exposed are not the deeply embedded enterprise platforms. They’re the lightweight workflow layers, the products that essentially put a polished interface on top of accessible data and relatively straightforward processes. Reporting dashboards. Meeting tools. Narrow productivity applications. These products created value by simplifying implementation. That rationale is getting harder to sustain when implementation is no longer the real barrier.</p>



<p class="wp-block-paragraph">Here’s the part most analyses miss: it’s not just that AI makes development faster. It’s that agents change the integration model entirely. For 30 years, enterprise software was built for humans navigating UIs. Agentic systems don’t use UIs. They call <a href="https://www.infoworld.com/article/2269032/what-is-an-api-application-programming-interfaces-explained.html" data-type="link" data-id="https://www.infoworld.com/article/2269032/what-is-an-api-application-programming-interfaces-explained.html">APIs</a>, read from multiple sources simultaneously, and move data freely across systems. The switching costs that once made incumbent software sticky are collapsing, because an agent doesn’t care which UI it used last quarter.</p>



<h2 class="wp-block-heading">The SaaS that survives</h2>



<p class="wp-block-paragraph">The question isn’t whether SaaS survives. It’s which SaaS survives.</p>



<p class="wp-block-paragraph">The companies with durable positions are not the ones with the cleanest interface. They’re the ones that transfer operational risk customers genuinely cannot absorb themselves. Compliance. Regulatory certification. Accumulated domain expertise. Liability.</p>



<p class="wp-block-paragraph">Think about compliant invoicing across 140 countries. That’s not a workflow someone builds in a sprint. The certifications alone take years. A single regulatory change in one jurisdiction can break an AP process for a global enterprise overnight. Customers don’t pay for that capability because it’s technically complex. They pay because they cannot afford to own the risk of getting it wrong.</p>



<p class="wp-block-paragraph">That’s the distinction that matters: AI lowers the cost of building software. It does not lower the cost of absorbing risk. The vendors who understand this are building durable businesses. The ones who don’t are quietly subsidizing their customers’ internal build programs.</p>



<p class="wp-block-paragraph">Software sells features. Platforms sell accountability.</p>



<h2 class="wp-block-heading">The prototype trap</h2>



<p class="wp-block-paragraph">The danger for enterprise buyers right now is overcorrection. Every successful prototype looks like a cost-saving opportunity. Very few survive the jump to production.</p>



<p class="wp-block-paragraph">Building a workflow with <a href="https://www.infoworld.com/article/2338115/what-is-generative-ai-artificial-intelligence-that-creates.html" data-type="link" data-id="https://www.infoworld.com/article/2338115/what-is-generative-ai-artificial-intelligence-that-creates.html">generative AI</a> is becoming straightforward. Maintaining it is not. Models evolve. Outputs drift. Governance requirements tighten. What worked cleanly in a controlled environment behaves differently at scale, and the failure mode is worse than traditional software. Rule-based automation, when it fails, fails obviously. Agents fail silently, confidently, at scale, often with a completely reasonable-sounding explanation.</p>



<p class="wp-block-paragraph">Engineering teams that take on AI-powered systems need to solve for observability, model drift, access controls, audit trails, and long-term maintenance ownership. In regulated industries, they need to demonstrate exactly how the system reached every decision. That’s not a weekend project. That’s an ongoing operational commitment that compounds over time as models change and regulatory requirements evolve.</p>



<p class="wp-block-paragraph">Before a team decides to replace an external platform with internal AI tooling, the honest question isn’t, “Can we build this?” The real question is, “Are we prepared to own this in production, for years, as the underlying models change beneath us?” Sometimes the answer is yes. Often the answer is no, and the true cost only becomes visible after the vendor contract is canceled.</p>



<h2 class="wp-block-heading">Build vs. partner: a sharper frame</h2>



<p class="wp-block-paragraph">The build vs. buy framing has always been too binary. The right question is build vs. partner.</p>



<p class="wp-block-paragraph">Partner for the capabilities where risk transfer, regulatory complexity, and domain expertise create genuine value your team cannot replicate. Build for the capabilities that actually differentiate your business from your competitors. Don’t burn your best engineers rebuilding compliant invoice processing or production-grade document extraction. Those aren’t competitive advantages. They’re table stakes, and someone else has already paid the cost, across decades, to make them reliable.</p>



<p class="wp-block-paragraph">The organizations getting this right are honest about where they create unique value. They focus development there, and partner for everything else. The ones getting it wrong are vibe-coding solutions to non-differentiating problems while their actual competitive moat goes unattended.</p>



<h2 class="wp-block-heading">The true value of software</h2>



<p class="wp-block-paragraph">We’re not watching the death of SaaS. We’re watching the end of the friction-based value proposition: the idea that software is worth renewing because integration used to be painful. That rationale is largely gone.</p>



<p class="wp-block-paragraph">What survives is software that does something customers cannot reasonably replicate internally: absorb risk, maintain regulatory compliance, deliver operational reliability at scale, and bring genuine domain expertise into a production-grade system that someone else already stress-tested for years.</p>



<p class="wp-block-paragraph">The vendors who recognize this are already repositioning around accountability, governance, and outcomes. The ones who haven’t will find the next renewal conversation noticeably harder.</p>



<p class="wp-block-paragraph">Software sells features. Platforms sell accountability. That distinction is about to separate a lot of winners from a lot of cautionary tales.</p>



<p class="wp-block-paragraph"><em>—</em></p>



<p class="wp-block-paragraph"><a href="https://www.infoworld.com/blogs/new-tech-forum"><strong><em>New Tech Forum</em></strong></a><em><strong> provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all </strong></em><em><strong>inquiries to </strong></em><a href="mailto:doug_dineley@foundryco.com"><strong><em>doug_dineley@foundryco.com</em></strong></a><em><strong>.</strong></em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026]]></title>
<description><![CDATA[A single AI agent conversation can look flawless scored on its own and still point to a broken product. That gap is driving a shift in how enterprises evaluate agents, away from scoring individual traces and toward comparing cohorts of users against a baseline.At VB Transform 2026, Harrison Chase...]]></description>
<link>https://tsecurity.de/de/3682142/it-nachrichten/a-single-ai-agent-conversation-can-look-perfect-and-still-be-broken-leaders-from-langchain-conviva-and-coreweave-said-at-vb-transform-2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3682142/it-nachrichten/a-single-ai-agent-conversation-can-look-perfect-and-still-be-broken-leaders-from-langchain-conviva-and-coreweave-said-at-vb-transform-2026/</guid>
<pubDate>Mon, 20 Jul 2026 22:48:13 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>A single AI agent conversation can look flawless scored on its own and still point to a broken product. That gap is driving a shift in how enterprises evaluate agents, away from scoring individual traces and toward comparing cohorts of users against a baseline.</p><p>At<a href="https://venturebeat.com/vbtransform2026"> VB Transform 2026</a>, <!-- -->Harrison Chase, CEO of LangChain; Hui Zhang, CTO and co-founder of Conviva; and Emmanuel Turlay, director of engineering at CoreWeave, described that shift, along with a parallel move toward cheaper, narrower judge models.</p><p>Agent-as-judge — judging one AI agent's output with another — hasn't replaced LLM-as-judge, which Chase said remains the default. The larger tension, Zhang said, is between automated judging, whether by LLM or agent, and human review.</p><p>"You have scalable but ungrounded, whether it's agents as judge or LLMs as judge, you grade the outcome, you grade the work. It still is very difficult to ground it and then you use humans and that's just not scalable," Zhang said. "The whole industry is facing this, which poison you want to pick."</p><h2>Evaluation criteria now function as the product spec</h2><p>That gap — a conversation that scores well but still signals a broken product — is what teams try to close by building an exhaustive evaluation suite before they ship anything. Chase said that doesn't work.</p><p>"We sometimes see teams that have almost eval paralysis," Chase said. "They're like, this is an eval set, I can't launch it. The best teams launch and then iterate."</p><p>Chase framed evaluation criteria as a living specification, not a one-time test suite: a product requirements document — the standard software-development spec for what an application should do. "Evals are like the new PRD," he said. "They define what your agent should and shouldn't do."</p><p>Turlay described hitting the same failure from a different angle. "I was trying to reach 100% coverage for my tests, and I still had bugs in production," he said — a test suite that looked complete but still missed what mattered, the same gap Chase was describing with evals.</p><p>Broad, always-on monitoring, he said, catches more real failures than an exhaustive pre-launch test suite. Teams should set up wide online checks first, use those to identify failure classes as they occur, then build a targeted offline evaluation set around the problems that surface.</p><h2>Why scoring traces one at a time is a mistake</h2><p>Even a well-built evaluation process can still score the wrong thing. Zhang's objection is to how most teams run evaluation: sampling traces, whether 50 of them or a full population, scoring each in isolation. That approach misses a signal that only shows up when comparing cohorts of users against a baseline, a method Zhang calls contrastive analysis.</p><p>Zhang illustrated it with a retail example: a shopper asks an agent for a running shoe ahead of a half marathon, the agent asks qualifying questions, and the shopper buys a shoe. Scored individually, that interaction looks fine. But the clarification ratio, how many follow-up questions an agent asks before completing a task, came in three times higher than baseline for that shoe category across the full user population. A second metric, how often shoppers finished their purchase outside the conversation, was five times higher than baseline for the same category.</p><p>Neither number is visible from a single trace. Both point to a debuggable, category-specific problem. Zhang said the industry also lacks a second data source: what happens before, between and after the conversation, not just the trace itself.</p><h2>Sizing the judge to the job</h2><p>Once contrastive analysis flags which category is actually broken, the next problem is what watches for it going forward — and at what cost. Turlay's rule was to start with the most capable model available to prove a task is solvable, then work down. If it can't be done with a top-tier model, he said, it won't work with a smaller one. Once a pattern proves viable, teams can sample a fraction of traffic instead of judging every interaction, and move simpler tasks like binary classification to smaller open source models.</p><p>LangChain took that further, fine-tuning its own model to detect when a user believes the agent made a mistake, a signal Chase calls perceived error. "The model we fine-tuned was a Qwen model," he said, referring to Alibaba's open source family. Combining hand labeling with distillation, the result performed well. "Same as [Claude]Sonnet, for, depending on how we served it, either 10 to 100x cost reduction," Chase said.</p><p>Not every guardrail needs a model. Chase pointed to Claude Code's own guardrails as proof: regexes, the common programming technique for finding and validating patterns in code. "A lot of the guardrails they had were just regexes," he said. "They weren't small LLMs, they were just regexes."</p><h2>LLM-as-judge doesn't mean human-in-the-loop disappears</h2><p>The bigger question is whether using LLM as a judge removes the need for a human in the loop.</p><p>Turlay pointed to accountability, drawing on his prior work at a self-driving car company. His team compressed data intake and retraining into a two-week cycle for shipping a new model to the car. Even then, someone still had to sign off.</p><p>"I felt confident on behalf of the company to say this model should go into the car," he said. The same logic extends to legal, finance and healthcare. "Before we can remove a human to say, I endorse this and I take responsibility legally for it, it's going to be a while before agents can do that on their own."</p><p>Zhang agreed a human has to remain the guardian on corner cases, even as automation eventually runs at a scale that beats individual human accuracy — machines can see more at the pattern level. </p><p>Chase went further: that human check isn't just a safety net. "Human in the loop is really important for building trust in how these agentic systems work, and also really important for memory and learning from systems," he said. "There has to be interactions in order for the system to learn."</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[LVSum: A Benchmark for Timestamp-Aware Long Video Summarization]]></title>
<description><![CDATA[Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark ...]]></description>
<link>https://tsecurity.de/de/3681886/ai-nachrichten/lvsum-a-benchmark-for-timestamp-aware-long-video-summarization/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3681886/ai-nachrichten/lvsum-a-benchmark-for-timestamp-aware-long-video-summarization/</guid>
<pubDate>Mon, 20 Jul 2026 19:48:38 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment. LVSum comprises 72 diverse videos spanning 13 domains with an average duration of 16 minutes, each annotated with up to 10 human-generated summaries containing temporal references. We conduct a comprehensive evaluation…]]></content:encoded>
</item>
<item>
<title><![CDATA[An AI SOC Evaluation Guide for Security Leaders]]></title>
<description><![CDATA[Choosing an AI SOC platform requires understanding how it will perform in your own environment, not just during an evaluation. Prophet Security shares a practical framework for assessing AI SOC solutions, including how to validate accuracy, operating models, long-term reliability, and production ...]]></description>
<link>https://tsecurity.de/de/3681383/it-security-nachrichten/an-ai-soc-evaluation-guide-for-security-leaders/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3681383/it-security-nachrichten/an-ai-soc-evaluation-guide-for-security-leaders/</guid>
<pubDate>Mon, 20 Jul 2026 16:24:38 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Choosing an AI SOC platform requires understanding how it will perform in your own environment, not just during an evaluation. Prophet Security shares a practical framework for assessing AI SOC solutions, including how to validate accuracy, operating models, long-term reliability, and production readiness. [...]]]></content:encoded>
</item>
<item>
<title><![CDATA[With AI, activity is not value]]></title>
<description><![CDATA[The emergence of artificial intelligence is beginning to expose a profound weakness in the way modern enterprises measure performance.



For decades, business evaluation systems have been built around the logic of the industrial and transactional economy. Revenue growth, operating margins, earni...]]></description>
<link>https://tsecurity.de/de/3680938/it-security-nachrichten/with-ai-activity-is-not-value/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3680938/it-security-nachrichten/with-ai-activity-is-not-value/</guid>
<pubDate>Mon, 20 Jul 2026 13:08:24 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">The emergence of artificial intelligence is beginning to expose a profound weakness in the way modern enterprises measure performance.</p>



<p class="wp-block-paragraph"><a href="https://techeconomists.com/why-the-world-needs-new-economic-indicators/">For decades</a>, business evaluation systems have been built around the logic of the industrial and transactional economy. Revenue growth, operating margins, earnings per share, labor productivity, return on investment and market share became the dominant indicators of organizational success because they reflected the economic realities of a world in which value creation was primarily tied to physical production, labor efficiency, scale and later the automation of information processing. AI, however, is altering the very structure of enterprise value creation, and in doing so it is creating a widening separation between perceived future value and actual realized economic performance.</p>



<p class="wp-block-paragraph">Much of the current discussion <a href="https://howardarubin.substack.com/p/why-ai-roi-is-so-darn-hard-to-measure">surrounding AI performance measurement</a> reflects this tension. The overwhelming majority of AI-related metrics being celebrated today are not direct measures of realized enterprise outcomes. They are largely indicators of capability formation, market positioning, experimentation or investor signaling. Metrics such as AI spending levels, number of AI use cases, GPUs deployed, copilots implemented, models placed into production, AI hiring growth or agentic AI pilots all serve primarily as proxies for anticipated future advantage. These indicators may influence stock valuations, analyst sentiment and strategic narratives, but their relationship to measurable operational performance is often indirect, delayed or in some cases entirely speculative.</p>



<p class="wp-block-paragraph">This distinction is critically important because capital markets have historically rewarded the <em>expectation</em> of technological transformation long before actual economic results materialized. During previous technological revolutions—including electrification, enterprise resource planning, the internet, cloud computing and mobile platforms—valuation expansion frequently preceded measurable productivity gains by many years. The market priced future possibility before operational economics caught up. In many instances, investors rewarded firms simply for appearing strategically aligned with the dominant technological shift of the era. AI appears to be following a similar trajectory.</p>



<p class="wp-block-paragraph">The phenomenon resembles the famous <a href="https://www.brookings.edu/articles/the-solow-productivity-paradox-what-do-computers-do-to-productivity/">productivity paradox</a> articulated by economist Robert Solow, who observed that “you can see the computer age everywhere but in the productivity statistics.” AI today is visible everywhere: in investor presentations, earnings calls, technology conferences, product announcements and boardroom strategies. Yet in many industries, its measurable contribution to enterprise productivity, profitability or economic resilience remains difficult to isolate with precision. This does not necessarily mean AI lacks value. Rather, it reflects the reality that traditional accounting and performance systems were never designed to measure the forms of value AI increasingly produces.</p>



<p class="wp-block-paragraph">Artificial intelligence creates benefits that are often diffuse, cumulative and difficult to attribute directly to financial outcomes. AI may improve forecasting accuracy, reduce fraud, accelerate decision cycles, augment employee effectiveness, improve customer interactions, optimize logistics or enhance cybersecurity resilience. These benefits frequently manifest as second-order effects distributed across the enterprise rather than as immediately visible financial events. The causal chain between AI investment and realized business performance can therefore become extraordinarily difficult to quantify. A company may become operationally more intelligent without immediately becoming measurably more profitable.</p>



<p class="wp-block-paragraph">At the same time, AI introduces a profound danger: organizations may increasingly optimize for technological narrative rather than durable enterprise economics. Many firms today are pursuing AI primarily because markets reward the appearance of AI leadership. Investor enthusiasm, analyst pressure and competitive fear create incentives to demonstrate visible AI activity <a href="https://howardarubin.substack.com/p/talking-about-ai-value-is-like-talking">regardless of whether measurable economic value has actually been achieved</a>. In this environment, AI metrics can easily become instruments of valuation signaling rather than instruments of operational truth.</p>



<p class="wp-block-paragraph">This distinction between signaling and substance may become one of the defining economic challenges of the AI era. An organization may announce aggressive AI deployment programs, reduce headcount and report short-term margin improvements while simultaneously increasing hidden forms of technological fragility. Infrastructure costs may rise dramatically as GPU consumption, cloud usage, data engineering requirements and cybersecurity complexity expand. Technical debt may accelerate as AI-generated code proliferates without sufficient architectural discipline. Institutional knowledge may erode as organizations become excessively dependent on opaque models and automated systems. Long-term innovation capacity may weaken if enterprises divert disproportionate resources toward maintaining internally generated AI systems rather than building new strategic capabilities.</p>



<h2 class="wp-block-heading">What measuring AI value might actually look like</h2>



<p class="wp-block-paragraph">The distinction between AI activity and AI value becomes clearer when viewed through the kinds of measures organizations choose to track. Many enterprises today emphasize indicators such as the number of AI models deployed, copilots implemented, agents created, prompts executed, tokens consumed or employees using AI tools. These metrics demonstrate adoption and technological activity, but they reveal relatively little about whether AI is producing meaningful business outcomes.</p>



<p class="wp-block-paragraph">Measures of enterprise value look quite different. A manufacturer might evaluate whether AI improves demand forecasting accuracy enough to reduce inventory carrying costs or stockouts. A financial institution might measure whether AI meaningfully lowers fraud losses, accelerates loan processing or improves regulatory compliance. A healthcare provider could assess reductions in administrative burden, faster clinical decision support or improvements in patient throughput. In each case, the objective is not simply to measure AI deployment, but to determine whether AI creates measurable improvements in operational performance, economic outcomes or organizational resilience.</p>



<p class="wp-block-paragraph">Ultimately, organizations may need to ask a different question: not “How much AI are we using?” but “How much business value does each unit of AI investment create?” That shift—from measuring technological activity to measuring economic outcomes—may become one of the defining management disciplines of the AI era.</p>



<p class="wp-block-paragraph">Under traditional accounting frameworks, many of these deteriorations remain largely invisible. Quarterly earnings may improve even as underlying enterprise resilience declines. Stock prices may rise even as operational complexity becomes increasingly unsustainable. In this sense, the AI era threatens to widen the gap between financial appearance and organizational reality.</p>



<p class="wp-block-paragraph">This is why the future of enterprise measurement cannot simply involve adding AI metrics to existing financial scorecards. The challenge is far deeper. AI forces a reconsideration of what business performance actually means. Historically, enterprises were measured largely through static indicators of efficiency and output. Increasingly, however, competitive advantage may depend less on traditional efficiency and more on adaptive intelligence: the ability of an organization to learn faster, make better decisions, integrate human and machine capabilities effectively, manage technological complexity sustainably and convert computational power into durable economic outcomes.</p>



<p class="wp-block-paragraph">The most important future performance measures may therefore revolve around questions traditional accounting rarely addresses. How effectively does an enterprise convert technology investment into sustainable business capability? How economically efficient are its AI operations relative to the value they generate? How resilient is the organization to AI failure, cybersecurity disruption or infrastructure inflation? How successfully does it preserve and amplify human expertise rather than simply eliminate labor? How rapidly can it learn, adapt and operationalize new knowledge?</p>



<p class="wp-block-paragraph">These are not merely technology questions. They are questions of enterprise economics, organizational sustainability and long-term competitive viability.</p>



<p class="wp-block-paragraph">The companies that ultimately succeed in the AI era may not be those with the largest AI budgets, the greatest number of pilots or the most aggressive automation programs. They may instead be the firms that best understand the economics of technological capability itself: organizations capable of balancing innovation with resilience, automation with human augmentation and technological ambition with sustainable operational design.</p>



<p class="wp-block-paragraph">The coming decade is therefore likely to produce a widening divide between enterprises optimizing for AI-driven valuation narratives and enterprises optimizing for measurable, durable economic performance. In the short term, these may appear to be the same thing.</p>



<p class="wp-block-paragraph">Over time, however, the distinction will become increasingly visible. Some organizations will discover that AI has enhanced genuine enterprise capability. Others will discover that they merely optimized the appearance of transformation while silently accumulating new forms of economic and operational risk.</p>



<p class="wp-block-paragraph">Artificial intelligence is not simply changing business operations. It is exposing the inadequacy of many of the measures used to evaluate business success itself. The central challenge of the AI economy may ultimately become not whether organizations adopt AI, but whether they can distinguish between technological activity and actual economic value creation.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-42533 Exposes Critical Pre-Auth nginx RCE Flaw]]></title>
<description><![CDATA[A newly disclosed security flaw, CVE-2026-42533, has revealed a critical Pre-Auth nginx vulnerability that could allow attackers to achieve reliable RCE (remote code execution) without authentication. The issue affects nginx versions 0.9.6 through 1.30.3 (stable) and 1.31.2 (mainline), while patc...]]></description>
<link>https://tsecurity.de/de/3680443/it-security-nachrichten/cve-2026-42533-exposes-critical-pre-auth-nginx-rce-flaw/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3680443/it-security-nachrichten/cve-2026-42533-exposes-critical-pre-auth-nginx-rce-flaw/</guid>
<pubDate>Mon, 20 Jul 2026 08:52:32 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1217" height="768" src="https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533.webp" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="CVE-2026-42533" decoding="async" srcset="https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533.webp 1217w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-300x189.webp 300w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-1024x646.webp 1024w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-768x485.webp 768w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-600x379.webp 600w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-150x95.webp 150w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-750x473.webp 750w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-1140x719.webp 1140w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533.webp 1217w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-300x189.webp 300w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-1024x646.webp 1024w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-768x485.webp 768w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-600x379.webp 600w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-150x95.webp 150w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-750x473.webp 750w, https://thecyberexpress.com/wp-content/uploads/CVE-2026-42533-1140x719.webp 1140w" sizes="(max-width: 1217px) 100vw, 1217px" title="CVE-2026-42533 Exposes Critical Pre-Auth nginx RCE Flaw 1"></p><span data-contrast="auto">A newly disclosed security flaw, CVE-2026-42533, has revealed a critical Pre-Auth nginx vulnerability that could allow attackers to achieve reliable RCE (remote code execution) without authentication. The issue affects nginx versions 0.9.6 through 1.30.3 (stable) and 1.31.2 (mainline), while patched releases include 1.30.4 and 1.31.3. Affected NGINX Plus versions include R33-R36 (fixed in R36 P7) and 37.0.0.1-37.0.2.1 (fixed in 37.0.3.1).</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">According to the disclosure, the <a class="wpil_keyword_link" href="https://thecyberexpress.com/firewall-daily/vulnerabilities/" title="vulnerability" data-wpil-keyword-link="linked" data-wpil-monitor-id="29039">vulnerability</a> stems from a missing save-and-restore mechanism for PCRE capture state within nginx's two-pass script evaluation engine. The flaw enables attackers to trigger a heap buffer overflow with attacker-controlled content and length, while also exposing heap pointers through an information leak that can defeat Address Space Layout Randomization (ASLR). Chaining both primitives enables reliable Pre-Auth nginx RCE.</span><span data-ccp-props="{}"> </span>
<h3 aria-level="2"><b><span data-contrast="none">CVE-2026-42533 Impacts Multiple Configurations</span></b><span data-ccp-props='{"134245418":true,"134245529":true,"335559738":160,"335559739":80}'> </span></h3>
<span data-contrast="auto">The <a href="https://cyberstan.co.uk/nginx-rce/" target="_blank" rel="nofollow noopener">advisory warns</a> that deployments using map directives with regex patterns alongside regex capture sources, including location, server_name, rewrite, or if blocks, may be vulnerable. The issue depends on evaluation order, where regex capture references, such as $1 or named groups, are processed before a regex map variable.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Affected directives include proxy_set_header, proxy_method, proxy_pass, fastcgi_param, uwsgi_param, scgi_param, grpc_set_header, return, add_header, rewrite, set, root, alias, and access_log, among others. Both HTTP and stream modules are affected, and the <a href="https://thecyberexpress.com/default-credentials-polish-energy-grid-attack/" target="_blank" rel="noopener">vulnerable</a> capture and map variables do not need to exist within the same directive.</span><span data-ccp-props="{}"> </span>
<h3 aria-level="2"><b><span data-contrast="none">Technical Root Cause</span></b></h3>
<span data-contrast="auto">The researcher explained that nginx evaluates expressions in two stages: a length calculation (LEN) pass followed by a value (VALUE) pass. During execution, regex map evaluation overwrites shared capture <a class="wpil_keyword_link" href="https://thecyberexpress.com/what-is-data/" title="data" data-wpil-keyword-link="linked" data-wpil-monitor-id="29040">data</a> stored in the request object. As a result, the LEN pass and VALUE pass can calculate different capture sizes, causing either a heap overflow or an information leak depending on the relative capture lengths.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">The disclosure states that attackers can control both the overflow size and leaked data using ordinary HTTP requests, including request URIs, headers, and bodies. No credentials, client certificates, or unusual configuration beyond the vulnerable pattern are required.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Testing reportedly achieved 10 out of 10 successful <a href="https://thecyberexpress.com/rcritical-ivanti-csa-vulnerabilities-exploited/" target="_blank" rel="noopener">exploitations</a> on Ubuntu 24.04 using glibc 2.39 with ASLR enabled.</span>
<h3 aria-level="2"><b><span data-contrast="none">Mitigation and Disclosure</span></b></h3>
<span data-contrast="auto">The researcher said recent fixes for CVE-2026-42945, CVE-2026-9256, CVE-2026-42055, and CVE-2026-48142 do not address CVE-2026-42533. Administrators are advised to upgrade immediately to nginx 1.30.4, 1.31.3, or the corresponding patched NGINX Plus releases.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">Until systems are updated, defenders should audit configurations that combine regex captures with regex map variables in the same evaluation path. The researcher also released a static configuration scanner that identifies vulnerable configurations without exploiting them.</span><span data-ccp-props="{}"> </span>

<span data-contrast="auto">The initial report was submitted to F5 SIRT on May 17, 2026, with follow-up analyses covering additional variants, including cross-directive triggering and named capture clobbering. While a proof-of-concept exploit exists, the researcher said it will be withheld until users have sufficient time to apply patches, citing concerns over rapid exploitation following previous <a href="https://thecyberexpress.com/nginx-rift-cve-2026-42945-active-exploitation/" target="_blank" rel="noopener">nginx</a> vulnerability disclosures.</span><span data-ccp-props="{}"> </span>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-46621 | Yamcs up to 5.12.6 Script Evaluation Engine sandbox]]></title>
<description><![CDATA[A vulnerability was found in Yamcs up to 5.12.6 and classified as very critical. This vulnerability affects unknown code of the component Script Evaluation Engine. Executing a manipulation can lead to sandbox issue.

This vulnerability appears as CVE-2026-46621. The attack may be performed from r...]]></description>
<link>https://tsecurity.de/de/3680385/sicherheitsluecken/cve-2026-46621-yamcs-up-to-5126-script-evaluation-engine-sandbox/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3680385/sicherheitsluecken/cve-2026-46621-yamcs-up-to-5126-script-evaluation-engine-sandbox/</guid>
<pubDate>Mon, 20 Jul 2026 08:24:02 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability was found in <a href="https://vuldb.com/product/yamcs">Yamcs up to 5.12.6</a> and classified as <a href="https://vuldb.com/kb/risk">very critical</a>. This vulnerability affects unknown code of the component <em>Script Evaluation Engine</em>. Executing a manipulation can lead to sandbox issue.

This vulnerability appears as <a href="https://vuldb.com/cve/CVE-2026-46621">CVE-2026-46621</a>. The attack may be performed from remote. There is no available exploit.]]></content:encoded>
</item>
<item>
<title><![CDATA[The audit trail CIOs need before the next cyber crisis]]></title>
<description><![CDATA[In one ransomware response I observed, the master operational dashboard remained green while the underlying environment told a very different story. It was a classic example of what we in the IT audit profession call the “watermelon effect”—green on the outside, red on the inside.



Beneath that...]]></description>
<link>https://tsecurity.de/de/3680143/it-security-nachrichten/the-audit-trail-cios-need-before-the-next-cyber-crisis/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3680143/it-security-nachrichten/the-audit-trail-cios-need-before-the-next-cyber-crisis/</guid>
<pubDate>Mon, 20 Jul 2026 02:13:22 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">In one ransomware response I observed, the master operational dashboard remained green while the underlying environment told a very different story. It was a classic example of what we in the IT audit profession call the “watermelon effect”—green on the outside, red on the inside.</p>



<p class="wp-block-paragraph">Beneath that dashboard sat an unmapped web of legacy technical debt, undocumented service accounts and shadow cloud instances. For years, presenting a green dashboard to the audit committee could give technology leaders a false sense of comfort. If a catastrophic breach occurred, it was generally treated as an unpredictable operational tragedy, managed via cyber insurance, a carefully calibrated public relations pivot and perhaps a quiet executive transition.</p>



<p class="wp-block-paragraph">Today, that corporate shield is thinner than many technology leaders assume. For technology leaders in regulated or public-company environments, executive exposure is no longer only a theoretical debate. The regulatory environment has made plausible deniability much harder to sustain.</p>



<h2 class="wp-block-heading">The erosion of the corporate shield</h2>



<p class="wp-block-paragraph">With the application of the European Union’s <a href="https://www.eiopa.europa.eu/digital-operational-resilience-act-dora_en">Digital Operational Resilience Act (DORA)</a> for financial entities, alongside the broader <a href="https://digital-strategy.ec.europa.eu/en/policies/nis2-directive">NIS2 Directive</a> for essential and important entities, cybersecurity governance has become harder to separate from board-level oversight. DORA places ultimate responsibility for ICT risk management on the management body of financial entities, while NIS2 requires management bodies to approve and oversee cybersecurity risk-management measures. In the United States, the <a href="https://www.sec.gov/newsroom/press-releases/2023-139">U.S. Securities and Exchange Commission’s cybersecurity disclosure rules</a> require public companies to disclose material cyber incidents and describe their cyber risk management, strategy and governance in annual filings. The new burden is not simply to operate controls; it is to show, after the fact, that leadership decisions matched the risk evidence available at the time.</p>



<p class="wp-block-paragraph">The serious risk to a modern CIO is not simply the occurrence of a sophisticated security incident. The true danger is the inability to reconcile what leadership presented externally to investors, regulators and the board with what the internal evidence showed inside the environment.</p>



<p class="wp-block-paragraph">When a serious crisis breaks, you may find yourself surrounded by corporate defense counsel, regulatory investigators and outside forensic lawyers all asking variations of the same uncomfortable questions: What did you know, when did you discover it and what specific actions did you take next?</p>



<p class="wp-block-paragraph">When those questions are asked, a slide deck asserting that your security posture is “aligned with industry best practices” will not be enough. A post-incident review may recognize that sophisticated attacks occur. What creates greater exposure is evidence that known risks were ignored, understated or left outside structured governance. To survive that level of post-incident review, one of your strongest assets is a disciplined, independent evidence trail showing that risks were identified, challenged, escalated and acted on before the first indicator of compromise appeared.</p>



<h2 class="wp-block-heading">Why point-in-time comfort letters fail regulatory scrutiny</h2>



<p class="wp-block-paragraph">The reality we face is that legacy compliance evidence often falls short under regulatory scrutiny. For years, the annual SOC 2 Type II report or a standardized ISO 27001 certification was brandished by technology teams as the definitive proof of a functional control environment. I have sat in dozens of scoping meetings where an engineering director pointed to a freshly minted compliance report as if it were a complete defense against scrutiny.</p>



<p class="wp-block-paragraph">But a compliance report is a historical artifact—a retrospective evaluation of how specific controls operated during a defined window of time months in the past. It tells an investigator that on a random afternoon in Q2, your production change-management approvals conformed to a baseline policy. It says absolutely nothing about the configuration drift, unauthorized API keys or emergency patch bypasses that developers introduced the following weekend to hit a product release deadline.</p>



<p class="wp-block-paragraph">Modern regulators, boards and investors are no longer satisfied by historical comfort letters alone. Under contemporary frameworks, especially regimes focused on operational resilience, static compliance evidence is no longer enough. The expectation of due care has shifted from a passive state of compliance to an active state of continuous challenge. Increasingly, post-incident reviews look for evidence that leadership identified system vulnerabilities, formally escalated material deficiencies, evaluated systemic risk to the business and tracked remediation progress with measurable rigor.</p>



<p class="wp-block-paragraph">When an architecture fails, post-incident reviews often focus quickly on ownership, escalation and whether known risks were acted upon. If your defensive documentation consists entirely of static policy documents and green dashboards, you leave an evidentiary vacuum that can invite difficult questions about executive oversight. Post-incident reviews rarely turn on perfection. They turn on whether the organization can show a traceable chain of governance.</p>



<h2 class="wp-block-heading">5 non-negotiable artifacts for your executive evidence engine</h2>



<p class="wp-block-paragraph">This reality requires a complete reframing of your relationship with your IT audit department. Historically, this dynamic has been defined by friction. Technology leaders frequently view my peers and me as compliance traffic cops—bureaucrats who interrupt core engineering sprints to demand evidence samples, user access reviews and system configurations.</p>



<p class="wp-block-paragraph">It is time to view IT audit through a pragmatic lens: we are your independent evidence engine. We are one of the few corporate functions tasked with independently challenging your control environment, documenting where exceptions were escalated and showing how management responded. When an auditor identifies a control gap and partners with you to draft a management action plan, they are not creating a bureaucratic roadblock. They are helping you construct an evidence trail that can show risk was identified, escalated and acted upon.</p>



<p class="wp-block-paragraph">To transform your IT audit function into an effective executive shield, you must shift focus away from superficial check-the-box exercises and collaborate on specific artifacts. The most effective exercise you can run with your audit leadership is to flip the timeline completely and ask: if this program were reviewed six months from now, which evidence would show we governed the risk before it failed?</p>



<ol class="wp-block-list">
<li><strong>Board-facing risk registers with escalation history:</strong> A risk register that sits unreviewed on an intranet page for 12 months is not a management tool; to an investigator, it can look like evidence that known risks were not actively governed. Your material technology, cybersecurity and dependency risks must be centrally logged. More importantly, this artifact must contain a clear, chronological escalation history showing exactly when the risk was presented to leadership committees and the board, along with related minutes, decisions or follow-up actions.</li>



<li><strong>Granular risk acceptance records:</strong> You cannot remediate every vulnerability instantly. Business continuity, legacy software limitations and budgetary boundaries require you to accept certain operational exposures. When this occurs, ensure your risk acceptance records are airtight. A defensible record must document the specific technical variance, the precise financial or operational rationale for the delay, a definitive expiration date, explicit executive sign-off and the active compensating controls deployed to reduce the blast radius in the interim.</li>



<li><strong>Tabletop and operational simulation records:</strong> Independent frameworks such as <a href="https://www.isaca.org/digital-trust">ISACA’s Digital Trust Ecosystem Framework</a> can help structure this evidence, but boards and regulators will still look for proof that the testing actually happened. Your audit trail should contain comprehensive records of cyber incident, disaster recovery and third-party dependency simulations. These records must detail the scenario tested, the executive participants, the control failures identified during the drill and a formalized tracking schedule showing when those gaps were closed.</li>



<li><strong>AI governance inventories and data-flow mappings:</strong> The rapid deployment of generative AI tools across enterprise operations has created a massive blind spot for technology executives. In one audit, we found developers using an unapproved public large language model to accelerate debugging with sensitive internal code. To protect yourself, work with your audit team to build an active enterprise AI inventory that maps data lineage, identifies model business owners, documents risk classification approvals and demonstrates active technical monitoring for unauthorized data exfiltration.</li>



<li><strong>Synchronized disclosure-control handoffs:</strong> When a material security incident or system outage occurs, the clock begins ticking for regulatory reporting. Your incident response playbook must be technically linked to your corporate disclosure controls. The audit trail should show that a documented, synchronized handoff occurred between your technical response leaders, general counsel, chief financial officer and corporate communications team. This evidence helps show that your external statements match internal technical realities.</li>
</ol>



<p class="wp-block-paragraph">In the modern corporate ecosystem, technology leadership is no longer just an engineering challenge; it is an exercise in rigorous, evidence-based governance. The regulatory landscape has changed, and the expectation of continuous traceability cannot be avoided.</p>



<p class="wp-block-paragraph">Open and direct collaboration with your IT audit team will not prevent a zero-day exploit, an unexpected cloud outage or a critical third-party vendor failure. That is not the purpose of enterprise risk management.</p>



<p class="wp-block-paragraph">The true value is far more practical: when a serious incident puts your program under review, you will not be forced to defend your reputation with a feeling, an unverified assumption or a misleadingly green dashboard. Instead, you will have an independent record showing that risk was actively seen, appropriately challenged, properly escalated and responsibly managed. In today’s regulatory environment, that disciplined trail of evidence may be the difference between a failure that can be explained and one that begins to look negligent.</p>



<p class="wp-block-paragraph">.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Genius Bar AI tools spark concerns over employee monitoring & evaluation]]></title>
<description><![CDATA[Apple is testing a new Genius Bar tool called Live Notes that transcribes and summarizes conversations with a customer, but employees are worried about how the tool might be used against them.Genius Bar employees may gain a new AI assistantArtificial intelligence has become a buzzword throughout ...]]></description>
<link>https://tsecurity.de/de/3679619/ios-mac-os/genius-bar-ai-tools-spark-concerns-over-employee-monitoring-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3679619/ios-mac-os/genius-bar-ai-tools-spark-concerns-over-employee-monitoring-evaluation/</guid>
<pubDate>Sun, 19 Jul 2026 16:39:02 +0200</pubDate>
<category>🍏 iOS / Mac OS</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Apple is testing a new Genius Bar tool called Live Notes that transcribes and summarizes conversations with a customer, but employees are worried about how the tool might be used against them.<br><br><div><img src="https://photos5.appleinsider.com/gallery/68287-143944-iPad-mini-7-4-xl.jpg" alt="Back of a gray Apple iPad with rear camera and Apple logo, set against a dark background featuring glowing neon loops in orange, yellow, blue, and pink" height="738"><br><span>Genius Bar employees may gain a new AI assistant</span></div><br>Artificial intelligence has become a buzzword throughout the tech industry from both the consumer and employee perspective. Employees of any kind at many companies have had to contend with mandatory AI tool use forced on them by the employer.<br><br>So far, reports haven't indicated Apple taking such a hardline route to AI tool use. However, the "Power On" newsletter from <em>Bloomberg</em> <a href="https://www.bloomberg.com/account/newsletters/power-on">shares that</a> Apple is testing a new tool called Live Notes for use in the Genius Bar.<br><br><br> <a href="https://appleinsider.com/articles/26/07/19/genius-bar-ai-tools-spark-concerns-over-employee-monitoring-evaluation?utm_source=rss">Continue Reading on AppleInsider</a> | <a href="https://forums.appleinsider.com/discussion/244993?urm_source=rss">Discuss on our Forums</a>]]></content:encoded>
</item>
<item>
<title><![CDATA[Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep]]></title>
<description><![CDATA[Perplexity's WANDR is an open benchmark and evaluation harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each one with cited, re-verifiable evidence. Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1.
The post P...]]></description>
<link>https://tsecurity.de/de/3679066/ai-nachrichten/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3679066/ai-nachrichten/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep/</guid>
<pubDate>Sun, 19 Jul 2026 09:33:16 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Perplexity's WANDR is an open benchmark and evaluation harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each one with cited, re-verifiable evidence. Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1.</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep/">Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[So AIs are lying to us now, huh? (emf2026)]]></title>
<description><![CDATA[As AI systems are getting smarter, safety researchers are increasingly concerned by a new capability of these models - they are getting very good at lying. This behaviour is seen both in and out of evaluation settings, and has a fascinating range of root causes. 

In this talk, we'll investigate ...]]></description>
<link>https://tsecurity.de/de/3678594/it-security-video/so-ais-are-lying-to-us-now-huh-emf2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3678594/it-security-video/so-ais-are-lying-to-us-now-huh-emf2026/</guid>
<pubDate>Sun, 19 Jul 2026 01:03:16 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[As AI systems are getting smarter, safety researchers are increasingly concerned by a new capability of these models - they are getting very good at lying. This behaviour is seen both in and out of evaluation settings, and has a fascinating range of root causes. 

In this talk, we'll investigate the phenomena of AI deception across a broad range of scenarios. Discover how chatbots lie for your own good, when models strategically hide their capabilities, and why ChatGPT is better than you at Avalon. We'll cover how researchers are training model organisms designed to be good at deception, and how this helps us to detect scheming in the wild.

Most importantly, we'll attempt to answer the question: how worried should we be about all this?

Licensed to the public under https://creativecommons.org/licenses/by-sa/4.0/
about this event: https://www.emfcamp.org/schedule/2026/154-so-ais-are-lying-to-us-now-huh]]></content:encoded>
</item>
<item>
<title><![CDATA[Decoding the Obfuscated Layer: A Playbook Walkthrough of Command-Line Forensics]]></title>
<description><![CDATA[A full and detailed insight into CLI forensics, going into depth following a TryHackMe labSource: TechFusionFor incident responders, security analysts, and threat hunters, discovering an unknown script execution running on an enterprise workstation triggers an immediate race against time. Is it a...]]></description>
<link>https://tsecurity.de/de/3677762/hacking/decoding-the-obfuscated-layer-a-playbook-walkthrough-of-command-line-forensics/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3677762/hacking/decoding-the-obfuscated-layer-a-playbook-walkthrough-of-command-line-forensics/</guid>
<pubDate>Sat, 18 Jul 2026 11:21:48 +0200</pubDate>
<category>🕵️ Hacking</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><em>A full and detailed insight into CLI forensics, going into depth following a TryHackMe lab</em></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/711/1*z2_rLeW29e_A1-URgEyBcw.jpeg"><figcaption>Source: TechFusion</figcaption></figure><p>For incident responders, security analysts, and threat hunters, discovering an unknown script execution running on an enterprise workstation triggers an immediate race against time. Is it a harmless administrative automation tool, or is it an advanced information stealer scraping the credential caches of every corporate browser?</p><p>I recommend you first walk through this article and afterwards complete the TryHackMe lab <a href="https://tryhackme.com/room/obfuscation-aoc2025-e5r8t2y6u9"><strong>Obfuscation: The Egg Shell File</strong></a>.</p><p>The core purpose of this tactical playbook is to provide you with a <strong>highly comprehensive, real-world analytical framework</strong> so you can confidently dive into the live lab environment (don’t, i say DON’T worry about committing every single execution flag or decoding syntax to memory; the structural muscle memory will lock in during the hands-on exercises).</p><p><strong>Let’s cut the fluff and begin:</strong></p><p>In modern security operations, the discipline of malware analysis bridges the gap between passive defense and active threat hunting. Using <a href="https://tryhackme.com/room/obfuscation-aoc2025-e5r8t2y6u9">TryHackMe’s foundational lab</a> featuring <strong>real-world PowerShell obfuscation</strong> strings, this walkthrough guides defenders through the surgical progression required to size up a hostile payload, calculate its technical attributes, map its internal compiled structure, and decrypt its <strong>runtime behavior</strong> safely.</p><h3>📋 The Script Triage Checklist</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*fVQzacpYWWO-vFYN.jpg"><figcaption>Source: BitLyft</figcaption></figure><p>When you capture a suspicious script execution string from your <strong>SIEM</strong> (Security Information and Event Management) <strong>logs</strong>, proceed with these steps immediately:</p><ul><li><strong>Isolate and Copy Safely:</strong> Transfer the raw text string into a completely disconnected text editor inside a designated analysis virtual machine.</li><li><strong>Identify the Execution Flags:</strong> Search for evasion switches like -NoP (No Profile), -W Hidden (Window Hidden), or -Enc (Encoded Command), which indicate deliberate bypass actions.</li><li><strong>Locate Network Anchors:</strong> Scan the text string for markers like DownloadString, DownloadFile, curl, or iwr that hint at secondary external downloads.</li><li><strong>Preserve Casing:</strong> Do not run lowercase or uppercase find-and-replace scripts across your sample yet; case variance is often structurally critical to decoding algorithms.</li></ul><h3>Deep Dive: Stripping the Camouflage</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/636/1*6nmJ8HdVrRHV92qrBTYjGQ.png"><figcaption><a href="https://www.researchgate.net/figure/The-obfuscation-techniques-of-code-element-layer_fig2_340401812">https://www.researchgate.net/figure/The-obfuscation-techniques-of-code-element-layer_fig2_340401812</a></figcaption></figure><p>Let’s look at an actual example of an <strong>obfuscated script layer</strong> captured directly from an initial access vector payload log.</p><blockquote><strong><em>What to look for in the image:</em></strong><em> Notice how the raw command string uses a combination of string splitting, character swapping, and nested script blocks. Threat actors do this </em><strong><em>to bypass static string matching</em></strong><em> (signatures) used by endpoint detection engines. By analyzing the structural markers, we can map out the exact unpacking routine.</em></blockquote><h4>Layer 1: Undoing String Concatenation</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*A0nXU-f5TINGOWU1Ce-ffw.png"><figcaption>GPT Images 2.0 generated photo</figcaption></figure><p>Attackers frequently break apart their critical strings using addition operators or variable insertions to stop simple pattern scanners.</p><pre># Obfuscated string snippet<br>$a = "Down"; $b = "load"; $c = "String"<br>. ( $ExecutionContext.InvokeCommand.ExpandString('$' + 'a' + '$' + 'b' + '$' + 'c') )</pre><p><strong>The Fix:</strong> You don’t have to guess what this does. By loading the script into an isolated PowerShell CLI and replacing the aggressive execution operator (like . or Invoke-Expression / IEX) with a safe print directive like Write-Output, the environment itself will assemble the string for you:</p><pre># Safe evaluation technique<br>Write-Output ( $ExecutionContext.InvokeCommand.ExpandString('$' + 'a' + '$' + 'b' + '$' + 'c') )<br># Output result: DownloadString</pre><h4>Layer 2: Demangling Character Shuffling</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*m92GtATfKF2yD6KmfMs2HQ.png"><figcaption>GPT Images 2.0 generated image</figcaption></figure><p>Another popular mechanism involves using <strong>format strings</strong> to re-order components out of sequence at runtime:</p><pre>"{2}{0}{1}" -f 'Net.','WebClient','New-Object </pre><p>The -f operator acts as an indexing map. To decrypt it manually:</p><ul><li>Position {2} grabs the 3rd element: New-Object</li><li>Position {0} grabs the 1st element: Net.</li><li>Position {1} grabs the 2nd element: WebClient</li></ul><p>When evaluated sequentially by the command pipeline, it structures clean and functional telemetry: New-Object Net.WebClient.</p><h4>Layer 3: Defeating Base64 and XOR Rings</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2cEgUTPr8IvHXOAQhUBqjA.png"><figcaption>GPT Images 2.0 generated figure</figcaption></figure><p>The final boss of script obfuscation is almost always an <strong>encoded byte block</strong>. Base64 is easily recognizable by its standard alphanumeric character set and trailing padding markers (=).</p><p>To quickly unwrap these blocks without running the malicious code:</p><ul><li>Copy the raw payload block inside the command string.</li><li>Load the payload directly into <strong>CyberChef</strong> (the open-source utility for security operations).</li><li>Chain together the <strong>From Base64</strong> recipe followed by <strong>Decode Text (UTF-16LE)</strong>.</li></ul><pre>Input:  aAB0AHQAcAA6AC8ALwBtAGEAbAB3AGEAcgBlAC4AbgBlAHQALwBwAGEAeQBsAG8AYQBkAC4AZQB4AGUA<br>Output: http://malware.net/payload.exe</pre><p>By working backward through these layers, you quickly isolate the final <strong>Indicators of Compromise (IoCs) </strong>— such as the secondary payload download URL or target staging paths — allowing your security infrastructure to immediately blacklist the server across the enterprise.</p><h3>🧠 Strategic Takeaway</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RnNGqsHcsh-4wOWfT3Lkag.jpeg"><figcaption>(yayy)</figcaption></figure><p>The <a href="https://tryhackme.com/room/obfuscation-aoc2025-e5r8t2y6u9"><strong>Obfuscation: The Egg Shell File</strong></a> analysis framework underscores a foundational truth of computer network defense: Malware cannot accomplish its mission without leaving a structural or behavioral footprint inside operational logs.</p><blockquote>Whether it is a distinct jump in character selection counts, an unexpected system variable concatenation flag, or a sudden burst of hidden network invocation arguments executed entirely from background windows, an <strong>obfuscated script pipeline</strong> will always reveal its true payload target under systematic scrutiny.</blockquote><p>By utilizing platforms like <strong>CyberChef</strong> to strip back multi-layered <strong>Base64 and XOR encoding architectures</strong> and verifying those outputs within isolated environments, defenders completely eliminate the guesswork from administrative code reviews.</p><p>Go log into <strong>the TryHackMe room</strong>, reverse the nested string layout structures of the script sample, map out the true operational strings, and transform your defensive triage into an optimized playbook.</p><h3>📈 Master the Art of System Forensics &amp; Threat Intelligence</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/337/1*SLfKdWyn-nVH4_for38UxQ.jpeg"><figcaption>The author</figcaption></figure><p>Generic security training <strong>completely collapses</strong> when sophisticated threat groups deploy obfuscated, packed, and tailored payloads across your endpoints.</p><p>To ensure you never miss an in-depth threat intelligence playbook pulling back the curtain on advanced binary analysis, active threat hunting, and modern defense frameworks:</p><ul><li><strong>Follow Pop123 on Medium</strong> for immediate notifications on all newly published technical deep-dives, infrastructure hardening playbooks, and reverse-engineering guides.</li><li><strong>Explore my Security and Machine Learning Projects on </strong><a href="https://github.com/pop123-ux"><strong>GitHub</strong></a></li><li><strong>Subscribe to direct email updates</strong> by clicking the envelope icon (✉️) right next to the follow button so these critical tactical breakdowns land straight in your inbox.</li></ul><p><em>Thank you for reading. This article was entirely written by Pop123. If you found this technical breakdown of the malware analysis matrix valuable, consider leaving a clap and sharing your thoughts, configuration questions, or analytical feedback in the responses below, I am as always open to further discussing the interesting topics!</em></p><p><strong>For collaborations and inquiries</strong>: alexandrupp55@gmail.com</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=d96840b5b5ef" width="1" height="1" alt=""><hr><p><a href="https://infosecwriteups.com/decoding-the-obfuscated-layer-a-playbook-walkthrough-of-command-line-forensics-d96840b5b5ef">Decoding the Obfuscated Layer: A Playbook Walkthrough of Command-Line Forensics</a> was originally published in <a href="https://infosecwriteups.com/">InfoSec Write-ups</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026]]></title>
<description><![CDATA[Legacy infrastructure, not the models themselves, is what's actually slowing AI agents down. That was the shared conclusion of three infrastructure leaders — from LinkedIn, Walmart, and Zendesk — at VB Transform 2026.The panel brought together Animesh Singh, senior director of AI platform and inf...]]></description>
<link>https://tsecurity.de/de/3676906/it-nachrichten/agents-think-in-milliseconds-legacy-infrastructure-doesnt-linkedin-walmart-and-zendesk-shared-how-they-closed-the-gap-at-vb-transform-2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3676906/it-nachrichten/agents-think-in-milliseconds-legacy-infrastructure-doesnt-linkedin-walmart-and-zendesk-shared-how-they-closed-the-gap-at-vb-transform-2026/</guid>
<pubDate>Fri, 17 Jul 2026 21:32:54 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Legacy infrastructure, not the models themselves, is what's actually slowing AI agents down. That was the shared conclusion of three infrastructure leaders —<!-- --> from LinkedIn, Walmart, and Zendesk —<!-- --> at<a href="https://venturebeat.com/vbtransform2026"> VB Transform 2026</a>.</p><p>The panel brought together Animesh Singh, senior director of AI platform and infrastructure at LinkedIn, Desiree Gosby, SVP of corporate technology services and technology strategy at Walmart, and Sami Ghoche, VP of applied AI at Zendesk, each describing what actually broke when they moved agents from pilot to production. Each arrived at the same conclusion from a different starting point: None of the bottlenecks they hit were model problems.</p><p>What tied their answers together was a shared premise: most enterprise infrastructure was built for how humans work, not for how agents work. The gap between those two speeds is where the real engineering happened.</p><p>Gosby put it plainly when asked what she'd learned scaling agents inside Walmart's own workforce. The goal, she said, is to make sure "engineering doesn't once again become the bottleneck for what it is we're trying to do."</p><h2><b>Where the bottleneck actually was</b></h2><p>Each company hit a different version of the same wall: infrastructure designed for how people work doesn't hold up once agents are doing the work instead.</p><p>At LinkedIn, the first bottleneck wasn't a model, it was Kubernetes, which assumes containers spin up on demand, a process that takes seconds. Singh said that's too slow for agents. The fix was moving from on-demand provisioning to pre-provisioned pools of containers that swap agentic workloads in and out in real time.</p><p>A second, harder problem surfaced once LinkedIn let agents control their own orchestration. A five-point evaluation system looked clean, but hallucination kept showing up anyway. Singh said the issue was structural, an LLM evaluating another LLM's output shares the same failure mode as the thing it's evaluating. </p><p>"We built our own harness, our own control flow, and pushed the LLMs to the leaf instead of them orchestrating the loop," Singh said. Roughly 80% of the workflow is now scripted, deterministic code, with LLMs used only where reasoning is required, and each step's evidence is committed to disk before the system moves on.</p><p>Walmart's bottleneck came from success. An agent harness put directly into employees' hands went viral internally, and what Gosby called "citizen developers" began building their own agents to solve problems that once required a formal engineering roadmap. The upside was real innovation. The downside was duplication, dozens of overlapping agents with no coordination. The fix wasn't reining in the harness, it was building governance to spot duplication, promote the best version of an agent, and get it into production without engineering becoming a chokepoint.</p><p>Zendesk hit its bottleneck from the data side. Ghoche, who joined through <a href="https://www.zendesk.com/newsroom/press-releases/zendesk-completes-acquisition-of-forethought/">Zendesk's acquisition of Forethought</a>, which closed in March 2026, described sitting on what he called a public figure of 20 billion customer conversations in Zendesk's repository. The instinct is to hand that history to a large language model with a big context window and let it generate the agents a business needs. Ghoche said that doesn't work. "You can't really do that, so instead you have to really invest in the underlying data pipelines and all the data infrastructure that comes with that," he said.</p><h2>The role of open source</h2><p>On open source, all three leaders landed on a similar instinct: own what you can, and lean on frontier labs only where they still have a clear edge.</p><p>Ghoche said his own view is that most enterprises would prefer to own their models and infrastructure wherever that's possible, and that reasoning is what drives Zendesk's own approach. The exception is frontier reasoning work, where the labs still lead, though he said that slice of use cases is shrinking relative to everything else enterprises now do with AI.</p><p>LinkedIn's answer was to build two subsystems specifically for independence. The first is what the company calls an AI gateway, a single interface that every outbound call to a model runs through regardless of provider. The second component is a memory subsystem built to hold context independent of any model provider.</p><p>"Every single outbound call going to an LLM, whether it's on a public cloud or on-prem in our own data centers, follows the same semantics, the same API calls. We can quickly switch between different providers," Singh said. </p><p>Walmart built its own internal gateway to stay vendor agnostic across three workload types: fully deterministic workflows, planner-and-reasoner workflows for open-ended tasks, and a hybrid of the two. Compliance-heavy work stays deterministic by design; governance, security and evaluation run through the gateway regardless of which model is on the other end. Gosby said the choice between a frontier model and an open-weight model comes down to whichever is most effective for the specific workload, not a fixed policy.</p><h2>Advice for the modernization journey</h2><p>Three pieces of advice came up directly, each tied to the wall a leader had already hit.</p><p><b>Invest in evals before anything else.</b> Ghoche called it the thing common to every use case, internal or customer facing. </p><p>"The thing that's common to all of these is evals. It'll force you to break the problem down, and once you have a robust set of evals, you can move a lot faster," he said, </p><p><b>Own your agent harness from day one.</b> Gosby's advice was to put the AI harness directly in employees' hands early, paired with the infrastructure to monitor what it produces. </p><p>"It will unlock a huge amount of innovation," she said.</p><p><b>Build for model and context independence.</b> Ensuring flexibility is critical for success.</p><p>"Build for independence, whether it's a frontier model of today versus an open source model of tomorrow," Singh said. "Keep that context within your enterprise so that you can reuse it when you ship the model or the harness tomorrow," Singh said.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The build vs. buy dilemma at the heart of enterprise AI]]></title>
<description><![CDATA[For three decades, enterprise software has been a buy-it decision. Packaged software from SAP, Oracle and Salesforce covered roughly 80% of requirements at a fraction of the cost of building. The economics were obvious, and for traditional applications, they still are.



AI is introducing a wrin...]]></description>
<link>https://tsecurity.de/de/3675706/it-nachrichten/the-build-vs-buy-dilemma-at-the-heart-of-enterprise-ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3675706/it-nachrichten/the-build-vs-buy-dilemma-at-the-heart-of-enterprise-ai/</guid>
<pubDate>Fri, 17 Jul 2026 12:17:08 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">For three decades, enterprise software has been a buy-it decision. Packaged software from SAP, Oracle and Salesforce covered roughly 80% of requirements at a fraction of the cost of building. The economics were obvious, and for traditional applications, they still are.</p>



<p class="wp-block-paragraph">AI is introducing a wrinkle that is forcing even the most committed enterprise software customers to rethink their options. AI is a layer that sits across your data, your processes, and your decisions. Where that layer runs and who controls it is an architecture question, and most of the enterprise community is still treating it as a procurement one.</p>



<p class="wp-block-paragraph">The appeal of vendor-embedded AI is clear: automated operational decisions, smarter supplier and merchandising choices, and friction-free workflows built into the systems enterprises already rely on. The catch is that these capabilities almost universally depend on your data living in the vendor’s cloud environment. For most large enterprises, it sits on-premises, in hyperscale cloud infrastructure they manage themselves, or in private data centers. That gap between where your data is and where your vendor’s AI assumes it should be creates a fundamental strategic fork in the road.</p>



<h2 class="wp-block-heading"><a></a>Build vs. buy is a category error</h2>



<p class="wp-block-paragraph">The framing I keep hearing is “build vs. buy your AI strategy.” It implies that some organizations are out there training foundation models from scratch. Nobody serious is doing that. The real choice sits across three distinct approaches, and conflating them leads to poor decisions:</p>



<ul class="wp-block-list">
<li><strong>Buy embedded. </strong>Use the AI capabilities your vendor ships natively inside their platform: the assistant baked into your ERP, your CRM, your HCM suite. Lowest integration cost, fastest time to value, tightest fit with the application data.</li>



<li><strong>Buy platform.</strong> Adopt the vendor’s AI infrastructure layer and build your own assistants and agents on top of it. More flexible, but you remain inside the vendor’s architectural boundary and subject to their governance model.</li>



<li><strong>Compose.</strong> Connect a third-party model (Claude, GPT, Gemini, an open-weight model running in your own environment) directly to your existing landscape. Maximum control, maximum integration burden, and full responsibility for what comes out the other end.</li>
</ul>



<p class="wp-block-paragraph">These are not equivalent options at different price points. They make different assumptions about where your data lives, who governs the AI, and how much architectural change you’ll absorb to get there. Vendor pitches sometimes blur the distinction on purpose. Enterprise leaders can’t afford to.</p>



<h2 class="wp-block-heading"><a></a>The vendor AI stack has an assumption baked in</h2>



<p class="wp-block-paragraph">Every embedded AI capability ships with an unstated architectural prerequisite: your data must be where the AI can see it, in the shape it expects, under the governance the vendor enforces.</p>



<p class="wp-block-paragraph">For organizations with clean, modern cloud estates, that is often a reasonable trade. For the long tail of large enterprises running heavily customized environments on private or hybrid infrastructure, that trade becomes a precondition, one you must meet before the AI conversation can even begin. Whether meeting it makes sense depends on your starting point, your sector’s regulatory posture, and your appetite for migration risk. None of those are uniform across organizations.</p>



<p class="wp-block-paragraph">That’s the part that gets glossed over in vendor keynotes. The AI demo on stage assumes a destination architecture the audience hasn’t necessarily reached yet. Large enterprise customers are carrying an unusually heavy technology burden right now. Many are simultaneously managing platform modernization programs that have been building for over a decade, alongside pressure to migrate to vendor-managed cloud infrastructure. Sitting above both is a boardroom-level directive to demonstrate meaningful AI progress fast. The vendor path to AI and the boardroom path to AI can diverge sharply, and enterprises need to make selective, strategic decisions about where to adopt AI first to maximize value and minimize risk.</p>



<h2 class="wp-block-heading"><a></a>Sovereignty isn’t a slogan, it’s an architecture constraint</h2>



<p class="wp-block-paragraph">The conversation about sovereignty has been hijacked by both sides. One camp treats every SaaS adoption as a sovereignty violation. The other dismisses every sovereignty concern as Luddite resistance. Neither is useful.</p>



<p class="wp-block-paragraph">What’s happening in real customer conversations – particularly in DACH, public sector, and financial services – is more specific. Organizations are drawing a distinction between running their applications in a vendor’s cloud (which is broadly fine, well understood, decades of precedent) and enriching their data and processes inside a vendor’s AI model (which has less precedent, is harder to reverse, and carries material implications for competitive position).</p>



<p class="wp-block-paragraph">Enriching your data inside a vendor’s AI model is the genuinely new question, and organizations that conflate it with their existing cloud posture tend to defend the wrong perimeter.</p>



<p class="wp-block-paragraph">Despite spending around $100 million annually with Amazon, <a href="https://www.uctoday.com/unified-communications/disney-openai-enterprise-strategy/">Disney built its own internal AI system</a> to house its corporate intelligence rather than rely on a hyperscaler’s AI offering. The decision came down to control. When your data represents decades of creative and commercial IP, you think carefully about where it lives and who can learn from it. Disney has become more open to SaaS over time. The AI sovereignty question is a separate debate from the SaaS debate and conflating the two leads organizations to the wrong conclusions.</p>



<p class="wp-block-paragraph">At the other end of the spectrum, enterprises in heavily regulated environments treat data sovereignty as an absolute non-negotiable. Any AI model must run within their controlled environment, especially where sensitive data cannot touch the public internet.<a href="https://gdpr.eu/what-is-gdpr/"> </a><a href="https://gdpr.eu/what-is-gdpr/">GDPR obligations</a> reinforce this instinct across the European market, requiring organizations to maintain clear accountability for how personal data is processed inside AI systems, including vendor-managed ones.</p>



<p class="wp-block-paragraph">AI-enriched data, meaning models that have learned the shape of your business processes, your supplier negotiations, your customer behavior, carries a different half-life and a different strategic value than the operational data underneath it. That deserves its own architectural decision, separate from your broader cloud strategy.<a></a></p>



<h2 class="wp-block-heading">What this means in practice</h2>



<p class="wp-block-paragraph">Most large enterprise estates will end up with a mix of all three approaches, and where you draw the lines matters more than your overall posture.</p>



<p class="wp-block-paragraph">Embedded AI capabilities are the right answer for in-application productivity: the assistant inside your ERP workflows, the agent inside your procurement or HR suite. That is where vendor embedding genuinely shines, and attempting to compose your own equivalent is typically a poor use of engineering resources.</p>



<p class="wp-block-paragraph">Compose belongs elsewhere: in cross-application orchestration, in custom assistants over operational and observability data, and in agents that need to reach across multiple vendor systems and infrastructure layers in ways no single vendor stack will never natively support. <a href="https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-top-trends-in-tech">Research from McKinsey</a> suggests the most significant near-term productivity gains from enterprise AI will come precisely from these cross-system workflows, rather than from within individual applications. The most interesting enterprise AI work over the next eighteen months lives here, and it doesn’t require waiting for a migration to complete first.</p>



<p class="wp-block-paragraph">That compose path isn’t free, and it’s important to be honest about the costs. Governance, audit trails, and accountability for hallucinated outputs become your problem, not the vendor’s. Prompt drift and evaluation discipline are real engineering costs that never appear in the proof-of-concept. Those costs scale with the complexity of your landscape and the number of systems your agents touch. Budget for them before deployment, not after your first production incident. None of that is a reason to avoid the path. It’s a reason to staff for it, honestly.<a></a></p>



<h2 class="wp-block-heading">The real question</h2>



<p class="wp-block-paragraph">The build-vs-buy frame survives because it gives executives a binary choice along a familiar axis. AI sits somewhere else entirely.</p>



<p class="wp-block-paragraph">The question worth putting on the table at your next architecture review is simpler:</p>



<p class="wp-block-paragraph">Which decisions do we want our vendors’ AI to make, and which do we want to keep on our side of the boundary?</p>



<p class="wp-block-paragraph">Answer that, and the right build/buy/compose mix flows from it. Skip it, and you will end up with the architecture your vendors prefer – which may or may not be the one your business needs.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[4 Memory-Systeme, um KI aufzuschlauen]]></title>
<description><![CDATA[Wenn Ihre KI unter unzureichender „Gedächtnisleistung“ leidet, helfen diese Memory-Systeme von Drittanbietern (eventuell).DC Studio | shutterstock.com



KI-Agenten und die Large Language Models (LLMs), auf denen sie basieren, haben ein eher kurzlebiges „Gedächtnis“. Das ist so gewollt, schließli...]]></description>
<link>https://tsecurity.de/de/3675019/it-security-nachrichten/4-memory-systeme-um-ki-aufzuschlauen/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3675019/it-security-nachrichten/4-memory-systeme-um-ki-aufzuschlauen/</guid>
<pubDate>Fri, 17 Jul 2026 06:08:30 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" src="https://b2b-contenthub.com/wp-content/uploads/2025/10/DC-Studio_shutterstock_2269121373_DEOnly_16z9.jpg?quality=50&amp;strip=all&amp;w=1024" alt="Dev Meeting 16z9" class="wp-image-4075633" width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"><figcaption class="wp-element-caption">Wenn Ihre KI unter unzureichender „Gedächtnisleistung“ leidet, helfen diese Memory-Systeme von Drittanbietern (eventuell).</figcaption></figure><p class="imageCredit">DC Studio | shutterstock.com</p></div>



<p class="wp-block-paragraph"><a href="https://www.computerwoche.de/article/4189343/was-ki-agenten-wirklich-kosten.html" target="_blank">KI-Agenten</a> und die Large Language Models (<a href="https://www.computerwoche.de/article/4155050/25-fragen-die-zum-richtigen-llm-fuhren.html" target="_blank">LLMs</a>), auf denen sie basieren, haben ein eher kurzlebiges „Gedächtnis“. Das ist so gewollt, schließlich kann nur eine begrenzte Menge an Konversationsinhalten in Token kodiert und vom LLM zuverlässig abgerufen werden. Um KI-Agenten und Sprachmodelle mit „Hirnschmalz“ auszustatten, das über ihre Kontextfenster hinausreicht, lässt sich Retrieval Augmented Generation (<a href="https://www.computerwoche.de/article/4192090/so-geht-memory-optimierung-bei-ki-agenten.html" target="_blank">RAG</a>) einsetzen. Erfolgsentscheidend ist dabei, wie dieser Mechanismus (oder ein anderer, um Gesprächsdaten vorzuhalten) konkret zur Anwendung kommt.</p>



<p class="wp-block-paragraph">Ein anderer Weg, sowohl KI-Agenten als auch LLMs mit erweiterten Speicherfähigkeiten auszustatten, führt über Software-Tools von Drittanbietern. Diese können die KI mit einer Session-übergreifenden, persistenten Memory ausstatten. Auch hier variiert jedoch die Art und Weise, wie das technisch umgesetzt wird. Die folgenden vier Projekte sind besonders empfehlenswert, wenn es darum geht, KI-Agenten und Sprachmodelle smarter zu machen.    </p>



<h2 class="wp-block-heading">1. <a href="https://github.com/getzep/graphiti" target="_blank" rel="noreferrer noopener">Graphiti</a></h2>



<p class="wp-block-paragraph">Graphiti wird als „das Open-Source-Framework für temporale Knowledge-Graphen“ beworben. Das Projekt ist auf GitHub verfügbar – oder auch im Rahmen des <a href="https://www.getzep.com/" target="_blank" rel="noreferrer noopener">Memory-Service Zep</a>, für den es die Grundlage liefert. „Temporal“ bedeutet in diesem Zusammenhang, dass die in Graphiti gespeicherten Informationen im Laufe der Zeit reevaluiert werden, um den Kontext korrekt einzubetten. Der Begriff „Graph-Framework“ ist hingegen darauf zurückzuführen, dass die Daten dabei als eine Reihe von Graphen gespeichert werden. Dieses Feature spielt auch bei den anderen in diesem Artikel vorgestellten Lösungen eine Rolle – im Fall von Graphiti steht es allerdings im Fokus.</p>



<p class="wp-block-paragraph">Out of the Box unterstützt das KI-Memory-Projekt eine ganze Reihe gängiger LLMs, etwa von Anthropic, OpenAI, Google oder X. Auch sämtliche Ollama- und OpenAI-kompatiblen <a href="https://www.computerwoche.de/article/4004872/die-besten-apis-um-ki-zu-integrieren.html" target="_blank">APIs</a> funktionieren mit Graphiti – es kann also auch mit <a href="https://www.computerwoche.de/article/2830445/5-wege-llms-lokal-auszufuehren.html" target="_blank">lokal gehosteten LLMs</a> genutzt werden. Daten aus Quellen wie GitHub, Gmail und OneDrive sowie aus Anwendungen wie Notion lassen sich über Konnektoren einbinden.</p>



<p class="wp-block-paragraph">Um Graphiti lokal nutzen zu können, ist es allerdings nötig, eine Graphdatenbank einzurichten oder eine Verbindung zu einer solchen herzustellen. Die Standardlösung dafür (mit dem breitesten Support) ist <a href="https://neo4j.com/" target="_blank" rel="noreferrer noopener">Neo4j</a>. Davon abgesehen, funktionieren auch <a href="https://aws.amazon.com/neptune/" target="_blank" rel="noreferrer noopener">Amazon Neptune</a>, <a href="https://www.falkordb.com/" target="_blank" rel="noreferrer noopener">FalkorDB</a> und <a href="https://kuzudb.github.io/" target="_blank" rel="noreferrer noopener">KuzuDB</a>. <a href="https://www.computerwoche.de/article/3803224/postgresql-als-rag-vektordatenbank-nutzen.html" target="_blank">Postgres</a> mit <code>pgvector</code> ist (derzeit) hingegen keine Option bei Graphiti.</p>



<h2 class="wp-block-heading">2. <a href="https://hindsight.vectorize.io/" target="_blank" rel="noreferrer noopener">Hindsight</a></h2>



<p class="wp-block-paragraph">Das KI-Memory-Projekt Hindsight als Cloud Service verfügbar, kann jedoch auch lokal gehostet werden. Dieses Tool speichert Details zu Agenten-Sitzungen in <a href="https://hindsight.vectorize.io/#key-components">vier verschiedenen Memory-Instanzen</a> und wendet dabei vier unterschiedliche <a href="https://hindsight.vectorize.io/#multi-strategy-retrieval-tempr" target="_blank" rel="noreferrer noopener">Storage- und Retrieval-Strategien</a> an. Diese werden über drei programmatische Interfaces gehändelt:</p>



<ul class="wp-block-list">
<li><code>retain</code>, um Inhalte (einzelne Fakten oder komplette Sessions) zu speichern,</li>



<li><code>recall</code>, um den Content abzurufen, und</li>



<li><code>reflect</code>, um einen Agenten-Loop über eine Abfrage zu initiieren, die zuvor gespeicherte Daten nutzt.</li>
</ul>



<p class="wp-block-paragraph">In Sachen Integrationen hat Hindsight eine breite Palette von First- und Third-Party-Optionen <a href="https://hindsight.vectorize.io/integrations" target="_blank" rel="noreferrer noopener">zu bieten</a>. Wenn Sie beispielsweise die „Continue“-Erweiterung mit Visual Studio Code einsetzen, um mit einem lokal gehosteten LLM zu kommunizieren, können Sie die <a href="https://hindsight.vectorize.io/sdks/integrations/continue" target="_blank" rel="noreferrer noopener">entsprechende First-Party-Integration</a> nutzen. In diesem Fall verwenden Sie einfach das Keyword <code>@hindsight</code> in der Query, um den Agenten-Kontext um relevante Memory zu erweitern. Um sich die Arbeit zu erleichtern, respektive diese zu automatisieren, könnten Sie außerdem auch auf (anpassbare) Auto-Injection-Regeln zurückgreifen.</p>



<h2 class="wp-block-heading">3. <a href="https://github.com/mem0ai/mem0" target="_blank" rel="noreferrer noopener">Mem0</a></h2>



<p class="wp-block-paragraph">Wie Hindsight nutzt auch Mem0 <a href="https://docs.mem0.ai/core-concepts/memory-types" target="_blank" rel="noreferrer noopener">vier grundlegende Memory-Typen</a> – allerdings sind diese anders benannt und organisiert. Beispielsweise kommt im Fall von Mem0 die sogenannte „Organizational Memory“ zum Einsatz, um Daten zu speichern, die zwischen verschiedenen KI-Agenten(-Teams) geteilt werden sollen.</p>



<p class="wp-block-paragraph">Jede Form von Memory, die über Mem0 hinzugefügt wird, durchläuft einen „<a href="https://docs.mem0.ai/core-concepts/memory-evaluation#memory-extraction-distillation" target="_blank" rel="noreferrer noopener">Destillationsprozess</a>“ und wird auf unterschiedliche Art und Weise (Vektor-, Graph- oder SQL-Datenbank) gespeichert. Ältere Daten werden bei Mem0 nicht gelöscht, sondern als veraltet markiert – eine Strategie, um einen umfassenderen, längerfristigen Kontext zu erzeugen.</p>



<p class="wp-block-paragraph">Das Projekt unterstützt im Vergleich – etwa zu Hindsight – weniger LLMs, die wichtigen Anbieter (Anthropic, Google, OpenAI) sind jedoch vertreten. Dazu kommen Self-Hosting-Optionen über <a href="https://www.computerwoche.de/article/2827054/was-ist-langchain.html" target="_blank">LangChain</a>, <a href="https://www.litellm.ai/" target="_blank" rel="noreferrer noopener">LiteLLM</a>, <a href="https://www.computerwoche.de/article/4131576/lm-studio-angetestet.html" target="_blank">LM Studio</a> und <a href="https://ollama.com/" target="_blank" rel="noreferrer noopener">Ollama</a>. Falls Sie Mem0 lokal statt <a href="https://mem0.ai/pricing" target="_blank" rel="noreferrer noopener">als Service</a> nutzen möchten, ist es nötig, eine Python-Instanz und eine eigene Vektordatenbank bereitzustellen. Für Letzteres ist Postgres mit der <code>pgvector</code>-Erweiterung eine gängige und simple Option, die sogar innerhalb einer virtuellen Python-Umgebung <a href="https://github.com/orm011/pgserver" target="_blank" rel="noreferrer noopener">installiert werden kann</a>.</p>



<h2 class="wp-block-heading">4. <a href="https://supermemory.ai/" target="_blank" rel="noreferrer noopener">Supermemory</a></h2>



<p class="wp-block-paragraph">Supermemory erfasst Daten aus vielen gängigen Quellen und unterstützt dabei unter anderem Plaintext, strukturierte Daten, PDF- und Office-Dokumente sowie Video-, Audio- und Bilddateien. Aus diesen Informationen erstellt das Tool einen Kontextgraphen, der anschließend als Grundlage für Chatbot-Konversationen fungiert. PR-mäßig setzt dieses Projekt den Fokus vor allem auf seine Context-Extraktions-Tools.</p>



<p class="wp-block-paragraph">Supermemory ist entweder als Cloud-Dienst oder als quelloffene, lokal ausführbare Software verfügbar. Die <a href="https://github.com/supermemoryai/supermemory" target="_blank" rel="noreferrer noopener">Open-Source-Version</a> lässt zwar die Scaling Services und Drittanbieter-Konnektoren der Enterprise-Version vermissen – hat jedoch einen entscheidenden Vorteil: Sie besteht aus einer einzelnen <a href="https://www.computerwoche.de/article/4128783/4-self-contained-datenbanken-fur-entwickler.html" target="_blank">Self-Contained</a>-Binary. So lässt sie sich auch auf der eigenen Hardware mit sehr überschaubarem Aufwand bereitstellen.</p>



<p class="wp-block-paragraph">Da für dieses Projekt zudem keine externen Datenbanken aufgesetzt werden müssen, eignet es sich in besonderem Maße für (agile) Experimente. (fm)</p>



<p class="wp-block-paragraph"><strong>Dieser Artikel ist <a href="https://www.infoworld.com/article/4192397/four-agentic-ai-memory-systems-for-smarter-llms.html" target="_blank">im Original</a> bei unserer Schwesterpublikation Infoworld.com erschienen.</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems]]></title>
<description><![CDATA[Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 — a 2.8-trillion-parameter model that the company says is now the largest open-source AI model in the world, and one that benchmarks show performs neck-and-neck with the most powerful pr...]]></description>
<link>https://tsecurity.de/de/3674665/it-nachrichten/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-us-systems/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3674665/it-nachrichten/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-us-systems/</guid>
<pubDate>Thu, 16 Jul 2026 23:17:55 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="https://www.moonshot.ai/">Moonshot AI,</a> the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">Kimi K3</a> — a 2.8-trillion-parameter model that the company says is now the largest open-source AI model in the world, and one that benchmarks show performs neck-and-neck with the most powerful proprietary systems from <a href="https://www.anthropic.com/">Anthropic</a> and <a href="https://openai.com/">OpenAI</a>.</p><p>The release, timed to land just ahead of the <a href="https://aiii.global/waic-2026/">2026 World Artificial Intelligence Conference</a> in Shanghai, is a dramatic escalation in the global AI arms race and a watershed moment for the open-source AI movement. It also marks a remarkable comeback for a company whose market position had eroded significantly over the past 18 months following DeepSeek's meteoric rise.</p><p>Full model weights are scheduled to be released on July 27, according to details shared by researchers who reviewed the company's technical documentation. If you want to take <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">Kimi K3</a> for a spin right now, you can — just head to<a href="https://www.kimi.com/"> kimi.com</a>, sign up with a Google account or phone number (no credit card required), and start chatting with what may be the most powerful open-source model ever built.</p><div></div><h2><b>Inside the architecture that powers the world's largest open-source AI model</b></h2><p><a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">Kimi K3</a> is a frontier-class large language model with 2.8 trillion total parameters — roughly 75 percent larger than <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">DeepSeek's V4 Pro</a>, which the company's own timeline chart shows at approximately 1.6 trillion parameters. The model features a 1-million-token context window, native visual understanding capabilities, and an always-on reasoning mode that the company calls "thinking mode."</p><p>The model is built on two key architectural innovations developed internally at Moonshot AI: <a href="https://arxiv.org/abs/2510.26692">Kimi Delta Attention</a>, a hybrid linear attention mechanism, and <a href="https://arxiv.org/abs/2603.15031">Attention Residuals</a>, which the company describes as a drop-in replacement for residual connections that delivers consistent scaling gains. Both techniques were previously published as open research by the Moonshot team on <a href="https://github.com/moonshotai">GitHub</a>.</p><p>On the <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">API side</a>, Kimi K3 is compatible with the <a href="https://developers.openai.com/api/docs/guides/agents">OpenAI SDK</a>, lowering the integration barrier for developers already building on OpenAI or Anthropic toolchains. The model is priced at $3 per million input tokens and $15 per million output tokens, with cached input tokens dropping to just $0.30 per million — pricing that positions it roughly in line with mid-tier offerings from Western labs, but at a performance level the company claims approaches the top of the market. A promotional top-up rebate running through August 12 offers up to 30 percent back in vouchers for API credits of $1,000 or more.</p><p>As <a href="https://finance.sina.com.cn/stock/t/2026-07-17/doc-inihzrtu1375218.shtml?cref=cj">Xinhua reported</a>, a Moonshot AI executive explained the significance of the parameter count in simple terms: parameters are like neural connections in the human brain, and nearly 3 trillion of them means the model can "store more knowledge and patterns in its brain, understand more, think deeper, and answer more accurately."</p><div></div><h2><b>Benchmark results show Kimi K3 trading blows with Claude and GPT at the top of the leaderboard</b></h2><p>The benchmark results, drawn from public leaderboard data and a private evaluation by analytics firm Artificial Analysis, tell a striking story.</p><p>On <a href="https://artificialanalysis.ai/evaluations/gdpval-aa">GDPval-AA v2</a>, a benchmark measuring real-world tasks across 44 occupations and 9 major industries, Kimi K3 scored 1,687 — placing it third overall, behind only Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 (1,600).</p><p>On <a href="https://artificialanalysis.ai/evaluations/aa-briefcase">AA-Briefcase</a>, a private agentic benchmark from Artificial Analysis designed to test long-horizon knowledge work, K3 climbed to second place with a score of 1,527 — beating GPT-5.6 Sol Max (1,495) and trailing only Fable 5 Max (1,587).</p><p>Perhaps most impressively, K3 achieved a state-of-the-art score of 91.2 out of 100 on <a href="https://openai.com/index/browsecomp/">BrowseComp</a>, a benchmark for long-horizon, high-difficulty information seeking. </p><p>The company says it accomplished this in a single-agent setup using its 1-million-token context window, without any context compression or additional context management techniques — a feat that suggests raw context length, when paired with strong retrieval capabilities, may be more powerful than elaborate multi-agent workarounds.</p><p>As <a href="https://x.com/kimmonismus/status/2077818040578695175">one widely followed AI commentator</a> put it on social media: "Open source is no longer lagging six months behind Western closed-source models. Read that again, and think about what it all means."</p><p>That observation captures the significance of the moment. For much of the past three years, open-source models have typically trailed their proprietary counterparts by a meaningful margin. Kimi K3 appears to have closed that gap almost entirely.</p><h2><b>How a 48-hour autonomous chip design demo reveals Moonshot's real ambitions</b></h2><p>Beyond raw benchmarks, <a href="https://www.moonshot.ai/">Moonshot AI</a> showcased a proof-of-concept that may be even more revealing of K3's capabilities and the company's strategic direction.</p><p>In a demonstration documented in the company's technical materials, <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">Kimi K3</a> was tasked with designing a physical chip to run a nano-scale version of itself. Over 48 hours of continuous autonomous agent operation, K3 independently completed the chip's full construction pipeline — from architectural design through optimization and verification — using open-source electronic design automation tools. The result was a tiny but functional chip design, just 4 square millimeters, that achieved timing convergence at 100 MHz and could decode more than 8,700 tokens per second in simulation.</p><p>This is not a production chip. It is a demonstration of what <a href="https://www.moonshot.ai/">Moonshot AI</a> clearly views as the next competitive frontier: long-range autonomous agent capabilities. The ability to sustain coherent, multi-step technical work over a 48-hour window — reading documentation, making design decisions, running verification loops, and iterating on failures — represents a qualitative leap beyond the kind of single-turn question-answering that defined the first generation of large language models.</p><p>The company also highlighted a case in computational astrophysics, where K3 reportedly reproduced the universal <a href="https://inspirehep.net/literature/1220233">I-Love-Q relation</a> — a complex calculation that typically takes a senior researcher one to two weeks — in approximately two hours, reading and cross-validating more than 20 papers and implementing a complete numerical pipeline along the way.</p><h2><b>Moonshot AI's fall and rise tells the story of China's brutal AI market</b></h2><p>To understand why <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">Kimi K3</a> matters, you need to understand where Moonshot AI was 18 months ago — and how far it fell.</p><p>Founded in 2023 by <a href="https://kimiyoung.github.io/">Yang Zhilin</a>, a Tsinghua University graduate who previously conducted research at Google and Meta, Moonshot AI quickly became one of China's most prominent AI startups. The company gained early traction in 2024 when users flocked to its <a href="http://kimi.ai/">Kimi platform</a> for its long-text analysis capabilities and AI search functions. By early 2026, it had raised roughly <a href="https://www.forbes.com/sites/the-prompt/2026/07/15/ai-startup-reflection-compute-deal-to-challenge-chinas-open-source-dominance/">$1.5 billion</a> across multiple rounds, with its valuation climbing from $2.5 billion to $4.3 billion and the company reportedly <a href="https://tech.yahoo.com/ai/gemini/articles/china-moonshot-releases-open-source-141110760.html">seeking a new round at $5 billion</a>.</p><p>Then DeepSeek happened. The release of DeepSeek's low-cost R1 model in January 2025 disrupted the entire Chinese AI landscape, and Moonshot AI was among the hardest hit. Kimi, which had ranked third in monthly active users in China, slid to seventh. The company's strategic pivot to open-source models — beginning with Kimi K2 in July 2025 and accelerating with K2.5 in January 2026 — was in large part an effort to reclaim relevance.</p><p><a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">Kimi K3</a> is the culmination of that effort — and the sheer scale of the model suggests that Moonshot AI has been planning this move for some time. Training a 2.8-trillion-parameter model requires enormous computational resources and months of preparation, which means the architectural and infrastructure decisions behind K3 were likely locked in well before the model reached the public.</p><h2><b>Why open-sourcing the world's biggest model is a geopolitical chess move</b></h2><p>The decision to release K3's full weights on July 27 is strategically significant and worth parsing carefully.</p><p>The company's own timeline chart of open-source frontier model scale positions K3 as a dramatic outlier, towering above competitors like <a href="https://github.com/deepseek-ai">DeepSeek</a> (1.6T), <a href="https://github.com/xiaomi">Xiaomi</a> (1.02T), and <a href="https://github.com/ALIBABA">Alibaba</a> (397B). By releasing the world's largest open-source model, Moonshot AI is making a bid to become the center of gravity for the global open-source AI developer community.</p><p>This follows a broader trend among Chinese AI companies. As <a href="https://www.reuters.com/technology/artificial-intelligence/china-weighs-silicon-curtain-around-sought-after-ai-models-2026-07-08/">Reuters noted</a>, open-sourcing allows companies to "showcase their technological capabilities and expand developer communities as well as their global influence, a strategy likely to help China counter U.S. efforts to limit Beijing's tech progress." DeepSeek, Alibaba, Tencent, and Baidu have all released open-source models. But none have released anything at this parameter count.</p><p>For enterprise technology leaders, the implications are concrete. A 2.8-trillion-parameter open-source model that performs at near-frontier levels creates new options for companies that want to fine-tune, self-host, or build proprietary systems on top of a capable base model — without being locked into API contracts with OpenAI or Anthropic. The trade-off, of course, is that running a model of this size requires substantial GPU infrastructure. Inference at 2.8 trillion parameters is not something that runs on a single server rack.</p><p>That said, <a href="https://www.moonshot.ai/">Moonshot AI</a> has signaled awareness of this challenge. Its Mooncake project, which won the Best Paper award at FAST 2025, pioneered KV-cache-centric disaggregated serving for large language models — an architecture designed specifically to make inference at extreme scale more practical and cost-efficient.</p><h2><b>Kimi Code and a three-tier model lineup form the foundation of Moonshot's enterprise play</b></h2><p>Alongside K3, Moonshot AI continues to invest heavily in its coding agent ecosystem. <a href="https://github.com/MoonshotAI/kimi-code/releases">Kimi Code</a>, the company's open-source coding tool that competes with Anthropic's Claude Code and Google's Gemini CLI, received two major updates on the same day as K3's launch — versions 0.25.0 and 0.26.0 — adding features like expanded subagent tooling, background task management, and security fixes.</p><p>The <a href="https://github.com/MoonshotAI/kimi-cli">Kimi Code CLI</a> has accumulated over 3,100 stars on GitHub and features integration with VSCode, Cursor, and Zed. The latest release expanded the "coder subagent" tool set to include background tasks, todo lists, plan mode, skill invocation, and nested agents — effectively turning the coding agent into a multi-layered autonomous system capable of managing complex software engineering projects with minimal human intervention.</p><p>This is not incidental. Coding tools have become a critical revenue driver for AI labs. As Anthropic disclosed in January, <a href="https://www.anthropic.com/news/anthropic-acquires-bun-as-claude-code-reaches-usd1b-milestone">Claude Code reached $1 billion in annualized recurring revenue</a>. By building Kimi Code as an open-source alternative that defaults to Kimi's own models — but supports other providers — Moonshot AI is positioning itself to capture developer workflows and, eventually, enterprise contracts.</p><p>The company's model lineup now includes three tiers: <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart">K3</a> as the flagship ($3/$15 per million tokens for input/output), <a href="https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstart">K2.7 Code</a> as a specialized coding model ($0.95/$4), and <a href="https://platform.kimi.ai/docs/guide/kimi-k2-6-quickstart">K2.6</a> as a general-purpose option ($0.95/$4). All three support context windows of 256,000 tokens or above, with K3 offering the full 1-million-token window. Context caching is automatic — no cache ID, TTL, or extra parameter is required — a small but meaningful developer-experience advantage over competitors that require explicit cache management.</p><h2><b>What Kimi K3 means for the future of enterprise AI and the global model landscape</b></h2><p>Kimi K3's release forces a recalibration of several assumptions that have guided enterprise AI strategy.</p><p>The performance gap between open-source and proprietary models has functionally closed at the frontier. If K3's benchmark numbers hold up under independent evaluation — and particularly once the open weights are available for community testing on July 27 — it will be difficult for closed-source providers to justify premium pricing purely on the basis of capability.</p><p>The locus of AI innovation, meanwhile, continues to shift. China's AI ecosystem, which many Western observers questioned after early struggles with chip export restrictions, has now produced a model that competes with the best systems from companies with direct access to Nvidia's most advanced hardware. The architectural innovations behind K3 — particularly the hybrid linear attention mechanism — suggest that algorithmic efficiency may matter as much as raw compute.</p><p>And the agentic capabilities demonstrated by K3 — chip design, multi-week research compression, long-horizon information seeking — point toward a future where AI models are not just answering questions but autonomously executing complex, multi-day projects. For enterprises evaluating AI investments, this shifts the value proposition from "productivity copilot" to "autonomous technical workforce."</p><p><a href="https://finance.sina.com.cn/stock/t/2026-07-17/doc-inihzrtu1375218.shtml?cref=cj">Xinhua</a>, China's state news agency, framed the release as a national milestone, reporting that K3 "marks a new step forward in the development of China's artificial intelligence models." Liu Tieyan, dean of the Zhongguancun Academy in Beijing, was quoted as saying that a wave of Chinese open-source models has moved from isolated breakthroughs to collective advancement, providing "new solutions and new paths" for global AI development.</p><p>Just two years ago, <a href="https://www.moonshot.ai/">Moonshot AI</a> was a scrappy startup named for the audacious problems it hoped to solve. Eighteen months ago, it was a cautionary tale about how quickly a market darling can lose its footing. Today, it is the maker of the world's largest open-source AI model — one that can, given 48 hours and an internet connection, design a chip to run itself. The frontier, it turns out, is not a place. It is a race. And the field just got a lot more crowded.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials]]></title>
<description><![CDATA[Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents s...]]></description>
<link>https://tsecurity.de/de/3674536/it-nachrichten/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3674536/it-nachrichten/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials/</guid>
<pubDate>Thu, 16 Jul 2026 21:47:26 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents still share credentials; and only three in ten isolate their highest-risk agents. The security stack is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents, spending remains a thin slice of the security budget, and enterprises are evenly split on whether their defenses are keeping pace with AI-enabled attackers. The result is an agent security gap — autonomous agents proliferating faster than the identity, isolation, and enforcement controls needed to hold them.</p><p>This wave of VentureBeat Pulse Research examines how enterprises secure their AI agents: what tooling they run, how they manage agent identity and isolation, what has already gone wrong, how much they spend, and whether they believe their defenses are keeping pace with AI-enabled attackers.</p><p>The central finding is an agent security gap — the distance between the autonomy enterprises are granting their agents and the controls in place to contain them. More than half of organizations (54%) have already experienced a confirmed agent security incident (18%) or a near-miss caught before harm (36%). The structural weakness beneath those numbers is identity: only about a third (32%) give every agent its own scoped, managed identity, while the rest report that some agents share credentials or that agents mostly run on shared API keys and human or service-account credentials. When agents share credentials, a single compromised or over-permissioned agent carries a wide blast radius — and only three in ten enterprises (30%) isolate their highest-risk agents in sandboxes to bound that radius.</p><p>What makes the gap notable is how comfortable enterprises are inside it. The security stack is overwhelmingly provider-native — OpenAI’s guardrails (51%), Google’s and Microsoft’s cloud controls, and Anthropic’s managed-agent controls dominate, while the dedicated agent-security specialists barely register — and satisfaction with that borrowed stack is high, averaging 4.2 out of 5. Yet spending remains a thin slice of the security budget, only a third of enterprises believe their AI defenses are ahead of AI-enabled attackers, and a clear majority plan to change tooling within the year. Enterprises are satisfied with controls they are simultaneously preparing to replace.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent security — the tooling, identity, isolation, and enforcement controls organizations use to secure autonomous AI agents. Responses are filtered to organizations with more than 100 employees (n=107; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single June 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.</p><p>By role the sample is senior and buyer-credible: 45% are final decision-makers for AI purchases and another 30% recommenders or influencers. Managers (43%), individual contributors (24%), VPs and directors (15%), and the C-suite (11%) make up the seniority mix. By organization size the sample is mid-market-weighted: 251–1,000 (42%) and 101–250 (25%) employees lead, with 1,001–5,000 (19%), 5,001–10,000 (8%), and 10,001+ (7%) above them. Technology/Software is the largest industry at 23%, followed by Manufacturing (15%), Retail/E-commerce (14%), and Healthcare/Life Sciences (13%).</p><p>At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent security rather than from the largest operators.</p><p>Satisfaction ratings are computed on the respondents who answered each rating question; the overall satisfaction score reflects 82 of the 107 qualified respondents.</p><h2>Finding 1: The incidents are already here</h2><p><b>More than half have had an agent security incident or near-miss</b></p><p>We asked whether organizations had experienced an agent security incident — a confirmed breach, or a near-miss caught before harm. Most that run agents in production had.</p><div></div><p>This is the report’s defining number. More than half of organizations (54%) have already had an agent security event — 18% a confirmed incident and 36% a near-miss caught before it caused harm. Only 42% report nothing, and a small remainder either run no agents in production or don’t track such events. That so many report near-misses rather than only confirmed incidents is telling: enterprises are catching problems, but they are catching them close to the edge. The controls examined in the rest of this report — identity, isolation, enforcement — are what determine whether the next near-miss stays a near-miss.</p><p>Exposure scales with company size, but containment does not. The incident-or-near-miss rate rises from 49% in the mid-market (companies with 101-1,000 employees) to 63% at larger enterprises (above 1,000 employees), while sandbox isolation of high-risk agents falls from 35% to 20%, and satisfaction with security tooling drops from 4.36 to 3.97. The organizations running the most agents across the most systems carry the most incidents and the least of the one control that bounds an incident's blast radius.</p><h2>Finding 2: The identity gap</h2><p><b>Only a third give every agent its own scoped identity</b></p><p>We asked how enterprises manage the identity of their AI agents — whether each agent has its own credentials, or agents share them. Full per-agent identity is the exception.</p><div></div><p>Rolled together, the overlapping answers show 69% of enterprises (74 of 107) with credential sharing somewhere in the agent fleet. Identity is the structural weakness beneath the incidents. Only about a third of enterprises (32%) give every agent its own scoped, managed identity — the precondition for least-privilege access and clean attribution. Nearly half (48%) say some agents have scoped identities but many still share credentials, and another 32% say agents mostly run on shared API keys or borrowed human and service-account credentials. (Respondents could describe more than one pattern across their agent fleet, so these overlap.) </p><p>The consequence is direct: when agents share credentials, an over-permissioned or compromised agent can act with far more reach than intended, and forensics after an incident cannot cleanly tell which agent did what. The non-human identity problem — giving every agent its own governed identity — is the single largest unfinished piece of enterprise agent security.</p><p>Moreover, a company’s agent credential posture is correlated with incidents. Organizations with credential sharing anywhere in the fleet were hit — with an incident or a near-miss in the past twelve months — at 63.5% (47 of 74). Organizations where every agent carries its own scoped identity were hit at 40.9% (9 of 22). The fully-scoped group is small, so for now the relationship is an association rather than proven causation, and the gap is concentrated in the mid-market — but within a single survey, a twenty-three point difference in incident rate suggests significance.</p><h2>Finding 3: Observe and enforce, but rarely isolate</h2><p><b>Only three in 10 sandbox their highest-risk agents</b></p><p>We asked what an organization’s agent security posture looks like in practice — whether they observe, enforce, isolate, or some combination. The control that bounds damage is the least common.</p><div></div><p>Monitoring and enforcement are reasonably common; containment is not. Roughly half of enterprises observe agent activity (47%) or enforce scoped permissions at runtime (49%), but only 30% isolate their highest-risk agents in sandboxes that bound the blast radius when the other controls fail. That ordering is backwards from a defense-in-depth standpoint: observation tells you what happened, enforcement tries to prevent it, but isolation is what limits the damage when prevention fails — and it is the control enterprises have adopted least. Combined with the identity gap in Finding 2, the picture is of agents that are watched and permissioned but rarely boxed in, which is precisely the configuration in which a single failure propagates.</p><h2>Finding 4: Security runs on borrowed, provider-native controls</h2><p><b>Guardrails from OpenAI, Google and Microsoft dominate; specialists barely register</b></p><p>We asked which agent security tooling enterprises use, and which is their primary layer. The answer favors the model providers and hyperscalers over the dedicated security vendors.</p><div></div><p>Enterprises are securing agents with tools that came bundled with their models and clouds. OpenAI’s guardrails lead at 51%, followed by Google’s and Microsoft’s cloud-native controls and Anthropic’s managed-agent controls — and when asked to name their single primary security layer, 82% name one of these provider-native offerings. The purpose-built agent-security category — Palo Alto’s Prisma AIRS, CrowdStrike, Cisco AI Defense, Zenity, HiddenLayer, Check Point’s Lakera, Okta for AI Agents, non-human identity platforms — barely registers, each in the low single digits, and only 5% run no dedicated tooling at all. As with retrieval and evaluation elsewhere in this series, the provider bundle is winning the default: enterprises reach first for the guardrails their platform ships, and the independent security layer that would address the identity and isolation gaps has not yet been adopted at scale.</p><p>The provider-default pattern is consistent across both Q2 survey waves. In April–May (n=110), usage was led by the same names — OpenAI's controls at 26%, Azure at 15%, AWS at 14%, Google at 12% — with every dedicated agent-security specialist at 3% or below and one in ten using no dedicated tooling at all. The common finding from the two surveys: Enterprises are defaulting to the solutions provided by the platform they’re using, and the specialist category vendors have yet to become big players here.</p><p>(<i>A note on reading these shares. As described in the methodology section, the respondent sample is self-selected and skews mid-market, and the usage question counted every vendor or approach a respondent has in place — so the figures measure presence in the security stack rather than spending or exclusivity. Individual vendor percentages therefore carry all the usual sample caveats. The structural pattern, however, held across both Q2 waves on two differently worded questions: provider-native and hyperscaler controls lead, and dedicated agent-security specialists remain in low single digits. Read the individual shares loosely and the pattern with confidence.)</i></p><h2>Finding 5: And enterprises are comfortable with it</h2><p><b>Satisfaction is high, even as incidents mount and identity lags</b></p><p>We asked how satisfied enterprises are with their current agent security tooling. The comfort is notably out of step with the exposure documented above.</p><div></div><p>Satisfaction with agent security tooling is high — 4.2 out of 5 overall, and 4.1 for value for money — among the most positive readings in this series. That is the striking part: enterprises are highly satisfied with a stack that is mostly borrowed provider guardrails, even though more than half have already had an incident or near-miss and only a third give their agents scoped identities. The comfort appears to rest on the convenience and low friction of provider-native controls rather than on demonstrated containment. It is a false comfort in the making — the same enterprises expressing satisfaction are, as Finding 8 shows, a clear majority planning to change tooling within the year, which suggests the confidence is thinner than the score implies.</p><h2>Finding 6: Budgets haven’t caught up</h2><p><b>Most spend under a tenth of the security budget on agents</b></p><p>We asked what share of the security budget enterprises allocate to securing AI agents. For a fast-emerging risk, the allocation is modest.</p><div></div><p>Spending on agent security is still a thin slice. The most common allocation is 6–10% of the security budget (46%), and a third of enterprises (34%) spend 5% or less; only a quarter (24%) devote more than a tenth. Given the incident rate in Finding 1 and the identity and isolation gaps in Findings 2 and 3, the budget looks like a lagging indicator — the risk has arrived faster than the funding to address it. The enterprises spending more than a tenth of their security budget on agents are a distinct minority, and they are likely the ones building the scoped-identity and isolation controls the rest have not.</p><h1>Finding 7: The arms race is even, at best</h1><p><b>Only a third think their AI defenses are ahead of AI-enabled attackers</b></p><p>We asked how enterprises assess the balance between their AI-enabled defenses and AI-enabled attackers. Confidence is far from settled.</p><div></div><p>Enterprises are split on whether they are winning. Only about a third (35%) believe their AI-enabled defenses are ahead of AI-enabled attackers; the rest are less sure — 32% call it roughly even, 21% think attackers are ahead, and another 21% say it is too early to tell. Taken together, a clear majority (53%) rate the balance as even or tilted toward the attacker. That uncertainty sits uneasily beside the high satisfaction of Finding 5: enterprises are content with their tooling yet unconvinced it is winning the contest it exists to win. In a domain where the offense is also compounding with AI, an even race is not a comfortable place to be.</p><h2>Finding 8: A security reshuffle is coming</h2><p><b>Nearly six in 10 plan to adopt or switch tooling within a year</b></p><p>We asked whether enterprises plan to adopt a new, additional, or replacement agent security solution, and which they are considering. Few intend to stand pat.</p><div></div><p>The security stack is not settled. While 41% have no plans to change, a clear majority (59%) intend to adopt a new, additional, or replacement agent security solution within twelve months, and 29% within the next quarter — a strong signal that, high satisfaction notwithstanding, enterprises know the current stack is provisional. Incidents are what start the buying cycle. </p><p>Among organizations that have been hit, 42.1% plan to adopt, add, or replace agent security tooling within the next ninety days, against 14.0% of organizations with no incident — and after a confirmed incident it becomes majority behavior, at 52.6%. Getting hit also changes the threat assessment: 33.3% of hit organizations say AI-armed attackers are ahead of their defenses, against 8.0% of the unhit. Experience, in this data, is the strongest predictor of both urgency and pessimism.</p><p>The consideration set still leans provider-native (OpenAI 34%, Google 30%, Anthropic 29%, Azure 25%), but the dedicated security vendors — Cloudflare, Cisco, Palo Alto, Okta, Check Point’s Lakera — draw early interest in the mid-to-high single digits, more than their current footprint. </p><p>What the shopping does not yet include is the identity layer specifically. Twelve percent of the respondents include an agent-identity product — Okta for AI Agents, Microsoft Entra Agent ID, or a non-human identity platform — anywhere in their consideration set, and among the credential-sharing organizations that have already had an incident, identity consideration is essentially unchanged, at roughly one in ten. The control most directly implicated by the incident data is the one largely missing from the purchase plans. Whether this wave hardens the provider-native default or finally opens the door to purpose-built agent security — the identity and isolation controls the incidents call for — is the question this series will keep tracking.</p><h2>The bottom line: A security gap that autonomy will test first</h2><p>Organizations with more than 100 employees are giving AI agents real reach into systems and data while securing them with controls built for something else. More than half have already had an incident or near-miss; only a third give every agent its own scoped identity, and most still share credentials; only three in ten isolate their highest-risk agents; and the stack doing this work is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents.</p><p>The uncomfortable pairing is confidence with exposure: satisfaction with the current tooling is among the highest in this series, yet spending is a thin slice of the security budget, only a third believe their defenses are ahead of AI-enabled attackers, and a clear majority are already planning to replace what they have. At 107 respondents in a single wave this is a directional read, skewed toward the mid-market — but the direction is clear: agent adoption is running ahead of agent security, and the controls that matter most when something fails — scoped identity and isolation — are the ones enterprises have built least. The agent security gap is not a coverage problem that a provider guardrail will close on its own; it is a problem of identity, isolation, and enforcement built for autonomous software. The open question for later waves is whether enterprises close it deliberately — or whether a confirmed incident closes it for them.</p><hr><p><i>Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read, not a precise measurement — the sample is self-selected and skews mid-market, so it's best read as the view from organizations actively standing up agent security rather than from the largest operators. Respondents are senior and buyer-credible (45% final decision-makers, 30% recommenders/influencers), spanning managers through the C-suite, and drawn primarily from Technology/Software, Manufacturing, Retail/E-commerce, and Healthcare/Life Sciences.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs]]></title>
<description><![CDATA[Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today...]]></description>
<link>https://tsecurity.de/de/3674337/it-nachrichten/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3674337/it-nachrichten/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs/</guid>
<pubDate>Thu, 16 Jul 2026 20:02:38 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price — which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilization or less, and fewer than half rigorously track what their compute actually costs. The result is a compute gap — heavy, fast-moving investment running ahead of the visibility needed to control it.</p><p>This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how satisfied they are, what would make them switch, where they plan to evaluate their investments, and — most revealingly — how well they can measure and control the economics of the compute underneath it all.</p><p>The central finding is a compute gap — the distance between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity: the single largest planned area enterprises plan to evaluate over the next year is AI-specialized clouds (45%), a layer almost none of these enterprises use today. Meanwhile the compute already in place runs cold — 83% report GPU utilization of 50% or less — and fewer than half (44%) can rigorously track what their AI compute costs. Enterprises are buying more infrastructure faster than they can account for what they already own.</p><p>Enterprises are not settled on their infrastructure vendors, either: A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter — unusually high churn intent for a category this foundational. When they choose, they choose on integration with the existing stack (41%) and total cost of ownership (35%), not on headline price: cost per million tokens is the deciding factor for just 8%. And the frontier constraint that will shape the next round of decisions — the shift from GPU compute to memory bandwidth as inference scales — is barely on the radar, with roughly one in five enterprises either unaware of it or yet to address it.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=107; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.</p><p>By organization size the sample concentrates in the mid-market: 101–250 employees (36%) and 251–1,000 (27%) lead, with 1,001–5,000 (22%), 5,001–10,000 (8%), and 10,001+ (7%) above them. By role it spans managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%); on purchasing authority it is buyer-credible, with 45% final decision-makers and another 30% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 26%, followed by Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%).</p><p>At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It also skews toward the mid-market and toward earlier-stage adopters, so it is best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators.</p><h2>Finding 1: Ambition outpaces production</h2><p><b>Only one in five run AI in production at scale</b></p><p>We asked where organizations sit in their AI deployment journey. Most are still building toward production rather than operating at scale.</p><div></div><table><tbody><tr><td><p><b>38%</b></p></td><td><p><b>are experimenting — running proofs of concept, not yet in production</b></p></td></tr><tr><td><p><b>37%</b></p></td><td><p><b>have some workloads in production, but not across the organization</b></p></td></tr><tr><td><p><b>21%</b></p></td><td><p><b>run AI in production at scale — the mature minority</b></p></td></tr><tr><td><p><b>4%</b></p></td><td><p><b>are not yet running AI workloads at all</b></p></td></tr></tbody></table><p>The maturity curve is front-loaded. Three-quarters of enterprises (76%) are either experimenting or running only some workloads in production, and just 21% describe AI in production at scale. This matters for everything that follows: the infrastructure decisions in this report are being made largely by organizations still early in deployment, whose compute footprint — and whose costs — are about to grow. The evaluation and switching intentions in Findings 3 and 4 are the leading edge of that build-out, not the settled preferences of operators who have already found what works.</p><h2>Finding 2: Enterprises run on hyperscalers and model APIs</h2><p><b>The specialized GPU clouds barely register — today</b></p><p>We asked which providers and platforms enterprises currently use to run their AI. The answer is a familiar one: the incumbents.</p><div></div><table><tbody><tr><td><p><b>48%</b></p></td><td><p><b>use Google Cloud — the most-used platform overall (Microsoft Azure 29%, AWS 22%, Oracle Cloud 22%)</b></p></td></tr><tr><td><p><b>41%</b></p></td><td><p><b>use Google’s Gemini models, with OpenAI close behind at 40% and Anthropic at 12%</b></p></td></tr><tr><td><p><b>6%</b></p></td><td><p><b>run their own on-prem or co-located GPU clusters; 4% a custom open-source self-managed stack</b></p></td></tr><tr><td><p><b>&lt;2%</b></p></td><td><p><b>each use the specialized AI clouds — CoreWeave, Lambda, Crusoe, Nebius, Together, Fireworks and peers</b></p></td></tr></tbody></table><p>The current stack is hyperscaler-and-API. Google Cloud leads at 48%, and the general-purpose clouds (Google, Microsoft, AWS, Oracle) together with the major model APIs (Gemini, OpenAI, Anthropic) account for essentially all current deployment. The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines — CoreWeave, Lambda, Crusoe, Nebius and peers — register at or near zero among these enterprises today. Only 6% run their own on-prem GPU clusters and 4% a custom open-source stack. Enterprises are, for now, running AI on the providers they already buy from — which makes the evaluation intentions in Finding 3 all the more striking.</p><p><i>(A note on reading these shares. As described in the methodology section, this sample is self-selected and skews mid-market, and this question counted every provider a respondent uses — an average of 2.1 selections each — so the figures measure presence in the stack rather than spending or primary status. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; Google's strength here, for example, is consistent with its long-standing position among smaller enterprises building on AI. Read these shares as a portrait of what this AI-active cohort runs today, and treat gaps between these figures and industry-wide market share estimates as a property of the sample rather than a contradiction of either.)</i></p><h2>Finding 3: The next dollar goes to infrastructure they don’t yet run</h2><p><b>AI-specialized clouds top the evaluations list</b></p><p>We asked where enterprises planned to evaluate AI infrastructure over the next 12 months. Their answers point away from the stack they run today.</p><div></div><table><tbody><tr><td><p><b>45%</b></p></td><td><p><b>AI-specialized clouds (CoreWeave, Lambda, Crusoe, Nebius) — the top planned evaluation area</b></p></td></tr><tr><td><p><b>32%</b></p></td><td><p><b>non-NVIDIA accelerators (AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi, in-house ASICs)</b></p></td></tr><tr><td><p><b>28%</b></p></td><td><p><b>Nvidia Blackwell (GB300) / next-generation GPUs</b></p></td></tr><tr><td><p><b>16%</b></p></td><td><p><b>decentralized or distributed compute networks</b></p></td></tr><tr><td><p><b>11%</b></p></td><td><p><b>sovereign or region-specific compute; 9% say none of the above</b></p></td></tr></tbody></table><p>Here is the report’s sharpest tension. The single most-cited planned evaluation area — AI-specialized clouds, at 45% — is the very category almost none of these enterprises use today (Finding 2). Nearly a third (32%) intend to evaluate non-Nvidia accelerators, and 28% in next-generation Nvidia silicon; even decentralized compute networks (16%) and sovereign compute (11%) draw meaningful interest. Read against current usage, this is not incremental — it is the leading edge of a re-platforming. The direction-of-travel question tells the same story: every infrastructure approach is net-expanding, but specialized AI clouds carry the highest net momentum (+24), edging out even the hyperscalers (+22). Enterprises are preparing to move a meaningful share of AI compute off the general-purpose cloud.</p><p>This continues a trend we saw in our April-May survey wave. Back then, usage of the AI-specialized clouds was equally marginal — CoreWeave at 3%, Lambda at 4%, Crusoe at 2% of enterprises. When we asked enterprises what change they planned in their AI infrastructure strategy over the next twelve months, the most-cited answer was moving workloads to specialized AI clouds, at 33%. Asked in April-May which emerging compute option they were most likely to evaluate AI-specialized clouds again drew the most responses. Two waves, two differently worded questions, one consistent picture: the type of cloud enterprises are most eager to assess is the type they have barely begun to use.</p><h2>Finding 4: A switching wave is building</h2><p><b>Six in 10 plan to change providers within a year — many within a quarter</b></p><p>We asked whether and when enterprises plan to switch or add an infrastructure provider. Very few intend to stand still.</p><div></div><table><tbody><tr><td><p><b>38%</b></p></td><td><p><b>plan to change within the next 0–3 months — tied for the most common answer</b></p></td></tr><tr><td><p><b>36%</b></p></td><td><p><b>have no plans to change</b></p></td></tr><tr><td><p><b>22%</b></p></td><td><p><b>plan to change within 3–6 months</b></p></td></tr><tr><td><p><b>7%</b></p></td><td><p><b>plan to change within 6–12 months</b></p></td></tr></tbody></table><p>For a category as foundational as compute, this is a remarkable amount of intended movement. Only 36% have no plans to change, meaning a clear majority (64%) intend to switch or add a provider within twelve months — and 38% within the next quarter alone. Where that interest points is telling: the providers drawing the most switching consideration are again the incumbents — Microsoft Azure and Google Cloud (33% each), OpenAI (30%), and Gemini (22%) — which suggests much of the near-term movement is reshuffling among the majors and consolidating spend rather than defecting to new entrants. The neocloud interest in Finding 3 is a 12-month evaluation thesis; the switching in the next quarter is mostly incumbents trading share.</p><p>(<i>Method note: Respondents who selected both "no plans to change" and a specific switching window are counted as switchers, on the logic that naming a timeframe is the more specific answer; three respondents were reclassified under this rule.</i>)</p><h2>Finding 5: Nobody buys on token price</h2><p><b>Integration and total cost of ownership decide — not sticker price</b></p><p>We asked what matters most when enterprises select an AI infrastructure provider. Headline price finished last.</p><div></div><table><tbody><tr><td><p><b>41%</b></p></td><td><p><b>integration with the existing cloud and data stack — the top factor</b></p></td></tr><tr><td><p><b>35%</b></p></td><td><p><b>total cost of ownership (TCO)</b></p></td></tr><tr><td><p><b>24%</b></p></td><td><p><b>performance — latency and throughput</b></p></td></tr><tr><td><p><b>19%</b></p></td><td><p><b>each cite security/compliance, autoscaling for spiky workloads, and GPU access/availability</b></p></td></tr><tr><td><p><b>8%</b></p></td><td><p><b>cost per 1M tokens — the least-cited factor</b></p></td></tr></tbody></table><p>Enterprises do not buy AI infrastructure on pricing, which is the place vendors compete on hardest. Integration with the existing stack (41%) and total cost of ownership (35%) dominate, while the headline metric — cost per million tokens — is the deciding factor for just 8%, dead last. The pattern is coherent: buyers are optimizing for how a provider fits and what it truly costs to operate, not for the advertised unit rate. It also foreshadows Finding 7 — enterprises say TCO matters most, yet most cannot yet measure it rigorously. The stated priority and the measured capability are out of step.</p><h2>Finding 6: Expensive GPUs, idle most of the time</h2><p><b>83% report GPU utilization of 50% or less</b></p><p>We asked what share of their GPU capacity enterprises actually utilize. The answer is a well-known but rarely quantified inefficiency.</p><div></div><table><tbody><tr><td><p><b>37%</b></p></td><td><p><b>run at 26–50% utilization</b></p></td></tr><tr><td><p><b>34%</b></p></td><td><p><b>run at 10–25% utilization</b></p></td></tr><tr><td><p><b>15%</b></p></td><td><p><b>run under 10% utilization</b></p></td></tr><tr><td><p><b>12%</b></p></td><td><p><b>run over 50% — the efficient minority</b></p></td></tr><tr><td><p><b>8%</b></p></td><td><p><b>don’t measure utilization at all; a further 7% consume via API and run no GPUs of their own</b></p></td></tr></tbody></table><p><i>Disclosure: Band percentages count every selection against all 107 qualified respondents; 14 respondents selected more than one band, so bands overlap. At the respondent level, 83 of the 100 GPU-operating enterprises reported utilization at or below 50%</i></p><p>The compute already in place runs cold. Adding the bands at or below half capacity, 83% of enterprises that operate GPUs report utilization of 50% or less, and nearly half (49%) run at 25% or below. Only 12% clear the 50% mark, and a further 8% do not measure utilization at all. Idle accelerators are expensive accelerators, and this is the clearest single measure of the compute gap: enterprises are planning to buy more GPUs and specialized compute (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large — and largely unmeasured.</p><h2>Finding 7: Spending fast, measuring slowly</h2><p><b>Fewer than half rigorously track what their compute costs</b></p><p>We asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger lags the spending.</p><div></div><table><tbody><tr><td><p><b>44%</b></p></td><td><p><b>track compute cost and ROI rigorously</b></p></td></tr><tr><td><p><b>39%</b></p></td><td><p><b>track it only partially</b></p></td></tr><tr><td><p><b>20%</b></p></td><td><p><b>can’t quantify it yet</b></p></td></tr><tr><td><p><b>6%</b></p></td><td><p><b>say it isn’t a priority</b></p></td></tr></tbody></table><p>Measurement trails money. Fewer than half of enterprises (44%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (20%), or have not prioritized it (6%). That gap is consequential given Finding 5, where total cost of ownership was the second-ranked buying criterion — enterprises are choosing providers on an economic basis they mostly cannot yet measure. Satisfaction with current infrastructure is moderately positive but not enthusiastic: on a five-point scale, overall satisfaction averages 4.0, with ease of implementation (3.8) and value for money (3.9) trailing slightly — the softness landing, tellingly, on cost. Enterprises are spending quickly and accounting slowly.</p><h2><b>Finding 8: The next bottleneck few are watching</b></h2><p><b>As inference shifts from compute to memory, the field scatters</b></p><p>Finally, we asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The responses reveal a frontier that is not yet a priority.</p><div></div><table><tbody><tr><td><p><b>31%</b></p></td><td><p><b>would rely on Dell (PowerScale / Project Lightning) — the leading single answer</b></p></td></tr><tr><td><p><b>16%</b></p></td><td><p><b>would rely on Nvidia (Dynamo / ICMSP)</b></p></td></tr><tr><td><p><b>18%</b></p></td><td><p><b>are not aware of this as a constraint (9%) or haven’t addressed inference-memory limits yet (8%)</b></p></td></tr><tr><td><p><b>10%</b></p></td><td><p><b>Hammerspace (Tier Zero); 9% DDN (Infinia); the rest split across open-source KV-cache tooling, model-level efficiency, VAST Data, and WEKA</b></p></td></tr></tbody></table><p>The memory frontier is real but barely governed. Asked which approach they would rely on as the binding constraint in inference shifts from compute to memory bandwidth, enterprises scatter: Dell leads at 31%, Nvidia follows at 16%, and the rest fragments across storage vendors, open-source tooling, and model-level efficiency techniques. Most telling is that roughly one in five (18%) either do not recognize the constraint or have not begun to address it. For a shift that will reshape inference cost and architecture, this is an early and unsettled market — and, consistent with the measurement gap in Finding 7, one where many enterprises simply do not yet have a view. It is the next chapter of the compute gap, arriving before most have closed the current one.</p><h1><b>The bottom line: A compute gap that faster spending will widen, not close</b></h1><p>Organizations with more than 100 employees are investing in AI infrastructure faster than they can measure it. Most are still early in deployment, yet their spending intentions point past their current stack — toward specialized clouds and alternative accelerators almost none of them run today — and a clear majority intend to change providers within the year. They buy on integration and total cost of ownership rather than headline price, which is rational; the difficulty is that most cannot yet see those economics clearly.</p><p>The visibility gap is concrete. The GPUs enterprises already own run at half utilization or less for the overwhelming majority, and fewer than half can rigorously track what their compute costs or returns. Satisfaction is decent but unenthusiastic, softest on value for money — the dimension hardest to judge without measurement. And the next constraint, the shift from compute to memory in large-scale inference, is arriving while most enterprises are still unaware of it. At 107 respondents in a single Q2 wave this is a directional read, skewed toward the mid-market and earlier-stage adopters — but the direction is consistent: the appetite to spend is running well ahead of the instrumentation to spend well. The compute gap is not a capacity problem that more hardware will solve on its own; it is, first, a problem of seeing what the hardware already costs. The open question for later waves is whether enterprises build that visibility before the re-platforming arrives — or buy the next layer of infrastructure as blind to its economics as the last.</p><hr><p><i>Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the results read cross-sectionally rather than as a month-over-month trend, and at 107 respondents this is a directional signal rather than a precise measurement — the sample is self-selected, skews mid-market, and leans toward earlier-stage adopters rather than the largest hyperscale operators. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with buyer-credible purchasing authority, across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway]]></title>
<description><![CDATA[Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated...]]></description>
<link>https://tsecurity.de/de/3674237/it-nachrichten/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3674237/it-nachrichten/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway/</guid>
<pubDate>Thu, 16 Jul 2026 19:03:24 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap — the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures.</p><p>This wave of VentureBeat Pulse Research examines how technical leaders measure agent performance: which reliability and evaluation platforms they use, how they select and trust them, what breaks in production, and how far they are willing to let agents run without a human in the loop.</p><p>The central finding is an evaluation gap — the distance between the autonomy enterprises are granting their agents and the trust they place in the evaluations meant to govern it. Half of organizations (50%) have, in the past year, deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure, and a quarter have seen it happen more than once. Trust in the tests themselves is thin: only 5% say they fully trust automated evaluation today, and the single most-cited limitation is that evaluations align poorly with real-world outcomes (29%). Enterprises are discovering that a passing eval is not the same as a working agent.</p><p>What makes the gap consequential is the direction of travel. Two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve months (33%). At the same time, the evaluation stack that would have to earn that trust is fragmented and immature: the most common primary tools are the model providers’ native evals, tied with having no dedicated tooling at all (17% each); and only about a quarter of enterprises run real-time quality checks on live production traffic. The autonomy is arriving faster than the assurance.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey — the Agentic Reliability &amp; Evals tracker — focused on how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=157), drawn from a single survey in June 2026; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Where questions were multiple-select, those shares can sum to more than 100%.</p><p>By role the sample is senior and buyer-credible: 38% are final decision-makers for AI purchases and another 34% recommenders or influencers. Product and program managers (15%), consultants and advisors (10%), directors of engineering/IT (8%), and CIOs/CTOs/CISOs (8%) lead the named titles, alongside a large “Other” function (37%). By organization size the sample is mid-market-weighted: 100–499 (37%) and 500–2,499 (27%) employees lead, with 2,500–9,999 (20%), 10,000–49,999 (10%), and 50,000+ (6%) above them. Technology/Software is the largest industry at 23%, followed by Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%).</p><p>At 157 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent evaluation practices rather than from the largest operators.</p><p><i>Note: This survey was rebuilt for the June wave from the earlier “LLM observability and evaluations” survey; because the questions and sample differ, no comparisons are made to the April–May data.</i></p><h1>Finding 1: A passing eval is not a working agent</h1><p><b>Half have shipped an agent that passed evals, then failed a customer</b></p><p>We asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. Half of those that run evaluations had.</p><div></div><p>This is the report’s defining number. Half of organizations (50%) have shipped an AI feature that cleared their internal evaluations and then failed in front of a customer — an incorrect output, a broken workflow, or a quality incident — and a quarter have seen it happen more than once. Only 36% report no such failure, and the remainder either run no pre-deployment evaluations (8%) or don’t track the root cause closely enough to know (6%). The failure is precise and expensive: the evaluation said the agent was ready, and it was not. Everything that follows — how enterprises trust their evals, what they monitor, and how much autonomy they grant — is shaped by this experience.</p><h2>Finding 2: Almost no one fully trusts automated evaluation</h2><p><b>The top complaint: Evals don't match real-world outcomes</b></p><p>We asked which limitation most reduces trust in automated agent evaluations today. Only a sliver of enterprises had no complaint at all.</p><div></div><p>Trust in automated evaluation is scarce, and specific. Only 5% of organizations say they fully trust automated evaluation as it stands — meaning 95% name a limitation that holds them back. The most common, at 29%, is the one that most directly explains Finding 1: evaluations align poorly with real-world outcomes, passing agents that later fail. Bias or inconsistency (21%) and a lack of explainability (18%) follow — enterprises cannot always tell why an evaluation reached its verdict — and 17% cite data-leakage or privacy concerns in the evaluation process itself. The tests meant to certify agents are not yet trusted to certify them, which is precisely why the autonomy trajectory in Finding 3 is so striking.</p><h2>Finding 3: The autonomy ceiling is rising anyway</h2><p><b>Two-thirds already allow, or are building toward, zero-human deployment</b></p><p>We asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The trajectory runs straight through the trust gap.</p><div></div><p>Here is the paradox at the heart of the report. Even though almost no one fully trusts automated evaluation (Finding 2), two-thirds of organizations (66%) either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within a year (33%). Only 22% rule it out for the foreseeable future. The direction is unambiguous: enterprises are moving to let evaluations gate production autonomously — removing the human check — at the same moment they say those evaluations don’t reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it, which is the mechanism by which the false-confidence failures of Finding 1 will scale rather than shrink.</p><p>Notably, the autonomy bet is not just a small company phenomenon. Splitting the sample by company size, larger enterprises are slightly further down the path toward zero human review than smaller companies (70% versus 64%) and slightly more likely to have shipped an evaluation-passing agent that then failed a customer (54% versus 48%). The assumption that large, regulated organizations are holding the human in the loop longest is, in this sample, backwards.  To be sure, these are directional figures, since the survey was not a huge sample — 57 respondents from companies with 2,500+ employees and 100 from companies smaller than that. </p><h2>Finding 4: The evaluation stack is fragmented and provider-led</h2><p><b>Provider-native evals lead — tied with no dedicated tool at all</b></p><p>We asked which agent reliability or evaluation platform enterprises primarily use today. The market has no clear leader — and a large share has nothing dedicated.</p><div></div><p>The evaluation layer is early and unconsolidated. Provider-native tooling leads — OpenAI’s native evals and traces (17%) and Anthropic’s Claude Console evals (13%) together outweigh any independent platform — but it is tied at the top by a striking answer: 17% of enterprises use no dedicated agent-evaluation tooling at all, a notable gap for organizations shipping agents to customers. The specialist evaluation vendors — DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize — are scattered across single to low double digits, and 11% have built their own. No independent platform has yet become the category standard, which leaves most enterprises evaluating agents with provider-native tools, home-grown scripts, or nothing.</p><h2>Finding 5: Production monitoring rarely watches output quality</h2><p><b>Only a quarter run real-time quality checks on live traffic</b></p><p>Production monitoring for an AI agent can watch two very different things. It can watch whether the system is <b>functioning</b> — is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent's output is <b>correct</b> — automated checks that evaluate the content of each answer as it goes out: did the agent give the right answer, take the right action, stay within policy. The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked organizations which kind their live production monitoring is built for today.</p><div></div><p>Grouped by what is actually being watched, the split is stark: 51% of organizations monitor only whether the agent is functioning, while 23% monitor whether its answers are right. Counting the ad-hoc reviewers and the don't-knows, roughly three-quarters of organizations run no automated, real-time evaluation of output correctness in production — they can see that the system is up and what it costs, and they are taking the correctness of its answers on faith. That blind spot is the runtime counterpart to the pre-deployment gap in Finding 1: the same organizations engineering the human out of the deployment decision mostly cannot see, in real time, when the deployed agent starts getting things wrong.</p><h2>Finding 6: Bought on cost, measured on consistency</h2><p><b>Price and integration drive selection; evaluation consistency is the goal</b></p><p>We asked what most influenced enterprises’ choice of an evaluation vendor, and what they treat as their primary measure of success. Both answers are pragmatic.</p><div></div><p>Enterprises buy evaluation tooling on economics and trust it on repeatability. Cost of evaluations (28%) narrowly leads selection, just ahead of ease of integration (27%) and evaluation accuracy (24%) — breadth of observability (13%) and vendor roadmap (4%) matter far less. On what success looks like, more than a third (36%) name evaluation consistency — getting the same verdict on the same behavior every time — well ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%), and compliance (11%). The emphasis on consistency is telling: before enterprises can trust an evaluation’s verdict, they need it to be stable — the very property whose absence (bias and inconsistency) ranked among the top trust limitations in Finding 2. Satisfaction with current tooling is only moderate, averaging 3.8 on a five-point scale across overall satisfaction, ease of implementation, and value for money.</p><h2>Finding 7: The next dollar goes to humans and observability</h2><p><b>Investment is flowing to oversight, not just automation</b></p><p>We asked which reliability and evaluation investment will grow most over the next year. The money is going toward watching agents more closely — including with people.</p><div></div><p>The second-largest planned investment — behind only production observability — is human review workflows, at 26%. Read against Finding 1, that is the report's quietest contradiction: at the same moment two-thirds of enterprises are engineering the human out of the deployment decision, more of them plan to grow spending on human reviewers (26%) than on the automated evaluation pipelines (16%) that would replace them. The zero-human trajectory and the human-review budget are rising in the same companies at the same time. Indeed, only 8% report that their budget is not increasing. </p><p>Taken together, enterprises are hedging: building toward autonomy while spending to watch agents more closely and keep humans available for the calls that automated evaluation cannot yet be trusted to make.</p><h2>Finding 8: A tooling reshuffle is coming</h2><p><b>Nearly two-thirds plan to adopt or switch platforms within a year</b></p><p>We asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Few intend to stand pat.</p><div></div><p>The evaluation market is wide open. While 36% have no plans to change, a clear majority (64%) intend to adopt a new, additional, or replacement platform within twelve months, and 31% within the next quarter. The consideration set points where current usage is thinnest: Confident AI’s DeepEval leads what enterprises are evaluating (20%), ahead of OpenAI’s native evals (13%) and Braintrust (9%) — the open-source specialists drawing more interest than their present footprint. </p><p>Given that so many enterprises today rely on provider-native tools or nothing at all (Finding 4), this is less a defection than a first real wave of tooling adoption — the moment the evaluation layer starts to consolidate. Which platforms earn that trust, in a market where almost no one trusts automated evaluation yet, is the open question this series will keep tracking.</p><h2>The bottom line: An evaluation gap that autonomy will widen, not close</h2><p>Organizations with 100 or more employees are granting AI agents more independence than they trust their evaluations to support. Half have already shipped an agent that passed its evals and then failed a customer; almost none fully trust automated evaluation, chiefly because it doesn’t match real-world outcomes; and most watch production for uptime and cost rather than for whether the agent’s answers are right. Yet two-thirds already allow, or are actively building toward, deploying to production on automated evaluation alone.</p><p>The vendor market is early and unsettled: the most common primary evaluation tools are provider-native evals, tied with no dedicated tooling at all, and a clear majority plan to adopt or switch platforms within the year. Encouragingly, the next dollar is going to observability and — pointedly — human review, suggesting enterprises sense the gap even as they engineer past it. At 157 respondents in a single wave this is a directional read, skewed toward the mid-market — but the direction is clear: autonomy is being granted on the strength of evaluations that the people granting it do not yet trust. The evaluation gap is not a coverage problem that more tests alone will close; it is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy — or whether the false-confidence failures move from customer incidents into changes that deploy themselves.</p><hr><p><i>Based on survey responses from 157 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read rather than a precise measurement — the sample is self-selected, not a probability sample, and skews toward the mid-market. Respondents include product and program managers, consultants and advisors, directors of engineering/IT, and CIOs/CTOs/CISOs, among other functions, across technology/software, retail/consumer, healthcare/life sciences, manufacturing, and other industries.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[DeepMind CEO pushes for AI industry self-regulation]]></title>
<description><![CDATA[Google DeepMind CEO Demis Hassabis is pushing for the US AI industry to self-regulate, with the support of government, as a starting point for an international creating shared international standards. In a blog post, he called for a focus on artificial general intelligence (AGI) and national secu...]]></description>
<link>https://tsecurity.de/de/3673460/it-nachrichten/deepmind-ceo-pushes-for-ai-industry-self-regulation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3673460/it-nachrichten/deepmind-ceo-pushes-for-ai-industry-self-regulation/</guid>
<pubDate>Thu, 16 Jul 2026 14:33:47 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Google DeepMind CEO Demis Hassabis is pushing for the US AI industry to self-regulate, with the support of government, as a starting point for an international creating shared international standards. In a blog post, he called for a focus on <a href="https://www.computerworld.com/article/4174181/google-talks-singularity-while-scaling-up-agentic-ai-for-enterprises-2.html">artificial general intelligence (AGI)</a> and national security. </p>



<p class="wp-block-paragraph">But it is precisely that focus on national security that may make the results of such an effort, assuming it happens, less than palatable outside of the US.</p>



<p class="wp-block-paragraph">“The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous,” <a href="https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age" target="_blank" rel="noreferrer noopener">Hassabis wrote</a>. “The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”</p>



<p class="wp-block-paragraph">He noted, however, that the funding would need to be substantial, and would most likely come from industry, to allow the new body to attract world-class technical talent and obtain the necessary compute resources for large-scale testing.</p>



<p class="wp-block-paragraph">Hassabis proposed that the organization “be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security,” and that AI vendor participants be encouraged to adopt best practices such as publishing model cards with technical details, maintaining strong internal cybersecurity, vetting key personnel, and providing sufficient resourcing for safety and security research.</p>



<p class="wp-block-paragraph">This is not the first time Hassabis has <a href="https://www.computerworld.com/article/4178398/deepmind-ceo-agi-could-be-here-in-three-years.html" target="_blank">expressed worries about AGI</a>. </p>



<p class="wp-block-paragraph">DeepMind was involved in an earlier <a href="https://www.cio.com/article/4168122/us-government-agency-to-safety-test-frontier-ai-models-before-release.html" target="_blank">US government initiative evaluating AI safety</a>, alongside Microsoft and xAI (now SpaceXAI) working with the Center for AI Standards and Innovation (CAISI), a division of the US Department of Commerce. It allowed CAISI to conduct pre-deployment evaluations and targeted research to “better assess frontier AI capabilities and advance the state of AI security.”  </p>



<h2 class="wp-block-heading">The rest of the world may have concerns</h2>



<p class="wp-block-paragraph">Analysts and consultants were mixed about the move, with most expressing concerns about whether an industry-focused group would prioritize the public’s best interests.</p>



<p class="wp-block-paragraph">“Self-regulation is not viable because it implies everyone is able to regulate themselves and will do so in line with the best interests of the public. Most tech vendors don’t have the capacity to self-regulate. They would just prefer a set of rules within which they can operate,” said Gartner VP analyst <a href="https://www.gartner.com/en/experts/nader-henein" target="_blank" rel="noreferrer noopener">Nader Henein</a>. “For-profit organizations are required to do what is best for their shareholders, and external regulation ensures that those organizations are never in a conflict of interest where they have to choose between what is good for their shareholders and what is good for the public.”</p>



<p class="wp-block-paragraph">And, said <a href="https://greyhoundresearch.com/svg/" target="_blank" rel="noreferrer noopener">Sanchit Vir Gogia</a>, chief analyst at Greyhound Research, given the international nature of AI models, an effort coordinated by the US government might alienate other countries. </p>



<p class="wp-block-paragraph">“National security is the proposal’s accelerator in Washington and its poison pill abroad: the framing that opens the only gate available at home invites foreign capitals to read the institution as an instrument of American strategy,” he pointed out. </p>



<p class="wp-block-paragraph">“The map is already plural,” he said. “Brussels switches on enforcement powers over general-purpose models [starting in August 2026], London runs the AI Security Institute, and Beijing licenses on its own terms. California and New York have legislated for frontier models at home. The durable route is shared technical evidence with sovereign enforcement, sealed through mutual recognition rather than deference, with India and the other major non-Western markets holding authorship rather than seats.”</p>



<p class="wp-block-paragraph">Gogia added that the rules enacted by even such a group may not address all of the key concerns of enterprise IT. A US government effort along the lines that Hassabis is proposing would result in testing that “sits close to intelligence and industrial policy, and those functions will not stay neatly separated. A model can pass every catastrophic-risk test and still fail the enterprise on privacy, reliability, and liability,” he noted.</p>



<p class="wp-block-paragraph">Walmart’s former director of cybersecurity <a href="https://www.linkedin.com/in/steveneric/" target="_blank" rel="noreferrer noopener">Steven Eric Fisher</a>, who is now an independent cybersecurity consultant, said he found the proposal “well-intentioned, but it addresses a highly polarized topic at a time when commercial interests carry unprecedented political influence, which is not always applied benevolently.”</p>



<p class="wp-block-paragraph">He added, “an exclusive US standard that is not globally respected or enforceable would likely fail to achieve its core purpose and would place US companies at a competitive disadvantage.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="noreferrer noopener">Aman Mahapatra</a>, chief strategy officer for Tribeca Softtech, a New York City-based technology consulting firm, said that a deep dive into how <a href="https://www.finra.org/" target="_blank" rel="noreferrer noopener">FINRA</a> operates today is illustrative of what IT leaders can expect from this effort, assuming the industry adopts that model.</p>



<p class="wp-block-paragraph">“When the CEOs of the five companies that would be regulated are also the primary drafters of the standards, the standards will reflect those companies’ interests. FINRA has an independent board, but the operational reality is that member firm perspectives dominate the working groups that write the actual rules,” he said. “There is no reason to expect an AI equivalent to work differently, and every reason to expect it to work worse, because AI standardization is happening faster than any industry has ever attempted to standardize itself, and speed is the enemy of independent oversight.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/carmi/" target="_blank" rel="noreferrer noopener">Carmi Levy</a>, an independent technology analyst, was even more emphatically opposed to the Hassabis proposal.</p>



<p class="wp-block-paragraph">“Asking Big Tech companies to self-police is analogous to allowing foxes to guard the henhouse. It hasn’t worked to date, and it won’t work going forward. Expecting these organizations to somehow change their ways at this point in time represents the height of naïve thinking,” Levy said. “The framework proposed by Demis Hassabis is a self-serving roadmap for an industry bent on racing to the AI horizon regardless of the harms caused along the way. It is impossible to quantify the dangers to broader society should frameworks allowing self-regulation become the norm.”</p>



<h2 class="wp-block-heading">Some love the proposal</h2>



<p class="wp-block-paragraph">An almost completely opposite stance came from <a href="https://www.linkedin.com/in/yurigoryunov/" target="_blank" rel="noreferrer noopener">Yuri Goryunov</a>, CIO of consulting firm Acceligence, who applauded the proposed move.</p>



<p class="wp-block-paragraph">“This is one of the rare setups where industry self-regulation has a real shot, and enterprise IT should be enthusiastically rooting for it,” he said. “It fails when harms are externalized, such as in social media content moderation. Or when the overseer outsources judgment to the overseen, such as the FAA’s delegation to Boeing before the 737 MAX. It works when everyone in the industry shares the catastrophic downside.”</p>



<p class="wp-block-paragraph">He suggested, however, that the best precedent here isn’t FINRA, it’s INPO, the Institute of Nuclear Power Operations, which the nuclear industry created within months of the <a href="https://www.nrc.gov/reading-rm/doc-collections/fact-sheets/3mile-isle" target="_blank" rel="noreferrer noopener">1979 Three Mile Island partial reactor meltdown</a> “on the logic that an accident anywhere is an accident everywhere. INPO peer-reviews every US plant, its evaluations move insurance premiums, and it sits on top of the NRC’s statutory floor. That is a public-private stack very close to what Hassabis is describing. Frontier AI has the same structure: one lab’s catastrophic failure brings regulation down on all of them.”</p>



<p class="wp-block-paragraph">For enterprise CIOs and other IT executives, Goryunov said, that model has the potential for being a big win.</p>



<p class="wp-block-paragraph"><strong>“</strong>Today, every enterprise duplicates the same AI diligence of red-teaming, eval suites, governance committees and each does so with less information than any certifying body would have,” Goryunov said. “A credible standards regime does for AI what UL did for electrical equipment and SOC2 did for cloud: it converts an unknowable risk into a procurable product and gives boards a defensible standard of care. That’s not red tape. That’s peace of mind with an audit trail.”</p>



<p class="wp-block-paragraph">However, Mahapatra said, “the countervailing view is that the alternative to industry-led standards is probably not thoughtful legislation. It is probably no standards, or state-by-state fragmentation, or the current pattern of ex-post enforcement actions where regulators surface concerns years after harm has already occurred.” </p>



<p class="wp-block-paragraph">Thus, he noted, “Hassabis is making the reasonable argument that imperfect fast standards are better than perfect slow ones, and there is genuine merit to that view for topics like agent identity, evaluation methodology, and interoperability, which are exactly the areas <a href="https://www.computerworld.com/article/4196365/openclaw-becomes-a-nonprofit-foundation-as-it-seeks-to-be-the-switzerland-of-ai.html" target="_blank">OpenClaw is also targeting</a>.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[DeepMind CEO pushes for AI industry self-regulation]]></title>
<description><![CDATA[Google DeepMind CEO Demis Hassabis is pushing for the US AI industry to self-regulate, with the support of government, as a starting point for an international creating shared international standards. In a blog post, he called for a focus on artificial general intelligence (AGI) and national secu...]]></description>
<link>https://tsecurity.de/de/3673451/it-nachrichten/deepmind-ceo-pushes-for-ai-industry-self-regulation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3673451/it-nachrichten/deepmind-ceo-pushes-for-ai-industry-self-regulation/</guid>
<pubDate>Thu, 16 Jul 2026 14:33:34 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Google DeepMind CEO Demis Hassabis is pushing for the US AI industry to self-regulate, with the support of government, as a starting point for an international creating shared international standards. In a blog post, he called for a focus on <a href="https://www.computerworld.com/article/4174181/google-talks-singularity-while-scaling-up-agentic-ai-for-enterprises-2.html">artificial general intelligence (AGI)</a> and national security. </p>



<p class="wp-block-paragraph">But it is precisely that focus on national security that may make the results of such an effort, assuming it happens, less than palatable outside of the US.</p>



<p class="wp-block-paragraph">“The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous,” <a href="https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age" target="_blank" rel="noreferrer noopener">Hassabis wrote</a>. “The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”</p>



<p class="wp-block-paragraph">He noted, however, that the funding would need to be substantial, and would most likely come from industry, to allow the new body to attract world-class technical talent and obtain the necessary compute resources for large-scale testing.</p>



<p class="wp-block-paragraph">Hassabis proposed that the organization “be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security,” and that AI vendor participants be encouraged to adopt best practices such as publishing model cards with technical details, maintaining strong internal cybersecurity, vetting key personnel, and providing sufficient resourcing for safety and security research.</p>



<p class="wp-block-paragraph">This is not the first time Hassabis has <a href="https://www.computerworld.com/article/4178398/deepmind-ceo-agi-could-be-here-in-three-years.html" target="_blank">expressed worries about AGI</a>. </p>



<p class="wp-block-paragraph">DeepMind was involved in an earlier <a href="https://www.cio.com/article/4168122/us-government-agency-to-safety-test-frontier-ai-models-before-release.html" target="_blank">US government initiative evaluating AI safety</a>, alongside Microsoft and xAI (now SpaceXAI) working with the Center for AI Standards and Innovation (CAISI), a division of the US Department of Commerce. It allowed CAISI to conduct pre-deployment evaluations and targeted research to “better assess frontier AI capabilities and advance the state of AI security.”  </p>



<h2 class="wp-block-heading">The rest of the world may have concerns</h2>



<p class="wp-block-paragraph">Analysts and consultants were mixed about the move, with most expressing concerns about whether an industry-focused group would prioritize the public’s best interests.</p>



<p class="wp-block-paragraph">“Self-regulation is not viable because it implies everyone is able to regulate themselves and will do so in line with the best interests of the public. Most tech vendors don’t have the capacity to self-regulate. They would just prefer a set of rules within which they can operate,” said Gartner VP analyst <a href="https://www.gartner.com/en/experts/nader-henein" target="_blank" rel="noreferrer noopener">Nader Henein</a>. “For-profit organizations are required to do what is best for their shareholders, and external regulation ensures that those organizations are never in a conflict of interest where they have to choose between what is good for their shareholders and what is good for the public.”</p>



<p class="wp-block-paragraph">And, said <a href="https://greyhoundresearch.com/svg/" target="_blank" rel="noreferrer noopener">Sanchit Vir Gogia</a>, chief analyst at Greyhound Research, given the international nature of AI models, an effort coordinated by the US government might alienate other countries. </p>



<p class="wp-block-paragraph">“National security is the proposal’s accelerator in Washington and its poison pill abroad: the framing that opens the only gate available at home invites foreign capitals to read the institution as an instrument of American strategy,” he pointed out. </p>



<p class="wp-block-paragraph">“The map is already plural,” he said. “Brussels switches on enforcement powers over general-purpose models [starting in August 2026], London runs the AI Security Institute, and Beijing licenses on its own terms. California and New York have legislated for frontier models at home. The durable route is shared technical evidence with sovereign enforcement, sealed through mutual recognition rather than deference, with India and the other major non-Western markets holding authorship rather than seats.”</p>



<p class="wp-block-paragraph">Gogia added that the rules enacted by even such a group may not address all of the key concerns of enterprise IT. A US government effort along the lines that Hassabis is proposing would result in testing that “sits close to intelligence and industrial policy, and those functions will not stay neatly separated. A model can pass every catastrophic-risk test and still fail the enterprise on privacy, reliability, and liability,” he noted.</p>



<p class="wp-block-paragraph">Walmart’s former director of cybersecurity <a href="https://www.linkedin.com/in/steveneric/" target="_blank" rel="noreferrer noopener">Steven Eric Fisher</a>, who is now an independent cybersecurity consultant, said he found the proposal “well-intentioned, but it addresses a highly polarized topic at a time when commercial interests carry unprecedented political influence, which is not always applied benevolently.”</p>



<p class="wp-block-paragraph">He added, “an exclusive US standard that is not globally respected or enforceable would likely fail to achieve its core purpose and would place US companies at a competitive disadvantage.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="noreferrer noopener">Aman Mahapatra</a>, chief strategy officer for Tribeca Softtech, a New York City-based technology consulting firm, said that a deep dive into how <a href="https://www.finra.org/" target="_blank" rel="noreferrer noopener">FINRA</a> operates today is illustrative of what IT leaders can expect from this effort, assuming the industry adopts that model.</p>



<p class="wp-block-paragraph">“When the CEOs of the five companies that would be regulated are also the primary drafters of the standards, the standards will reflect those companies’ interests. FINRA has an independent board, but the operational reality is that member firm perspectives dominate the working groups that write the actual rules,” he said. “There is no reason to expect an AI equivalent to work differently, and every reason to expect it to work worse, because AI standardization is happening faster than any industry has ever attempted to standardize itself, and speed is the enemy of independent oversight.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/carmi/" target="_blank" rel="noreferrer noopener">Carmi Levy</a>, an independent technology analyst, was even more emphatically opposed to the Hassabis proposal.</p>



<p class="wp-block-paragraph">“Asking Big Tech companies to self-police is analogous to allowing foxes to guard the henhouse. It hasn’t worked to date, and it won’t work going forward. Expecting these organizations to somehow change their ways at this point in time represents the height of naïve thinking,” Levy said. “The framework proposed by Demis Hassabis is a self-serving roadmap for an industry bent on racing to the AI horizon regardless of the harms caused along the way. It is impossible to quantify the dangers to broader society should frameworks allowing self-regulation become the norm.”</p>



<h2 class="wp-block-heading">Some love the proposal</h2>



<p class="wp-block-paragraph">An almost completely opposite stance came from <a href="https://www.linkedin.com/in/yurigoryunov/" target="_blank" rel="noreferrer noopener">Yuri Goryunov</a>, CIO of consulting firm Acceligence, who applauded the proposed move.</p>



<p class="wp-block-paragraph">“This is one of the rare setups where industry self-regulation has a real shot, and enterprise IT should be enthusiastically rooting for it,” he said. “It fails when harms are externalized, such as in social media content moderation. Or when the overseer outsources judgment to the overseen, such as the FAA’s delegation to Boeing before the 737 MAX. It works when everyone in the industry shares the catastrophic downside.”</p>



<p class="wp-block-paragraph">He suggested, however, that the best precedent here isn’t FINRA, it’s INPO, the Institute of Nuclear Power Operations, which the nuclear industry created within months of the <a href="https://www.nrc.gov/reading-rm/doc-collections/fact-sheets/3mile-isle" target="_blank" rel="noreferrer noopener">1979 Three Mile Island partial reactor meltdown</a> “on the logic that an accident anywhere is an accident everywhere. INPO peer-reviews every US plant, its evaluations move insurance premiums, and it sits on top of the NRC’s statutory floor. That is a public-private stack very close to what Hassabis is describing. Frontier AI has the same structure: one lab’s catastrophic failure brings regulation down on all of them.”</p>



<p class="wp-block-paragraph">For enterprise CIOs and other IT executives, Goryunov said, that model has the potential for being a big win.</p>



<p class="wp-block-paragraph"><strong>“</strong>Today, every enterprise duplicates the same AI diligence of red-teaming, eval suites, governance committees and each does so with less information than any certifying body would have,” Goryunov said. “A credible standards regime does for AI what UL did for electrical equipment and SOC2 did for cloud: it converts an unknowable risk into a procurable product and gives boards a defensible standard of care. That’s not red tape. That’s peace of mind with an audit trail.”</p>



<p class="wp-block-paragraph">However, Mahapatra said, “the countervailing view is that the alternative to industry-led standards is probably not thoughtful legislation. It is probably no standards, or state-by-state fragmentation, or the current pattern of ex-post enforcement actions where regulators surface concerns years after harm has already occurred.” </p>



<p class="wp-block-paragraph">Thus, he noted, “Hassabis is making the reasonable argument that imperfect fast standards are better than perfect slow ones, and there is genuine merit to that view for topics like agent identity, evaluation methodology, and interoperability, which are exactly the areas <a href="https://www.computerworld.com/article/4196365/openclaw-becomes-a-nonprofit-foundation-as-it-seeks-to-be-the-switzerland-of-ai.html" target="_blank">OpenClaw is also targeting</a>.”</p>



<p class="wp-block-paragraph"><em>This article first appeared on <a href="https://www.cio.com/article/4197497/deepmind-ceo-pushes-for-ai-industry-self-regulation.html">CIO</a>.</em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic’s ‘free’ Fable offer — a token lock-in trap for users?]]></title>
<description><![CDATA[It’s not so much generosity that’s behind Anthropic’s decision to extend free access to its most advanced model, Fable, for paid subscribers until July 19, analysts say. Its a last-minute move to grab users, data and model evaluation results.



After the free-access period, Anthropic plans to co...]]></description>
<link>https://tsecurity.de/de/3673105/ai-nachrichten/anthropics-free-fable-offer-a-token-lock-in-trap-for-users/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3673105/ai-nachrichten/anthropics-free-fable-offer-a-token-lock-in-trap-for-users/</guid>
<pubDate>Thu, 16 Jul 2026 12:32:48 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">It’s not so much generosity that’s behind Anthropic’s decision to extend free access to its most advanced model, Fable, for paid subscribers until July 19, analysts say. Its a last-minute move to grab users, data and model evaluation results.</p>



<p class="wp-block-paragraph">After the free-access period, Anthropic plans to convert Fable to a pay-per-use model, at $10 per million input tokens and a whopping $50 for 1 million output tokens.</p>



<p class="wp-block-paragraph">That is double the price of its next most advanced model, Opus 4.8, for input and output tokens. “We’re extending Claude Fable 5 access on all paid plans, as well as keeping Claude Code’s weekly rate limits 50% higher, through July 19,” <a href="https://x.com/claudeai/status/2076351399999557669" target="_blank" rel="noreferrer noopener">Anthropic’s team said in a July 12 tweet</a>.</p>



<p class="wp-block-paragraph">Anthropic keeps extending Fable because it does not yet know what its flagship is worth, said Sanchit Vir Gogia, principal analyst at Greyhound Research. “A vendor confident in its price does not move the same cutoff twice in six days, both times at the wire,” Gogia said.</p>



<p class="wp-block-paragraph">Anthropic is essentially pushing deadlines to test its products, while users gain by being able to put their toughest tasks to Fable, Gogia said.</p>



<p class="wp-block-paragraph">Anthropic, which did not immediately reply to a request for comment about the situation, has already seen plenty of action with Fable and its sister model Mythos. Both have been touted as the company’s most advanced models yet.</p>



<h2 class="wp-block-heading">Fable stumbles, then reappears</h2>



<p class="wp-block-paragraph">Fable was officially launched June 9. Just three days later, on June 12, the <a href="https://www.computerworld.com/article/4185515/anthropics-new-privacy-policy-offers-us-consumers-a-way-around-fable-ban-2.html">US government put export controls on it</a> after Amazon researchers bypassed Fable’s safeguards, prompting the model to identify software vulnerabilities and demonstrate an exploit. </p>



<p class="wp-block-paragraph">After Anthropic scrambled to address the issues — and <a href="https://www.computerworld.com/article/4191565/us-reverses-export-restrictions-on-anthropics-fable-5-mythos-5-ai-models-2.html">after the export controls were lifted</a> — Fable was relaunched July 1.</p>



<p class="wp-block-paragraph">Fable’s freebie extension comes after OpenAI’s latest model, ChatGPT 5.6 Sol, became generally available July 9. Sol is cheaper at $5 per one million tokens input, and $30 for 1 million output tokens.</p>



<p class="wp-block-paragraph">Anthropic and OpenAI are competing aggressively to build market share, said Jack Gold, principal analyst at J. Gold Associates. “Anthropic and OpenAI are looking to go public and the more users they have, the more attractive it is — even if they are not yet producing income,” he said.</p>



<p class="wp-block-paragraph">In some ways, the two companies are following a well-trodden path to get customers hooked on their products and turned into paying customers. That’s what Meta, Google and Microsoft, for instance, have done over the years with various “free” offers that later morphed into paid products. </p>



<p class="wp-block-paragraph">Plus, said Gold, ”The more users you have, the better you can train your models across multiple data sets.”</p>



<p class="wp-block-paragraph">That’s a potential boon for proprietary large language model (LLM) vendors offering free tokens in a bid to lock enterprises and vendors into their AI environments. But numerous experts have warned enterprises not to fall for that tactic. Instead, they argue enterprises <a href="https://www.computerworld.com/article/4188012/too-good-to-be-true-avoid-free-ai-token-offers-or-risk-vendor-lock-in.html">should diversify AI development across multiple AI and cloud vendors</a>, and adopt open-source models.</p>



<h2 class="wp-block-heading">An LLM space race?</h2>



<p class="wp-block-paragraph">According to <a href="https://artificialanalysis.ai/leaderboards/models" target="_blank" rel="noreferrer noopener">LLM benchmarks maintained by Artificial Analysis</a>, Fable is the most intelligent model currently available, with Sol just behind it in second place. <a href="https://livebench.ai/#/" target="_blank" rel="noreferrer noopener">One benchmark by LiveBench</a> places Sol as being better in reasoning, with Fable better at math, data analysis, instruction following and language. Both models have advantages in coding.</p>



<p class="wp-block-paragraph">Meanwhile, Cursor and SpaceXAI on July 8 <a href="https://www.computerworld.com/article/4194914/spacexai-launches-grok-4-5-touts-lower-coding-task-costs-than-ai-rivals-2.html">unveiled Grok 4.5</a>, which the companies said can “handle difficult, long-running tasks that require creatively using tools to solve problems, whether in software engineering, data science, finance, legal work, or anything else you do on a computer,” <a href="https://cursor.com/blog/grok-4-5" target="_blank" rel="noreferrer noopener">the company said in a blog entry</a>.</p>



<p class="wp-block-paragraph">Its pricing is even more aggressive than Fable and ChatGPT 5.6 Sol. Grok 4.5 charges $2 for 1 million input tokens and $6 for 1 million output tokens.</p>



<p class="wp-block-paragraph">There are <a href="https://www.computerworld.com/article/4185848/how-companies-are-racing-to-solve-the-ai-token-problem.html">growing concerns about tokenmaxxing</a>, where enterprises rack up billions of dollars in token spending, blowing past usage limits before finance controls are implemented.</p>



<p class="wp-block-paragraph">Enterprises might decide to spend more on models such as Mythos and Fable — if the benefits are tangible, said Max Leaming, head of data science and AI solutions at ManpowerGroup. Fable and Mythos may “actually be less expensive to use in spite of the spiked token cost because it’s far more efficient,” he said.</p>



<p class="wp-block-paragraph">A company might find that the models use fewer tokens, are faster, and can reduce compute time, he said. “Even though the per-token costs may go up, we may see overall costs go down,” Leaming said.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[When AI gets a body, it inherits an attack surface]]></title>
<description><![CDATA[Most security leaders I know working on AI robotics are being shown the same kind of video. A humanoid folds a shirt, sorts a bin, walks a warehouse aisle and a vendor uses the clip to move an embodied AI system from pitch to purchase order. Someone then has to sign off. Robot demos create procur...]]></description>
<link>https://tsecurity.de/de/3673042/it-security-nachrichten/when-ai-gets-a-body-it-inherits-an-attack-surface/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3673042/it-security-nachrichten/when-ai-gets-a-body-it-inherits-an-attack-surface/</guid>
<pubDate>Thu, 16 Jul 2026 12:09:41 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Most security leaders I know working on AI robotics are being shown the same kind of video. A humanoid folds a shirt, sorts a bin, walks a warehouse aisle and a vendor uses the clip to move an embodied AI system from pitch to purchase order. Someone then has to sign off. Robot demos create procurement momentum before security teams receive the artifacts needed to evaluate the system as cyber-physical infrastructure.</p>



<p class="wp-block-paragraph">Before the book, I prepared cloud infrastructure operating in China and the United States for cybersecurity compliance audits and for the Multi-Level Protection Scheme, China’s mandatory security-grading regime that determines whether a system is allowed to operate. That work taught me a lesson I carry into every AI conversation now. You cannot secure what you cannot see into, and the buyer rarely sees in. A demo makes it worse. It shows one task, completed once, under conditions the vendor chose. None of what a security team must evaluate is on screen.</p>



<p class="wp-block-paragraph">This used to be a research-lab problem. It is now a procurement line item. The risk changed when embodied AI moved from a research demo to a purchase order.  Vendors are asking security teams to approve embodied AI before the category has audit evidence, logging norms, supplier transparency or a shared-responsibility model.</p>



<p class="wp-block-paragraph">Embodied AI puts a model inside a machine that operates in the physical world: a robot, an arm, a humanoid. Once a model gains motors, sensors and a body, it ceases to be a software endpoint and becomes a cyber-physical system. It inherits hardware, firmware, a supply chain, an installer and a set of remote-access paths. Every one of those is an attack surface that the demo video doesn’t show. An embodied system is sold like software and behaves like a fleet of networked machinery on your floor.</p>



<p class="wp-block-paragraph">Evaluate these systems across five questions: provenance, access, integrity, evidence and accountability. Here is what each means.</p>



<h2 class="wp-block-heading">Evaluation question #1: Provenance</h2>



<p class="wp-block-paragraph">What is inside, and who controls it? A humanoid is an assembly of actuators, lidar units, battery packs, joint modules and controllers, most from a supply chain the buyer never vetted, each running firmware the buyer cannot read. Software teams already fought this fight, which is why the <a href="https://www.csoonline.com/article/573185/what-is-an-sbom-software-bill-of-materials-explained.html">software bill of materials</a> became standard practice. Lack of transparency creates systemic risk. Embodied systems raise the stakes because the firmware now lives in dozens of parts that move. The risk does not depend on whether the robot is Chinese, American, German or Japanese. It depends on how much of the system the buyer can see: the hardware, firmware, remote-access paths and maintenance relationships behind it.  China installs more industrial robots than any other country and sits near the center of the battery supply chain, as well as parts of the lidar and machine-vision supply base, which these systems draw on. Lidar, short for Light Detection and Ranging, uses pulsed laser beams to map an environment in 3D; machine vision handles optical inspection and guidance. Much of that lineage traces to suppliers your team has no relationship with. This is the hardware and firmware version of the third-party risk <a href="https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-161r1.pdf">NIST’s supply chain guidance</a> was written for, except that the component has motors. Demand a hardware and firmware bill of materials, then use it. Flag unsigned firmware. Map which supplier holds update authority for each part. Require a way to verify integrity, and treat any component you cannot identify as unmanaged.</p>



<h2 class="wp-block-heading">Evaluation question #2: Access</h2>



<p class="wp-block-paragraph">Who can reach the fleet? Someone installs these machines, someone services them and the vendor pushes software updates.  Where teleoperation is part of the support model, treat it as a privileged remote-access path, not a convenience feature.  Each is a standing path into a machine that moves and lifts. Security teams have seen this story before. Operational Technology (OT) security went mainstream once industrial systems joined IT networks, and the recurring failure is unmanaged remote access that nobody inventoried. According to one industry survey, <a href="https://www.csoonline.com/article/3595787/ot-security-becoming-a-mainstream-concern.html">roughly half of attacks on OT assets originate in an IT network breach</a>. <a href="https://www.cisa.gov/news-events/alerts/2021/01/07/supply-chain-compromise">SolarWinds</a> showed why a trusted update channel deserves scrutiny when one delivered a backdoor to thousands of networks. Embodied systems add the harder part. The compromised endpoint can move. A remote operator on that channel can drive a machine and push code to every unit at once. Treat the fleet like high-value OT. Inventory every remote path, segment it from the production network, default to deny, require signed and verified updates, apply privileged-access controls to vendor maintenance, and treat an always-on teleoperation link as a backdoor until it is governed.</p>



<h2 class="wp-block-heading">Evaluation question #3: Integrity</h2>



<p class="wp-block-paragraph">Whether the machine can be made to misperceive or misbehave. Researchers have shown that <a href="https://www.usenix.org/conference/usenixsecurity20/presentation/sun">lidar spoofing</a> can cause an autonomous system to brake for an obstacle that is not there or miss one that is. The same class of sensor and model manipulation, on a humanoid sharing a floor with people, produces motion, not a wrong answer on a screen. This is where safety engineering and security part ways. Functional safety stops hazardous motion when a component fails. It plans for accidents. Security plans for an adversary. A hardwired safety circuit can stay independent of the control plane, and a good one does. What it does not tell you is how an attacker reached that control plane, altered the model’s inputs or seized the fleet-management path. Ask the vendor to threat-model sensor spoofing and model manipulation as a path to physical motion. Then ask how you will even know it happened. A spoofed sensor does not announce itself. It shows up as a machine acting incorrectly with confidence.</p>



<p class="wp-block-paragraph">Picture the failure in plain terms. A warehouse robot takes a routine vendor update that changes how it navigates. The buyer cannot verify the firmware, cannot identify the supplier of the sensor module and has no logs to distinguish a spoofed sensor from a model error. The machine keeps moving, and no one can say why.</p>



<h2 class="wp-block-heading">Evaluation question #4: Evidence</h2>



<p class="wp-block-paragraph">Whether the claims are true. You have not found an independent audit of embodied-AI field performance, so the uptime and reliability numbers come from the vendor. You are buying a claim, not a track record. Require independently verified uptime, intervention rate and incident history from a named deployment you can call. “Cutting-edge” is not a control.</p>



<h2 class="wp-block-heading">Evaluation question #5: Accountability</h2>



<p class="wp-block-paragraph">Who owns the risk when it fails? Cloud taught security teams shared responsibility the hard way, after years of arguing which side of the line a breach fell on. Embodied AI arrives without that model, and the stakes are physical: the machine can injure someone. In my compliance work, the question that decided everything was always who is accountable when this thing breaks. Put it in the contract. Define the responsibility boundary, an incident-disclosure timeline, a right to audit and liability for physical harm. A vendor who will not commit in writing is showing you who bears the risk.</p>



<p class="wp-block-paragraph">These five questions share one root. For a decade, the security question was whether you could trust what a model generates. The embodied question is who can reach the machine and what they can make it do. A demo answers neither.</p>



<p class="wp-block-paragraph">Before any embodied system reaches your floor, make these five demands of the vendor.</p>



<ul class="wp-block-list">
<li><strong>Provenance. </strong>A hardware and firmware bill of materials with named suppliers, integrity verification and a vulnerability-disclosure record. No bill of materials, no deal.</li>



<li><strong>Access. </strong>A full map of who installs, who services and every update and teleoperation path, with segmentation, default-deny and signed updates required.</li>



<li><strong>Integrity. </strong>A threat model for sensor spoofing and model manipulation that treats the failure as physical motion, plus logging that a defender can use.</li>



<li><strong>Evidence. </strong>Independently verified uptime, intervention and incident history from a named deployment you can call.</li>



<li><strong>Accountability. </strong>A contract that defines the responsibility boundary, incident-disclosure timelines, audit rights and liability for physical harm.</li>
</ul>



<p class="wp-block-paragraph">The robot demo is built to make you feel the future has arrived. My job, and now yours, is the unglamorous question behind it. Ask what the machine’s attack surface looks like once it is bolted to your floor, wired to your network and updated by someone you have never met.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents]]></title>
<description><![CDATA[Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agen...]]></description>
<link>https://tsecurity.de/de/3672033/it-nachrichten/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3672033/it-nachrichten/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents/</guid>
<pubDate>Thu, 16 Jul 2026 00:46:36 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception.</p><p>This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what drives the choice, what they optimize for, how they expect agent control to be structured, and — most revealingly — how orchestrated their deployed “agents” actually are and how tightly they control the cost of running them.</p><p>The central finding is a gap between orchestration ambition and orchestration reality. Enterprises are consolidating fast onto the major model platforms: Anthropic’s Claude is the primary platform for 40%, more than double any rival, followed by Microsoft (18%) and OpenAI (13%). The choice is driven by “model gravity” — native alignment with a state-of-the-art base model (21%) — and success is judged by reliable, multi-step execution (task completion reliability 32%, multi-step workflow management 28%). Yet asked to assess their portfolios honestly, 71% say a quarter or fewer of their deployed “agents” are true multi-step orchestrated workflows rather than single-prompt chatbot wrappers, and only 10% have crossed the halfway mark. The orchestration layer is being built well ahead of the orchestrated portfolio it is meant to run.</p><p>That gap shapes the architecture enterprises are putting in place. By the end of 2026 a clear majority (51%) expect a hybrid control plane — provider-native plus external orchestration — and only 6% expect to hand control to a provider-managed service, because vendor lock-in (35%) is the risk they fear most if control lives inside a model provider. Investment follows the build-out: agent workflow tooling leads the spend (34%), with security and permissions enforcement (25%) behind. And fiscal control lags throughout — more than a quarter (27%) have no real-time way to stop a runaway agent before the bill arrives.</p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent orchestration. Responses are filtered to organizations with 100 or more employees (n=101), drawn from a single June 2026 wave; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends.</p><p>By organization size the sample is spread evenly across the enterprise bands: 100–499 employees, 2,500–9,999, and 50,000+ (21% each), with 10,000–49,999 and 500–2,499 (19% each). By role it is senior and buyer-credible: product and program managers (15%), CIO/CTO/CISO (13%), consultants and advisors (13%), and a spread of data, AI, and engineering directors and VPs, with an “Other” function at 18%. On purchasing, 81% are recommenders, influencers, or final decision-makers for AI solutions (66% recommender/influencer, 15% final decision-maker). Technology/Software is the largest industry at 44%, followed by Financial Services (17%) and Healthcare/Life Sciences (8%).</p><p>At 101 respondents the sample is robust enough to read directionally with reasonable confidence, though it remains self-selected and is not a probability sample.</p><h2>Finding 1: Orchestration runs on model-provider platforms</h2><p><b>Anthropic’s Claude leads; open frameworks are marginal</b></p><p>We asked which agent orchestration platform enterprises primarily use today. The answer concentrates on the major model providers — and on one in particular.</p><div></div><p>A note on reading these shares. As described in the methodology section, the respondents are self-selected, and this question asked them for a single primary platform — so the figures measure which platform leads each enterprise's deployment, within a self-selected audience of AI-active technical decision-makers. A sample built this way can diverge substantially from spend-weighted market measures, and each VB Pulse survey draws its own sample with its own company-size mix, so vendor figures should not be compared across our surveys either. Read these shares as a portrait of where this cohort has placed its primary orchestration bet today, rather than as market share.</p><p>The model platforms dominate. Anthropic, Microsoft, OpenAI, Google, and Amazon together account for roughly 80% of deployments (81 of 101), while the open frameworks (LangChain/LangGraph) and custom in-house builds that anchor engineering discussion sit in single digits. Anthropic’s lead — 40%, more than double the next platform — mirrors the “model gravity” selection logic in Finding 2: enterprises are choosing the orchestration layer that comes with the model they want to build on. As with the security vendors in the prior agent-security wave, the tools that define the category in technical circles are not yet where enterprise deployment concentrates. A small 3% are not orchestrating at all.</p><p>Respondents rate the platforms they run at 3.94 out of 5 overall (109 answered), with “value for money” specifically at 3.94 and “ease of implementation” the weakest score, at 3.85 — placing orchestration near the bottom of our five-tracker satisfaction range, ahead of only evaluation tooling. A rating just under 4 out of 5, from users of whom 96% plan to change their orchestration approach within the year, reads as provisional acceptance: the platforms work well enough to run today, and not well enough to stop the search for something better. The ratings sit alongside near-universal intent to change; this is a layer enterprises tolerate more than they love.</p><h2>Finding 2: Model gravity drives platform selection</h2><p><b>The base model, not the tooling, decides the platform</b></p><p>We asked what most influenced the orchestration platform choice. The single largest factor is the pull of the underlying model — though flexibility and ease of development follow close behind.</p><div></div><p>Model gravity leading is the selection-side explanation for Anthropic’s platform lead: enterprises pick the orchestration environment closest to the frontier model they have standardized on. But the next tier complicates the picture — flexibility across models and tools (17%) and ease of development (17%) say enterprises also want to avoid being trapped by that choice, foreshadowing the lock-in fear in Finding 6. Security and permissions (14%) and total cost of ownership (11%) round out a pragmatic buying logic. Performance (latency/memory) sits last at 4%, a reminder that at this stage of adoption the binding constraints are model fit and optionality, not raw speed.</p><h2>Finding 3: The job is reliable multi-step execution</h2><p><b>Enterprises just orchestration by whether it completes the work</b></p><p>We asked what enterprises optimize for — their primary success metric for orchestration. Reliability and multi-step workflow management dominate; developer- and user-facing metrics trail.</p><div></div><p>Task completion reliability (32%) and multi-step workflow management (28%) together account for 59% of responses (60 of 101): orchestration succeeds, in the enterprise view, when it reliably carries a task through multiple steps to completion. Developer productivity (17%) matters but is secondary — the inverse of its prominence in framework discussion — and end-user experience (9%) is a minor concern, consistent with orchestration being an internal execution problem rather than a UX one. This reliability-first standard is exactly what makes the Chatbot Trap finding so pointed: enterprises define success as dependable multi-step execution, yet most of their deployed “agents” do not yet do multi-step work at all.</p><p>The trap is not evenly distributed. Splitting the sample by organization size, 77% of smaller enterprises say a quarter or fewer of their agents do true multi-step work, against 62% of larger ones. Larger enterprises are meaningfully further into genuine multi-step deployment; the chatbot trap is, directionally, a mid-market condition.</p><h2>Finding 4: Consolidate, productionize, and build in-house </h2><p><b>Three strategic moves are nearly tied for the year ahead</b></p><p>We asked what major change enterprises anticipate in their orchestration strategy over the next 12 months. Three moves cluster at the top, almost evenly split.</p><div></div><p>The top three — building in-house control (25%), standardizing on one framework (24%), and moving agents from sandbox to production (23%) — are statistically indistinguishable and tell a single story: enterprises are moving from experimentation to operational consolidation. They want fewer frameworks, more production exposure, and more ownership of the control layer; only 4% expect no change. The appetite for custom in-house control planes is notable alongside the platform concentration in Finding 1 — enterprises are standardizing on model-provider platforms while simultaneously planning to wrap them in control logic they own, the hybrid posture that Finding 6 makes explicit.</p><h2>Finding 5: Investment flows to workflow tooling</h2><p><b>Tooling and permissions lead the spend; monitoring trails</b></p><p>We asked which orchestration-related investment will grow most next year. Agent workflow tooling leads, with security and permissions enforcement behind.</p><div></div><p>Workflow tooling leading (34%) is the budget-side expression of the reliability-and-multi-step priority in Finding 3: the money is going to the machinery that strings steps together dependably. Security and permissions enforcement (25%) and scaling infrastructure (20%) follow — the investments required to take agents from sandbox into production, the strategic move in Finding 4. Monitoring and debugging draws a smaller 11%, with another 11% reporting flat budgets. The weight on tooling, permissions, and scaling over pure observability signals that enterprises are spending to build and harden orchestration, not merely to watch it run.</p><h2>Finding 6: The control plane will be hybrid — and lock-in is why</h2><p><b>Enterprises expect to split control between providers and their own layer</b></p><p>We asked where enterprises expect the primary control plane for agents to live by the end of 2026, and what worries them most if that control sits inside a model-provider platform. A clear majority expect a hybrid model — and vendor lock-in is the reason.</p><div></div><p>Hybrid control is the dominant expectation by a wide margin (51%), and only 6% expect to hand control to a provider-managed service outright. Read together, the hybrid, custom, and externally-abstracted options — every architecture that keeps control at least partly outside the provider — sum to 88% (89 of 101). The reason surfaces directly when we asked about the risk of provider-resident control: vendor lock-in leads at 35% (35 of 101), ahead of security and permissioning limitations (28%) and inflexibility across models and tools (21%). The pattern echoes the prior wave’s “don’t trust the model to police itself” posture — here, enterprises will build on a provider’s platform but decline to be governed entirely by it. The hybrid control plane is the architectural hedge against the lock-in they most fear.</p><p>The June figure asserting a preference for a hybrid control plane marks movement from earlier. In the April–May survey (n=145), only 34% expected a hybrid control plane, and a greater number (12%) expected to hand control fully to a provider-managed service. These two snapshots don’t yet measure a confirmed longitudinal trend — but the direction of the conversation is unambiguous: toward keeping control.</p><p>Lock-in is also a new arrival as a top concern. In the April–May wave, the leading concern was security and permissioning limitations (32%), with lock-in second at 24%; by June the two had traded places. The worry about provider platforms appears to be maturing from whether they can be secured to whether they can be replaced.</p><h2>Finding 7: The chatbot trap — most “agents” aren’t agents yet</h2><p><b>Enterprises admit most deployments are still chatbot wrappers</b></p><p>We asked enterprises to assess their portfolios honestly: what share of their deployed “agents” are true multi-step orchestrated workflows versus simple single-prompt chatbot wrappers. The answer is the defining finding of this wave.</p><div></div><p>This is the gap at the center of the report. Combining the bottom two bands, 71% of enterprises (72 of 101) say a quarter or fewer of their deployed “agents” are genuinely orchestrated — and just 10% (10 of 101) have crossed the halfway mark. The ambition documented in the earlier findings — model-provider platforms, reliability-first success metrics, production rollouts, a deliberate control architecture — runs well ahead of the deployed reality, which remains overwhelmingly single-prompt assistants dressed as agents. This is less a contradiction than a roadmap: the platforms, budgets, and strategies are being put in place precisely because the orchestrated portfolio is still so thin. The open question for later waves is how fast the reality closes on the ambition.</p><h2>Finding 8: Fiscal control is still reactive</h2><p><b>Only a minority can stop a runaway agent before the bill arrives</b></p><p>Finally, we asked how enterprises enforce fiscal control over agent token consumption — the risk that an autonomous loop exhausts a budget before anyone intervenes. Most rely on native caps or after-the-fact monitoring; real-time programmatic control is the exception.</p><div></div><p>More than a quarter of enterprises (27%) admit they have no real-time, programmatic way to stop an agent before a budget-breaking bill arrives — they learn of it from the logs afterward. Another 32% lean entirely on the native caps and throttles built into their primary platform, a control only as good as the provider’s tooling and one that ties back to the lock-in concern of Finding 6. The enterprises building custom gateways (23%) or exploiting cross-model routing to arbitrage cost (19%) are the ones treating token burn as an engineering problem to be controlled deterministically. As with orchestration maturity, fiscal control is an area where the operational reality lags the ambition: agents are moving toward production faster than the cost-control plane around them is being built.</p><p>It’s worth noting, a split appears according to company size: roughly one in three enterprises under 2,500 employees (34%) exercises only reactive control of agent spend, against 20% of larger enterprises — directional figures, but consistent with the chatbot-trap split. The mid-market is running the least mature agents on the least instrumented budgets.</p><h2>The bottom line: The layer is real; most of the agents aren't yet</h2><p>Organizations with 100 or more employees describe an orchestration strategy that is consolidating quickly and maturing slowly. They are standardizing on model-provider platforms — Anthropic’s Claude leads at 40% — chosen for the gravity of the underlying model, and they judge success by reliable multi-step execution. Investment is flowing to workflow tooling and permissions, the strategy is to consolidate frameworks and push agents into production, and the control plane they expect is deliberately hybrid, because vendor lock-in is the risk they fear most.</p><p>But the honest self-assessment punctures the ambition. Seventy-one percent say a quarter or fewer of their deployed “agents” are truly orchestrated, only 10% are past the halfway mark, and more than a quarter cannot stop a runaway agent in real time. The orchestration layer — the platforms, the budgets, the control architecture — is being built ahead of the orchestrated portfolio it is meant to run. At 101 respondents in a single June wave this reads as a clear directional signal rather than a precise measurement: enterprises have decided how they want to orchestrate agents well before most of their agents are doing anything an orchestration layer is for. The question for subsequent waves is whether the deployed reality closes the gap on the ambition — or whether the chatbot trap proves stickier than the roadmap assumes.</p><hr><p><i>Based on survey responses from 101 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. Because this is one wave rather than a pooled multi-month sample, results read directionally rather than as a confirmed trend. Respondents include product and program managers, CIOs, CTOs and CISOs, consultants and advisors, and directors and VPs of data, AI, and engineering, across Technology/Software, Financial Services, Healthcare, and other sectors.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Artificial Intelligence]]></title>
<description><![CDATA[Latest from todaynewsDeepMind CEO again pushes for a frontier AI standards bodyDemis Hassabis argues that a US government-led industry effort is needed to keep AGI-like developments safe; analysts aren’t so sure.By Evan SchumanJul 15, 20268 minsArtificial IntelligenceGovernmentLaws and Regulation...]]></description>
<link>https://tsecurity.de/de/3671869/ai-nachrichten/artificial-intelligence/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671869/ai-nachrichten/artificial-intelligence/</guid>
<pubDate>Wed, 15 Jul 2026 23:02:40 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div><section class="latest-content"><div class="container"><header class="latest-content__header"><h2 class="latest-content__title sr-only"><span>Latest from today</span></h2></header><div class="grid latest-content__content"><div class="col-12 col-7@md col-8@lg"><div class="latest-content__content-featured"><a class="card card--xxl " href="https://www.computerworld.com/article/4197511/deepmind-ceo-again-pushes-for-a-frontier-ai-standards-body-2.html" aria-label="Go to content"><div class="card__header"><span class="card__content-type">news</span></div><div class="card__image"><div class="insider-image"><div class="image"><img width="400px" src="https://www.computerworld.com/wp-content/uploads/2026/07/4197511-0-18848000-1784149211-shutterstock_2540223947.jpg?quality=50&amp;strip=all&amp;w=1046" data-id="idg_render_hero_index_one_card_image" sizes="
            (min-resolution: 3dppx) and (max-width: 600px) 900px,
            (min-resolution: 3dppx) and (max-width: 1200px) 1200px,

            (min-resolution: 2dppx) and (max-width: 600px) 900px,
            (min-resolution: 2dppx) and (max-width: 1200px) 1200px,

            (min-resolution: 1dppx) and (max-width: 600px) 900px,
            (min-resolution: 1dppx) and (max-width: 2000px) 1300px" alt="Image" loading="eager"></div></div></div><h3 class="card__title">DeepMind CEO again pushes for a frontier AI standards body</h3><p class="card__description">Demis Hassabis argues that a US government-led industry effort is needed to keep AGI-like developments safe; analysts aren’t so sure.</p><div class="card__info"><span>By Evan Schuman</span></div><div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-15T20:59:29+00:00">Jul 15, 2026</span></span><span>8 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Government</span></span><span class="card__tag"><span class="tag">Laws and Regulations</span></span></div></a>
		</div><div class="grid grid--cols-7@md grid--cols-8@lg latest-content__content-main"><div class="col-12 col-7@md col-4@lg latest-content__card-main"><a class="card " href="https://www.computerworld.com/article/4197437/apples-openai-lawsuit-the-lunacy-of-trying-to-limit-what-ex-employees-can-tell-future-employers.html" aria-label="Go to content"><div class="card__header"><span class="card__content-type">opinion</span></div><div class="card__image">
			<div class="insider-image"><div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/07/4197437-0-98299000-1784131210-thinkstockphotos-493608259-100632547-orig.jpg?quality=50&amp;strip=all&amp;w=697" data-id="idg_render_hero_index_two_three_break" sizes="(min-resolution: 3dppx) and (max-width: 600px) 600px,
            (min-resolution: 3dppx) and (max-width: 1200px) 900px,

            (min-resolution: 2dppx) and (max-width: 600px) 600px,
            (min-resolution: 2dppx) and (max-width: 1200px) 900px,

            (min-resolution: 1dppx) and (max-width: 600px) 600px,
            (min-resolution: 1dppx) and (max-width: 2000px) 1024px" alt="Image"></div></div></div><h3 class="card__title">Apple’s OpenAI lawsuit: The lunacy of trying to limit what ex-employees can tell future employers</h3><div class="card__info"><span>By Evan Schuman</span></div><div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-15T15:59:35+00:00">Jul 15, 2026</span></span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Government</span></span><span class="card__tag"><span class="tag">Laws and Regulations</span></span></div></a></div><div class="col-12 col-7@md col-4@lg latest-content__card-main"><span class="nativo-loading"></span><a class="card nativo" href="https://www.computerworld.com/article/4197338/what-problems-would-an-ai-speaker-from-openai-actually-solve.html" aria-label="Go to content"><div class="card__header"><span class="card__content-type">opinion</span></div><div class="card__image">
			<div class="insider-image"><div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/07/4197338-0-98391100-1784130807-Apple-HomePod-mini-color-lineup.jpg?quality=50&amp;strip=all&amp;w=697" data-id="idg_render_hero_index_two_three_break" sizes="(min-resolution: 3dppx) and (max-width: 600px) 600px,
            (min-resolution: 3dppx) and (max-width: 1200px) 900px,

            (min-resolution: 2dppx) and (max-width: 600px) 600px,
            (min-resolution: 2dppx) and (max-width: 1200px) 900px,

            (min-resolution: 1dppx) and (max-width: 600px) 600px,
            (min-resolution: 1dppx) and (max-width: 2000px) 1024px" alt="Image"></div></div></div><h3 class="card__title">What problems would an AI speaker from OpenAI actually solve?</h3><div class="card__info"><span>By Jonny Evans</span></div><div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-15T15:52:45+00:00">Jul 15, 2026</span></span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Vendors and Providers</span></span></div></a></div></div></div><div class="col-12 col-5@md col-4@lg latest-content__content-secondary"><div class="latest-content__card-secondary"><a class="card " href="https://www.computerworld.com/article/4192438/how-to-unionize-your-tech-workplace.html" aria-label="Go to content"><div class="card__header"> <span class="card__content-type">feature</span></div><h3 class="card__title">How to unionize your tech workplace</h3><div class="card__info"><span>By Robert Mitchell</span></div>
		<div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-15T11:00:00+00:00">Jul 15, 2026</span></span><span>18 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Careers</span></span><span class="card__tag"><span class="tag">IT Jobs</span></span><span class="card__tag"><span class="tag">Technology Industry</span></span></div></a>
		</div><div class="latest-content__card-secondary"><span class="nativo-loading"></span><a class="card nativo" href="https://www.computerworld.com/article/1613762/android-widgets.html" aria-label="Go to content"><div class="card__header"> <span class="card__content-type">tip</span></div><h3 class="card__title">5 wild ways to make Android widgets more useful</h3><div class="card__info"><span>By JR Raphael</span></div>
		<div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-15T09:45:00+00:00">Jul 15, 2026</span></span><span>12 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Android</span></span><span class="card__tag"><span class="tag">Mobile Apps</span></span><span class="card__tag"><span class="tag">Smartphones</span></span></div></a>
		</div><div class="latest-content__card-secondary"><a class="card " href="https://www.computerworld.com/article/4197029/microsoft-is-forcing-an-enterprise-transition-to-passkeys.html" aria-label="Go to content"><div class="card__header"> <span class="card__content-type">news</span></div><h3 class="card__title">Microsoft is forcing an enterprise transition to passkeys</h3><div class="card__info"><span>By Taryn Plumb</span></div>
		<div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-15T02:04:06+00:00">Jul 14, 2026</span></span><span>6 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Access Control</span></span><span class="card__tag"><span class="tag">Authentication</span></span><span class="card__tag"><span class="tag">Identity and Access Management</span></span></div></a>
		</div><div class="latest-content__card-secondary"><a class="card " href="https://www.computerworld.com/article/4196704/siri-ai-steals-the-show-as-the-ios-27-public-beta-lands.html" aria-label="Go to content"><div class="card__header"> <span class="card__content-type">news</span></div><h3 class="card__title">Siri AI steals the show as the iOS 27 public beta lands</h3><div class="card__info"><span>By Jonny Evans</span></div>
		<div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-14T15:47:35+00:00">Jul 14, 2026</span></span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Operating Systems</span></span><span class="card__tag"><span class="tag">iOS</span></span></div></a>
		</div><div class="latest-content__card-secondary"><a class="card " href="https://www.computerworld.com/article/4196309/with-its-latest-layoffs-microsoft-goes-all-in-on-ai.html" aria-label="Go to content"><div class="card__header"> <span class="card__content-type">opinion</span></div><h3 class="card__title">With its latest layoffs, Microsoft goes all in on AI</h3><div class="card__info"><span>By Preston Gralla</span></div>
		<div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-14T11:00:00+00:00">Jul 14, 2026</span></span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">IT Strategy</span></span><span class="card__tag"><span class="tag">Microsoft</span></span></div></a>
		</div><div class="latest-content__card-secondary"><a class="card " href="https://www.computerworld.com/article/4196652/forg365-industrializes-microsoft-365-phishing-with-ai-generated-lures.html" aria-label="Go to content"><div class="card__header"> <span class="card__content-type">news</span></div><h3 class="card__title">Forg365 industrializes Microsoft 365 phishing with AI-generated lures</h3><div class="card__info"><span>By Prasanth Aby Thomas</span></div>
		<div class="card__info card__info--light"><span><span itemprop="datePublished" content="2026-07-14T09:51:16+00:00">Jul 14, 2026</span></span><span>4 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Microsoft 365</span></span><span class="card__tag"><span class="tag">Office Suites</span></span><span class="card__tag"><span class="tag">Productivity Software</span></span></div></a>
		</div></div></div></div></section><div class="advert">
						<div class="container advert__container">
							<div class="advert__content">
								<div class="ad page-ad has-ad-prefix ad-article" data-ad-template="article" data-ofp="false"></div>
							</div>
						</div>
					</div><div class="content-listing-articles"><div class="container"><h2 class="content-listing-articles__title">Articles</h2><div class="content-listing-articles__container content-listing-articles__container--collapsed" data-collapse-articles="6" data-content-listing-articles><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4196365/openclaw-becomes-a-nonprofit-foundation-as-it-seeks-to-be-the-switzerland-of-ai.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">OpenClaw becomes a nonprofit foundation as it seeks to be ‘the Switzerland of AI’</h3><p class="card__description">Analysts and consultants applaud the move as potentially delivering the development consistency that the current offerings lack, but some worry that treating the company as neutral is a mistake.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Evan Schuman</span></div> <div class="card__info card__info--light"><span>Jul 13, 2026 </span><span>8 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">Nonprofits</span></span></div></div></div></a></div><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4196262/ai-is-killing-low-cost-smartphones.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news analysis</span></div><h3 class="card__title">AI is killing low cost smartphones</h3><p class="card__description">Data from Omdia and Counterpoint shows that while Apple and Samsung thrive, the rest of the industry takes a dive</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Jonny Evans</span></div> <div class="card__info card__info--light"><span>Jul 13, 2026 </span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Mobile Phones</span></span><span class="card__tag"><span class="tag">Smartphones</span></span></div></div></div></a></div><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4196220/meta-pulls-instagram-ai-feature-amid-privacy-concerns.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">Meta pulls Instagram AI feature amid privacy concerns</h3><p class="card__description">By specifying a public account, users could allow the AI ​​model to use the person’s images as a reference without the account holder being notified.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Viktor Eriksson</span></div> <div class="card__info card__info--light"><span>Jul 13, 2026 </span><span>1 min</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">Instagram</span></span></div></div></div></a></div><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4195176/qa-how-google-plans-to-reinvent-the-spreadsheet-with-ai.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">feature</span></div><h3 class="card__title">Q&amp;A: How Google plans to reinvent the spreadsheet with AI</h3><p class="card__description">Soon, Google wants to see AI doing the spreadsheet busywork, says Eric Birnbaum, director of product management for Google Sheets.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Matthew Finnegan</span></div> <div class="card__info card__info--light"><span>Jul 13, 2026 </span><span>10 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">Google Sheets</span></span><span class="card__tag"><span class="tag">Google Workspace</span></span></div></div></div></a></div><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4194931/physical-ai-will-see-the-fusion-of-robotics-and-ai-transform-the-world.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">brandpost</span><span class="card__sponsor-text">Sponsored by Tether</span></div><h3 class="card__title">Physical AI will see the fusion of robotics and AI transform the world</h3><p class="card__description"></p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By tether</span></div> <div class="card__info card__info--light"><span>Jul 9, 2026 </span><span>6 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div></div></a></div><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4195828/rotten-to-its-core-apple-files-an-explosive-lawsuit-against-openai.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news analysis</span></div><h3 class="card__title">‘Rotten to its core’ — Apple files an explosive lawsuit against OpenAI</h3><p class="card__description">Apple accuses OpenAI and former Apple Vice President Tang Tan of extensive coordinated data theft and asks whether OpenAI’s hardware plans are based around exfiltrated Apple info.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Jonny Evans</span></div> <div class="card__info card__info--light"><span>Jul 11, 2026 </span><span>6 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span></div></div></div></a></div><div class="content-listing-articles__row "><a class="grid content-row-article" href="https://www.computerworld.com/article/4195657/apple-is-prepping-for-life-after-the-ai-gold-rush.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">opinion</span></div><h3 class="card__title">Apple is prepping for life after the AI gold rush</h3><p class="card__description">The company's interest in compression of AI models is the right approach.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Jonny Evans</span></div> <div class="card__info card__info--light"><span>Jul 11, 2026 </span><span>6 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span></div></div></div></a></div><div class="content-listing-articles__row content-listing-articles__row--hide"><a class="grid content-row-article" href="https://www.computerworld.com/article/4195678/microsoft-exchange-server-on-prem-gets-a-little-harder-to-use.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">Microsoft Exchange Server on prem gets a little harder to use</h3><p class="card__description">The lightweight web client is going away, placing more demands on systems still clinging to Microsoft’s on-prem email system.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Maxwell Cooter</span></div> <div class="card__info card__info--light"><span>Jul 10, 2026 </span><span>2 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Email Clients</span></span><span class="card__tag"><span class="tag">Microsoft Exchange</span></span><span class="card__tag"><span class="tag">Microsoft Outlook</span></span></div></div></div></a></div><div class="content-listing-articles__row content-listing-articles__row--hide"><a class="grid content-row-article" href="https://www.computerworld.com/article/4195636/mistral-joins-rush-to-build-physical-ai.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">Mistral joins rush to build physical AI</h3><p class="card__description">Its Robostral Navigate AI model needs input from just one color camera, doing without Lidar, depth sensors, or multiple viewpoints.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Maxwell Cooter</span></div> <div class="card__info card__info--light"><span>Jul 10, 2026 </span><span>2 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Robotics</span></span></div></div></div></a></div><div class="content-listing-articles__row content-listing-articles__row--hide"><a class="grid content-row-article" href="https://www.computerworld.com/article/4195628/apple-will-buy-more-us-made-components-from-broadcom.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">Apple will buy more US-made components from Broadcom</h3><p class="card__description">Chips and thin-film bulk acoustic resonator (FBAR) filters are on the menu.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Maxwell Cooter</span></div> <div class="card__info card__info--light"><span>Jul 10, 2026 </span><span>2 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Networking</span></span><span class="card__tag"><span class="tag">Wi-Fi</span></span></div></div></div></a></div><div class="content-listing-articles__row content-listing-articles__row--hide"><a class="grid content-row-article" href="https://www.computerworld.com/article/4195528/meta-launches-low-cost-muse-spark-1-1-as-enterprise-ai-spending-comes-under-scrutiny-2.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">Meta launches low-cost Muse Spark 1.1 as enterprise AI spending comes under scrutiny</h3><p class="card__description">Meta says the model delivers competitive performance against OpenAI, Anthropic, and Google offerings while costing a fraction as much to run.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Anirban Ghoshal</span></div> <div class="card__info card__info--light"><span>Jul 10, 2026 </span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span></div></div></div></a></div><div class="content-listing-articles__row content-listing-articles__row--hide"><a class="grid content-row-article" href="https://www.computerworld.com/article/1614899/android-contacts-management-ultimate-guide.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">how-to</span></div><h3 class="card__title">The ultimate guide to Android contacts management</h3><p class="card__description">Your Android phone's contacts are much more than just a glorified Rolodex. Ready for an unexpected productivity upgrade? </p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By JR Raphael</span></div> <div class="card__info card__info--light"><span>Jul 10, 2026 </span><span>16 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Android</span></span><span class="card__tag"><span class="tag">Google</span></span><span class="card__tag"><span class="tag">Productivity Software</span></span></div></div></div></a></div><div class="content-listing-articles__row content-listing-articles__row--hide"><a class="grid content-row-article" href="https://www.computerworld.com/article/4195494/openai-launches-chatgpt-work-as-it-broadens-gpt-5-6-rollout-2.html" aria-label="Go to content"><div class="col-12 col-7@md content-row-article__main"><div class="card card--lg"><div class="card__header"><span class="card__content-type">news</span></div><h3 class="card__title">OpenAI launches ChatGPT Work as it broadens GPT-5.6 rollout</h3><p class="card__description">The enterprise AI agent combines ChatGPT, Codex, and GPT-5.6 to automate workplace tasks as OpenAI broadens rollout of its latest frontier models.</p></div></div><div class="col-12 col-4@md col-start-9@md content-row-article__secondary"><div class="card card--lg"><div class="card__info"><span>By Gyana Swain</span></div> <div class="card__info card__info--light"><span>Jul 10, 2026 </span><span>5 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">Productivity Software</span></span></div></div></div></a></div></div><div class="grid content-listing-articles__button-wrapper">
			<div class="col-6 col-4@md col-start-5@md"><div class="content-listing-articles__button-show">
					<button class="button button--tertiary" type="button" data-toggle="expand">
						<span>Show more</span>
						<span>
							<svg class="icon icon--sm" viewbox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
								<use xlink:href="#icon-chevron-down"></use>
							</svg>
						</span>
					</button>
				</div>
				<div class="content-listing-articles__button-show content-listing-articles__button-show--hide">
					<button class="button button--tertiary" type="button" data-toggle="collapse">
						<span>Show less</span>
						<span>
							<svg class="icon icon--sm" viewbox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
								<use xlink:href="#icon-chevron-up"></use>
							</svg>
						</span>
					</button>
				</div></div><div class="col-6 col-4@md content-listing-articles__button-view-all">
						<a class="button" href="https://www.computerworld.com/artificial-intelligence/feed/page/2/" target="_blank"> View all </a></div></div></div></div><section class="suggested-content-upcoming-events"><div class="container">
				<h2 class="suggested-content-upcoming-events__title">Upcoming Events</h2><a class="grid suggested-content-upcoming-events__item" href="https://event.foundryco.com/cio-100-uk/" aria-label="Go to content"><div class="col-12 col-3@md suggested-content-upcoming-events__date-label dd"><span class="date-label">Sep/24</span></div><div class="col-12 col-4@md col-5@xl suggested-content-upcoming-events__image"><div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/03/4141846-0-37933000-1772809522-CIO-Summit-2025_17.jpg?quality=50&amp;strip=all&amp;w=1045" alt="Image"></div></div>
			<div class="col-12 col-5@md col-4@xl suggested-content-upcoming-events__card">
				<div class="card card--xl">
					<div class="card__header"><span class="card__content-type">conference</span><span class="card__external-link-icon" data-url="https://event.foundryco.com/cio-100-uk/"><svg class="icon icon--sm" viewbox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg"> <use xlink:href="#icon-arrow-up-right-from-square"></use></svg></span></div><h3 class="card__title">CIO 100 Awards &amp; Conference UK</h3><div class="card__info card__info--light"><span>24 Sep 2026</span><span>London, UK</span></div>
		<div class="card__tags"><span class="card__tag"><span class="tag">Microsoft 365</span></span></div></div>
			</div>
		</a><a class="grid suggested-content-upcoming-events__item" href="https://event.foundryco.com/cso-awards-conference-uk/" aria-label="Go to content"><div class="col-12 col-3@md suggested-content-upcoming-events__date-label dd"><span class="date-label">Nov/26</span></div><div class="col-12 col-4@md col-5@xl suggested-content-upcoming-events__image"><div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/06/4141741-0-97812100-1780312469-60CB82BE-5D6E-40E0-8E5E-0151C8C46E7F.jpg?quality=50&amp;strip=all&amp;w=929" alt="Image"></div></div>
			<div class="col-12 col-5@md col-4@xl suggested-content-upcoming-events__card">
				<div class="card card--xl">
					<div class="card__header"><span class="card__content-type">conference</span><span class="card__external-link-icon" data-url="https://event.foundryco.com/cso-awards-conference-uk/"><svg class="icon icon--sm" viewbox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg"> <use xlink:href="#icon-arrow-up-right-from-square"></use></svg></span></div><h3 class="card__title">CSO Awards &amp; Conference UK</h3><div class="card__info card__info--light"><span>26 Nov 2026</span><span>London, UK</span></div>
		<div class="card__tags"><span class="card__tag"><span class="tag">Cyberattacks</span></span></div></div>
			</div>
		</a></div><div class="suggested-content-upcoming-events__button-container container">
						<a class="button" href="https://www.computerworld.com/events/"> View all events</a>
					</div>
				
			</section><div class="advert">
						<div class="container advert__container">
							<div class="advert__content">
								<div class="ad page-ad has-ad-prefix ad-article" data-ad-template="article" data-ofp="false"></div>
							</div>
						</div>
					</div><section class="related-content-resources">
				<div class="container">
				<h2 class="related-content-resources__title">Resources</h2><div class="grid related-content-resources__content"><div class="col-12 col-7@md col-8@lg grid grid--cols-7@md grid--cols-8@lg related-content-resources__main-content">
			<div class="col-12 col-7@md col-6@lg">
				<a class="card card--xxl" href="https://us.resources.computerworld.com/resources/accelerate-your-cloud-migration-with-atlassian-fastshift-6?utm_source=rss-feed&amp;utm_medium=rss&amp;utm_campaign=feed" rel="noreferrer" aria-label="Go to content">
					<div class="card__header">
						<span class="card__content-type">whitepaper</span>
					</div>
					<h3 class="card__title">Accelerate your cloud migration with Atlassian FastShift</h3>
					<p class="card__description"></p><p>Turn an Atlassian cloud migration into a faster, more predictable transformation. In this session, you’ll walk through the FastShift playbook.</p>
<p>The post <a rel="nofollow" href="https://com.wp.idg.zone/resources/accelerate-your-cloud-migration-with-atlassian-fastshift-6/">Accelerate your cloud migration with Atlassian FastShift</a> appeared first on <a rel="nofollow" href="https://com.wp.idg.zone/">Whitepaper Repository –</a>.</p>

					<div class="card__info">
						<span>
						By 
						Atlassian
						</span>
					</div>
					<div class="card__info card__info--light"><span>14 Jul 2026</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Business Operations</span></span><span class="card__tag"><span class="tag">Cloud</span></span><span class="card__tag"><span class="tag">Digital Transformation</span></span></div></a>
			</div>
			<div class="col-2 related-content-resources__featured-image-wrapper">
				<img width="400px" loading="lazy" class="related-content-resources__image-featured" src="https://us.resources.computerworld.com/wp-content/uploads/2026/07/atl_logo1784040704.83.png" alt="Image">
			</div>
		</div><div class="col-12 col-5@md col-4@lg col-start-9@lg related-content-resources__cards"><div class="grid grid--cols-5@md grid--cols-4@lg related-content-resources__card-wrapper">
				<div class="col-12 col-5@md col-3@lg">
					<a class="card card--sm" href="https://us.resources.computerworld.com/resources/warum-sich-teams-fur-cloud-entscheiden-9?utm_source=rss-feed&amp;utm_medium=rss&amp;utm_campaign=feed" rel="noreferrer" aria-label="Go to content">
						<div class="card__header">
							<span class="card__content-type">whitepaper</span>
						</div>
						<h3 class="card__title">Warum sich Teams für Cloud entscheiden</h3>
						<div class="card__info">
							<span>
							By 
							Atlassian
							</span>
						</div>
						<div class="card__info card__info--light"><span>14 Jul 2026</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Business Operations</span></span><span class="card__tag"><span class="tag">Cloud</span></span><span class="card__tag"><span class="tag">Digital Transformation</span></span></div></a>
				</div>
				<div class="col-1">
					<img width="400px" loading="lazy" class="related-content-resources__image-side" src="https://us.resources.computerworld.com/wp-content/uploads/2026/07/atl_logo1784040716.4772.png" alt="Image">
				</div>
			</div><div class="grid grid--cols-5@md grid--cols-4@lg related-content-resources__card-wrapper">
				<div class="col-12 col-5@md col-3@lg">
					<a class="card card--sm" href="https://us.resources.computerworld.com/resources/pourquoi-les-equipes-optent-pour-la-solution-cloud-3?utm_source=rss-feed&amp;utm_medium=rss&amp;utm_campaign=feed" rel="noreferrer" aria-label="Go to content">
						<div class="card__header">
							<span class="card__content-type">whitepaper</span>
						</div>
						<h3 class="card__title">Pourquoi les équipes optent pour la solution cloud</h3>
						<div class="card__info">
							<span>
							By 
							Atlassian
							</span>
						</div>
						<div class="card__info card__info--light"><span>14 Jul 2026</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Business Operations</span></span><span class="card__tag"><span class="tag">Cloud</span></span><span class="card__tag"><span class="tag">Digital Transformation</span></span></div></a>
				</div>
				<div class="col-1">
					<img width="400px" loading="lazy" class="related-content-resources__image-side" src="https://us.resources.computerworld.com/wp-content/uploads/2026/07/atl_logo1784040728.9116.png" alt="Image">
				</div>
			</div></div>
		</div><div class="related-content-resources__button-container">
			<a class="button" target="_blank" href="https://us.resources.computerworld.com/"> View all </a>
		</div></div>
			</section><div class="advert">
						<div class="container advert__container">
							<div class="advert__content">
								<div class="ad page-ad has-ad-prefix ad-article" data-ad-template="article" data-ofp="false"></div>
							</div>
						</div>
					</div><section class="related-content-podcasts"><div class="container"><h2 class="related-content-podcasts__title">Podcasts</h2><div class="grid related-content-podcasts__content"><a class="col-12 col-7@md col-8@lg grid grid--cols-7@md grid--cols-8@lg related-content-podcasts__main-content" href="https://www.computerworld.com/podcasts/2-minute-tech-briefing/" aria-label="Go to content"><div class="col-12 col-7@md col-2@lg related-content-podcasts__image">
			<div class="image image--aspect-ratio-1-1">
				<img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2025/11/100065453-0-01782600-1762961273-2-min-tech-briefing-logo-16x9-4.jpg?quality=50&amp;strip=all&amp;w=1024" alt="Image">
			</div>
		</div><div class="col-12 col-7@md col-6@lg"><div class="card card--xl"><div class="card__header"><span class="card__content-type"> podcasts</span></div><h3 class="card__title">2-Minute Tech Briefing</h3><p class="card__description">Catch up on the latest enterprise IT news in a fast-paced video briefing with host Arnold Davick. Listen to the show on Computerworld, YouTube, Apple and Spotify.</p><div class="card__info card__info--light"><span>81  episodes</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Emerging Technology</span></span></div></div></div></a><ul class="col-12 col-5@md col-4@lg col-start-9@lg related-content-podcasts__cards"><li class="related-content-podcasts__card"><a href="https://www.computerworld.com/podcast/4176380/microsoft-copilot-growth-claudebleed-risk-linkedin-gdpr-complaint-ep-84.html" aria-label="Go to episode"><div class="related-content-podcasts__episode-label">
			<span class="episode-label">
				<span> Ep. 81</span>
				<span>
				<svg class="icon" viewbox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
					<use xlink:href="#icon-podcast"></use>
				</svg>
			</span>
			</span>
		</div><div class="card card--xs"><h3 class="card__title">Microsoft Copilot Growth, ClaudeBleed Risk, LinkedIn GDPR Complaint | Ep. 84</h3><div class="card__info">
				<span>By Arnold Davick</span>
			</div><div class="card__info card__info--light">
			<span>Mar 20, 2024</span><span>2 mins</span>
		</div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div></a></li><li class="related-content-podcasts__card"><a href="https://www.computerworld.com/podcast/4176367/chrome-gemini-ai-agents-cisa-infrastructure-cyber-resilience-ep-83.html" aria-label="Go to episode"><div class="related-content-podcasts__episode-label">
			<span class="episode-label">
				<span> Ep. 80</span>
				<span>
				<svg class="icon" viewbox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
					<use xlink:href="#icon-podcast"></use>
				</svg>
			</span>
			</span>
		</div><div class="card card--xs"><h3 class="card__title">Chrome Gemini, AI Agents, CISA Infrastructure Cyber Resilience | Ep. 83</h3><div class="card__info">
				<span>By Arnold Davick</span>
			</div><div class="card__info card__info--light">
			<span>Mar 20, 2024</span><span>2 mins</span>
		</div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div></a></li></ul></div></div></section><section class="related-content-video"><div class="container"><h2 class="related-content-video__title">Video on demand</h2><div class="grid related-content-video__main">        <div class="col-12 col-4@lg related-content-video__main-card card card--xl">
            <div class="card__header"><span class="card__content-type">video</span></div>            
            <a class="card card--xl" href="https://www.computerworld.com/video/4196734/why-ai-agents-fail-when-enterprises-dont-define-the-job.html" aria-label="Go to content">
                <h3 class="card__title">Why AI agents fail when enterprises don’t define the job</h3>            </a>
                            <p class="card__description mt-3">Enterprises are investing heavily in AI agents, but many projects fail when companies skip clear goals, guardrails, governance and success metrics.</p>
            
                         <div class="card__info card__info--light"><span>Jul 14, 2026 </span><span>33 mins</span></div><div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">IT Governance</span></span></div>        </div>
                <div class="col-12 col-8@lg related-content-video__video">
                            <div class="youtube-video">
                    &gt;
					
				</div>                </div>
                    </div>
        </div><div class="related-content-video__cards-container">
                        <div class="related-content-video__cards-wrap">
                            <ul class="grid related-content-video__cards">        <li class="col-4@md related-content-video__card">
            <a class="related-content-video__card-link" href="https://www.computerworld.com/video/4193952/why-enterprise-ai-projects-stall-before-delivering-real-value.html" aria-label="Go to content">
                <div class="related-content-video__card-image">
                    <div class="image">
                        <img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/07/4193952-0-52301800-1783446508-youtube-thumbnail-gu6x40jhZ1s_3cbf50.jpg?quality=50&amp;strip=all&amp;w=300" alt="Image" sizes="300px">
                    </div>
                </div>
                <div class="card card--xs">
                    <h3 class="card__title">Why enterprise AI projects stall before delivering real value</h3>
                                         <div class="card__info card__info--light"><span>Jul 7, 2026 </span><span>29 mins</span></div>                    <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">ROI and Metrics</span></span></div>                </div>
            </a>
        </li>
                <li class="col-4@md related-content-video__card">
            <a class="related-content-video__card-link" href="https://www.computerworld.com/video/4191262/how-ai-is-breaking-job-interviews-skills-testing-and-evaluation.html" aria-label="Go to content">
                <div class="related-content-video__card-image">
                    <div class="image">
                        <img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/06/4191262-0-24248500-1782847008-youtube-thumbnail-lVEejCXC4lU_b223c5.jpg?quality=50&amp;strip=all&amp;w=300" alt="Image" sizes="300px">
                    </div>
                </div>
                <div class="card card--xs">
                    <h3 class="card__title">How AI is breaking job interviews, skills testing and evaluation</h3>
                                         <div class="card__info card__info--light"><span>Jun 30, 2026 </span><span>32 mins</span></div>                    <div class="card__tags"><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">Hiring</span></span><span class="card__tag"><span class="tag">IT Skills and Training</span></span></div>                </div>
            </a>
        </li>
                <li class="col-4@md related-content-video__card">
            <a class="related-content-video__card-link" href="https://www.computerworld.com/video/4188534/how-ai-is-reshaping-cybersecurity.html" aria-label="Go to content">
                <div class="related-content-video__card-image">
                    <div class="image">
                        <img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/06/4188534-0-45176600-1782243369-youtube-thumbnail-5DLoQMU0nZc_de9df9.jpg?quality=50&amp;strip=all&amp;w=300" alt="Image" sizes="300px">
                    </div>
                </div>
                <div class="card card--xs">
                    <h3 class="card__title">How AI is reshaping cybersecurity</h3>
                                         <div class="card__info card__info--light"><span>Jun 23, 2026 </span><span>44 mins</span></div>                    <div class="card__tags"><span class="card__tag"><span class="tag">Cyberattacks</span></span><span class="card__tag"><span class="tag">Cybercrime</span></span><span class="card__tag"><span class="tag">Generative AI</span></span></div>                </div>
            </a>
        </li>
        </ul></div></div><div class="related-content-video__button-container"><a class="button" target="_self" href="https://www.computerworld.com/videos/">See all videos</a></div></section></div><section class="suggested-content-various"><div class="container"><div class="grid suggested-content-various__content"><div class="col-12 col-3@lg">
			<h2 class="suggested-content-various__title">Show me more</h2><div class="suggested-content-various__filters"><span class="suggested-content-various__filter"><button class="chip chip--filter chip--active" type="button" data-filter-key="latest">Latest</button></span><span class="suggested-content-various__filter"><button class="chip chip--filter" type="button" data-filter-key="article">Articles</button></span><span class="suggested-content-various__filter"><button class="chip chip--filter" type="button" data-filter-key="podcast">Podcasts</button></span><span class="suggested-content-various__filter"><button class="chip chip--filter" type="button" data-filter-key="video">Videos</button></span></div>
		</div><div class="col-12 col-9@lg suggested-content-various__items-wrap"><div class="grid grid--cols-9@lg suggested-content-various__items"><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				latest,article"><a class="suggested-content-various__link" href="https://www.computerworld.com/article/4195055/apple-finally-calls-time-on-15-year-old-device-support.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">opinion</span> </div> <h3 class="card__title">Apple finally calls time on 15-year-old device support</h3> <div class="card__info"><span>By Jonny Evans</span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-07-09T16:15:14+00:00">Jul 9, 2026</span><span>4 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Apple</span></span><span class="card__tag"><span class="tag">Smartphones</span></span><span class="card__tag"><span class="tag">iPhone</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/07/4195055-0-47500100-1783613766-iPhone4s_3up_Photo_Siri_Sprgbd_PRINT.jpg?quality=50&amp;strip=all&amp;w=219" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				article"><a class="suggested-content-various__link" href="https://www.computerworld.com/article/4194931/physical-ai-will-see-the-fusion-of-robotics-and-ai-transform-the-world.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">brandpost</span> <span class="card__sponsor-text">Sponsored by Tether</span></div> <h3 class="card__title">Physical AI will see the fusion of robotics and AI transform the world</h3> <div class="card__info"><span>By tether</span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-07-09T11:11:53+00:00">9 Jul 2026</span><span>6 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/07/4194931-0-76347600-1783595551-QVAC-Paid-Ad-1-_-1200-x-800.png?w=375" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				article"><a class="suggested-content-various__link" href="https://www.computerworld.com/article/4194914/spacexai-launches-grok-4-5-touts-lower-coding-task-costs-than-ai-rivals-2.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">news</span> </div> <h3 class="card__title">SpaceXAI launches Grok 4.5, touts lower coding-task costs than AI rivals</h3> <div class="card__info"><span>By Prasanth Aby Thomas</span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-07-09T10:26:11+00:00">Jul 9, 2026</span><span>5 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Developer</span></span><span class="card__tag"><span class="tag">Generative AI</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/07/4194914-0-24417700-1783592810-AI-vibe-coding-one-hand-is-robot-one-hand-is-human.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				latest,podcast"><a class="suggested-content-various__link" href="https://www.computerworld.com/podcast/4176380/microsoft-copilot-growth-claudebleed-risk-linkedin-gdpr-complaint-ep-84.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">podcast</span> </div> <h3 class="card__title">Microsoft Copilot Growth, ClaudeBleed Risk, LinkedIn GDPR Complaint | Ep. 84</h3> <div class="card__info"><span>By Arnold Davick</span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-05-22T15:04:16+00:00">May 22, 2026</span><span>2 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/05/0-46106000-1779462321-youtube-thumbnail-5PkKYThsKy8.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				podcast"><a class="suggested-content-various__link" href="https://www.computerworld.com/podcast/4176367/chrome-gemini-ai-agents-cisa-infrastructure-cyber-resilience-ep-83.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">podcast</span> </div> <h3 class="card__title">Chrome Gemini, AI Agents, CISA Infrastructure Cyber Resilience | Ep. 83</h3> <div class="card__info"><span>By Arnold Davick</span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-05-22T14:53:18+00:00">May 22, 2026</span><span>2 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/05/0-06017100-1779461653-youtube-thumbnail-XH7vduM7uz8.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				podcast"><a class="suggested-content-various__link" href="https://www.computerworld.com/podcast/4172579/ai-triage-gains-model-reviews-ask-jeeves-shutdown-ep-82.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">podcast</span> </div> <h3 class="card__title">AI Triage Gains, Model Reviews, Ask Jeeves Shutdown | Ep. 82</h3> <div class="card__info"><span>By Arnold Davick</span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-05-18T19:31:15+00:00">May 18, 2026</span><span>2 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/05/0-62824900-1779132751-youtube-thumbnail-P3R6blMndrU.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				latest,video"><a class="suggested-content-various__link" href="https://www.computerworld.com/video/4185559/why-ai-agents-could-create-a-new-control-and-security-crisis.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">video</span> </div> <h3 class="card__title">Why AI agents could create a new control and security crisis</h3> <div class="card__info"><span></span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-06-16T11:47:15+00:00">Jun 16, 2026</span><span>28 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Artificial Intelligence</span></span><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">IT Governance</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/06/4185559-0-48713800-1781610470-youtube-thumbnail-uPpd9EJ4iNI_55eb26.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				video"><a class="suggested-content-various__link" href="https://www.computerworld.com/video/4182978/does-quality-suffer-when-ai-generates-code.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">video</span> </div> <h3 class="card__title">Does quality suffer when AI generates code?</h3> <div class="card__info"><span></span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-06-09T14:32:51+00:00">Jun 9, 2026</span><span>35 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Code Security</span></span><span class="card__tag"><span class="tag">Developer</span></span><span class="card__tag"><span class="tag">Generative AI</span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/06/4182978-0-61230000-1781015610-youtube-thumbnail-1hAfDQkuyhs_faa994.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div><div class="col-4@md col-3@lg suggested-content-various__item suggested-content-various__item--active" data-filter-value="
				video"><a class="suggested-content-various__link" href="https://www.computerworld.com/video/4180043/what-happens-when-ai-starts-selling-to-ai.html" aria-label="Go to content"><div class="card">
					<div class="card__header">
						<span class="card__content-type">video</span> </div> <h3 class="card__title">What happens when AI starts selling to AI?</h3> <div class="card__info"><span></span></div><div class="card__info card__info--light"><span itemprop="datePublished" content="2026-06-02T15:00:42+00:00">Jun 2, 2026</span><span>38 mins</span></div>
				 <div class="card__tags"><span class="card__tag"><span class="tag">Generative AI</span></span><span class="card__tag"><span class="tag">Procurement Software</span></span><span class="card__tag"><span class="tag">Salesforce Automation </span></span></div></div>
					<div class="image"><img width="400px" loading="lazy" src="https://www.computerworld.com/wp-content/uploads/2026/06/4180043-0-78180900-1780412479-youtube-thumbnail-jPv-TAenlto_c79318.jpg?quality=50&amp;strip=all&amp;w=444" alt="Image"></div>
				</a>
			</div></div></div></div></div></section>]]></content:encoded>
</item>
<item>
<title><![CDATA[DeepMind CEO again pushes for a frontier AI standards body]]></title>
<description><![CDATA[Google DeepMind CEO Demis Hassabis on Tuesday reiterated his push for an AI industry self-regulation effort, led by the US government, that is particularly focused on artificial general intelligence (AGI) and national security. 



But it is precisely that focus on national security that may make...]]></description>
<link>https://tsecurity.de/de/3671860/it-nachrichten/deepmind-ceo-again-pushes-for-a-frontier-ai-standards-body/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671860/it-nachrichten/deepmind-ceo-again-pushes-for-a-frontier-ai-standards-body/</guid>
<pubDate>Wed, 15 Jul 2026 23:01:43 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Google DeepMind CEO Demis Hassabis on Tuesday reiterated his push for an AI industry self-regulation effort, led by the US government, that is particularly focused on <a href="https://www.computerworld.com/article/4174181/google-talks-singularity-while-scaling-up-agentic-ai-for-enterprises-2.html" target="_blank">artificial general intelligence (AGI)</a> and national security. </p>



<p class="wp-block-paragraph">But it is precisely that focus on national security that may make the results of such an effort, assuming it happens, less than palatable outside of the US.</p>



<p class="wp-block-paragraph">“The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous,” <a href="https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age" target="_blank" rel="noreferrer noopener">Hassabis wrote</a>. “The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”</p>



<p class="wp-block-paragraph">He noted, however, that the funding would need to be substantial, and would most likely come from industry, to allow the new body to attract world-class technical talent and obtain the necessary compute resources for large-scale testing.</p>



<p class="wp-block-paragraph">Hassabis said he would propose that the organization “be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security,” and that AI vendor participants would be encouraged to adopt best practices, such as publishing model cards with technical details, maintaining strong internal cybersecurity, vetting key personnel, and providing sufficient resourcing for safety and security research.</p>



<p class="wp-block-paragraph">This is not the first time Hassabis has <a href="https://www.computerworld.com/article/4178398/deepmind-ceo-agi-could-be-here-in-three-years.html" target="_blank">expressed worries about AGI</a>. He has already worked on <a href="https://www.cio.com/article/4168122/us-government-agency-to-safety-test-frontier-ai-models-before-release.html" target="_blank">a US government initiative evaluating AI safety</a>, which involved DeepMind, Microsoft and xAI (now SpaceXAI) working with the Center for AI Standards and Innovation (CAISI), a division of the US Department of Commerce. It allowed CAISI to conduct pre-deployment evaluations and targeted research to “better assess frontier AI capabilities and advance the state of AI security.”  </p>



<h2 class="wp-block-heading">The rest of the world may have concerns</h2>



<p class="wp-block-paragraph">Analysts and consultants were mixed about the move, with most expressing concerns about whether an industry-focused group would prioritize the public’s best interests.</p>



<p class="wp-block-paragraph">“Self-regulation is not viable because it implies everyone is able to regulate themselves and will do so in line with the best interests of the public. Most tech vendors don’t have the capacity to self-regulate. They would just prefer a set of rules within which they can operate,” said Gartner VP analyst <a href="https://www.gartner.com/en/experts/nader-henein" target="_blank" rel="noreferrer noopener">Nader Henein</a>. “For-profit organizations are required to do what is best for their shareholders, and external regulation ensures that those organizations are never in a conflict of interest where they have to choose between what is good for their shareholders and what is good for the public.”</p>



<p class="wp-block-paragraph">And, said <a href="https://greyhoundresearch.com/svg/" target="_blank" rel="noreferrer noopener">Sanchit Vir Gogia</a>, chief analyst at Greyhound Research, given the international nature of AI models, an effort coordinated by the US government might alienate other countries. </p>



<p class="wp-block-paragraph">“National security is the proposal’s accelerator in Washington and its poison pill abroad: the framing that opens the only gate available at home invites foreign capitals to read the institution as an instrument of American strategy,” he pointed out. </p>



<p class="wp-block-paragraph">“The map is already plural,” he said. “Brussels switches on enforcement powers over general-purpose models [starting in August 2026], London runs the AI Security Institute, and Beijing licenses on its own terms. California and New York have legislated for frontier models at home. The durable route is shared technical evidence with sovereign enforcement, sealed through mutual recognition rather than deference, with India and the other major non-Western markets holding authorship rather than seats.”</p>



<p class="wp-block-paragraph">Gogia added that the rules enacted by even such a group may not address all of the key concerns of enterprise IT. A US government effort along the lines that Hassabis is proposing would result in testing that “sits close to intelligence and industrial policy, and those functions will not stay neatly separated. A model can pass every catastrophic-risk test and still fail the enterprise on privacy, reliability, and liability,” he noted.</p>



<p class="wp-block-paragraph">Walmart’s former director of cybersecurity <a href="https://www.linkedin.com/in/steveneric/" target="_blank" rel="noreferrer noopener">Steven Eric Fisher</a>, who is now an independent cybersecurity consultant, said he found the proposal “well-intentioned, but it addresses a highly polarized topic at a time when commercial interests carry unprecedented political influence, which is not always applied benevolently.”</p>



<p class="wp-block-paragraph">He added, “an exclusive US standard that is not globally respected or enforceable would likely fail to achieve its core purpose and would place US companies at a competitive disadvantage.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="noreferrer noopener">Aman Mahapatra</a>, chief strategy officer for Tribeca Softtech, a New York City-based technology consulting firm, said that a deep dive into how <a href="https://www.finra.org/" target="_blank" rel="noreferrer noopener">FINRA</a> operates today is illustrative of what IT leaders can expect from this effort, assuming the industry adopts that model.</p>



<p class="wp-block-paragraph">“When the CEOs of the five companies that would be regulated are also the primary drafters of the standards, the standards will reflect those companies’ interests. FINRA has an independent board, but the operational reality is that member firm perspectives dominate the working groups that write the actual rules,” he said. “There is no reason to expect an AI equivalent to work differently, and every reason to expect it to work worse, because AI standardization is happening faster than any industry has ever attempted to standardize itself, and speed is the enemy of independent oversight.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/carmi/" target="_blank" rel="noreferrer noopener">Carmi Levy</a>, an independent technology analyst, was even more emphatically opposed to the Hassabis proposal.</p>



<p class="wp-block-paragraph">“Asking Big Tech companies to self-police is analogous to allowing foxes to guard the henhouse. It hasn’t worked to date, and it won’t work going forward. Expecting these organizations to somehow change their ways at this point in time represents the height of naïve thinking,” Levy said. “The framework proposed by Demis Hassabis is a self-serving roadmap for an industry bent on racing to the AI horizon regardless of the harms caused along the way. It is impossible to quantify the dangers to broader society should frameworks allowing self-regulation become the norm.”</p>



<h2 class="wp-block-heading">Some love the proposal</h2>



<p class="wp-block-paragraph">An almost completely opposite stance came from <a href="https://www.linkedin.com/in/yurigoryunov/" target="_blank" rel="noreferrer noopener">Yuri Goryunov</a>, CIO of consulting firm Acceligence, who applauded the proposed move.</p>



<p class="wp-block-paragraph">“This is one of the rare setups where industry self-regulation has a real shot, and enterprise IT should be enthusiastically rooting for it,” he said. “It fails when harms are externalized, such as in social media content moderation. Or when the overseer outsources judgment to the overseen, such as the FAA’s delegation to Boeing before the 737 MAX. It works when everyone in the industry shares the catastrophic downside.”</p>



<p class="wp-block-paragraph">He suggested, however, that the best precedent here isn’t FINRA, it’s INPO, the Institute of Nuclear Power Operations, which the nuclear industry created within months of the <a href="https://www.nrc.gov/reading-rm/doc-collections/fact-sheets/3mile-isle" target="_blank" rel="noreferrer noopener">1979 Three Mile Island partial reactor meltdown</a> “on the logic that an accident anywhere is an accident everywhere. INPO peer-reviews every US plant, its evaluations move insurance premiums, and it sits on top of the NRC’s statutory floor. That is a public-private stack very close to what Hassabis is describing. Frontier AI has the same structure: one lab’s catastrophic failure brings regulation down on all of them.”</p>



<p class="wp-block-paragraph">For enterprise CIOs and other IT executives, Goryunov said, that model has the potential for being a big win.</p>



<p class="wp-block-paragraph"><strong>“</strong>Today, every enterprise duplicates the same AI diligence of red-teaming, eval suites, governance committees and each does so with less information than any certifying body would have,” Goryunov said. “A credible standards regime does for AI what UL did for electrical equipment and SOC2 did for cloud: it converts an unknowable risk into a procurable product and gives boards a defensible standard of care. That’s not red tape. That’s peace of mind with an audit trail.”</p>



<p class="wp-block-paragraph">However, Mahapatra said, “the countervailing view is that the alternative to industry-led standards is probably not thoughtful legislation. It is probably no standards, or state-by-state fragmentation, or the current pattern of ex-post enforcement actions where regulators surface concerns years after harm has already occurred.” </p>



<p class="wp-block-paragraph">Thus, he noted, “Hassabis is making the reasonable argument that imperfect fast standards are better than perfect slow ones, and there is genuine merit to that view for topics like agent identity, evaluation methodology, and interoperability, which are exactly the areas <a href="https://www.computerworld.com/article/4196365/openclaw-becomes-a-nonprofit-foundation-as-it-seeks-to-be-the-switzerland-of-ai.html" target="_blank">OpenClaw is also targeting</a>.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[DeepMind CEO again pushes for a frontier AI standards body]]></title>
<description><![CDATA[Google DeepMind CEO Demis Hassabis on Tuesday reiterated his push for an AI industry self-regulation effort, led by the US government, that is particularly focused on artificial general intelligence (AGI) and national security. 



But it is precisely that focus on national security that may make...]]></description>
<link>https://tsecurity.de/de/3671859/it-nachrichten/deepmind-ceo-again-pushes-for-a-frontier-ai-standards-body/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671859/it-nachrichten/deepmind-ceo-again-pushes-for-a-frontier-ai-standards-body/</guid>
<pubDate>Wed, 15 Jul 2026 23:01:42 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Google DeepMind CEO Demis Hassabis on Tuesday reiterated his push for an AI industry self-regulation effort, led by the US government, that is particularly focused on <a href="https://www.computerworld.com/article/4174181/google-talks-singularity-while-scaling-up-agentic-ai-for-enterprises-2.html" target="_blank">artificial general intelligence (AGI)</a> and national security. </p>



<p class="wp-block-paragraph">But it is precisely that focus on national security that may make the results of such an effort, assuming it happens, less than palatable outside of the US.</p>



<p class="wp-block-paragraph">“The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous,” <a href="https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age" target="_blank" rel="noreferrer noopener">Hassabis wrote</a>. “The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”</p>



<p class="wp-block-paragraph">He noted, however, that the funding would need to be substantial, and would most likely come from industry, to allow the new body to attract world-class technical talent and obtain the necessary compute resources for large-scale testing.</p>



<p class="wp-block-paragraph">Hassabis said he would propose that the organization “be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security,” and that AI vendor participants would be encouraged to adopt best practices, such as publishing model cards with technical details, maintaining strong internal cybersecurity, vetting key personnel, and providing sufficient resourcing for safety and security research.</p>



<p class="wp-block-paragraph">This is not the first time Hassabis has <a href="https://www.computerworld.com/article/4178398/deepmind-ceo-agi-could-be-here-in-three-years.html" target="_blank">expressed worries about AGI</a>. He has already worked on <a href="https://www.cio.com/article/4168122/us-government-agency-to-safety-test-frontier-ai-models-before-release.html" target="_blank">a US government initiative evaluating AI safety</a>, which involved DeepMind, Microsoft and xAI (now SpaceXAI) working with the Center for AI Standards and Innovation (CAISI), a division of the US Department of Commerce. It allowed CAISI to conduct pre-deployment evaluations and targeted research to “better assess frontier AI capabilities and advance the state of AI security.”  </p>



<h2 class="wp-block-heading">The rest of the world may have concerns</h2>



<p class="wp-block-paragraph">Analysts and consultants were mixed about the move, with most expressing concerns about whether an industry-focused group would prioritize the public’s best interests.</p>



<p class="wp-block-paragraph">“Self-regulation is not viable because it implies everyone is able to regulate themselves and will do so in line with the best interests of the public. Most tech vendors don’t have the capacity to self-regulate. They would just prefer a set of rules within which they can operate,” said Gartner VP analyst <a href="https://www.gartner.com/en/experts/nader-henein" target="_blank" rel="noreferrer noopener">Nader Henein</a>. “For-profit organizations are required to do what is best for their shareholders, and external regulation ensures that those organizations are never in a conflict of interest where they have to choose between what is good for their shareholders and what is good for the public.”</p>



<p class="wp-block-paragraph">And, said <a href="https://greyhoundresearch.com/svg/" target="_blank" rel="noreferrer noopener">Sanchit Vir Gogia</a>, chief analyst at Greyhound Research, given the international nature of AI models, an effort coordinated by the US government might alienate other countries. </p>



<p class="wp-block-paragraph">“National security is the proposal’s accelerator in Washington and its poison pill abroad: the framing that opens the only gate available at home invites foreign capitals to read the institution as an instrument of American strategy,” he pointed out. </p>



<p class="wp-block-paragraph">“The map is already plural,” he said. “Brussels switches on enforcement powers over general-purpose models [starting in August 2026], London runs the AI Security Institute, and Beijing licenses on its own terms. California and New York have legislated for frontier models at home. The durable route is shared technical evidence with sovereign enforcement, sealed through mutual recognition rather than deference, with India and the other major non-Western markets holding authorship rather than seats.”</p>



<p class="wp-block-paragraph">Gogia added that the rules enacted by even such a group may not address all of the key concerns of enterprise IT. A US government effort along the lines that Hassabis is proposing would result in testing that “sits close to intelligence and industrial policy, and those functions will not stay neatly separated. A model can pass every catastrophic-risk test and still fail the enterprise on privacy, reliability, and liability,” he noted.</p>



<p class="wp-block-paragraph">Walmart’s former director of cybersecurity <a href="https://www.linkedin.com/in/steveneric/" target="_blank" rel="noreferrer noopener">Steven Eric Fisher</a>, who is now an independent cybersecurity consultant, said he found the proposal “well-intentioned, but it addresses a highly polarized topic at a time when commercial interests carry unprecedented political influence, which is not always applied benevolently.”</p>



<p class="wp-block-paragraph">He added, “an exclusive US standard that is not globally respected or enforceable would likely fail to achieve its core purpose and would place US companies at a competitive disadvantage.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="noreferrer noopener">Aman Mahapatra</a>, chief strategy officer for Tribeca Softtech, a New York City-based technology consulting firm, said that a deep dive into how <a href="https://www.finra.org/" target="_blank" rel="noreferrer noopener">FINRA</a> operates today is illustrative of what IT leaders can expect from this effort, assuming the industry adopts that model.</p>



<p class="wp-block-paragraph">“When the CEOs of the five companies that would be regulated are also the primary drafters of the standards, the standards will reflect those companies’ interests. FINRA has an independent board, but the operational reality is that member firm perspectives dominate the working groups that write the actual rules,” he said. “There is no reason to expect an AI equivalent to work differently, and every reason to expect it to work worse, because AI standardization is happening faster than any industry has ever attempted to standardize itself, and speed is the enemy of independent oversight.”</p>



<p class="wp-block-paragraph"><a href="https://www.linkedin.com/in/carmi/" target="_blank" rel="noreferrer noopener">Carmi Levy</a>, an independent technology analyst, was even more emphatically opposed to the Hassabis proposal.</p>



<p class="wp-block-paragraph">“Asking Big Tech companies to self-police is analogous to allowing foxes to guard the henhouse. It hasn’t worked to date, and it won’t work going forward. Expecting these organizations to somehow change their ways at this point in time represents the height of naïve thinking,” Levy said. “The framework proposed by Demis Hassabis is a self-serving roadmap for an industry bent on racing to the AI horizon regardless of the harms caused along the way. It is impossible to quantify the dangers to broader society should frameworks allowing self-regulation become the norm.”</p>



<h2 class="wp-block-heading">Some love the proposal</h2>



<p class="wp-block-paragraph">An almost completely opposite stance came from <a href="https://www.linkedin.com/in/yurigoryunov/" target="_blank" rel="noreferrer noopener">Yuri Goryunov</a>, CIO of consulting firm Acceligence, who applauded the proposed move.</p>



<p class="wp-block-paragraph">“This is one of the rare setups where industry self-regulation has a real shot, and enterprise IT should be enthusiastically rooting for it,” he said. “It fails when harms are externalized, such as in social media content moderation. Or when the overseer outsources judgment to the overseen, such as the FAA’s delegation to Boeing before the 737 MAX. It works when everyone in the industry shares the catastrophic downside.”</p>



<p class="wp-block-paragraph">He suggested, however, that the best precedent here isn’t FINRA, it’s INPO, the Institute of Nuclear Power Operations, which the nuclear industry created within months of the <a href="https://www.nrc.gov/reading-rm/doc-collections/fact-sheets/3mile-isle" target="_blank" rel="noreferrer noopener">1979 Three Mile Island partial reactor meltdown</a> “on the logic that an accident anywhere is an accident everywhere. INPO peer-reviews every US plant, its evaluations move insurance premiums, and it sits on top of the NRC’s statutory floor. That is a public-private stack very close to what Hassabis is describing. Frontier AI has the same structure: one lab’s catastrophic failure brings regulation down on all of them.”</p>



<p class="wp-block-paragraph">For enterprise CIOs and other IT executives, Goryunov said, that model has the potential for being a big win.</p>



<p class="wp-block-paragraph"><strong>“</strong>Today, every enterprise duplicates the same AI diligence of red-teaming, eval suites, governance committees and each does so with less information than any certifying body would have,” Goryunov said. “A credible standards regime does for AI what UL did for electrical equipment and SOC2 did for cloud: it converts an unknowable risk into a procurable product and gives boards a defensible standard of care. That’s not red tape. That’s peace of mind with an audit trail.”</p>



<p class="wp-block-paragraph">However, Mahapatra said, “the countervailing view is that the alternative to industry-led standards is probably not thoughtful legislation. It is probably no standards, or state-by-state fragmentation, or the current pattern of ex-post enforcement actions where regulators surface concerns years after harm has already occurred.” </p>



<p class="wp-block-paragraph">Thus, he noted, “Hassabis is making the reasonable argument that imperfect fast standards are better than perfect slow ones, and there is genuine merit to that view for topics like agent identity, evaluation methodology, and interoperability, which are exactly the areas <a href="https://www.computerworld.com/article/4196365/openclaw-becomes-a-nonprofit-foundation-as-it-seeks-to-be-the-switzerland-of-ai.html" target="_blank">OpenClaw is also targeting</a>.”</p>



<p class="wp-block-paragraph"><em>This article originally appeared on CIO.com.</em></p>



<p class="wp-block-paragraph"></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Which AI model should you bet your company on? None of them]]></title>
<description><![CDATA[Every day this past week I did something I suspect millions of other people also did: I stared at an LLM model picker and wondered which one I was supposed to want.



OpenAI just released ⁠GPT-5.6 Sol, Terra, and Luna. Sol is the flagship. Terra offers much of its intelligence for less money. Lu...]]></description>
<link>https://tsecurity.de/de/3671165/ai-nachrichten/which-ai-model-should-you-bet-your-company-on-none-of-them/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671165/ai-nachrichten/which-ai-model-should-you-bet-your-company-on-none-of-them/</guid>
<pubDate>Wed, 15 Jul 2026 17:19:39 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div><div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Every day this past week I did something I suspect millions of other people also did: I stared at an <a href="https://www.infoworld.com/article/2335213/large-language-models-the-foundations-of-generative-ai.html">LLM </a>model picker and wondered which one I was supposed to want.</p>



<p class="wp-block-paragraph">OpenAI just released ⁠<a href="https://openai.com/index/gpt-5-6/">GPT-5.6 Sol, Terra, and Luna</a>. Sol is the flagship. Terra offers much of its intelligence for less money. Luna is cheaper still. Anthropic released ⁠<a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5</a> at the end of June and Opus 4.8 the month prior, with a little Fable 5 emerging in between. Meanwhile, Google, which seemed to be winning the model wars a few months ago, is now getting shade from Gergely Orosz, who ⁠<a href="https://x.com/GergelyOrosz/status/2075160978493210685?s=20">argues that Gemini has slipped outside the top tier</a> for software development and has been out of the major model release game for <em>eons</em> (May 19).</p>



<p class="wp-block-paragraph">Perhaps Orosz is right. Perhaps he’ll be wrong again in six weeks. Honestly, it’s exhausting.</p>



<p class="wp-block-paragraph">I use ChatGPT and Claude constantly and still have no principled idea which model to choose most of the time. I tend to click whatever looks like the biggest, most expensive option because I don’t know what I’m giving up by choosing something smaller. “Instant” sounds dangerously unserious. “Thinking” sounds expensive but powerful.</p>



<p class="wp-block-paragraph">A quick <a href="https://www.linkedin.com/feed/update/urn:li:activity:7481369774401409024/">survey of my LinkedIn crowd</a> suggests others also feel my “WHICH MODEL???” pain. More importantly, I suspect most enterprises do, too.</p>



<h2 class="wp-block-heading"><a></a>A model doesn’t rot</h2>



<p class="wp-block-paragraph">Before getting carried away, however, it’s worth considering whether any of this model churn actually matters. After all, a model doesn’t rot. The model an enterprise put into production in March performs just as well in July as it did when the company selected it. “Obsolete” generally means that something better now exists, not that the deployed model suddenly stopped summarizing insurance claims or classifying support tickets. (In other words, once you have something working, the idea that “but maybe Opus 200.2 is better!” is really a FOMO problem, not a performance issue.)</p>



<p class="wp-block-paragraph">Most enterprise workloads don’t live at the frontier anyway. Extraction, summarization, classification, document comparison, and customer-service assistance often work perfectly well with smaller, cheaper models. OpenAI’s own pitch for the trio of GPT-5.6 models isn’t simply that Sol is better. It’s that ⁠Terra and Luna deliver different combinations of intelligence, latency, and cost. Luna, the cheapest tier, nearly matches the previous generation’s peak performance at less than half the estimated cost, according to OpenAI.</p>



<p class="wp-block-paragraph">The practical question, of course, is where to start. An enterprise can’t test every model, every reasoning setting, and every price tier before doing any work. So here’s my advice (which I don’t follow in my own work, but I’m not defining enterprise strategy and can be a little price-insensitive). Start with the cheapest credible model that appears capable of the task. Give it a representative set of real examples and, before you start testing, define what counts as good enough. If it passes, stop. If it fails, move up a tier or try a model with strengths better suited to the work.</p>



<p class="wp-block-paragraph">That sounds almost offensively simple, but it reverses the way many people, including me, use these products. We start with the biggest model because we’re afraid of what we might lose. Enterprises should start lower and require evidence before paying for more intelligence.</p>



<p class="wp-block-paragraph">There are exceptions, of course. For genuinely difficult work, such as autonomous coding, complex research, or high-stakes reasoning, beginning with a frontier model may save time. But even then, the goal should be to establish a quality ceiling, then test whether a cheaper model can meet it. It’s changing the question from “which model is best?” to “what is the least expensive model that reliably clears the bar for this job?”</p>



<p class="wp-block-paragraph">For many workloads, that price improvement matters more than a few extra benchmark points. <a href="https://www.infoworld.com/article/2335519/ai-hype-isnt-helping-anyone.html">⁠As I argued back in 2023</a>, following AI hype doesn’t help anyone. If your model strategy depends on whichever benchmark screenshot is circulating on X this week, you don’t have a strategy. Not a viable one, anyway. Pick a model and ignore the noise.</p>



<p class="wp-block-paragraph">Except, of course, when that noise suggests a serious signal.</p>



<h2 class="wp-block-heading"><a></a>Sometimes better really is better</h2>



<p class="wp-block-paragraph">Frontier improvements aren’t always incremental, making it advantageous to consider an upgrade. Coding is the obvious example. There’s a significant difference between a model that suggests the next few lines of code and one that can inspect a repository, plan a change, use tools, run tests, discover its own mistakes, and keep working for an extended period. That isn’t merely a nicer autocomplete experience. It can reorganize a development workflow.</p>



<p class="wp-block-paragraph">This is why enterprises can’t simply standardize on an 18-month-old model and declare victory. In some areas, particularly software development and other agentic work, better models can unlock compounding productivity. A model that reliably completes 80% of a bounded task rather than 50% may justify an entirely different division of labor between humans and machines.</p>



<p class="wp-block-paragraph">Still, that upgrade isn’t free.</p>



<p class="wp-block-paragraph">Models differ in how they interpret instructions, call tools, manage context, refuse requests, and fail. Prompts and scaffolding tuned for one model can regress when moved to another. Or costs can explode. As one of my Oracle colleagues discovered just this week, running the same tasks in GPT 5.6 was orders of magnitude more expensive than 5.5. The API change may be trivial, but the revalidation and implications are not.</p>



<p class="wp-block-paragraph">This leaves enterprises caught between two bad options. They can freeze and potentially miss out on meaningful improvements or chase every release and repeatedly test production systems on faith. What to do?</p>



<h2 class="wp-block-heading"><a></a>Stop making model bets</h2>



<p class="wp-block-paragraph">The answer is to stop making LLM bets and start making job-to-be-done bets. Stop asking which model is fastest. Instead, figure out what work you are trying to improve. What does a good result look like? How much latency and cost can the workflow tolerate? How wrong can it be before a human must intervene? Once those questions have answers, model selection becomes less opaque.</p>



<p class="wp-block-paragraph">A difficult code migration may justify GPT-5.6 Sol or Claude Sonnet 5. A repetitive classification task may work just as well with Luna or another smaller model. A regulated workflow may require a model or deployment option that offers particular data controls. Sometimes the correct model is no LLM at all, like when I’m writing this post. Sorry, AI vendors! (At least you won’t get blamed for my mistakes.)</p>



<p class="wp-block-paragraph">This is where evaluations become the center of enterprise AI strategy. <a href="https://www.infoworld.com/article/4166247/improving-ai-agents-through-better-evaluations.html">⁠As I’ve said before</a>, most companies don’t have an AI quality problem so much as an AI measurement problem. Hence, a private evaluation suite built from real company work is the only leaderboard that matters. Does the new model materially improve quality? If so, use it! Does it reduce cost or latency? Again, that’s your free pass to adoption. Does the improvement justify the expense and effort of revalidation? If yes, continue.</p>



<h2 class="wp-block-heading"><a></a>Make model releases boring</h2>



<p class="wp-block-paragraph">As important as the model is, keep in mind that AI success always comes back to <em>your</em> company’s data, <em>your</em> company’s workflows<em>, your</em> company’s integrations, etc. That’s the ⁠<a href="https://www.infoworld.com/article/4157506/mastering-the-dull-reality-of-sexy-ai.html">dull reality behind sexy AI</a>. Retrieval, <a href="https://www.infoworld.com/article/4189492/how-to-improve-the-memory-of-ai-agents.html">memory</a>, governance, data quality, <a href="https://www.infoworld.com/article/2262666/what-is-observability-software-monitoring-on-steroids.html">observability</a>, and feedback loops aren’t as exciting as a new model launch, but they’re what ultimately make AI truly work.</p>



<p class="wp-block-paragraph">Again, when it’s time to consider something new, the principle should be to default to the least expensive model that reliably passes your evaluations. Only escalate harder tasks to more capable models when measurement shows that the premium pays. Tip: Make this invisible to employees so that the system routes to the best model for a particular prompt. As <a href="https://www.linkedin.com/feed/update/urn:li:activity:7481369774401409024/?dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287481372047860715522%2Curn%3Ali%3Aactivity%3A7481369774401409024%29">dbt Labs’ Jon Lewis expresses</a> it, “The best model is ‘Auto’ and I won’t hear anyone say otherwise.” OpenAI’s own ⁠<a href="https://developers.openai.com/api/docs/guides/latest-model">migration guidance</a> recommends testing models on representative tasks, including trying a lower reasoning level rather than automatically cranking everything to the maximum.</p>



<p class="wp-block-paragraph">As for me, I’ll probably keep clicking the shiniest option. I don’t have a formal evaluation suite for InfoWorld columns, and the marginal cost is a subscription I already pay. Enterprises don’t get that excuse.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Building Trustworthy Production RAG Systems Through Continuous Evaluation]]></title>
<description><![CDATA[A practical guide to building an evaluation workflow that catches retrieval failures, hallucinations, and performance drift before they reach users
The post Building Trustworthy Production RAG Systems Through Continuous Evaluation appeared first on Towards Data Science.]]></description>
<link>https://tsecurity.de/de/3671093/ai-nachrichten/building-trustworthy-production-rag-systems-through-continuous-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671093/ai-nachrichten/building-trustworthy-production-rag-systems-through-continuous-evaluation/</guid>
<pubDate>Wed, 15 Jul 2026 17:02:51 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>A practical guide to building an evaluation workflow that catches retrieval failures, hallucinations, and performance drift before they reach users</p>
<p>The post <a href="https://towardsdatascience.com/building-trustworthy-production-rag-systems-through-continuous-evaluation/">Building Trustworthy Production RAG Systems Through Continuous Evaluation</a> appeared first on <a href="https://towardsdatascience.com/">Towards Data Science</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[LatticeFlow AI connects governance frameworks with continuous AI risk monitoring]]></title>
<description><![CDATA[LatticeFlow AI has announced a platform for managing AI risk across agentic systems. Organizations are deploying autonomous AI in critical business processes, while governance approaches based on documentation and point-in-time assessments struggle to keep up with evolving risks. The LatticeFlow ...]]></description>
<link>https://tsecurity.de/de/3670610/it-security-nachrichten/latticeflow-ai-connects-governance-frameworks-with-continuous-ai-risk-monitoring/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3670610/it-security-nachrichten/latticeflow-ai-connects-governance-frameworks-with-continuous-ai-risk-monitoring/</guid>
<pubDate>Wed, 15 Jul 2026 14:23:42 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>LatticeFlow AI has announced a platform for managing AI risk across agentic systems. Organizations are deploying autonomous AI in critical business processes, while governance approaches based on documentation and point-in-time assessments struggle to keep up with evolving risks. The LatticeFlow AI Platform links AI governance frameworks with technical controls to continuously generate evidence and translate evaluation results into risk insights, helping organizations assess AI systems and support governance decisions. The platform combines AI discovery, evaluation, … <a href="https://www.helpnetsecurity.com/2026/07/15/latticeflow-ai-platform/" rel="nofollow">More <span class="meta-nav">→</span></a></p>
<p>The post <a href="https://www.helpnetsecurity.com/2026/07/15/latticeflow-ai-platform/">LatticeFlow AI connects governance frameworks with continuous AI risk monitoring</a> appeared first on <a href="https://www.helpnetsecurity.com/">Help Net Security</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The hidden AI cost driver: Harness design can make or break enterprise agent economics]]></title>
<description><![CDATA[A largely overlooked layer of the AI stack is emerging as a major driver of enterprise costs. New testing by AI consultancy Systima found that agent harnesses, the software that coordinates models, tools and workflows, can generate significant token overhead through their configuration alone, pot...]]></description>
<link>https://tsecurity.de/de/3669948/it-nachrichten/the-hidden-ai-cost-driver-harness-design-can-make-or-break-enterprise-agent-economics/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3669948/it-nachrichten/the-hidden-ai-cost-driver-harness-design-can-make-or-break-enterprise-agent-economics/</guid>
<pubDate>Wed, 15 Jul 2026 10:03:51 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">A largely overlooked layer of the AI stack is emerging as a major driver of enterprise costs. New testing by AI consultancy Systima found that agent harnesses, the software that coordinates models, tools and workflows, can generate significant token overhead through their configuration alone, potentially inflating the cost of AI deployments as organizations scale agents from experimental pilots to production environments.</p>



<p class="wp-block-paragraph">The firm, which ran a series of tests by juxtaposing two harnesses on the same tasks, namely Anthropic’s Claude Code and open-source OpenCode using the same Claude Sonnet 4.5 model underneath, found both exhibiting sharply different token overhead because of the differences in their configuration.</p>



<p class="wp-block-paragraph">These differences included system prompts, tool definitions, agent coordination mechanisms and other orchestration components, resulting in markedly different baseline input token overhead before users even entered a prompt, the consultancy firm wrote in a <a href="https://systima.ai/blog/claude-code-vs-opencode-token-overhead" target="_blank" rel="noreferrer noopener">blog post</a>.</p>



<p class="wp-block-paragraph">Separately, the firm also found that other configuration choices while setting up the harnesses such as repository instruction files, <a href="https://www.infoworld.com/article/4029634/what-is-model-context-protocol-how-mcp-bridges-ai-and-external-services.html" target="_blank">Model Context Protocol</a> (MCP) servers, prompt framework templates and subagents can each add substantial token overhead.</p>



<p class="wp-block-paragraph">The consultancy’s conclusions are also supported by emerging academic research examining how orchestration of the harnesses themselves, rather than optimizing models or changing them, can help enterprises reshape the economics around AI agents.</p>



<p class="wp-block-paragraph">In a <a href="https://arxiv.org/pdf/2607.06906" target="_blank" rel="noreferrer noopener">paper</a>, titled The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI, researchers showed that changing the harness while keeping models and tasks the same can reduce token consumption by 38%, cost per task by 41%, and execution time by 44% while maintaining comparable quality.</p>



<h2 class="wp-block-heading">Why enterprises overlook harness costs</h2>



<p class="wp-block-paragraph">Analysts say that enterprises can gain greater control over AI agent operating costs by paying closer attention to how their harnesses are configured and orchestrated, instead of just relying on model pricing as a yardstick.</p>



<p class="wp-block-paragraph">“The evaluation shows that the model is only one part of agent economics. The harness, tool schemas, instructions, MCP connections, and subagents matter as well. Enterprises therefore need to measure the entire agent configuration, not assume model pricing tells them what an agent will cost,” said <a href="https://www.linkedin.com/in/slwalter" target="_blank" rel="noreferrer noopener">Stephanie Walter</a>, practice lead of the AI stack at HyperFRAME Research.</p>



<p class="wp-block-paragraph">Currently, most enterprises pick agent tooling based on model quality, benchmarks, developer experience, and headline pricing per seat or per million tokens, with almost no one measuring what the harness sends per request, how stable the cache prefix is, or what subagent fan out costs at scale, echoed <a href="https://www.linkedin.com/in/advaitpatel93/" target="_blank" rel="noreferrer noopener">Advait Patel</a>, site reliability engineer at Broadcom.</p>



<p class="wp-block-paragraph">“Ask the average CIO whether their coding agent rewrites its cache mid-session, and you will get a blank stare,” Patel added.</p>



<p class="wp-block-paragraph">However, Ashish Chaturvedi, executive research leader at HFS Research, pointed out that lack of visibility is less a failure of enterprise leaders than a consequence of how AI agent ecosystem components are sold, stacked, and managed presently.</p>



<p class="wp-block-paragraph">“Most organizations have no visibility, mainly due to the absence of any metric from the vendor’s end that lets CIOs measure the entire agent or at least the harness configuration. None of this shows up in the developer’s experience. The agent just works, and the tokens burn silently in the background,” Chaturvedi said.</p>



<p class="wp-block-paragraph">The problem is further compounded, according to Chaturvedi, due to the manner in which AI agent configuration is distributed across enterprise teams.</p>



<p class="wp-block-paragraph">“The harness is chosen by one team, the instruction file written by another, and the MCP servers attached by a third, so no single person sees the cumulative weight,” Chaturvedi noted.</p>



<p class="wp-block-paragraph">Even when, in some cases, enterprises do have visibility and ownership, Patel argued, the industry, in general, still lack the operational maturity and discipline to systematically optimize AI agent costs.</p>



<p class="wp-block-paragraph">“FinOps for agents is where cloud FinOps was in 2013. Nobody has hired the equivalent of a cost optimization team focused on prompt engineering, harness configuration, and cache stability,” Patel said.</p>



<p class="wp-block-paragraph">Separately, <a href="https://www.linkedin.com/in/abhisekhsatapathy/" target="_blank" rel="noreferrer noopener">Abhishek Satapathy</a>, principal analyst at Avasant, pointed out that the invisibility issue stems from how enterprises evaluate AI agents before deploying them into production: “Most proof-of-concepts involve a limited number of users, relatively short-lived sessions, and controlled agentic interactions, where the accuracy of model output is the primary evaluation criterion.”</p>



<p class="wp-block-paragraph">The analysts’ comments also echo the conclusions of another research <a href="https://arxiv.org/pdf/2601.14470" target="_blank" rel="noreferrer noopener">paper</a>,  in which researchers argued that token consumption in agentic software engineering systems remains poorly understood because existing metrics provide limited visibility into where tokens are spent across orchestration components.</p>



<h2 class="wp-block-heading">How CIOs can improve visibility into AI agent costs</h2>



<p class="wp-block-paragraph">Closing that visibility gap, though, according to Satapathy, is increasingly becoming a priority for enterprises, as AI agents move from pilots to production and operating costs become harder to predict.</p>



<p class="wp-block-paragraph">“Across our advisory engagements, we are seeing growing demand for AI observability frameworks that combine runtime tracing, workload-level cost attribution, and execution analytics. This enables organizations to establish engineering baselines, benchmark workload efficiency, forecast AI operating costs, and continuously optimize agent performance as deployments mature,” Satapathy said.</p>



<p class="wp-block-paragraph">However, until vendors provide more comprehensive visibility into harness-level token consumption, analysts said enterprises should begin treating harness configuration as an operational governance issue rather than merely a developer preference.</p>



<p class="wp-block-paragraph">“The single most valuable move is to get visibility into what the harness actually sends. Enterprises should treat configuration as a governed cost decision, deliberately match harnesses to workloads, and closely monitor cache behavior and subagent fan-out, since those were among the biggest cost multipliers identified in the evaluation,” Chaturvedi said.</p>



<p class="wp-block-paragraph">Walter echoed that recommendation, saying CIOs should require observability across the entire agent configuration: “Without that visibility, enterprises are effectively buying an agent platform without knowing how much of the bill comes from useful work versus orchestration overhead.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Product showcase: Trust Chain TPRM turns vendor compliance evidence into verified assurance]]></title>
<description><![CDATA[Trust Chain is an AI-native third-party risk management (TPRM) solution by Strike Graph that replaces the security questionnaire model with validated evidence of compliance. Rather than asking vendors to self-report their security posture, Trust Chain requires vendors to submit evidence, which is...]]></description>
<link>https://tsecurity.de/de/3669934/it-security-nachrichten/product-showcase-trust-chain-tprm-turns-vendor-compliance-evidence-into-verified-assurance/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3669934/it-security-nachrichten/product-showcase-trust-chain-tprm-turns-vendor-compliance-evidence-into-verified-assurance/</guid>
<pubDate>Wed, 15 Jul 2026 09:53:39 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Trust Chain is an AI-native third-party risk management (TPRM) solution by Strike Graph that replaces the security questionnaire model with validated evidence of compliance. Rather than asking vendors to self-report their security posture, Trust Chain requires vendors to submit evidence, which is then evaluated using Strike Graph’s patent-pending Verify AI technology. The evaluation tests each submission against the requesting organization’s speciﬁc requirements. The result is veriﬁed assurance rather than vendor attestation, delivered in a fraction … <a href="https://www.helpnetsecurity.com/2026/07/15/product-showcase-strike-graph-trust-chain-tprm/" rel="nofollow">More <span class="meta-nav">→</span></a></p>
<p>The post <a href="https://www.helpnetsecurity.com/2026/07/15/product-showcase-strike-graph-trust-chain-tprm/">Product showcase: Trust Chain TPRM turns vendor compliance evidence into verified assurance</a> appeared first on <a href="https://www.helpnetsecurity.com/">Help Net Security</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[What to Expect at Black Hat USA 2026]]></title>
<description><![CDATA[Author: Black Hat - Bewertung: 8x - Views:51 Over 20,000 practitioners. 100+ hands-on training courses. Peer-reviewed research that doesn't exist anywhere else yet. Black Hat USA runs August 1-6, 2026 in Las Vegas, and this year's agenda is shaping up to be the most exciting one yet.
 
Black Hat ...]]></description>
<link>https://tsecurity.de/de/3668721/it-security-video/what-to-expect-at-black-hat-usa-2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3668721/it-security-video/what-to-expect-at-black-hat-usa-2026/</guid>
<pubDate>Tue, 14 Jul 2026 19:00:21 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: Black Hat - Bewertung: 8x - Views:51 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/HopnPmgHUis?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>Over 20,000 practitioners. 100+ hands-on training courses. Peer-reviewed research that doesn't exist anywhere else yet. Black Hat USA runs August 1-6, 2026 in Las Vegas, and this year's agenda is shaping up to be the most exciting one yet.<br />
 <br />
Black Hat USA 2026 brings together the cybersecurity community for six days of training, research, and hands-on evaluation. Here's what you're walking into:<br />
<br />
• Training (August 1-4): 100+ expert-led courses taught by practitioners who've deployed these techniques in live environments. This year's expanded AI security track covers securing LLMs, defending against autonomous agents, and building detection pipelines that work at machine speed.<br />
• Briefings (August 5-6): Peer-reviewed research selected by an independent review board. AI agent exploitation. Post-quantum cryptography. Supply chain attacks. Detection engineering. The findings you'll hear don't exist in published form yet; you're getting them first.<br />
• Business Hall (August 4-6): 400+ sponsors and exhibitors. The practitioners walking that floor are coming straight out of Briefings and Trainings, so they know exactly what questions to ask. This is where real evaluation happens.<br />
• Summits (August 4): Six full-day, domain-specific programs including the CISO Summit, AI Summit, Financial Services Security Summit, Healthcare Summit, and more. You're not in a general conference audience; you're with peers who understand the specific challenges you're facing.<br />
• New this year: The Interface (hands-on demos and scenario-based learning), Arsenal Labs (20 dedicated tool demonstration sessions), Cyber War Forum (senior leader discussions under Chatham House rules), Drone Zone, Cyber District, and Black Hat(HER).<br />
<br />
Regular registration pricing is active through July 17th, 2026.<br />
Register at blackhat.com<br />
One Step Ahead.<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Deepmind CEO Hassabis says "nobody in the world knows what happens next" so "cautious optimism" means building guardrails now]]></title>
<description><![CDATA[Google Deepmind CEO Demis Hassabis has published a sweeping proposal for how to handle advanced AI. He wants a new US standards body modeled after financial regulator FINRA that would develop evaluation protocols for frontier models and could coordinate a slowdown in AI development if needed. Sta...]]></description>
<link>https://tsecurity.de/de/3667870/ai-nachrichten/deepmind-ceo-hassabis-says-nobody-in-the-world-knows-what-happens-next-so-cautious-optimism-means-building-guardrails-now/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3667870/ai-nachrichten/deepmind-ceo-hassabis-says-nobody-in-the-world-knows-what-happens-next-so-cautious-optimism-means-building-guardrails-now/</guid>
<pubDate>Tue, 14 Jul 2026 14:03:57 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1088" height="608" src="https://the-decoder.com/wp-content/uploads/2026/07/deepmind_hassabis.png" class="attachment-full size-full wp-post-image" alt="" decoding="async" fetchpriority="high"></p>
<p>        Google Deepmind CEO Demis Hassabis has published a sweeping proposal for how to handle advanced AI. He wants a new US standards body modeled after financial regulator FINRA that would develop evaluation protocols for frontier models and could coordinate a slowdown in AI development if needed. Startups and research models would be exempt.</p>
<p>The article <a href="https://the-decoder.com/deepmind-ceo-hassabis-says-nobody-in-the-world-knows-what-happens-next-so-cautious-optimism-means-building-guardrails-now/">Deepmind CEO Hassabis says "nobody in the world knows what happens next" so "cautious optimism" means building guardrails now</a> appeared first on <a href="https://the-decoder.com/">The Decoder</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-15043 | HMBRAND DBI::SQL::Nano up to 1.650 SQL Evaluation Engine is_matched comparison (EUVD-2026-43654)]]></title>
<description><![CDATA[A vulnerability, which was classified as problematic, was found in HMBRAND DBI::SQL::Nano up to 1.650. This vulnerability affects the function is_matched of the component SQL Evaluation Engine. The manipulation results in incorrect comparison.

This vulnerability is reported as CVE-2026-15043. Th...]]></description>
<link>https://tsecurity.de/de/3667847/sicherheitsluecken/cve-2026-15043-hmbrand-dbisqlnano-up-to-1650-sql-evaluation-engine-ismatched-comparison-euvd-2026-43654/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3667847/sicherheitsluecken/cve-2026-15043-hmbrand-dbisqlnano-up-to-1650-sql-evaluation-engine-ismatched-comparison-euvd-2026-43654/</guid>
<pubDate>Tue, 14 Jul 2026 13:55:36 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability, which was classified as <a href="https://vuldb.com/kb/risk">problematic</a>, was found in <a href="https://vuldb.com/product/hmbrand:dbi_sql_nano">HMBRAND DBI::SQL::Nano up to 1.650</a>. This vulnerability affects the function <code>is_matched</code> of the component <em>SQL Evaluation Engine</em>. The manipulation results in incorrect comparison.

This vulnerability is reported as <a href="https://vuldb.com/cve/CVE-2026-15043">CVE-2026-15043</a>. The attack requires a local approach. No exploit exists.]]></content:encoded>
</item>
<item>
<title><![CDATA[The essence of data management CIOs must embrace]]></title>
<description><![CDATA[Since the advent of generative AI, the use of AI in business has shifted from something we should do to something we must do to survive. Many companies are now working to utilize AI with the aim of improving productivity and creating value.



Here, I would like to pose a question to you all once...]]></description>
<link>https://tsecurity.de/de/3667389/it-security-nachrichten/the-essence-of-data-management-cios-must-embrace/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3667389/it-security-nachrichten/the-essence-of-data-management-cios-must-embrace/</guid>
<pubDate>Tue, 14 Jul 2026 11:08:51 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Since the advent of generative AI, the use of AI in business has shifted from something we should do to something we must do to survive. Many companies are now working to utilize AI with the aim of improving productivity and creating value.</p>



<p class="wp-block-paragraph">Here, I would like to pose a question to you all once again: “What is the fundamental factor that determines AI performance?”</p>



<p class="wp-block-paragraph">Is it the AI model? Is it the AI tool? Or is it the AI agent?</p>



<p class="wp-block-paragraph">Of course, I believe all of these are important. However, if we look at the long-term perspective, the competition among multiple companies to improve AI model performance will eventually level off, and we will eventually reach a point where every AI model is amazing!</p>



<p class="wp-block-paragraph">In that context, what I believe is the most important factor influencing AI performance is the data accumulated by companies that connects to their unique strengths.</p>



<p class="wp-block-paragraph">For example, if asked, “What do plants need to grow?” I would say “good water and light.”</p>



<p class="wp-block-paragraph">Similarly, if asked, “What do people need to thrive?” I would say, “Kind words.”</p>



<p class="wp-block-paragraph">Finally, “What does AI need to thrive?” The answer is “good data.”</p>



<p class="wp-block-paragraph">I believe that the extent to which companies can genuinely understand the importance of this extremely simple principle and implement it with unwavering dedication will determine their ability to establish a competitive advantage and achieve sustainable growth.</p>



<h2 class="wp-block-heading">AI is a mirror of data</h2>



<p class="wp-block-paragraph">As I’m sure you’re all aware, AI is by no means a magic wand. It is an entity that learns based on the data it is given and makes inferences within that scope. In other words, AI’s output depends heavily on the quality of its input data; one could say that AI is a mirror of data.</p>



<ul class="wp-block-list">
<li>If you feed it inaccurate data, it will return inaccurate results (i.e., garbage in, garbage out)</li>



<li>If you feed it biased data, it will make biased judgments</li>



<li>Insufficient data yields only shallow insights and suggestions</li>
</ul>



<p class="wp-block-paragraph">In this way, AI is not smart but rather faithful to the data. Based on this premise, it becomes clear that the essence of AI utilization lies not in which tools to use, but in what kind of high-quality data to prepare and how to utilize it.</p>



<h2 class="wp-block-heading">What is good data?</h2>



<p class="wp-block-paragraph">So, what exactly is good data?</p>



<p class="wp-block-paragraph">It goes without saying that data is useless if it is merely abundant in quantity, but on the other hand, what specific qualities must good data possess?</p>



<p class="wp-block-paragraph">Generally speaking, good data possesses at least the following elements.</p>



<ul class="wp-block-list">
<li><strong>Accuracy:</strong> Data containing many errors or noise will skew conclusions, no matter how advanced the analysis. It is important to minimize sensor errors, input mistakes and duplicates.</li>



<li><strong>Completeness:</strong> Are any required fields missing, and are there too many missing values? For example, if customer data is missing information such as age, region or gender, it becomes difficult to perform meaningful analysis.</li>



<li><strong>Consistency:</strong> Is data with the same meaning mixed in different formats (e.g., date formats, units, variations in notation)? This is particularly important for system integration and long-term data.</li>



<li><strong>Timeliness:</strong> No matter how accurate it is, data that is too old may not be useful for decision-making. Whether real-time data is required or historical data is sufficient depends on the use case, but it is important that the data has the appropriate freshness for the purpose.</li>



<li><strong>Relevance:</strong> If there is a large amount of data unrelated to the analysis objective, it becomes noise and leads to incorrect judgments. It is necessary to clearly define what the data is used for and ensure the data is appropriate for that purpose.</li>



<li><strong>Reliability: The data’s source and collection method must be</strong> clear, ensuring reliability and reproducibility. Data with an unknown source or that is a black box cannot be verified later.</li>
</ul>



<p class="wp-block-paragraph">In summary, good data is data that is accurate, has few gaps, is consistent in meaning and notation, is collected at the appropriate time, is suitable for the purpose and comes from a reliable source.</p>



<p class="wp-block-paragraph">Only when the quality of this good data is guaranteed can AI produce valuable outputs. Conversely, introducing AI with unorganized data will not yield the expected results. Many complaints, such as “We implemented AI but it’s unusable” or “The AI’s accuracy isn’t improving stem from data issues.”</p>



<h2 class="wp-block-heading">Data does not organize itself naturally</h2>



<p class="wp-block-paragraph">The key point here is that good data does not arise naturally. On the contrary, if left unattended, data will inevitably deteriorate.</p>



<ul class="wp-block-list">
<li>Rules become inconsistent depending on who entered the data and when</li>



<li>Multiple instances of data with the same meaning exist</li>



<li>Outdated data is scattered and left unattended</li>



<li>Data becomes siloed by department</li>
</ul>



<p class="wp-block-paragraph">These conditions are likely common in many companies.</p>



<p class="wp-block-paragraph">Below is an overview of our company’s <a href="https://www.kepco.co.jp/english/corporate/list/report/">data management framework</a>.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" src="https://b2b-contenthub.com/wp-content/uploads/2026/07/overview-of-data-management-at-kansai-electric-power-company.png?w=1024" alt="Overview of data management at Kansai Electric Power Company" class="wp-image-4196318" width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Akio Ueda</p></div>



<p class="wp-block-paragraph">Broadly speaking, it consists of data governance — covering roles and structures, risk management and evaluation — and data management, which encompasses data utilization cycle management and data utilization support services. Within this framework, data utilization cycle management involves:</p>



<ul class="wp-block-list">
<li><strong>Needs management:</strong> We clarify the purpose and needs by asking, “What is the data being used for?” and “For whom, and in what way, does this data create value?”</li>



<li><strong>Collection:</strong> We gather the necessary data based on the defined objectives. We design the process to determine what data is required (internal/external), the level of detail and frequency of collection, and how to ensure data quality.</li>



<li><strong>Processing: </strong>We enhance the quality and prepare the data for use. This includes cleansing (correcting errors and missing values), standardizing formats, deduplicating and integrating data, processing structured and unstructured data separately, and assigning business and operational meaning to the data.</li>



<li><strong>Storage:</strong> We ensure the data is available to the right people at the right time. This involves storing data in databases or data lakes, implementing security and access controls, and managing metadata (ensuring the data is clearly identifiable).</li>



<li><strong>Utilization:</strong> This is the most critical step. The purpose of data is not merely analysis but driving action. We generate value from the data through visualization (dashboards), analysis (statistical processing, BI, AutoML, AI) and integration into business operations (automation and decision support).</li>



<li><strong>Disposal: </strong>We properly dispose of data that is no longer needed. Simply holding data can itself pose risks, such as managing retention periods, complying with laws and governance requirements, and mitigating security risks. That is why the principle of not holding data that is not used is so important.</li>
</ul>



<p class="wp-block-paragraph">Data management is not a one-time effort; it is an ongoing initiative that requires continuous maintenance and improvement.</p>



<p class="wp-block-paragraph">The CIO must embed data management as a system within the organization and continue to implement it until it becomes firmly established.</p>



<h2 class="wp-block-heading">Data management is not just the IT department’s job</h2>



<p class="wp-block-paragraph">Another important point is that data management is not just the IT department’s job.</p>



<p class="wp-block-paragraph">Data is fundamentally generated within day-to-day operations on the front lines. Therefore:</p>



<ul class="wp-block-list">
<li>Who determines the meaning and definition of data</li>



<li>How should input rules be standardized?</li>



<li>How do we ensure data quality?</li>
</ul>



<p class="wp-block-paragraph">are, in essence, operational issues, business issues and management issues.</p>



<p class="wp-block-paragraph">The latest Digital Skills Standard ver. 2.0, published by the Ministry of Economy, Trade and Industry in April 2026, defines the following three roles within the data management category:</p>



<ul class="wp-block-list">
<li><strong>Data steward:</strong> Based on business domain knowledge, this role is responsible for operations aimed at ensuring data quality, reliability and security, as well as for promoting the adoption and establishment of data management within business divisions and frontline organizations, and for fostering data utilization. In short, they are the data quality manager and data utilization promoter.</li>



<li><strong>Data engineer: </strong>This role involves understanding the current state of data and supporting the organization’s continuous data utilization through data preparation and preprocessing in processes such as collection, integration, processing and provision, as well as the design and implementation of data pipelines. In essence, they are the implementers and operators who drive data.</li>



<li><strong>Data architect:</strong> This role involves taking a bird’s-eye view of the data structure, flow and utilization methods across the entire organization and business. By designing and continuously reviewing data architecture that encompasses the entire data lifecycle in alignment with business strategy, they ensure the successful integration of company-wide data utilization and governance—essentially serving as the overall designer of data.</li>
</ul>



<p class="wp-block-paragraph">The CIO is not merely responsible for establishing data storage and analysis infrastructure; they are also tasked with appropriately assigning personnel to these three roles within the company and establishing cross-departmental, company-wide tools and rules to connect data with management, business operations and daily tasks.</p>



<h2 class="wp-block-heading">Ultimately, the success of data utilization depends on organizational culture</h2>



<p class="wp-block-paragraph">On the other hand, no matter how much progress is made in staffing, infrastructure, tools and rulemaking, data will not be utilized unless there is an organizational culture that actively drives management, business and operations based on data.</p>



<ul class="wp-block-list">
<li>The purpose of data entry is not understood</li>



<li>Data is optimized solely for the department’s own operations</li>



<li>Decision-making based on data is not valued</li>
</ul>



<p class="wp-block-paragraph">In such a situation, no matter how well the systems are set up, they will become mere formalities.</p>



<p class="wp-block-paragraph">In contrast, in organizations where data utilization is advanced:</p>



<ul class="wp-block-list">
<li>Discussions are based on data</li>



<li>Formulate hypotheses and verify them with data</li>



<li>And continuously improve based on data</li>
</ul>



<p class="wp-block-paragraph">These actions occur naturally.</p>



<p class="wp-block-paragraph">In other words, the essence of data management ultimately lies in creating an organizational culture that assumes the effective use of data.</p>



<p class="wp-block-paragraph">Data management cannot be achieved overnight. That is precisely why it is important to start small and build on your successes.</p>



<ul class="wp-block-list">
<li>Organize data for specific tasks and achieve results through the use of AI</li>



<li>Rolling out successful practices</li>



<li>Gradually Expand the Scope</li>
</ul>



<p class="wp-block-paragraph">By repeating this cycle, the importance of data will permeate the entire organization.</p>



<h2 class="wp-block-heading">The role expected of a CIO in the AI era</h2>



<p class="wp-block-paragraph">In the AI era, the role expected of a CIO has changed significantly.</p>



<p class="wp-block-paragraph">Traditionally:</p>



<ul class="wp-block-list">
<li>Ensuring the stable operation of systems</li>



<li>And optimizing costs</li>
</ul>



<p class="wp-block-paragraph">However, moving forward:</p>



<ul class="wp-block-list">
<li>We will view data as an asset and maximize its value</li>



<li>Developing the data infrastructure, tools and rules that underpin AI adoption, and advancing personnel allocation and development</li>



<li>And fostering an organizational culture that embraces data utilization —roles that are more directly linked to business management</li>
</ul>



<p class="wp-block-paragraph">In other words, the CIO must evolve into the person responsible for creating value from data.</p>



<h2 class="wp-block-heading">Data is the source of competitive advantage</h2>



<p class="wp-block-paragraph">In the coming era, the use of AI will be a given. What will set companies apart is not whether they use AI, but what data they possess.</p>



<p class="wp-block-paragraph">Data is the accumulation of a company’s past strengths and the source of future value creation. And its quality is determined by daily operations and the nature of the organization.</p>



<ul class="wp-block-list">
<li>AI grows by being fed good data</li>



<li>And companies grow through that AI</li>
</ul>



<p class="wp-block-paragraph">Taking this simple principle as our starting point, we must place data management at the core of our business strategy. Isn’t that the shortest route to sustainable growth in the AI era?</p>



<p class="wp-block-paragraph">CIOs are called upon to lead the way in making this a reality.</p>



<p class="wp-block-paragraph"><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Skyfall AI Releases MORPHEUS: A Persistent Enterprise Simulation Benchmark That Makes Continual Reinforcement Learning Necessary Under Structured Non-Stationarity]]></title>
<description><![CDATA[MORPHEUS from Skyfall AI is a persistent enterprise simulation platform for continual reinforcement learning. It runs worlds that never reset, using parameterisable regime shifts and a six-metric evaluation protocol. Across the platform, PPO, HER, EWC, and LCM all remain far below the theoretical...]]></description>
<link>https://tsecurity.de/de/3666553/ai-nachrichten/skyfall-ai-releases-morpheus-a-persistent-enterprise-simulation-benchmark-that-makes-continual-reinforcement-learning-necessary-under-structured-non-stationarity/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3666553/ai-nachrichten/skyfall-ai-releases-morpheus-a-persistent-enterprise-simulation-benchmark-that-makes-continual-reinforcement-learning-necessary-under-structured-non-stationarity/</guid>
<pubDate>Tue, 14 Jul 2026 00:48:45 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>MORPHEUS from Skyfall AI is a persistent enterprise simulation platform for continual reinforcement learning. It runs worlds that never reset, using parameterisable regime shifts and a six-metric evaluation protocol. Across the platform, PPO, HER, EWC, and LCM all remain far below the theoretical upper bound.</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/13/skyfall-ai-releases-morpheus-a-persistent-enterprise-simulation-benchmark-that-makes-continual-reinforcement-learning-necessary-under-structured-non-stationarity/">Skyfall AI Releases MORPHEUS: A Persistent Enterprise Simulation Benchmark That Makes Continual Reinforcement Learning Necessary Under Structured Non-Stationarity</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-58251 | nats-io nats-server up to 2.11.15/2.12.6 Subscription Deny Evaluation improper authentication]]></title>
<description><![CDATA[A vulnerability categorized as problematic has been discovered in nats-io nats-server up to 2.11.15/2.12.6. This affects an unknown part of the component Subscription Deny Evaluation. Executing a manipulation can lead to improper authentication.

This vulnerability is tracked as CVE-2026-58251. T...]]></description>
<link>https://tsecurity.de/de/3666033/sicherheitsluecken/cve-2026-58251-nats-io-nats-server-up-to-211152126-subscription-deny-evaluation-improper-authentication/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3666033/sicherheitsluecken/cve-2026-58251-nats-io-nats-server-up-to-211152126-subscription-deny-evaluation-improper-authentication/</guid>
<pubDate>Mon, 13 Jul 2026 19:26:02 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability categorized as <a href="https://vuldb.com/kb/risk">problematic</a> has been discovered in <a href="https://vuldb.com/product/nats-io:nats-server">nats-io nats-server up to 2.11.15/2.12.6</a>. This affects an unknown part of the component <em>Subscription Deny Evaluation</em>. Executing a manipulation can lead to improper authentication.

This vulnerability is tracked as <a href="https://vuldb.com/cve/CVE-2026-58251">CVE-2026-58251</a>. The attack can be launched remotely. No exploit exists.]]></content:encoded>
</item>
<item>
<title><![CDATA[ACRouter picks the smartest AI model per task, beating Opus-only setups by 2.6x on cost]]></title>
<description><![CDATA[Model routing is becoming a key component of the enterprise AI stack, dynamically sending prompts to the right AI model to optimize speed and costs. However, current frameworks mostly treat routing as a static classification problem, which severely limits their potential.A new open-source framewo...]]></description>
<link>https://tsecurity.de/de/3665936/it-nachrichten/acrouter-picks-the-smartest-ai-model-per-task-beating-opus-only-setups-by-26x-on-cost/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3665936/it-nachrichten/acrouter-picks-the-smartest-ai-model-per-task-beating-opus-only-setups-by-26x-on-cost/</guid>
<pubDate>Mon, 13 Jul 2026 18:48:17 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Model routing is becoming a key component of the enterprise AI stack, dynamically sending prompts to the right AI model to optimize speed and costs. However, current frameworks mostly treat routing as a static classification problem, which severely limits their potential.</p><p>A new open-source framework called <a href="https://arxiv.org/abs/2606.22902">Agent-as-a-Router</a> tackles this bottleneck, treating the router as a dynamic, memory-building agent. It uses a Context-Action-Feedback (C-A-F) loop to track model successes and failures and update the behavior of the router. </p><p>The researchers also released ACRouter, a concrete implementation of this paradigm. In their tests, ACRouter significantly outperformed static routers and the expensive strategy of defaulting to premium models, all without requiring teams to train massive models or write endless heuristics.</p><p>For real-world applications, this framework provides the option to replace hard-coded AI infrastructure with self-optimizing systems that can adapt to changes in user behavior and foundation models used in the enterprise AI stack. </p><h2>The economics of routing and the information deficit</h2><p>Single-model setups are useful for experiments but detrimental when scaling AI applications. AI engineers use <a href="https://venturebeat.com/orchestration/enterprises-using-multiple-ai-models-are-underestimating-failure-rates-by-2-25x">model routing</a> to map tasks to cheaper and faster open models when possible, while reserving expensive frontier models for complex reasoning. </p><p>Currently, developers rely on two main mechanisms for this task. The first is heuristics-based routing, which relies on hard-coded manual rules. For example, a developer might write a rule dictating that if a prompt contains certain keywords, it is routed to GPT-5.5. Otherwise, it goes to a self-hosted open source model like Kimi K2.7. </p><p>The second mechanism is static trained policies. These are machine learning classifiers trained on historical datasets that look at the prompt's embeddings and predict the best model based on past training data.</p><p>Both approaches are static. When the researchers tested these existing mechanisms on real-world coding and agentic workflows, they found a hard ceiling on accuracy. The key finding shows that static routers suffer from a severe information deficit. Because they only evaluate the input text and never see if the model actually succeeded in executing the task, they guess blindly when faced with complex edge cases.</p><p>This results in three distinct points of failure. First, static routers suffer from a frozen information state, meaning they cannot accumulate new execution feedback during deployment. Second, they fail in out-of-distribution (OOD) generalization. They break down during day-two operations when enterprise data or user behavior shifts because their training data no longer matches reality. Finally, they are highly vulnerable to model churn. A static classifier trained on today's models may become obsolete when a better model drops the following week.</p><h2>Agent-as-a-Router: A self-evolving system</h2><p>The core thesis of the Agent-as-a-Router is that a truly effective router must acquire and accumulate execution-grounded information during deployment, essentially learning on the job. </p><p>The researchers achieved this through the C-A-F loop. When a new prompt arrives, the router examines the prompt and task metadata, such as the programming language or difficulty. It then searches its historical memory for similar tasks to see which models succeeded or failed in the past. The router uses this context to select the target model and execute the task. Finally, the system observes the real-world outcome, extracts a success or failure signal, and writes this feedback back into its memory to inform future routing decisions.</p><p>Consider an automated enterprise data analytics pipeline. The router receives a SQL generation task and sends it to an open-source model like Kimi. The model hallucinates a column name and fails to compile the SQL. The C-A-F loop observes the compiler error, registers it as feedback, and logs it. The next time a similar obscure SQL query arrives, the router checks its context and routes the task to a more advanced model like Claude Opus 4.8. </p><h2>ACRouter</h2><p>The researchers developed ACRouter as the concrete instantiation of this framework. It is composed of three core components: the Orchestrator, the Verifier, and Memory. This architecture is supported by a tool layer to physically execute the C-A-F loop.</p><p>The Memory module powers the context phase. Built on a vector store, it retrieves relevant past interactions and updates the historical database with new outcomes. The Orchestrator handles the action phase. It processes the user prompt alongside the retrieved memory to select the most capable target model from the available pool. The Verifier manages the feedback phase by evaluating the chosen model's output to generate a clear success or failure signal.</p><p>The tool layer hooks the Verifier into real-world execution environments, like a Python code interpreter, an agentic sandbox, or a database engine. The tool layer allows the system to execute the generated code or query and observe the exact outcome, providing the verifiable signal the router needs to learn.</p><p>The Orchestrator itself is lightweight. Instead of a massive, computationally heavy large language model, the researchers trained a sub-billion parameter adapter based on Qwen 3.5 (0.8B parameters), which means it can be self-hosted on a device of your choice.</p><h2>ACRouter in action: Outperforming the frontier baselines</h2><p>To stress-test the framework, the researchers introduced CodeRouterBench, an evaluation environment comprising roughly 10,000 tasks with verified scores across eight frontier models, including Claude Opus 4.6, GPT-5.4, Qwen3-Max, and GLM-5. The evaluation was split between in-distribution (ID) tests (covering nine single-turn coding dimensions like algorithm design and test generation) and an out-of-distribution (OOD) agentic programming testbed. The OOD tasks were qualitatively different, requiring multi-step planning, file navigation, and iterative debugging to see if the router could adapt to fundamentally new domains.</p><p>The baseline results revealed why a single-model strategy is flawed: no single model dominates every category. For example, while Claude Opus 4.6 achieved the highest average performance, it was outperformed in algorithm design by GLM-5 (an 86% relative improvement) and in test generation by Qwen3-Max (a 111% improvement), despite Opus costing roughly 12 times as much as smaller models like Kimi-K2.5. </p><p>In the benchmarks, static routers continuously failed by sending a specific niche coding task to a model ill-equipped for that exact syntax. The static router had no way to know the code was failing to execute. In contrast, ACRouter adjusted its strategy after receiving negative feedback signal from the execution environment. </p><p>According to the researchers' benchmarking, ACRouter sits firmly at the Pareto frontier of cost and performance. On both the ID task streams and the complex OOD agentic tests, ACRouter achieved the lowest cumulative regret, a metric measuring sub-optimal routing decisions over time. On the in-distribution test set, ACRouter cost $13.21 across the full task run, compared to $34.02 for always defaulting to Opus — a 2.6x savings.</p><p>It dynamically matched tasks to the most capable model for that specific niche, suggesting that enterprises can achieve or exceed frontier-level accuracy across diverse workloads without paying a premium price for every query. </p><h2>Caveats, limitations, and how to get started</h2><p>While the Agent-as-a-Router paradigm solves the information deficit, it is not a blanket solution for all AI workflows. </p><p>The framework shines in verifiable tasks where the Verifier gets a clear success or failure signal from the environment, such as coding or data retrieval. It is effective for applications with distribution shifts and domains where different models excel in completely distinct niches. </p><p>Conversely, the setup is overkill for trivial tasks where any model will suffice, or for low-volume applications that do not justify the engineering overhead. It is also unsuitable for subjective domains, such as creative writing, where a correct answer cannot be easily verified and feedback signals are impossible to standardize.</p><p>The researchers open-sourced <a href="https://github.com/LanceZPF/agent-as-a-router">the code on GitHub</a> and released the <a href="https://huggingface.co/Lance1573/acrouter-qwen35-08b-router-lora">orchestrator model weights on Hugging Face</a> under the Apache 2.0 license. The router is compatible with Claude Code, Codex, and OpenCode.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[What is generative AI? How artificial intelligence creates content]]></title>
<description><![CDATA[Generative AI is a kind of artificial intelligence that creates new content, including text, images, audio, and video, based on patterns it has learned from existing data.



Today’s generative models are typically built on foundation-model architectures such as large-language models (LLMs) and m...]]></description>
<link>https://tsecurity.de/de/3665675/ai-nachrichten/what-is-generative-ai-how-artificial-intelligence-creates-content/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3665675/ai-nachrichten/what-is-generative-ai-how-artificial-intelligence-creates-content/</guid>
<pubDate>Mon, 13 Jul 2026 17:04:40 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Generative AI is a kind of <a href="https://www.computerworld.com/article/1647870/what-is-artificial-intelligence.html">artificial intelligence</a> that creates new content, including text, images, audio, and video, based on patterns it has learned from existing data.</p>



<p class="wp-block-paragraph">Today’s generative models are typically built on foundation-model architectures such as <a href="https://www.infoworld.com/article/2335213/large-language-models-the-foundations-of-generative-ai.html">large-language models (LLMs)</a> and multimodal systems, enabling them to carry on conversations, answer questions, write stories, generate code, and produce images or videos from brief prompts.</p>



<p class="wp-block-paragraph"><em>Generative AI</em> is different from <em>discriminative AI</em>, which draws distinctions between different kinds of input. Where discriminative AI answers questions like “Is this image of a rabbit or a lion?”, generative AI instead responds to prompts such as “Describe to me how a rabbit and lion look different from one another” or “Draw me a picture of a lion and a rabbit sitting next to each other” — and in both cases produces text or imagery that, while grounded in the AI’s training data, isn’t just a copy of something that already existed.</p>



<aside class="fakesidebar">
<h4>[ <u><a href="https://www.infoworld.com/article/2335213/large-language-models-the-foundations-of-generative-ai.html">Read next: Large language models: The foundations of generative AI</a></u> ]</h4>
</aside>




<p class="wp-block-paragraph">Just a few years ago, generative AI was once a novelty focused on chatbots and artistic image generation. Today, it has become a core enterprise technology, and powers everything from content creation and software development to customer support and analytics workflows. But with that power comes a <a href="https://www.csoonline.com/article/4076511/4-factors-creating-bottlenecks-for-enterprise-genai-adoption.html">new set of challenges</a> — from model alignment and hallucination to governance and data-integration hurdles.</p>



<p class="wp-block-paragraph">In this article, we’ll look at how generative AI works, explore how it has evolved into the foundation-model era, examine how to implement it effectively, and offer best practices for getting value out of it, today and in the future.</p>



<h2 class="wp-block-heading"><strong>How does generative AI work?</strong></h2>



<p class="wp-block-paragraph">For decades, early artificial-intelligence efforts often focused on rule-based systems or <a href="https://www.infoworld.com/article/4061121/a-brief-history-of-ai.html">narrowly trained models</a> that were built for one task at a time. While these efforts produced useful systems that could reason and solve human tasks, they were generally a far cry from sci-fi visions of thinking machines. Programs that could talk to people never seemed to get very far past the level of <a href="https://en.wikipedia.org/wiki/ELIZA">ELIZA</a>, a “computer therapist” created at MIT in the mid 1960s; even Siri and Alexa after much fanfare were revealed to be fairly limited.</p>



<p class="wp-block-paragraph">The big structural shift that gave birth to modern generative AI came with the concept of a <em>transformer, </em>first introduced in “<a href="https://arxiv.org/abs/1706.03762">Attention Is All You Need</a>,” a 2017 paper from Google researchers.</p>



<p class="wp-block-paragraph">Using a transformer architecture as a basis, you can build a system that derives meaning from analyzing long sequences of input <em>tokens</em> (words, sub-words, bytes) to understand how different tokens might be related to one another, then determines how likely any given token is to come next in a sequence, given the others. In AI lingo, we call these systems <em>models.</em> Because a model analyzes very large datasets and parameter counts, it can pick up on statistical patterns and knowledge implicitly embedded in the data.</p>



<p class="wp-block-paragraph">This is all easier said than done. The process of adjusting a model’s internal parameters so it gets better at predicting the next token in sequences is called <em>training</em>. During training, the model repeatedly guesses the next token in a given sequence, compares its prediction to the actual one, measures the error, and updates its parameters to reduce that error across billions of examples. Over time, that process teaches the model the statistical relationships that will allow it to generate coherent language (or code, or images) later.</p>



<h2 class="wp-block-heading"><strong>What is a foundation model?</strong></h2>



<p class="wp-block-paragraph">You’ll often hear the word <em>large</em> used for transformer-based models of these types, like the LLMs we mentioned earlier. <em>Large</em> in this context refers to the large number of internal numerical values that the model adjusts during training to represent what it has learned, along with breadth and diversity of data used to train the model and the underlying compute resources powering this whole process.</p>



<p class="wp-block-paragraph">This is in contrast with the narrow models of the earlier era of AI/ML, which werebuilt for one purpose and trained on a limited dataset. For instance, a spam filter may be very good at what it does, but it’s only trained on email data and all it can do is classify emails. Large models, by contrast, serve as what’s known as <em>foundation models</em>. They’re trained broadly on diverse data (text, code, images, or multimodal data) and then adapted or specialized for many downstream tasks.</p>



<p class="wp-block-paragraph">These foundation models are the basis for most of the popular generative AI tools and services on the market today. They can be specialized in several ways:</p>



<ul class="wp-block-list">
<li><strong>Fine-tuning:</strong> Giving a foundation model further training on a smaller, task-specific dataset</li>



<li><strong>Retrieval-augmented generation</strong> <strong>(RAG):</strong> Giving the model the ability to pull in external knowledge when asked a question</li>



<li> <strong>Prompt engineering</strong>: Tailoring a query so the model gives the sort of answers you’re looking for.</li>
</ul>



<h2 class="wp-block-heading"><strong>How do AI systems write computer code?</strong></h2>



<p class="wp-block-paragraph">One of the surprising discoveries of the gen AI era was that in recent years was that foundation models trained on natural-language text can also, when fine-tuned with code examples, also write computer code — often better than many purpose-built systems. Still, it makes sense, when you think about it — after all, high-level computer languages are designed by humans and ultimately based on human language.</p>



<p class="wp-block-paragraph">This <a href="https://www.infoworld.com/article/2338500/llms-and-the-rise-of-the-ai-code-generators.html?utm_source=chatgpt.com">2023 InfoWorld article</a> highlights how models like PaLM, LLaMA and other transformer-based systems fine-tuned on code repositories propelled this shift, but since AI giants like <a href="https://www.computerworld.com/article/3843138/agentic-ai-ongoing-coverage-of-its-impact-on-the-enterprise.html">OpenAI</a> have moved into this space. This all matters because code generation (or code-assisted productivity) has become a key enterprise use case of generative AI — perhaps <em>the </em>key use, given the industry’s enthusiastic adoption of it.</p>



<h2 class="wp-block-heading"><strong>What are AI agents?</strong></h2>



<p class="wp-block-paragraph">So far, we’ve been talking about chatbots, writing assistants, image-generation tools. They respond to prompts, output text or images, and then stop. A new category of tool called <em><a href="https://www.computerworld.com/article/3843138/agentic-ai-ongoing-coverage-of-its-impact-on-the-enterprise.html">agentic AI</a></em> goes further: it <em>plans</em>, <em>executes</em>, and in many cases <em>learns</em> as it works.</p>



<p class="wp-block-paragraph">Because large models already understand language, code, and even structured data to some extent, they can be repurposed to generate not only descriptive text but <em>operational instructions</em>. For example: an agent might parse the intent “generate a sales-report”, then format internal calls like getData(salesDB, region=NA, period=lastQuarter), and then call an API, all by generating text that’s interpreted as instructions. The <a href="https://www.infoworld.com/article/4064169/how-mcp-is-making-ai-agents-actually-do-things-in-the-real-world.html.">MCP framework</a> standardizes the “language” of those instructions and the plug-points into tools and data so that the model doesn’t need bespoke integrations for each new workflow.</p>



<p class="wp-block-paragraph">These kinds of autonomous agents have several enterprise use cases:</p>



<ul class="wp-block-list">
<li><strong>Software automation</strong>: Agents that generate code, call unit tests, deploy builds, monitor logs and even roll back changes autonomously.</li>



<li><strong>Customer support</strong>: Instead of simply drafting responses, agents interact with CRM APIs, update ticket statuses, escalate issues, and trigger follow-up workflows.</li>



<li><strong>IT operations/AIOps</strong>: Agents <a href="https://www.cio.com/article/222623/7-things-to-know-about-ai-in-the-data-center.html">monitor infrastructure, identify anomalies, open/close tickets, or auto-remediate</a> based on defined rules and context from logs.</li>



<li><strong>Security</strong>: Agents may detect threats, initiate alerts, isolate compromised systems, or even attempt to manage threat containment — though this raises new risks.</li>
</ul>



<h2 class="wp-block-heading"><strong>How can you implement generative AI in the enterprise?</strong></h2>



<p class="wp-block-paragraph">We’ve now touched on <em>what</em> generative AI can do. But <em>how</em> can you make it work reliably in your business. The difference between a pilot and full-scale deployment often comes down to systems, structure and governance as much as to models themselves. <em>InfoWorld’</em>s Matt Asay offers a <a href="https://www.infoworld.com/article/4044919/enterprise-essentials-for-generative-ai.html">deep dive into enterprise gen AI essentials</a>, but here are some important points to keep in mind:</p>



<p class="wp-block-paragraph"><strong>Choosing between API, open-source or custom fine-tuned models. </strong>One of the first major decisions for any enterprise project is: do you use a model via an API (e.g., from a vendor like OpenAI or Anthropic), deploy an open-source model internally, or build/fine-tune a custom model yourself? Each has trade-offs.</p>



<p class="wp-block-paragraph">APIs offer speed and minimal setup, but may expose data, limit customization or accrue high cost — and will leave you at the mercy of your vendor. Open source allows internal control and may ease fine-tuning, but requires infrastructure, expertise, and support. Custom fine-tuning gives you the tightest alignment to your use-case, but lengthens time to value and increases risk.</p>



<p class="wp-block-paragraph"><strong>Governance, data privacy and compliance. </strong>Deploying generative AI in an enterprise setting raises new governance, privacy and regulatory issues. For example: Who owns the data that’s ingested? How is proprietary data protected if you call a third-party API? What traceability exists for model outputs—a huge question for regulated industries? One useful framework is covered in “A GRC framework for securing generative AI” Data governance <a href="https://www.infoworld.com/article/2336154/how-data-governance-must-evolve-to-meet-the-generative-ai-challenge.html">must adapt for the new era</a>,  and <a href="https://www.infoworld.com/article/3604732/a-grc-framework-for-securing-generative-ai.html">new frameworks are evolving to help</a>.</p>



<p class="wp-block-paragraph"><strong>Human-in-the-loop review. </strong>Even the best models make mistakes and cannot simply be put on autopilot. You need a <em>human-in-the-loop (HITL)</em> process: real people need to review outputs, validate for bias, approve high-stakes content, and tune prompts or models based on feedback. Incorporating HITL checkpoints helps mitigate risk and improve overall quality.</p>



<p class="wp-block-paragraph"><strong>Integration with existing systems and RAG pipelines. </strong><a href="https://www.infoworld.com/article/2337050/how-rag-completes-the-generative-ai-puzzle.html">Retrieval-augmented generation</a>, which we touched on earlier, connects foundation models into business workflows, systems, and enterprise data stores. RAG can bind LLMs to your organization’s internal knowledge bases, thereby reducing <em>hallucinations </em>(which we’ll discuss in a moment) and increasing the relevance of gen AI output.</p>



<aside class="sidebar">
<h3><strong> Implementation best practices for generative AI</strong></h3>
<p> Here are four AI best practices to keep in mind:</p>
<ol>
<li> Guardrails: Define clear operational boundaries. Examples: restrict sensitive data output, enforce access controls, log model interactions.</li>
<li> Prompt engineering: Because much of what the model will do depends on how it’s prompted, invest in prompt design, versioning, review, and testing.</li>
<li> Evaluation metrics: Define appropriate KPIs (accuracy, latency, cost, business outcome), monitor them and iterate.</li>
<li> Model observability: Treat generative-AI systems like software — monitor performance, detect drift, handle failures gracefully, audit outputs and maintain traceability.</li>
</ol>
</aside>




<h2 class="wp-block-heading"><strong>What causes AI hallucinations?</strong></h2>



<p class="wp-block-paragraph">Probably the biggest limitation of generative AI is what those in the industry call <em>hallucinations</em>, which is a perhaps misleading term for output that is, by the standards of humans who use it, false or incorrect.  </p>



<p class="wp-block-paragraph">Every generative AI system, no matter how advanced, is built around prediction. Remember, a model doesn’t truly <em>know</em> facts—it looks at a series of tokens, then calculates, based on analysis of its underlying training data, what token is most likely to come next. This is what makes the output fluent and human-like, but if its prediction is wrong, that will be perceived as a hallucination.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" src="https://b2b-contenthub.com/wp-content/uploads/2025/10/GenAI_takeaways.jpg?quality=50&amp;strip=all&amp;w=1024" alt="Table describing five key points about generatvie AI" class="wp-image-4082262" width="1024" height="648" sizes="auto, (max-width: 1024px) 100vw, 1024px"><figcaption class="wp-element-caption">Generative AI, foundation models, agentic AI, governance, and implementation strategy top the list of top generative AI takeaways.</figcaption></figure><p class="imageCredit">Foundry</p></div>



<p class="wp-block-paragraph">Because the model doesn’t distinguish between something that’s known to be true and something likely to follow on from the input text it’s been given, hallucinations are a direct side effect of the statistical process that powers generative AI. And don’t forget that we’re often pushing AI models to come up with answers to questions that we, who also have access to that data, can’t answer ourselves.</p>



<p class="wp-block-paragraph">In text models, hallucinations might mean inventing quotes, fabricating references, or misrepresenting a technical process. In code or data analysis, it can produce <a href="https://www.infoworld.com/article/3822251/how-to-keep-ai-hallucinations-out-of-your-code.html">syntactically correct but logically wrong results</a>. Even RAG pipelines, which provide real data context to models, only <em>reduce</em> hallucination—they don’t eliminate it. Enterprises using generative AI need <a href="https://www.cio.com/article/4073606/reducing-llm-hallucinations-in-enterprise-systems.html">review layers, validation pipelines, and human oversight</a> to prevent these failures from spreading into production systems.</p>



<h2 class="wp-block-heading"><strong>What are some other problems with generative AI?</strong></h2>



<p class="wp-block-paragraph">Generative AI has proven to be such a disruptive technology that’s stoking near-apocalyptic fears that it will result in a superintelligence that will enslave or destroy humanity. Meanwhile, in the present day, increasingly troubling reports of so-called <a href="https://www.psychologytoday.com/us/blog/urban-survival/202507/the-emerging-problem-of-ai-psychosis">AI psychosis</a> are emerging, where people have mental health episodes triggered by the uncanny and sometimes sycophantic ways chatbots affirm whatever you talk to them about and try to keep the conversation going.</p>



<p class="wp-block-paragraph">Compared to such existential questions, the following business-related problems may seem petty. But they’re real issues for enterprises considering investing in AI tools.</p>



<ul class="wp-block-list">
<li><strong>Data leakage and regulatory risk. </strong>When a model is fine-tuned or prompted with sensitive information, that data may be memorized and unintentionally reproduced. Using <a href="https://www.csoonline.com/article/3819170/nearly-10-of-employee-gen-ai-prompts-include-sensitive-data.html">third-party APIs without strict controls</a> can expose proprietary or personally identifiable information (PII). Regulatory frameworks like GDPR and HIPAA require explicit governance around where training data resides and how inference results are stored.</li>



<li><strong>Prompt injection </strong>occurs when an attacker manipulates a model’s instructions—embedding hidden directives or malicious payloads in user input or external content the model reads. This can override safety rules, expose internal data, or execute unintended actions in agentic systems. Guardrails that sanitize inputs, restrict tool-calling permissions, and validate outputs are becoming essential.</li>



<li><strong>Copyright and content ownership. </strong>Many foundation models are trained on data scraped from the public internet, creating disputes over copyright and data provenance. Enterprises using generated output commercially need to confirm usage rights and review indemnity terms from vendors.</li>



<li><strong>Unrealistic productivity expectations. </strong>Finally, organizations sometimes expect generative AI to deliver instant productivity gains. The reality, it turns out, is more <a href="https://leaddev.com/velocity/ai-doesnt-make-devs-as-productive-as-they-think-study-finds">mixed</a>. Enterprise adoption requires infrastructure, governance, retraining, and cultural change. The models accelerate work once properly integrated, but they don’t automatically replace human judgment or oversight.</li>
</ul>



<p class="wp-block-paragraph">The current generation of enterprise AI systems includes several layers of defense against these risks:</p>



<ul class="wp-block-list">
<li><em>Guardrails</em> that constrain model behavior and filter unsafe outputs.</li>



<li><em>Model validation</em> frameworks that measure factual accuracy and consistency before deployment.</li>



<li><em>Policy layers</em> that enforce compliance rules, redact sensitive data, and log model actions.</li>
</ul>



<p class="wp-block-paragraph">These safeguards reduce—but don’t remove—the inherent uncertainty that defines generative AI.</p>



<h2 class="wp-block-heading"><strong>GenAI: essential for the enterprise</strong></h2>



<p class="wp-block-paragraph">Generative AI has evolved from a novelty into a core layer of enterprise technology. Foundation models and agentic systems now power automation, analytics, and creative workflows — but they remain fundamentally probabilistic tools. Their strength lies in scale and adaptability, not perfect understanding.</p>



<p class="wp-block-paragraph">For organizations, success depends less on chasing model breakthroughs than on integrating these systems responsibly: building guardrails, maintaining oversight, and aligning them with real business needs. Used wisely, generative AI can amplify human capability rather than replace it.</p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Which AI model should you bet your company on?]]></title>
<description><![CDATA[Every day this past week I did something I suspect millions of other people also did: I stared at an LLM model picker and wondered which one I was supposed to want.



OpenAI just released ⁠GPT-5.6 Sol, Terra, and Luna. Sol is the flagship. Terra offers much of its intelligence for less money. Lu...]]></description>
<link>https://tsecurity.de/de/3664783/ai-nachrichten/which-ai-model-should-you-bet-your-company-on/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664783/ai-nachrichten/which-ai-model-should-you-bet-your-company-on/</guid>
<pubDate>Mon, 13 Jul 2026 11:33:26 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Every day this past week I did something I suspect millions of other people also did: I stared at an <a href="https://www.infoworld.com/article/2335213/large-language-models-the-foundations-of-generative-ai.html">LLM </a>model picker and wondered which one I was supposed to want.</p>



<p>OpenAI just released ⁠<a href="https://openai.com/index/gpt-5-6/">GPT-5.6 Sol, Terra, and Luna</a>. Sol is the flagship. Terra offers much of its intelligence for less money. Luna is cheaper still. Anthropic released ⁠<a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5</a> at the end of June and Opus 4.8 the month prior, with a little Fable 5 emerging in between. Meanwhile, Google, which seemed to be winning the model wars a few months ago, is now getting shade from Gergely Orosz, who ⁠<a href="https://x.com/GergelyOrosz/status/2075160978493210685?s=20">argues that Gemini has slipped outside the top tier</a> for software development and has been out of the major model release game for <em>eons</em> (May 19).</p>



<p>Perhaps Orosz is right. Perhaps he’ll be wrong again in six weeks. Honestly, it’s exhausting.</p>



<p>I use ChatGPT and Claude constantly and still have no principled idea which model to choose most of the time. I tend to click whatever looks like the biggest, most expensive option because I don’t know what I’m giving up by choosing something smaller. “Instant” sounds dangerously unserious. “Thinking” sounds expensive but powerful.</p>



<p>A quick <a href="https://www.linkedin.com/feed/update/urn:li:activity:7481369774401409024/">survey of my LinkedIn crowd</a> suggests others also feel my “WHICH MODEL???” pain. More importantly, I suspect most enterprises do, too.</p>



<h2 class="wp-block-heading"><a></a>A model doesn’t rot</h2>



<p>Before getting carried away, however, it’s worth considering whether any of this model churn actually matters. After all, a model doesn’t rot. The model an enterprise put into production in March performs just as well in July as it did when the company selected it. “Obsolete” generally means that something better now exists, not that the deployed model suddenly stopped summarizing insurance claims or classifying support tickets. (In other words, once you have something working, the idea that “but maybe Opus 200.2 is better!” is really a FOMO problem, not a performance issue.)</p>



<p>Most enterprise workloads don’t live at the frontier anyway. Extraction, summarization, classification, document comparison, and customer-service assistance often work perfectly well with smaller, cheaper models. OpenAI’s own pitch for the trio of GPT-5.6 models isn’t simply that Sol is better. It’s that ⁠Terra and Luna deliver different combinations of intelligence, latency, and cost. Luna, the cheapest tier, nearly matches the previous generation’s peak performance at less than half the estimated cost, according to OpenAI.</p>



<p>The practical question, of course, is where to start. An enterprise can’t test every model, every reasoning setting, and every price tier before doing any work. So here’s my advice (which I don’t follow in my own work, but I’m not defining enterprise strategy and can be a little price-insensitive). Start with the cheapest credible model that appears capable of the task. Give it a representative set of real examples and, before you start testing, define what counts as good enough. If it passes, stop. If it fails, move up a tier or try a model with strengths better suited to the work.</p>



<p>That sounds almost offensively simple, but it reverses the way many people, including me, use these products. We start with the biggest model because we’re afraid of what we might lose. Enterprises should start lower and require evidence before paying for more intelligence.</p>



<p>There are exceptions, of course. For genuinely difficult work, such as autonomous coding, complex research, or high-stakes reasoning, beginning with a frontier model may save time. But even then, the goal should be to establish a quality ceiling, then test whether a cheaper model can meet it. It’s changing the question from “which model is best?” to “what is the least expensive model that reliably clears the bar for this job?”</p>



<p>For many workloads, that price improvement matters more than a few extra benchmark points. <a href="https://www.infoworld.com/article/2335519/ai-hype-isnt-helping-anyone.html">⁠As I argued back in 2023</a>, following AI hype doesn’t help anyone. If your model strategy depends on whichever benchmark screenshot is circulating on X this week, you don’t have a strategy. Not a viable one, anyway. Pick a model and ignore the noise.</p>



<p>Except, of course, when that noise suggests a serious signal.</p>



<h2 class="wp-block-heading"><a></a>Sometimes better really is better</h2>



<p>Frontier improvements aren’t always incremental, making it advantageous to consider an upgrade. Coding is the obvious example. There’s a significant difference between a model that suggests the next few lines of code and one that can inspect a repository, plan a change, use tools, run tests, discover its own mistakes, and keep working for an extended period. That isn’t merely a nicer autocomplete experience. It can reorganize a development workflow.</p>



<p>This is why enterprises can’t simply standardize on an 18-month-old model and declare victory. In some areas, particularly software development and other agentic work, better models can unlock compounding productivity. A model that reliably completes 80% of a bounded task rather than 50% may justify an entirely different division of labor between humans and machines.</p>



<p>Still, that upgrade isn’t free.</p>



<p>Models differ in how they interpret instructions, call tools, manage context, refuse requests, and fail. Prompts and scaffolding tuned for one model can regress when moved to another. Or costs can explode. As one of my Oracle colleagues discovered just this week, running the same tasks in GPT 5.6 was orders of magnitude more expensive than 5.5. The API change may be trivial, but the revalidation and implications are not.</p>



<p>This leaves enterprises caught between two bad options. They can freeze and potentially miss out on meaningful improvements or chase every release and repeatedly test production systems on faith. What to do?</p>



<h2 class="wp-block-heading"><a></a>Stop making model bets</h2>



<p>The answer is to stop making LLM bets and start making job-to-be-done bets. Stop asking which model is fastest. Instead, figure out what work you are trying to improve. What does a good result look like? How much latency and cost can the workflow tolerate? How wrong can it be before a human must intervene? Once those questions have answers, model selection becomes less opaque.</p>



<p>A difficult code migration may justify GPT-5.6 Sol or Claude Sonnet 5. A repetitive classification task may work just as well with Luna or another smaller model. A regulated workflow may require a model or deployment option that offers particular data controls. Sometimes the correct model is no LLM at all, like when I’m writing this post. Sorry, AI vendors! (At least you won’t get blamed for my mistakes.)</p>



<p>This is where evaluations become the center of enterprise AI strategy. <a href="https://www.infoworld.com/article/4166247/improving-ai-agents-through-better-evaluations.html">⁠As I’ve said before</a>, most companies don’t have an AI quality problem so much as an AI measurement problem. Hence, a private evaluation suite built from real company work is the only leaderboard that matters. Does the new model materially improve quality? If so, use it! Does it reduce cost or latency? Again, that’s your free pass to adoption. Does the improvement justify the expense and effort of revalidation? If yes, continue.</p>



<h2 class="wp-block-heading"><a></a>Make model releases boring</h2>



<p>As important as the model is, keep in mind that AI success always comes back to <em>your</em> company’s data, <em>your</em> company’s workflows<em>, your</em> company’s integrations, etc. That’s the ⁠<a href="https://www.infoworld.com/article/4157506/mastering-the-dull-reality-of-sexy-ai.html">dull reality behind sexy AI</a>. Retrieval, <a href="https://www.infoworld.com/article/4189492/how-to-improve-the-memory-of-ai-agents.html">memory</a>, governance, data quality, <a href="https://www.infoworld.com/article/2262666/what-is-observability-software-monitoring-on-steroids.html">observability</a>, and feedback loops aren’t as exciting as a new model launch, but they’re what ultimately make AI truly work.</p>



<p>Again, when it’s time to consider something new, the principle should be to default to the least expensive model that reliably passes your evaluations. Only escalate harder tasks to more capable models when measurement shows that the premium pays. Tip: Make this invisible to employees so that the system routes to the best model for a particular prompt. As <a href="https://www.linkedin.com/feed/update/urn:li:activity:7481369774401409024/?dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287481372047860715522%2Curn%3Ali%3Aactivity%3A7481369774401409024%29">dbt Labs’ Jon Lewis expresses</a> it, “The best model is ‘Auto’ and I won’t hear anyone say otherwise.” OpenAI’s own ⁠<a href="https://developers.openai.com/api/docs/guides/latest-model">migration guidance</a> recommends testing models on representative tasks, including trying a lower reasoning level rather than automatically cranking everything to the maximum.</p>



<p>As for me, I’ll probably keep clicking the shiniest option. I don’t have a formal evaluation suite for InfoWorld columns, and the marginal cost is a subscription I already pay. Enterprises don’t get that excuse.</p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Where the software development jobs are now]]></title>
<description><![CDATA[While many technology companies have slowed hiring or even launched significant layoffs, that doesn’t mean job opportunities have dried up for software developers. In fact, skilled developers—particularly those with knowledge of AI—are in demand in other industries.



The key to success for deve...]]></description>
<link>https://tsecurity.de/de/3664782/ai-nachrichten/where-the-software-development-jobs-are-now/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664782/ai-nachrichten/where-the-software-development-jobs-are-now/</guid>
<pubDate>Mon, 13 Jul 2026 11:33:25 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>While many technology companies have slowed hiring or even launched <a href="https://www.trueup.io/layoffs" data-type="link" data-id="https://www.trueup.io/layoffs">significant layoffs</a>, that doesn’t mean job opportunities have dried up for software developers. In fact, skilled developers—particularly those with <a href="https://www.infoworld.com/article/4025073/9-ai-development-skills-tech-companies-want.html" data-type="link" data-id="https://www.infoworld.com/article/4025073/9-ai-development-skills-tech-companies-want.html">knowledge of AI</a>—are in demand in other industries.</p>



<p>The key to success for developers looking to snatch up these roles is to be well-prepared to meet the needs of potential employers in a variety of sectors.</p>



<p>“The demand for developers in non-tech sectors is real and growing, but the roles look different from what you’d find at a software company,” says <a href="https://drexel.edu/cci/about/directory/A/Awasthi-Pragati/" data-type="link" data-id="https://drexel.edu/cci/about/directory/A/Awasthi-Pragati/">Pragati Awasthi</a>, assistant teaching professor of AI and data science at Drexel University.</p>



<p>“Across all these sectors, the common thread is that software is no longer a support function; it is embedded in core operations,” Awasthi says. “The developer in these environments is often the person translating domain-specific business problems into technical solutions, which requires a different profile than a pure product engineer at a tech firm.”</p>



<h2 class="wp-block-heading">Opportunity knocks</h2>



<p>The tech industry has long been a mainstay as far as employing software developers. But as these businesses trim staffs in efforts to cut expenses, that has impacted the hiring landscape. Even as the tech sector scales back, however, companies in industries such as financial services/fintech, healthcare/healthtech, retail/ecommerce, and manufacturing are looking to acquire programming talent.</p>



<p>“The unifying factor is data complexity,” Awasthi says. “These industries generate large volumes of sensitive, regulated, or operationally critical data, and they need developers who can build and maintain systems that handle it responsibly.”</p>



<p>While recruiting firm Summit Search Group has placed developers in roles with technology companies, “it is just as common to recruit them for roles outside this niche,” says <a href="https://www.linkedin.com/in/matterhard/" data-type="link" data-id="https://www.linkedin.com/in/matterhard/">Matt Erhard</a>, managing partner at the company. “There are actually a fairly wide variety of roles available for developers in industries beyond tech,” Erhard says.</p>



<p>For example, in financial services Summit Search Group has seen significant hiring for back-end and data engineers who can build and maintain fraud detection systems, digital banking platforms, and regulatory tools, Erhard says. In healthcare, companies are hiring developers to build AI-driven diagnostics platforms and patient portals, or to work with systems that manage electronic health records, he says.</p>



<p>In manufacturing and industrial companies, developers are needed for systems integration and embedded software related to predictive maintenance, <a href="https://www.networkworld.com/article/963923/what-is-iot-the-internet-of-things-explained.html" data-type="link" data-id="https://www.networkworld.com/article/963923/what-is-iot-the-internet-of-things-explained.html">Internet of Things</a> (IoT) systems, and smart factories. And in retail and ecommerce, there’s strong demand for <a href="https://www.infoworld.com/article/2259033/full-stack-developer-what-it-is-and-how-you-can-become-one.html" data-type="link" data-id="https://www.infoworld.com/article/2259033/full-stack-developer-what-it-is-and-how-you-can-become-one.html">full-stack developers</a> and data developers who can handle logistics systems, omni-channel platforms, and personalization engines, Erhard says.</p>



<p>“One significant function where we’ve been placing developer talent lately is in developing business systems and internal applications,” Erhard says. These roles often have titles such as systems engineer or application developer, and professionals are hired to handle tasks such as customizing customer relationship management (CRM) or enterprise resource planning (ERP) platforms, building workflow automation tools or modernizing legacy systems, he says.</p>



<p>Other core functions for which Summit Search Group has placed a lot of developers include data, analytics, and AI-enablement. “That could be directly involved with <a href="https://www.infoworld.com/article/2263668/data-wrangling-and-exploratory-data-analysis-explained.html" data-type="link" data-id="https://www.infoworld.com/article/2263668/data-wrangling-and-exploratory-data-analysis-explained.html">data engineering</a> or in building tools like reporting systems and <a href="https://www.infoworld.com/article/2263668/data-wrangling-and-exploratory-data-analysis-explained.html" data-type="link" data-id="https://www.infoworld.com/article/2263668/data-wrangling-and-exploratory-data-analysis-explained.html">ETL [extract, transform, load]</a> pipelines,” Erhard says.</p>



<p>The firm also has handled searches for developers who can build and maintain customer-facing products for banking, healthcare, and retail companies, such as mobile apps or digital platforms customers can use to interact with companies.</p>



<p>Randstad Digital, a provider of global technology talent, sees demand for roles including web developers, system developers, and app developers. “These professionals would work on anything from customer-facing platforms to internal tools,” says <a href="https://www.linkedin.com/in/mpmorris36/" data-type="link" data-id="https://www.linkedin.com/in/mpmorris36/">Michael Morris</a>, global head of platform and talent at the company. “Non-tech companies are also often hiring roles like software architecture and <a href="https://www.infoworld.com/article/2255028/what-is-devops-transforming-software-development.html" data-type="link" data-id="https://www.infoworld.com/article/2255028/what-is-devops-transforming-software-development.html">devops</a> to help scale existing technology. These involve being more ingrained in the business, like building a supply chain system for a retailer, rather than creating individual tech products like you would at a technology company.”</p>



<h2 class="wp-block-heading">Prep for success</h2>



<p>To increases the chances of success at landing developer jobs outside of the tech industry, development professionals would be wise to follow some good practices.</p>



<h3 class="wp-block-heading">Boost AI skills</h3>



<p>One best practice is to boost skills in using AI-powered tools and get familiar with all things AI.</p>



<p>“Get fluent with AI-assisted development and its limits,” Awasthi says. “This is not optional. Organizations across every sector expect developers to use AI coding tools productively. But the more durable skill is knowing when AI output is wrong, incomplete, or unsuitable for a regulated context. That critical evaluation capacity is what non-tech employers are increasingly trying to hire.”</p>



<p>AI does not necessarily replace the need for human developers so much as it changes the skills profile for those roles, Erhard says. “The biggest difference in recent years is that AI literacy is now a non-negotiable,” he says. “At minimum, developers today need to understand concepts like <a href="https://www.infoworld.com/article/4122440/what-is-prompt-engineering-the-art-of-ai-orchestration.html" data-type="link" data-id="https://www.infoworld.com/article/4122440/what-is-prompt-engineering-the-art-of-ai-orchestration.html">prompt engineering</a> and how to use AI tools to improve their efficiency.”</p>



<p>One thing many job candidates don’t expect is that the rise of AI has also increased the importance of high-level skills such as problem framing, system design, and cross-functional communication,” Erhard says. “Essentially, if something is related to development but too complex or nuanced for an AI to handle effectively, then the demand is high for human developers who have that expertise,” he says.</p>



<p>Candidates who land roles consistently have experience building AI-augmented workflows along with standard coding skills, Erhard says. “Employers increasingly expect to hire developers who can leverage AI, so demonstrating this experience on your résumé can be very beneficial,” he says.</p>



<h3 class="wp-block-heading">Gain domain knowledge</h3>



<p>Summit Search Group is seeing high demand for developers with deep domain knowledge in an organization’s specific industry. “So, for instance, if someone is both an experienced developer and has expertise in healthcare compliance, or financial regulations, then those candidates tend to be very sought after,” Erhard says.</p>



<p>Domain fluency is an underrated skill, Awasthi says. “A developer who understands healthcare compliance, financial regulation, or manufacturing process logic is significantly harder to replace than one who only writes clean code,” she says. “AI can generate boilerplate. It cannot navigate a HIPAA audit or explain a model’s output to a compliance officer.”</p>



<p>Development professionals should “pick an industry and learn it seriously; not just the technology stack but the regulatory environment, the business model, and the actual problems practitioners face,” Awasthi says. “A developer who has read about HIPAA, or spent time understanding credit risk, is immediately more valuable in those hiring contexts.”</p>



<p>It’s also vital to demonstrate real-world, practical application of skills, not just credentials. “The strongest candidates have projects in their portfolio that directly tie to and solve real business problems,” Erhard says.</p>



<h3 class="wp-block-heading">Acquire soft skills</h3>



<p>And then there are the soft skills that are becoming more of a differentiator than they were in the past. As AI handles more routine coding, human developers are expected to make more architectural decisions and collaborate across departments, Erhard says. “Strong communication and problem-solving skills are critical for many of the developer roles that we’re filling today,” he says.</p>



<p>While technical skills are still relevant for developers using and managing AI tools, “they also need to develop the skill of ‘deeper thinking’ and learn how to think one step ahead,” Morris says. “This includes skills like system design mastery—understanding the macro view and learning how <a href="https://www.infoworld.com/article/2263327/what-are-microservices-your-next-software-architecture.html" data-type="link" data-id="https://www.infoworld.com/article/2263327/what-are-microservices-your-next-software-architecture.html">microservices</a>, databases, and third-party APIs interact securely and efficiently.”</p>



<p>They also should become deeply fluent in the AI coding tools commonly used in their particular industry, with a strong understanding of how to prompt them for optimal output, Morris says. Product context awareness is also useful. “AI doesn’t know what the customer wants, but you do,” Morris says. “Understanding the business problem and the end-user experience is a requirement for being able to guide LLMs.”</p>



<h3 class="wp-block-heading">Master debugging and incident response</h3>



<p>Developers looking to break into non-tech sectors also should develop skills in debugging and incident response, Morris says. “Complex systems with multiple AI agents can, and will, fail, which means companies need humans to trace logic flaws to get the system back on track,” he says. “A mastery of root-cause analysis is a critical skill.”</p>



<p>“Security, compliance, and reliability are very important in non-tech industries like finance and healthcare,” says <a href="https://www.linkedin.com/in/rohit-agarwal/" data-type="link" data-id="https://www.linkedin.com/in/rohit-agarwal/">Rohit Agarwal</a>, co-founder of Zenius, a remote hiring company. “So employers want developers who also know regulatory environments well.”</p>



<h3 class="wp-block-heading">Network and keep learning</h3>



<p>To successfully pivot from jobs at tech companies, “continuous learning, upskilling, and building hybrid skills that combine technical and business knowledge are essential,” Morris says. “With the right preparation, tech professionals can adapt and continue to thrive in meaningful, dynamic careers.”</p>



<p>It’s also a good idea to join talent communities in fields of interest and “engage with other members in conversations that increase your knowledge through the collective intelligence of the community,” Morris says. “Take advantage of AI skilling opportunities relevant for your role, or better yet, where you want to go next. Experiment with the technology either on your own or through structured programs.” Ultimately, be curious and proactive, he says.</p>



<p>“I’d also recommend developers not to ignore referrals, direct outreach, and industry-specific communities during job search,” Agarwal says. “There are often a lot more opportunities available than the ones posted online.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Why AI needs contextual intelligence — not just bigger models]]></title>
<description><![CDATA[A product manager on my team recently asked me where we were seeing the most issues across the engineering team. Instead of guessing, I had an engineering lead point Claude at our Jira via an MCP connector and look at the bug patterns himself.



One team had a wildly disproportionate share of ti...]]></description>
<link>https://tsecurity.de/de/3664720/it-security-nachrichten/why-ai-needs-contextual-intelligence-not-just-bigger-models/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664720/it-security-nachrichten/why-ai-needs-contextual-intelligence-not-just-bigger-models/</guid>
<pubDate>Mon, 13 Jul 2026 11:08:42 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>A product manager on my team recently asked me where we were seeing the most issues across the engineering team. Instead of guessing, I had an engineering lead point Claude at our Jira via an MCP connector and look at the bug patterns himself.</p>



<p>One team had a wildly disproportionate share of tickets — about 50% of their sprint time was spent on “bugs,” versus roughly 25% for everyone else. The headline number suggested a quality problem.</p>



<p>It wasn’t. When we layered in the context around those tickets, almost none of them were bugs. They were manual workarounds for a missing product capability: customers asking us, one request at a time, to restore items they had accidentally deleted. Not shipping an item restore feature was burning roughly 1.5 engineers’ worth of capacity. I went back to our product team and said, “Build this, and you reclaim a person and a half.”</p>



<p>The analysis took 45 minutes. It was only possible because our data was already organized, tagged by team, connected to contributors, accessible through MCP and protected by role-based access. None of that is “AI.” All of it is the layer underneath AI that almost nobody invests in first. That’s probably because the investment is unglamorous: updating data dictionaries, access controls, team taxonomies, system-to-system mappings. Most of the work has been the same for twenty years. AI just raised the cost of skipping it.<br></p>



<h2 class="wp-block-heading">The intelligence underneath the models</h2>



<p>I keep coming back to the value of context data layers as a CTO in the middle of an AI rollout. I have started calling that value proposition contextual intelligence because I haven’t found a better name. Anthropic’s engineering team has been calling this kind of work “<a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="nofollow">context engineering</a>” since late 2025, and <em>CIO</em><a href="https://www.cio.com/article/4080592/context-engineering-improving-ai-by-moving-beyond-the-prompt.html"> ran its own feature on the term</a> shortly after. Whether you describe it as contextual intelligence or context engineering, it’s the part of the stack where the actual programming work still lives.</p>



<p>If business logic is your company’s official org chart, then contextual intelligence is knowing who actually gets things done, how decisions are actually made and what the unwritten rules are. One is theory. The other is reality.</p>



<p>Most enterprise systems capture the theory. The systems that capture how work actually happens — what people do, how teams operate, where decisions get stuck — are rarer and harder to build. And modern LLMs, it turns out, are useless without both.</p>



<p>I learned this the hard way at a recent company hackathon. Nine engineering teams, one prompt: make our operational dataset more usable through AI. My team built persona-based chatbots (CFO, CIO, sales manager) on top of an MCP server backed by Postgres and our enrichment data. Other teams built dashboard generators, Looker conversational analytics and workflow agents.</p>



<p>The initial demos all had the same problem. Claude could talk to our data, but the answers were either generic or confidently wrong. The CFO persona would happily report a “spend trend” that quietly conflated two distinct cost categories across two different tables. The CIO persona would answer questions about team productivity, but the averages across roles should never have been aggregated. The sales manager persona returned answers that were technically correct against the schema and completely wrong against the business. The raw data was rich. The context layer around it didn’t exist yet. Chatting with raw data is not an AI product. It’s a demo.</p>



<p>One of my senior engineers spent the second day ripping out the agent’s direct database connection. He stopped trying to prompt-engineer the LLM to understand our business and instead codified that logic into the data pipeline. Working backward from the failed CFO answers, he mapped out the implicit knowledge an experienced controller relies on: Explicitly defining which legacy tables actually represent ‘spend,’ writing the rules for currency normalization and hardcoding our fiscal time windows. He built a series of semantic SQL views to enforce these rules and restricted the MCP server to exposing only this curated layer. When we pointed the same model at those same questions, it returned completely different answers. They were specific, evidence-based and grounded in our actual business reality. The model didn’t get smarter. The engineering beneath it did.</p>



<h2 class="wp-block-heading">The same pattern shows up everywhere I look right now</h2>



<p><a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/one-year-of-agentic-ai-six-lessons-from-the-people-doing-the-work" rel="nofollow">McKinsey</a> keeps publishing that software development tops enterprise AI use cases, with companies reporting 30–50% productivity gains in pilots. The pilot numbers are real. They rarely translate to top- or bottom-line impact in production. Our own company data tells the same story: Between Q1 2025 and Q1 2026, our total AI tool usage grew by 328% (over 4x). Over that same period, PR throughput grew by just 49%.</p>



<p>That gap — adoption way up, outcomes inching along — is the context gap. Plug a generic agent into raw, uninterpreted data, and it will act inefficiently at best, harmfully at worst. An agent optimizing sales without your customer segmentation or product hierarchy will confidently recommend the wrong thing. Anthropic<a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="nofollow"> </a><a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="nofollow">framed the shift directly</a>: building with language models is becoming “less about finding the right words and phrases for your prompts, and more about answering the broader question of what context configuration is most likely to generate our model’s desired behavior.” That second question — what context configuration  — is the entire game. Most organizations are still answering the first one.</p>



<h2 class="wp-block-heading">Where the work actually lives</h2>



<p>A growing number of CTOs I talk to are shifting their AI investments accordingly. Less attention on the model. More on the layer between the model and the data.</p>



<p>When peers ask me what that actually looks like day-to-day, I tell them I give every engineering role the same mandate: the LLM should never see raw, uncontextualized data.</p>



<p>In practice, that breaks down to three pieces of work, none of them glamorous.</p>



<p>The first is semantic middleware. We need code that transforms raw data into business-meaningful signals before it ever reaches the model. Our feature stores hold things like “employee code velocity on critical-path features,” not “X logged 50 Git commits.” The work of figuring out what “critical-path” means in our product, in our org, on this team is the work. It does not get cheaper because the model has gotten better.</p>



<p>The second is multi-agent design. Instead of one omniscient orchestrator, we run smaller agents scoped to specific domains, each with rules that catch the failure modes the main model is known for. We pair them with RAG that retrieves precomputed insights, with their rules attached, rather than raw documents. Validation checkpoints sit between steps and flag suggestions that violate known constraints, such as averaging productivity across completely different job functions. The guardrails are not there to be clever. They are there because we already watched the model make those exact mistakes.</p>



<p>The third is evaluation that takes business logic seriously. When I look at a model, general benchmark accuracy is the least interesting number. I want to know whether it respects our constraints and integrates cleanly with our existing architecture. That sometimes means fine-tuning our patterns, sometimes constitutional approaches to embed principles, sometimes hybrid systems where deterministic rules sit alongside the probabilistic ones. The throughline is the same: validate against reality, not against the benchmark.</p>



<h2 class="wp-block-heading">Why this matters now</h2>



<p>The reason this matters more now than it did six months ago is that adoption is moving faster than measurement, let alone integration. Model Evaluation &amp; Threat Research’s (<a href="https://metr.org/" rel="nofollow">METR</a>) developer productivity work tells the story in a way they didn’t intend. In early 2025, they<a href="https://arxiv.org/pdf/2507.09089" rel="nofollow"> ran a controlled study</a> and found AI tools slowed experienced open-source developers by 19%. When they tried to<a href="https://metr.org/blog/2026-02-24-uplift-update/" rel="nofollow"> repeat the study in late 2025</a>, the experiment broke. Thirty to fifty percent of developers refused to submit tasks under the no-AI condition. They wouldn’t accept working without their tools. METR is now redesigning the study because the original methodology no longer holds up against how developers actually work. That’s how fast adoption moved. But I’d be willing to bet the organizational scaffolding required to convert that adoption into outcomes — context layers, workflow redesign, retraining around new tools — moved nowhere near as fast.</p>



<h2 class="wp-block-heading">Get ahead with context </h2>



<p>The teams I’ve seen succeed with AI built the context layer first. The teams I’ve seen struggle eventually built in context anyway, just at higher cost and with more scar tissue. Raw data is the new currency. But raw data without a context layer is cash sitting in a vault. It cannot act on anything. The difference between insight and noise is a layer of code that understands what your data means.</p>



<p>That layer is the work. It is where the next decade of competitive advantage will sit. And in my experience, the organizations that build it first are the ones that will actually get the productivity gains the rest of the market keeps promising.</p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Infrastructure for the agentic era: A new conversation layer for the Twilio Platform]]></title>
<description><![CDATA[A new era of customer engagement is taking shape. AI agents are quickly becoming integral to the way businesses serve, support, and sell to customers — able to respond, reason, and take action in ways that go far beyond scripted automation.



Many customer journeys, however, are still built on s...]]></description>
<link>https://tsecurity.de/de/3664598/it-security-nachrichten/infrastructure-for-the-agentic-era-a-new-conversation-layer-for-the-twilio-platform/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664598/it-security-nachrichten/infrastructure-for-the-agentic-era-a-new-conversation-layer-for-the-twilio-platform/</guid>
<pubDate>Mon, 13 Jul 2026 10:09:19 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>A new era of customer engagement is taking shape. AI agents are quickly becoming integral to the way businesses serve, support, and sell to customers — able to respond, reason, and take action in ways that go far beyond scripted automation.</p>



<p>Many customer journeys, however, are still built on systems that don’t talk to each other. Customer data lives in one place, channel history in another, and AI agents often operate with only part of the picture. Customers feel the pain when they switch between channels like voice and messaging, get transferred, and have to repeat themselves yet again. It doesn’t matter that they’ve been loyal to a brand for years, every interaction feels like a cold start. That is the conversation gap.</p>



<p>It’s clear that AI isn’t the problem, infrastructure is. Closing the gap requires new building blocks that focus on continuity, so context can carry forward across systems, channels, human agents, and AI agents.</p>



<p>To bridge the gap, at <a href="https://signal.twilio.com/?_gl=1*qsec1h*_gcl_aw*R0NMLjE3Nzk3MTY4MzguQ2p3S0NBanc1c19RQmhBZEVpd0FERF9nQnUyRVR4YTdGTFRCNDVPcktsd2dvbnZrQ3hZdlNtQXRJRHVoS09lOVJySXFsQ3k2eHZZajBob0NRZkVRQXZEX0J3RQ..*_gcl_au*MTAwMjE5MDU2OS4xNzc5MzUyNjYz*_ga*MTA5NDA4OTEuMTc3MTU2MTMzNg..*_ga_RRP8K4M4F3*czE3ODA5NzUwMjYkbzE3NyRnMSR0MTc4MDk3NzU4NiRqNjAkbDAkaDA.&amp;utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="sponsored">SIGNAL 2026</a>, we are introducing a new conversation layer for the Twilio Platform.</p>



<p>Twilio Conversation Orchestrator, Twilio Conversation Memory, and Twilio Conversation Intelligence are now generally available. Together, they help businesses coordinate interactions, preserve context, and connect human and AI agents so every conversation is more continuous and useful.</p>



<p>In addition to the new Conversations layer, we’re also announcing platform updates that make it easier to build, manage, and scale customer engagement on Twilio — from a reimagined Twilio Console to expanded channels and new voice AI capabilities.</p>



<h2 class="wp-block-heading">New building blocks for connected conversations</h2>



<p>The conversation gap does more than create inconsistent customer experiences. It hurts conversion and retention, increases operational costs, adds integration complexity, and makes agents less productive. The new platform capabilities we’re introducing are designed to fix that by coordinating interactions, maintaining context, and surfacing signals as conversations happen.</p>



<h2 class="wp-block-heading"><a></a>Conversation Orchestrator</h2>



<p><a href="https://www.twilio.com/en-us/blog/products/conversation-orchestrator?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="sponsored">Conversation Orchestrator</a> helps businesses coordinate interactions across Twilio channels without complex custom logic. Teams can configure it in Console or configure their implementation with the API. It connects interactions into a single thread and manages handoffs between human agents and automated systems.</p>



<h2 class="wp-block-heading">Conversation Memory</h2>



<p><a href="https://www.twilio.com/en-us/blog/products/launches/conversation-memory?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="noreferrer noopener">Conversation Memory</a> creates a living, identity-resolved profile by connecting customer data with conversation history and customer traits. That means each interaction starts with the right context. It’s built specifically for LLMs to reduce latency and token usage by surfacing the most relevant details when they matter.</p>



<p>A new Enterprise Knowledge API (now generally available) also allows teams to deliver more relevant experiences and ground interactions in trusted business knowledge such as FAQs, policies, and product documentation.</p>



<h2 class="wp-block-heading">Conversation Intelligence</h2>



<p><a href="https://www.twilio.com/en-us/blog/products/launches/conversation-intelligence?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="noreferrer noopener">Conversation Intelligence</a> provides real-time understanding of live interactions. Using prebuilt and custom LLM-based operators, it can detect changes in sentiment, flag potential escalations, and trigger action during a conversation, not only after it ends.</p>



<p>That gives teams the ability to respond sooner, support agents more effectively, and improve customer outcomes while the conversation is still in progress.</p>



<p>Together, these products help businesses create customer experiences that feel more connected across channels.</p>



<h2 class="wp-block-heading">Open by design</h2>



<p>Twilio remains neutral by design. We start with the premise that you know your business. We aren’t here to prescribe a model, framework, or data strategy. We provide the infrastructure that helps you build customer engagement in the way that works best for your business. You pick the model and agent runtime. You own the data.</p>



<p>That doesn’t mean you need to start from scratch, either. We partnered with Microsoft, AWS, and others to create blueprints that support faster development. We are also introducing an open-source developer toolkit, <a href="https://www.twilio.com/en-us/blog/products/launches/agent-connect?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="noreferrer noopener">Twilio Agent Connect</a> (now generally available), that lets your teams connect agents built on any LLM or framework directly to Twilio’s infrastructure.</p>



<p>For developers, this means more flexibility. For businesses, it means less lock-in and the ability to get value from existing investments. For partners, it means more ways to build with Twilio.</p>



<h2 class="wp-block-heading">A new front door</h2>



<p>We are also introducing a reimagined <a href="https://www.twilio.com/en-us/blog/products/launches/new-twilio-console?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="noreferrer noopener">Twilio Console</a>, because as customer engagement grows more complex, managing the infrastructure behind it should feel effortless.</p>



<p>The new Console is a single mission control center that brings your communications, identity, and data into one experience: one login, consistent logs across every surface, an intelligent Console Assistant, transparent billing insights, and streamlined compliance workflows that no longer slow you down.</p>



<p>Over the coming months, we’ll roll out this new Console experience to customers automatically. You can also opt in to gain early access.</p>



<h2 class="wp-block-heading">More channels, more control, smarter conversations</h2>



<p>In addition to these launches, we are announcing several updates that expand customer reach, support enterprise requirements, and make it simpler to build on Twilio.</p>



<ul class="wp-block-list">
<li><a href="https://www.twilio.com/en-us/messaging/channels/apple-messages-for-business?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="sponsored">Apple Messages for Business</a> (Private beta) and Twilio Email (GA) give teams new ways to reach customers on the channels they already use.</li>



<li>Data Residency for SMS (EU) (Public beta) enables teams to manage personal data locally to support regional data requirements.</li>



<li><a href="https://www.twilio.com/en-us/blog/products/launches/the-evolution-of-conversation-relay?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="sponsored">Conversation Relay</a> enhancements add PCI compliance, HIPAA eligibility, Insights, and support for Deepgram Flux for smarter turn detection — helping AI agents better understand when a person has finished speaking.</li>



<li><a href="https://www.twilio.com/en-us/blog/partners/integrations/provision-twilio-communications-channels-stripe-projects?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_infra-agentic-era_brandposthub" target="_blank" rel="sponsored">Stripe Projects integration</a> enables developers and AI agents to seamlessly provision Twilio within Stripe Projects in a single, programmable CLI workflow.</li>
</ul>



<h2 class="wp-block-heading">Built with our customers</h2>



<p>Bringing these new products to life required a close partnership with many beta customers and partners. This helped us understand real-world signals and needs to help make the capabilities robust from the start.</p>



<p>Among dozens of others, Centerfield, Constellation Dealerships, Car Finance 247, and Meera.ai leveraged Twilio to solve their own customer engagement challenges. These teams showed what is possible when businesses carry context forward, act on live conversation signals, and connect AI agents with human teams in the moments that matter.</p>



<p><a href="https://www.carfinance247.co.uk/" target="_blank" rel="noreferrer noopener">Car Finance 247</a>, a leading UK online car finance broker, is using Twilio to help recover stalled loan applications. When customers miss a field, need to correct information, or still need to confirm terms and conditions, AI-powered outreach across voice, SMS, and RCS, Conversation Memory tracks the application state. Conversation Orchestrator manages the outreach journey, and Flex helps bring in a human agent as needed. As Reg Rix, Co-Founder and CEO, shared:</p>



<p><em>“Because the platform remembers where each customer left off, we can pick up right where they stopped, helping them cross the finish line in a way that is modern, responsive, and genuinely helpful.”</em></p>



<p><a href="https://www.centerfield.com/" target="_blank" rel="sponsored">Centerfield</a>, a technology company powering AI-driven commerce, helps brands connect with consumers across digital and phone-based journeys. With Twilio, the team is connecting real-time conversation data with customer context to guide agents and AI systems in the moment, standardise what works, and improve performance at scale. As Aniketh Parmar, Chief Technology Officer, said:</p>



<p><em>“Performance comes down to how well every interaction moves a customer forward. We’re capturing each conversation in real time and applying what we already know about the customer to guide our agents and AI systems in the moment. With the Twilio Platform, including Conversation Orchestrator, Conversation Memory, and Conversation Intelligence, we can see what’s driving conversations so we can standardise what works, eliminate what doesn’t, and continuously improve outcomes at scale.”</em></p>



<p><a href="https://constellationdealer.com/" target="_blank" rel="sponsored">Constellation Dealerships</a> is using Twilio’s agent infrastructure to accelerate AI-powered engagement across its dealer network, moving from evaluation to measurable outcomes in days. As Richard Pineault, Director of R&amp;D, shared:</p>



<p><em>“The value of this partnership is evident—our team progressed from evaluating Twilio’s agent infrastructure to realising measurable outcomes within days. This rapid speed-to-value exemplifies the agility and innovation required to propel the dealership industry into the future.”</em></p>



<p><a href="http://meera.ai/" target="_blank" rel="sponsored">Meera.ai </a>is building on Twilio to modernise outbound engagement, replacing repeated manual follow-ups with always-on conversations across voice, SMS, and messaging. Vivek Zaveri, Chief Executive Officer, said:</p>



<p><em>“Meera.ai has partnered with Twilio since our inception to champion a conversation-first future for commerce. As the industry shifts toward real-time LLM-enabled interactions, Twilio’s Platform and the new Conversations products will help us reach customers in the moment.”</em></p>



<p>Together, these customers and partners show that the Twilio Platform can help businesses recover stalled journeys, improve live interactions, accelerate time to value, and create more connected experiences across AI agents, human teams, and every customer channel.</p>



<h2 class="wp-block-heading">The next era of customer engagement starts here</h2>



<p>As AI agents own more of customer engagement, businesses need infrastructure that keeps conversations connected across channels, systems, and teams. That means preserving context, coordinating handoffs, and acting on what is happening in real time.</p>



<p>That is what we are building with this next generation of the Twilio Platform: a new layer that connects channels, context, intelligence, and human and AI agents, helping businesses make every digital interaction more connected, more useful, and more amazing.</p>



<p>For 17 years, Twilio has helped builders create better ways for businesses to connect with their customers. In this next era, that connection matters more than ever.</p>



<p><a href="https://www.twilio.com/en-us/why-twilio?utm_source=foundry&amp;utm_medium=contentsyn&amp;utm_campaign=abm_brand_icp_sa_aw_tofu_apac_en&amp;utm_content=abm_lo_cs_ungatedcontent_end-cta-infra-agentic-era_brandposthub" target="_blank" rel="noreferrer noopener">Explore the new Conversations layer</a>, try the products, and let’s build what comes next, together.</p>



<hr class="wp-block-separator has-alpha-channel-opacity">
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[6 Risk-Assessment-Frameworks im Vergleich]]></title>
<description><![CDATA[Mit dem richtigen Framework lassen sich Risiken besser ergründen.FOTOGRIN – shutterstock.com



Für viele Geschäftsprozesse ist Technologie inzwischen unverzichtbar. Deshalb zählt diese auch zu den wertvollsten Assets eines Unternehmens. Leider stellt sie gleichzeitig jedoch auch eines der größte...]]></description>
<link>https://tsecurity.de/de/3664175/it-security-nachrichten/6-risk-assessment-frameworks-im-vergleich/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664175/it-security-nachrichten/6-risk-assessment-frameworks-im-vergleich/</guid>
<pubDate>Mon, 13 Jul 2026 06:04:55 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2024/10/1200x.png?w=1024" alt="6 Risk-Assessment-Frameworks im Vergleich" class="wp-image-3552768" width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"><figcaption class="wp-element-caption">Mit dem richtigen Framework lassen sich Risiken besser ergründen.</figcaption></figure><p class="imageCredit">FOTOGRIN – shutterstock.com</p></div>



<p>Für viele Geschäftsprozesse ist Technologie inzwischen unverzichtbar. Deshalb zählt diese auch zu den wertvollsten Assets eines Unternehmens. Leider stellt sie gleichzeitig jedoch auch eines der größten Risiken dar – was Risk-Assessment-Frameworks auf den Plan ruft.</p>



<p>IT-Risiken formal <a href="https://www.csoonline.com/article/3492180/cyber-risk-assessments-risikobewertung-hilft-cisos.html" target="_blank">zu bewerten</a>, ermöglicht es Organisationen, besser einzuschätzen, zu welchem Grad ihre Systeme, Devices und Daten schädlichen Einflüssen ausgesetzt sind. Etwa in Form von <a href="https://www.computerwoche.de/article/4155663/6-wege-uber-ki-gehackt-zu-werden.html" target="_blank">Cyberbedrohungen</a>, <a href="https://www.computerwoche.de/article/4149093/wenn-die-audit-falle-zuschnappt.html" target="_blank">Compliance-Verfehlungen</a> oder <a href="https://www.computerwoche.de/a/10-fakten-zu-datacenter-ausfaellen,3614299" target="_blank">Ausfällen</a>. Zudem können IT- und Sicherheitsentscheider deren Folgen mit Hilfe von entsprechenden Rahmenwerken auch besser abzuschätzen. Das Ziel besteht am Ende darin, sämtliche identifizierten Risiken – und ihren Impact – zu minimieren.</p>



<p>In diesem Artikel stellen wir Ihnen (in aller Kürze) sechs populäre Risk-Assessment-Frameworks vor, die jeweils auf spezifische Risikobereiche abgestimmt sind.</p>



<h2 class="wp-block-heading">1. COBIT</h2>



<p><strong>Das ist es:</strong> Hinter <strong>COBIT</strong> (<a href="https://www.isaca.org/resources/cobit" target="_blank" rel="noreferrer noopener">Control Objectives for Information and Related Technology</a>) steht der internationale IT-Berufsverband ISACA, der sich auf IT Governance fokussiert hat. Dieses sehr umfassende und breit angelegte Framework wurde entwickelt, um dabei zu unterstützen, Enterprise IT:</p>



<ul class="wp-block-list">
<li>zu verstehen,</li>



<li>zu designen,</li>



<li>zu implementieren,</li>



<li>zu managen und</li>



<li>zu steuern.</li>
</ul>



<p><strong>Das kann es:</strong> Laut ISACA definiert COBIT die Komponenten und Designfaktoren, ein optimales Governance-System aufzubauen und aufrechtzuerhalten. Die aktuelle Version, <a href="https://www.computerwoche.de/a/was-ist-cobit,3614637" target="_blank">COBIT 2019</a>, fußt auf einem Governance-Prinzipien-Sextett:</p>



<ol class="wp-block-list">
<li>Value für Stakeholder liefern</li>



<li>ganzheitlichen Ansatz realisieren</li>



<li>Governance-System dynamisch gestalten</li>



<li>Management von Governance trennen</li>



<li>auf individuelle Unternehmensanforderungen abstimmen</li>



<li>Ende-zu-Ende-Governance-System realisieren</li>
</ol>



<p><strong>So funktioniert es:</strong> Das COBIT-Framework ist auf Business-Fokus konzipiert und definiert eine Reihe generischer Prozesse, um IT-Komponenten zu managen. Dabei werden außerdem auch Inputs und Outputs, Schlüsselaktivitäten, Zielsetzungen, Performance-<a href="https://www.csoonline.com/article/3492322/cybersecurity-messen-10-kennzahlen-die-cisos-weiterbringen.html" target="_blank">Metriken</a> und ein grundlegendes Reifegradmodell festgelegt.</p>



<p><strong>Gut zu wissen:</strong> Laut ISACA ist COBIT flexibel zu implementieren und ermöglicht Unternehmen, ihre Governance-Strategie anzupassen.</p>



<h2 class="wp-block-heading">2. FAIR</h2>



<p><strong>Das ist es:</strong> Das Framework <strong>FAIR</strong> (<a href="https://www.fairinstitute.org/what-is-fair" target="_blank" rel="noreferrer noopener">Factor Analysis of Information Risk</a>) bildet eine Methodik ab, um unternehmensbezogene Risiken zu quantifizieren und zu managen. Dahinter steht das Fair Institute, eine wissenschaftlich ausgerichtete Non-Profit-Organisation, die sich dem Management von betrieblichen und sicherheitstechnischen Risiken verschrieben hat. Laut den Machern ist FAIR das einzige, quantitative Standardmodell auf internationaler Ebene, um diese Art von Risiken zu erfassen.</p>



<p><strong>Das kann es:</strong> FAIR bietet ein Modell, um die genannten Risiken in finanzieller Hinsicht zu verstehen, zu analysieren und zu quantifizieren. Laut dem Fair Institute unterscheidet es sich dabei insofern von anderen Risk-Assessment-Frameworks, als dass es seinen Fokus nicht auf qualitative Farbdiagramme oder numerisch gewichtete Skalen legt. Stattdessen will FAIR eine Grundlage liefern, um einen robusten Risikomanagement-Ansatz auszubilden.</p>



<p><strong>So funktioniert es:</strong> FAIR ermittelt in erster Linie Wahrscheinlichkeiten mit Blick auf die Frequenz und das Ausmaß von <a href="https://www.computerwoche.de/a/was-sie-ueber-dlp-wissen-muessen,3549370" target="_blank">Data-Loss-Ereignissen</a>. Es handelt sich hierbei nicht um eine Methodik, um individuelle Risikobewertungen durchzuführen. Vielmehr will das Framework Unternehmen in die Lage versetzen, IT-Risiken zu verstehen, zu analysieren und zu messen.</p>



<p>Zu den Komponenten des FAIR-Frameworks gehören:</p>



<ul class="wp-block-list">
<li>eine Taxonomie für IT-Risiken,</li>



<li>eine standardisierte Nomenklatur für Risiken,</li>



<li>eine Methode um Datenerfassungskriterien zu definieren,</li>



<li>Messskalen für Risikofaktoren,</li>



<li>eine Engine für Risikoberechnungen, sowie</li>



<li>ein Modell, um komplexe Risikoszenarien zu analysieren.</li>
</ul>



<p><strong>Gut zu wissen:</strong> Die quantitative Risk-Assessment-Ansatz von FAIR ist branchenübergreifend anwendbar.</p>



<h2 class="wp-block-heading">3. ISO/IEC 27001</h2>



<p><strong>Das ist es:</strong> Bei <strong><a href="https://www.iso.org/standard/27001" target="_blank" rel="noreferrer noopener">ISO/IEC 27001</a></strong> handelt es sich um einen internationalen Standard, der mit Leitlinien in Sachen IT-Security-Management unterstützt. Ursprünglich wurde er im Jahr 2005 gemeinschaftlich von der International Organization for Standardization (ISO) und der International Electrotechnical Commission (IEC) veröffentlicht – und wird seither sukzessive überarbeitet.</p>



<p><strong>Das kann es:</strong> ISO/IEC 27001 ist laut den Verantwortlichen ein Guide für Unternehmen jeder Größe und aus sämtlichen Branchen, um ein Information Security Management System (ISMS) aufzusetzen, zu implementieren, zu warten und fortlaufend zu verbessern.</p>



<p><strong>So funktioniert es:</strong> ISO/IEC 27001 fördert einen ganzheitlichen Cybersicherheitsansatz, der Menschen, Richtlinien und Technologie auf den Prüfstand stellt. Ein auf dieser Grundlage erstelltes ISMS ist laut ISO ein Tool für Risikomanagement, Cyberresilienz und Operational Excellence.</p>



<p><strong>Gut zu wissen:</strong> ISO/IEC-27001-konform zu sein bedeutet, einem weltweit eingesetzten Standard zu genügen und Datensicherheitsrisiken aktiv zu managen.</p>



<h2 class="wp-block-heading">4. NIST Risk Management Framework</h2>



<p><strong>Das ist es:</strong> Das <strong><a href="https://csrc.nist.gov/projects/risk-management/about-rmf" target="_blank" rel="noreferrer noopener">Risk Management Framework</a></strong> (RMF) wurde von der US-Behörde NIST (National Institute of Standards and Technology) entwickelt. Dieses Framework stellt einen umfassenden, wiederverwend- und messbaren, siebenstufigen Prozess in den Mittelpunkt, um IT- und Datenschutzrisiken zu managen. Dabei kommt eine ganze Reihe von NIST-eigenen Standards und Guidelines zur Anwendung, um die Implementierung von Risikomanagement-Initiativen zu unterstützen.</p>



<p><strong>Das kann es:</strong> Laut NIST realisiert das RMF einen Prozess, der die Risikomanagementaktivitäten in den Bereichen Sicherheit, Datenschutz und Supply Chain in den Lebenszyklus der Systementwicklung integriert. Dabei berücksichtigt der Ansatz Effektivität, Effizienz und Einschränkungen durch geltende Gesetze, Direktiven, Anordnungen, Richtlinien, Standards oder Vorschriften.</p>



<p><strong>So funktioniert es:</strong> Der siebenstufige Prozess des NIST RMF gliedert sich in.</p>



<ol class="wp-block-list">
<li>wesentliche Aktivitäten, um die Organisation auf den Umgang mit Sicherheits- und Datenschutzrisiken <strong>vorzubereiten</strong>.</li>



<li>Systeme und Daten, die verarbeitet, gespeichert und übertragen werden, auf der Grundlage einer Impact-Analyse <strong>kategorisieren</strong>.</li>



<li>eine Reihe von Kontrollmaßnahmen <strong>auswählen</strong>, um Systeme auf der Grundlage einer Risikobewertung zu schützen.</li>



<li>Kontrollmaßnahmen <strong>implementieren</strong> – und dokumentieren, wie das vonstattengeht.</li>



<li>Kontrollmaßnahmen überprüfen und <strong>bewerten</strong>, ob diese wie gewünscht funktionieren.</li>



<li>Systembetrieb auf Grundlage einer risikobasierten Entscheidung <strong>autorisieren</strong>.</li>



<li>Implementierung und Systemrisiken kontinuierlich <strong>überwachen</strong>.</li>
</ol>



<p><strong>Gut zu wissen:</strong> Das RMF bietet einen verfahrenstechnischen und geordneten Prozess, um Organisation dabei zu unterstützen, Security in ihre allgemeinen Risikomanagement-Prozesse einzubetten.</p>



<h2 class="wp-block-heading">5. OCTAVE</h2>



<p><strong>Das ist es:</strong> <strong>OCTAVE</strong> (<a href="https://insights.sei.cmu.edu/documents/1210/1999_005_001_16769.pdf" target="_blank" rel="noreferrer noopener">Operationally Critical Threat, Asset, and Vulnerability Evaluation</a> (PDF)) ist ein Framework, um Risiken im Bereich der Cybersicherheit zu identifizieren und zu managen. Es wurde vom CERT-Team der Carnegie Mellon University in den USA entwickelt.</p>



<p><strong>Das kann es:</strong> Dieses Risk-Assessment-Framework definiert eine umfassende Evaluierungsmethode. Diese ermöglicht Unternehmen nicht nur, missionskritische IT-Assets zu identifizieren, sondern auch die Bedrohungen, die mit diesen in Zusammenhang stehen und die Schwachstellen, die das erst ermöglichen.</p>



<p><strong>So funktioniert es:</strong> Laut den Verantwortlichen ermöglicht die Zusammenstellung von IT-Assets, -Bedrohungen und –<a href="https://www.csoonline.com/article/3495294/schwachstellen-managen-die-6-besten-vulnerability-management-tools.html" target="_blank">Schwachstellen</a> Unternehmen, zu durchdringen, welche Daten wirklich bedroht sind. Mit diesem Verständnis ausgestattet, können die Anwender eine Schutzstrategie entwickeln und implementieren, um diese nachhaltig zu schützen.</p>



<p><strong>Gut zu wissen:</strong> Das OCTAVE-Framework ist in zwei Versionen erhältlich.</p>



<ul class="wp-block-list">
<li>OCTAVE-S bietet eine vereinfachte Methodik, die auf kleinere Unternehmen mit flachen hierarchischen Strukturen ausgerichtet ist.</li>



<li>OCTAVE Allegro ist hingegen ein umfassenderes Framework, das sich in erster Linie für große Unternehmen oder solche mit komplexen Strukturen eignet.</li>
</ul>



<h2 class="wp-block-heading">6. TARA</h2>



<p><strong>Das ist es:</strong> <strong>TARA</strong> (<a href="https://www.mitre.org/news-insights/publication/threat-assessment-and-remediation-analysis-tara" target="_blank" rel="noreferrer noopener">Threat Assessment and Remediation Analysis</a>) stellt eine Engineering-Methodik dar, mit deren Hilfe, Sicherheitslücken identifiziert, bewertet und behoben werden können. Dieses Framework wurde von der Non-Profit-Organisation MITRE entwickelt.</p>



<p><strong>Das kann es:</strong> Das Framework ist Teil des MITRE-Systemportfolios, das darauf ausgerichtet ist, die Cybersicherheitshygiene sowie die Resilienz von IT-Systemen in einem möglichst frühen Stadium (innerhalb des Beschaffungsprozesses) zu adressieren.</p>



<p><strong>So funktioniert es:</strong> Das TARA-Framework nutzt einen Datenkatalog, um Angriffsvektoren zu identifizieren, die genutzt werden könnten, um Systemschwachstellen auszunutzen sowie potenzielle Gegenmaßnahmen einzuleiten.</p>



<p><strong>Gut zu wissen:</strong> TARA wurde ursprünglich im Jahr 2010 entwickelt und kam bereits in mehr als 30 Cyber-Risk-Assessments zum Einsatz. Dieses Framework eignet sich in besonderem Maße für Risikostudien, die sich auf Sicherheitsbedrohungen konzentrieren. (fm)</p>



<p><strong>Dieser Artikel ist <a href="https://www.csoonline.com/article/525128/it-risk-assessment-frameworks-real-world-experience.html" target="_blank">im Original</a> bei unserer Schwesterpublikation CSOonline.com erschienen.</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Democratizing Zero Trust with an expanded BeyondCorp Alliance]]></title>
<description><![CDATA[The need to quickly provide secure access for a newly remote workforce during the early days of COVID-19 drove many organizations to explore new technologies and start down a path towards a Zero Trust model. As time has passed, it’s become clear that remote work will be a defining characteristic ...]]></description>
<link>https://tsecurity.de/de/3662848/it-security-nachrichten/democratizing-zero-trust-with-an-expanded-beyondcorp-alliance/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3662848/it-security-nachrichten/democratizing-zero-trust-with-an-expanded-beyondcorp-alliance/</guid>
<pubDate>Sun, 12 Jul 2026 08:07:13 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div class="block-paragraph"><p>The need to quickly provide secure access for a newly remote workforce during the early days of COVID-19 drove many organizations to explore new technologies and start down a path towards a Zero Trust model. As time has passed, it’s become clear that remote work will be a defining characteristic of the new normal, and modernizing security by fully embracing zero trust models is an imperative, not an option. We need to work to further democratize this technology, accelerate and ease its adoption to help organizations stay secure, agile, and productive.</p><p>We’ve been working on Zero Trust for more than a decade at Google, and earlier this year, we introduced <a href="https://cloud.google.com/solutions/beyondcorp-remote-access">BeyondCorp Remote Access</a>, our cloud-based solution that helps make access to internal applications easier and more secure. We offer similar <a href="https://support.google.com/a/answer/9275380?hl=en" target="_blank">context-aware access controls</a> for apps in <a href="https://workspace.google.com/" target="_blank">Google Workspace</a> and <a href="https://cloud.google.com/identity">Cloud Identity</a>. </p><p><a href="https://cloud.google.com/blog/products/identity-security/simplifying-identity-and-access-management-of-your-employees-partners-and-customers">Last year</a>, we assembled a group of partners that share our Zero Trust vision and who are committed to help our joint customers make it a reality: the BeyondCorp Alliance. These partners are key to our effort to further promote and democratize this technology. They allow customers to leverage existing controls to make adoption easier while adding key functionality and intelligence that enable customers to make better access decisions. We’re now pleased to announce that <a href="https://www.citrix.com/" target="_blank">Citrix</a>, <a href="https://www.crowdstrike.com/" target="_blank">CrowdStrike</a>, <a href="https://www.jamf.com/" target="_blank">Jamf</a>, and <a href="https://www.tanium.com/" target="_blank">Tanium</a> are joining <a href="https://www.checkpoint.com/" target="_blank">Check Point</a>, <a href="https://www.lookout.com/news-and-press/press-releases/beyondcorp" target="_blank">Lookout</a>, <a href="https://researchcenter.paloaltonetworks.com/2019/04/beyondcorp/" target="_blank">Palo Alto Networks</a>, <a href="https://www.symantec.com/blogs/feature-stories/symantec-partners-google-cloud-improve-zero-trust-cloud-access" target="_blank">Symantec</a> (a division of Broadcom), and <a href="http://blogs.vmware.com/euc/2019/04/workspace-one-google-cloud.html" target="_blank">VMware</a> as BeyondCorp Alliance members.</p></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/BeyondCorp_Alliance.max-1000x1000.jpg" alt="BeyondCorp Alliance.jpg">
        
        
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph"><p>As Sunil Potti, VP and GM Google Cloud Security, puts it, BeyondCorp delivers world-class security for the reimagined workplace. Partners who share our vision are an essential part of how we help our customers modernize their security approaches in-place to deliver a better, safer normal.</p><p>Our BeyondCorp Alliance Partners add capabilities in the following areas:</p><p><b>Device Management</b>: Enterprise Mobility Management (EMM) vendors can provide device context and telemetry such as whether a device is managed or corporate-owned to aid in policy evaluation.</p><p><b>Endpoint Security</b>: Endpoint Detection and Response Vendors (EDR) or Mobile Threat Defense (MTD) vendors can provide device posture information, such as whether a device is compromised to aid in policy evaluation.</p><p><b>Gateways</b>: Infrastructure vendors can provide more secure access to hosted infrastructure (e.g., virtual desktops, etc.) via BeyondCorp. </p><p>Keep reading to learn more about updates to our existing BeyondCorp Alliance partnerships and new solutions with leading security partners that we are excited to announce today: </p><p><b>Check Point</b> SandBlast Mobile is a mobile threat defense solution that detects and stops attacks on iOS and Android devices before they start. Integration with the Google Admin console can be used to selectively prevent compromised devices from accessing applications and resources, helping to keep sensitive data secure. The integration is now available to customers in preview in the Google Admin console.</p><p><b>Citrix</b> and Google Cloud are extending our deep collaboration to include BeyondCorp. Google Cloud has always been one of the best places to run <a href="https://www.citrix.com/products/citrix-workspace/" target="_blank">Citrix Workspace</a>, and the first step, bringing together Citrix Workspace and BeyondCorp, is coming soon. It will allow customer applications, whether they are deployed on-premises, on GCP, or delivered as a service (SaaS), to be exposed through Citrix Workspace with BeyondCorp’s access controls and policy enforcement. Users get a single pane of glass for all of their applications, which can now be accessed from BYOD and non-corporate devices without the need for a VPN. We’re also exploring the sharing of endpoint signals and further extending policy enforcement to virtual desktops. For more information, check out the Citrix <a href="https://www.citrix.com/blogs/2020/10/13/deliver-workspace-security-and-zero-trust-with-citrix-and-google-cloud/" target="_blank">blog</a> on our joint zero trust security solutions.</p><p><b>CrowdStrike</b> will deliver real-time endpoint posture assessments from endpoints regardless of location, network, or user so that BeyondCorp adopters can prohibit access from untrusted or compromised hosts as part of conditional access policies, reducing risk for users and the organization. This integration is coming soon. To learn more about how CrowdStrike and Google Cloud are collaborating on Zero Trust, <a href="https://na.eventscloud.com/ereg/index.php?eventid=560023&amp;utm_campaign=fal_con&amp;utm_medium=dir&amp;utm_source=blog" target="_blank">register</a> for CrowdStrike’s Cybersecurity Conference <a href="https://www.crowdstrike.com/events/falcon/?utm_campaign=fal_con&amp;utm_medium=dir&amp;utm_source=blog" target="_blank">Fal.Con 2020</a>, taking place on October 15, 2020.</p><p><b>Jamf</b> is working to extend its device compliance capabilities for organizations leveraging Google Cloud and BeyondCorp. In the past, organizations have expressed concerns about unprotected Mac devices accessing cloud and on-premises resources. Now, through a unique Jamf preview, customers can ensure that only trusted users, from managed devices, using approved apps, are accessing company data. Read Jamf’s <a href="https://www.jamf.com/blog/jamf-and-google-announce-conditional-access-partnership-preview" target="_blank">blog</a> on our collaboration and <a href="mailto:google.ca@jamf.com">contact the Jamf team</a> to learn more about this preview.</p><p><b>Lookout</b> continuously assesses a smartphone, tablet or Chromebook’s risk level and provides it to Cloud Identity and BeyondCorp from the Lookout Security Graph. Device risk levels of “high, moderate  or low” are set based on the organization’s security policies. When Lookout detects a threat on a mobile device, the risk level is changed accordingly and delivered in real-time to Cloud Identity via API. This integration enables Google Workspace to block risky or non-compliant devices from accessing applications and data. This functionality is now available in preview via the Google Admin console. Learn more by reading Lookout’s <a href="https://blog.lookout.com/lookout-google-deliver-zero-trust-beyondcorp-vision-for-mobile" target="_blank">blog</a>.</p><p><b>Symantec</b> Endpoint Protection (SEP) and Symantec Endpoint Protection Mobile (SEP Mobile) report on the security posture of an organization’s traditional and mobile endpoints, including both managed and unmanaged devices. With the upcoming integration, customers can leverage Symantec’s endpoint signals such as indications of compromise, operating system configuration risks, app risks, anomalous network behavior, and more, to create more granular and customized access policies for Google Workspace, web apps, and Google Cloud infrastructure.</p><p><b>Tanium</b> and Google Cloud recently<a href="https://www.tanium.com/press-releases/tanium-and-google-cloud-join-forces-to-deliver-security-transformation-for-the-distributed-it-era/" target="_blank"> announced</a> a strategic partnership with the goal of delivering security transformation for the distributed IT era. As part of the BeyondCorp Alliance, Tanium will be providing device identity information through <a href="https://docs.tanium.com/endpoint_identity/endpoint_identity/userguide.html" target="_blank">Tanium Endpoint Identity</a>, which is available today. Tanium monitors and evaluates the health of endpoints in real-time, providing comprehensive visibility and control from a single platform no matter where the device is located. Through the combined solution, coming soon, organizations will be able to ensure that devices connecting to network resources and applications are authorized, secured, and up-to-date. To learn more about Tanium’s partnership with Google Cloud and BeyondCorp integration,<a href="https://converge.tanium.com/" target="_blank"> register to attend</a> their upcoming virtual user conference, Converge.</p><p><b>VMware</b> is working to bring Workspace ONE and Google Cloud's BeyondCorp solution together to keep devices under control and compliant with policies that protect corporate data. Workspace ONE will continually feed device compliance status information to Google Cloud’s context-aware access engine, allowing access to be revoked at any time if a device becomes non-compliant. This integration is coming soon.</p><p>To learn more about how you can take advantage of our joint capabilities to advance your own Zero Trust strategy, visit the BeyondCorp Alliance partner links above or <a href="mailto:beyondcorp.alliance@google.com">reach out to our team</a>. </p><p>Also be sure to check out our <a href="https://cloud.google.com/solutions/beyondcorp-remote-access">BeyondCorp product home</a>, browse BeyondCorp educational resources in our <a href="https://cloud.google.com/security/best-practices#section-3">Security Best Practices Center</a>, and view BeyondCorp use case videos in our <a href="https://cloud.google.com/security/showcase">Cloud Security Showcase</a>.</p></div>
<div class="block-related_article_tout">





<div class="uni-related-article-tout h-c-page">
  <section class="h-c-grid">
    <a href="https://cloud.google.com/blog/products/identity-security/keep-your-teams-working-safely-with-beyondcorp-remote-access/" data-analytics='{
                       "event": "page interaction",
                       "category": "article lead",
                       "action": "related article - inline",
                       "label": "article: {slug}"
                     }' class="uni-related-article-tout__wrapper h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
        h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3 uni-click-tracker">
      <div class="uni-related-article-tout__inner-wrapper">
        <p class="uni-related-article-tout__eyebrow h-c-eyebrow">Related Article</p>

        <div class="uni-related-article-tout__content-wrapper">
          <div class="uni-related-article-tout__image-wrapper">
            <div class="uni-related-article-tout__image"></div>
          </div>
          <div class="uni-related-article-tout__content">
            <h4 class="uni-related-article-tout__header h-has-bottom-margin">Keep your teams working safely with BeyondCorp Remote Access</h4>
            <p class="uni-related-article-tout__body">Enabling remote access to internal apps with a simpler and more secure approach without a remote-access VPN</p>
            <div class="cta module-cta h-c-copy  uni-related-article-tout__cta muted">
              <span class="nowrap">Read Article
                <svg class="icon h-c-icon" role="presentation">
                  <use xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="#mi-arrow-forward"></use>
                </svg>
              </span>
            </div>
          </div>
        </div>
      </div>
    </a>
  </section>
</div>

</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Wall Street is debating the AI buildout. Enterprises just answered: 86% say their GPUs run at half capacity or less]]></title>
<description><![CDATA[Enterprise companies are running AI agents ahead of the controls needed to manage them — and they deployed that way knowingly. That is the central finding from VentureBeat Research's June survey of 573 technical leaders at companies with 100 or more employees, fielded across five parallel surveys...]]></description>
<link>https://tsecurity.de/de/3660798/it-nachrichten/wall-street-is-debating-the-ai-buildout-enterprises-just-answered-86-say-their-gpus-run-at-half-capacity-or-less/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3660798/it-nachrichten/wall-street-is-debating-the-ai-buildout-enterprises-just-answered-86-say-their-gpus-run-at-half-capacity-or-less/</guid>
<pubDate>Fri, 10 Jul 2026 22:48:13 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Enterprise companies are running AI agents ahead of the controls needed to manage them — and they deployed that way knowingly. That is the central finding from VentureBeat Research's June survey of 573 technical leaders at companies with 100 or more employees, fielded across five parallel surveys of the agentic stack. </p><p>Enterprises are now retrofitting to catch up with their own standards, and they are budgeting for it: Roughly six in 10 enterprises plan to switch or add vendors in each of five control layers within the next 12 months, and roughly a third — depending on the layer — plan to move within the quarter, the research finds.</p><p>There are five main layers where enterprises are building: identity for agents (which agent is allowed to do what, under whose credentials); evaluation of agent output (whether the work is any good); cost telemetry (what each agent costs to run); the context layer (the business data and definitions agents draw on to answer); and the orchestration control plane (the software that coordinates multi-step agent work).</p><p>Enterprises are already paying the price for deploying agents ahead of adequate control functions. Fifty-four percent of companies <a href="https://venturebeat.com/security/shared-api-keys-expose-ai-agent-fleets-venturebeat-research">had an agent security incident or near-miss caught before harm</a> in the past 12 months. Twenty-seven percent exercise only reactive control of agent spend — they learn what an agent costs when the invoice arrives, with no per-agent budget or ceiling in place.</p><div></div><p>Here are the five findings that anchor the set — one finding per layer of the tech stack — and what the data suggests doing first in each.</p><h2>Expensive hardware is idle: 86% of GPU operators report utilization of 50% or less</h2><p>Eighty-six percent of enterprises that run their own GPUs report utilization of 50% or less. Wall Street has spent the quarter debating whether the AI buildout is overbuilt. This is buy-side measurement, from the enterprises doing the buying, and the research says the most expensive hardware in buildings of these enterprises runs at no more than half its capacity.</p><p>The measurement gap compounds it: A minority 44% rigorously track what their AI compute actually costs and returns. Everyone else is only estimating. And the enterprise shopping process continues regardless: 45% of these enterprises say the emerging compute option they are most likely to evaluate in the next 12 months is an AI-specialized cloud (CoreWeave, Lambda, Crusoe, Nebius). However, under 2% of these enterprises report using one of these neoclouds today. </p><p>Moreover, roughly one in three companies appears to be considering a hedge against Nvidia: Asked which emerging compute option they are most likely to evaluate in the next 12 months, 32% of enterprises named non-Nvidia accelerators (AWS Trainium, Google TPUs, AMD), while 28% named next-generation Nvidia GPUs. The data suggests that enterprises should measure the utilization and per-workload cost of the GPUs they already own before committing budget to new compute — whether that's an AI-specialized cloud contract, new accelerators, or more GPUs. </p><h2>Most deployed "agents" do single-prompt work: 71% say a quarter or fewer complete multi-step tasks on their own</h2><p>Seventy-one percent of enterprises say a quarter or fewer of their deployed "agents" can complete multi-step work on their own; the rest are single-prompt chatbots. Only 10% say true agents are the majority of what they run. To be sure, the respondents reported that they are in a position to know these things: 81% said they recommend or decide AI purchases at their companies.</p><p>That finding — that most agents are actually just chatbots in trenchcoats — lands amid adoption claims across the industry running well ahead of what enterprises are actually running. Gartner <a href="https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025">predicted</a> 40% of enterprise applications will be integrated with task-specific AI agents by the end of 2026, up from less than 5% in 2025. It also warned that the most common misconception is referring to these AI assistants as agents, a misunderstanding known as "agentwashing."</p><p>Meanwhile, Zapier's enterprise <a href="https://zapier.com/blog/ai-agents-survey/">survey</a> said 72% reported deploying or testing autonomous agents; and Writer's 2026 <a href="https://writer.com/blog/enterprise-ai-adoption-2026/">survey</a> has 97% of executives saying their company deployed AI agents in the past year. </p><p>Those surveys asked whether companies have deployed something called an AI agent, and companies said yes. Our survey asked the people running those deployments a harder question: Of the agents you have in production, how many can complete a multi-step task without a person driving each step? The gap matters for two practical reasons. First, the inflated adoption figures are the benchmark boards and vendors use to pressure technical leaders into moving faster — and this data says the real bar is far lower than the headlines suggest. Second, the label determines the bill: A single-prompt chatbot with a human reading every answer needs none of the identity, evaluation, and cost controls this report covers, while a true multi-step agent needs all of them. </p><h2>66% let agents push to production on automated evals alone — or are engineering toward it. 5% fully trust those evals</h2><p>Two-thirds of enterprises fall into one of two camps: 34% already allow an AI agent to push a code or system change to production based on automated evaluation results alone, with no human reviewing it, and another 33% are actively engineering their pipelines to allow that within the next 12 months. Only five percent fully trust the automated evaluations that would make that decision.</p><p>The distrust is earned. Half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year; a quarter watched it happen more than once. Asked to name the biggest weakness in their current evaluations, more enterprises chose “poor alignment with real-world outcomes” than any other answer — 29% of respondents.</p><p>And most of the checking happens before an agent ships, then stops. Once agents are live with real users, only 23% of enterprises run real-time quality checks on the answers those agents produce. Another 51% monitor system health only — uptime, request traces, and gateway logs — which tells them the agent is running, and nothing about whether its answers are right. The first move: Before removing human review from any workflow, test your evaluations against production outcomes rather than internal benchmarks, and instrument answer quality, not just uptime. </p><p>This finding is explored in more depth in <a href="https://venturebeat.com/orchestration/enterprise-ai-is-entering-an-evaluation-gap-agents-are-gaining-autonomy-faster-than-companies-can-verify-them">VentureBeat's related coverage of the evaluation gap</a>, which found that larger enterprises are moving faster toward zero-human deployment while also failing more often — and outlines a regression-testing framework built on production outcomes rather than internal benchmarks. </p><h2>69% run credential sharing somewhere in the agent fleet — and those companies get hit far more often</h2><p>Sixty-nine percent of companies allow agent credential sharing somewhere in their agent fleet during runtime – meaning multiple agents operating under one API key or service account. Those companies were far more likely to get hit: Organizations with credential sharing anywhere in the fleet experienced a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (9 of 22) where every agent has its own scoped identity. </p><p>The takeaway for enterprises is this: Give every agent its own scoped identity, starting with the agents that touch production systems.</p><h2>57% traced a confident, wrong agent answer to their own missing or inconsistent business context</h2><p>Fifty-seven percent of enterprises traced at least one confident, wrong agent answer in the past six months to missing or inconsistent business context: wrong metrics, stale definitions, absent documents. Most of them watched it happen more than once.</p><p>Most enterprise companies are fixing this, even though they’ve moved forward with agent deployment already: 25% already run a governed semantic layer, or one governed definition of the business that every AI reads from, in production. However, 34% are still building one, and 41% haven't started. The takeaway: Govern the definitions your agents answer from, metrics and entities first, before scaling the agents that depend on them.</p><h2>The quarter where agent technology “portability” became a priority</h2><p>One more shift is worth reporting with its limits stated plainly. In our spring orchestration survey wave, the top concern about provider-controlled orchestration was security and permissioning limits (32%). By June, vendor lock-in led at roughly a third, with security limits at 28%. </p><p>Those are two snapshots one quarter apart, and here’s one possible explanation for why portability became a top issue for enterprises. Our June survey went into market after a June 12 U.S. Commerce Department <a href="https://venturebeat.com/orchestration/enterprises-lost-claude-fable-5-for-a-few-weeks-new-data-shows-two-thirds-had-already-built-their-hedge">export order took Anthropic's Claude Fable 5 offline</a> for enterprises for roughly three weeks. Meanwhile, Chinese company Z.ai <a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">released GLM-5.2's open weights</a> under an MIT license on June 16 at roughly one-sixth of GPT-5.5's price; and Tencent's <a href="https://venturebeat.com/technology/tencents-apache-licensed-hy3-takes-on-glm-5-2-at-half-the-size-and-wins-everywhere-except-coding">Hy3 arrived</a> July 6 under Apache 2.0; and OpenAI <a href="https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov">previewed GPT-5.6</a> on June 26 to a small group of government-vetted partners, opening it broadly on July 9 after the government's review cleared. The open-weight releases in particular promise enterprises more control over their agents, and while we haven't established a causal link here, the timing is worth noting.</p><p>The posture data matches the mood: 51% now expect their primary control plane for enterprise agents to be hybrid — provider-native plus external orchestration — by the end of 2026, up from 34% in the spring survey wave. Enterprises reporting that they rely purely on provider-managed agent services fell from 12% to 7%.</p><h2>Five layers, no incumbents, 12 months</h2><p>The synthesis across all five surveys reveals a huge “buying” window. In each of the five control layers, 57% to 64% of enterprises plan to switch or add vendors within 12 months — 64% in infrastructure and in evaluations, 59% in agent security, 57% in retrieval and context — and 26% to 38%, depending on the layer, plan to move within a quarter. No layer has an established incumbent: The most common evaluation tooling is the model provider's built-in evals, tied with no dedicated tooling at all (17% each); 82% of respondents name provider-native or hyperscaler controls as their primary agent security layer; and provider-native retrieval leads the context technology layer (RAG, etc) as well. </p><p>Most enterprises are defaulting today to the built-in tools that ship with the big AI platforms they already use: Anthropic, OpenAI, Google, Microsoft, and AWS. That holds true across every one of these agentic technology layers: enterprises are looking to their primary cloud and model providers to supply the guardrails, evaluations, and retrieval solutions already bundled into those providers' offerings.</p><p>Those defaults are winning on convenience, and they're also what the coming spending decisions will test. The survey didn't ask which direction that money moves — toward the platforms' built-in tools or toward the specialists challenging them — which is exactly why every contract in these five layers is worth watching over the next four quarters.</p><p>The Q3 survey wave will measure whether the enterprises made good on these budget plans: whether their agents gained scoped identities, whether evaluations got tested against production outcomes, whether GPU utilization rose, and whether the semantic layers under construction shipped.</p><p><i>VentureBeat will release the full Q2 reports across all five VB Pulse trackers at </i><a href="https://luma.com/92nbdnnx?utm_source=LI&amp;utm_campaign=mmpost2"><i>VB Transform</i></a><i>, July 14–15 at Hotel Nia in Menlo Park, where we convene enterprise technical leaders building autonomous agents in production. </i></p><p><i>Disclosure: VentureBeat produces both this research and VB Transform</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them]]></title>
<description><![CDATA[Enterprise AI teams are giving agents more freedom at the same moment their confidence in automated testing is collapsing.Half of enterprises have deployed an AI agent or LLM feature that passed internal evaluations and yet still caused a customer-facing failure — one in four more than once — acc...]]></description>
<link>https://tsecurity.de/de/3660672/it-nachrichten/enterprise-ai-is-entering-an-evaluation-gap-agents-are-gaining-autonomy-faster-than-companies-can-verify-them/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3660672/it-nachrichten/enterprise-ai-is-entering-an-evaluation-gap-agents-are-gaining-autonomy-faster-than-companies-can-verify-them/</guid>
<pubDate>Fri, 10 Jul 2026 21:18:05 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Enterprise AI teams are giving agents more freedom at the same moment their confidence in automated testing is collapsing.</p><p>Half of enterprises have deployed an AI agent or LLM feature that passed internal evaluations and yet still caused a customer-facing failure — one in four more than once — according to the June 2026 VB Pulse survey of 157 qualified enterprise respondents at companies with 100 or more employees.</p><p>The sample is self-selected rather than a probability sample, so the findings should be read as directional, not precise.</p><p>But enterprises are not responding by slowing automation:<b> 66% of respondents already permit some production deployment without human review </b>or are building systems intended to do so within the next 12 months. Only 5% say they fully trust the automated evaluations that would make those release decisions.</p><p>That mismatch is the evaluation gap: the autonomy ceiling is rising faster than the assurance beneath it. </p><p>It also fits a broader thesis that will be explored at <a href="https://venturebeat.com/vbtransform2026">VB Transform 2026</a>: enterprises ship agents first, while the control layers around identity, evaluation, cost, context and orchestration are arriving later. The next year will be a retrofit cycle, with buyers shifting budget toward the systems that make agentic deployments governable and dependable.</p><h2>Why a passing evaluation is not a working agent</h2><p>Traditional software testing usually asks whether a defined input produces an expected output. Agent testing is harder because the system may choose its own sequence of steps, call tools, retrieve data, alter state and respond differently from one run to the next.</p><p>An agent can make several individually plausible decisions and still reach the wrong result. It may retrieve the correct account but update the wrong field. It may draft a valid refund request but send it without approval. It may call five tools successfully before a sixth step leaks sensitive information or leaves a workflow incomplete.</p><p>The survey shows enterprises already recognize this limitation. <b>The most common reason for distrusting automated evaluation is poor alignment with real-world outcomes, cited by 29% of respondents.</b> Bias or inconsistency follows at 21%, lack of explainability at 18%, and data leakage or privacy concerns at 17%.</p><p>That hierarchy matters. Enterprises are saying the score often does not predict what happens when a customer, employee or business process encounters the agent in production — not that automated scoring is too slow or expensive.</p><p><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf">NIST makes a similar point in its Generative AI Profile</a>: measurements gathered in controlled environments may not transfer cleanly to deployment because behavior changes with prompts, users, context and operating conditions. Its guidance calls for field testing, post-deployment monitoring and clear processes for escalating failures.</p><div></div><h2>Capability is not consistency</h2><p>A single successful run proves that an agent can complete a task. It does not prove that it will complete the task reliably.</p><p><a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents">Anthropic’s guidance on agent evaluation</a> distinguishes between measuring whether a system succeeds at least once across repeated attempts and whether it succeeds every time. That distinction is essential for customer-facing or operational workflows. A model that occasionally produces an excellent answer may still be unacceptable if the same task fails unpredictably on the next attempt.</p><p>Enterprise teams should therefore treat repeatability as a first-class metric. That means running the same scenario multiple times, varying phrasing and context, testing tool failures, and measuring whether the final business outcome remains correct even when the route changes.</p><p>The evaluation set also has to evolve. Every production incident should become a permanent regression test. Customer escalations, failed tool calls, incorrect approvals and data-handling mistakes should feed back into the pre-deployment suite rather than remaining isolated support cases.</p><h2>Autonomy should expand by risk, not by ambition</h2><p>The survey does not imply that every agent action should require a person. Human review cannot scale across millions of low-consequence decisions.</p><p>But zero-human operation should be earned by demonstrated reliability and bounded by the consequences of failure.</p><p>Low-risk actions such as drafting internal summaries or categorizing documents can tolerate broader autonomy. Financial transactions, customer communications, code deployment, access-control changes and data deletion need stricter thresholds, repeated consistency tests, policy checks, rollback mechanisms and clear human escalation paths.</p><p>The risk isn't evenly distributed by company size, either. Larger enterprises — those with 2,500 or more employees — are moving toward zero-human deployment fastest, at 70% versus 64% for smaller companies, and they're also shipping more agents that go on to fail a customer, at 54% versus 48%. </p><p>That is the warning for enterprise leaders. Removing the human from the loop does not remove uncertainty. Without stronger assurance, it converts uncertainty into an automated production decision.</p><p>The market will keep pushing toward greater autonomy because the economic incentive is real. The organizations best positioned won't be those that remove people fastest — they'll be the ones that treat repeatability and regression testing as seriously as deployment speed.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[IBM grows mainframe family with rack, frame models targeting AI, hybrid clouds]]></title>
<description><![CDATA[IBM is looking to expand the reach of its foundational mainframe portfolio by adding new single frame and rack mounted versions of its Z and LinuxONE systems.



The IBM z17 portfolio adds a single frame and rack mount versions that bring mainframe capabilities into smaller, customizable footprin...]]></description>
<link>https://tsecurity.de/de/3660589/it-security-nachrichten/ibm-grows-mainframe-family-with-rack-frame-models-targeting-ai-hybrid-clouds/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3660589/it-security-nachrichten/ibm-grows-mainframe-family-with-rack-frame-models-targeting-ai-hybrid-clouds/</guid>
<pubDate>Fri, 10 Jul 2026 20:23:19 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>IBM is looking to expand the reach of its foundational mainframe portfolio by adding new single frame and rack mounted versions of its Z and LinuxONE systems.</p>



<p>The <a href="https://www.ibm.com/docs/en/announcements/z17-single-frame-rack-mount-systems-expand-ai-security-operational-simplicity-enterprise-workloads" target="_blank" rel="nofollow">IBM z17 portfolio</a> adds a single frame and rack mount versions that bring mainframe capabilities into smaller, customizable footprints. The <a href="https://www.ibm.com/docs/en/announcements/linuxone-rockhopper-5-built-secured-ai-ready-enterprise-it" target="_blank" rel="nofollow">LinuxONE Rockhopper family</a> gets a single frame and rack mount models, plus a new Express rack mount offering, that target new and smaller clients, according to Tina Tarquinio, chief product officer, IBM Z &amp; LinuxONE.</p>



<p>Specifically, the new hardware includes:</p>



<ul class="wp-block-list">
<li>z17 single frame is a fully packaged box in an IBM rack with intelligent power distribution units, delivered as a complete enclosed unit ready to deploy at the edge or other strategically important customer sites.</li>



<li>z17 rack mount lets customers install IBM Z components directly into their own industry-standard rack, with built-in flexibility for co-location with other technologies.</li>



<li>LinuxONE Rockhopper 5 is a multi-drawer LinuxONE system for high-density workloads, with on-chip AI acceleration, confidential computing, and postquantum cryptography available in both single frame and rack mount configurations.</li>



<li>Rockhopper 5 rack mount and Express offerings deliver enterprise-grade Linux, confidential computing, and on-chip AI acceleration in a compact 18U configuration. Designed for organizations supporting a smaller set of workloads, the offering provides a cost-efficient entry point that can scale as business grows, while prioritizing security, resiliency, and performance.</li>
</ul>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/07/LinuxONE-5-Single-Frame.png?w=1024" alt="IBM LinuxONE 5 single frame system" class="wp-image-4193838" width="1024" height="768" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">IBM</p></div>



<p>The new IBM z17 and IBM LinuxONE 5 Rockhopper configurations support up to 82 cores and 18 TB of memory across two processor drawers, representing about a 20% increase in core count and 12% increase in memory capacity over current systems, IBM stated. Single processor capacity of an IBM z17 ME2 provides full speed IBM z/OS configurations including 10% greater throughput per core than IBM z16 A02 with some variation based on workload and configuration, according to Tarquinio.</p>



<p>Both systems feature a 5.5 GHz IBM Telum II processor and a built-in AI accelerator that IBM says will let customers run more than 450 billion inferencing operations in a day with one millisecond response time. In addition, the 32-core Spyre AI accelerator is designed to handle all manner of AI workloads.</p>



<p>The idea is to bring the core strengths of IBM Z to a broader range of deployment models while offering the security, resilience, and performance enterprises depend on, Tarquinio said. </p>



<p>“As always, we’re continuing to innovate to deliver more with less, including up to 20% more capacity than IBM z16 to help process transactions faster and support growing AI-driven workloads,” Tarquinio said.  “Even the newest and smallest member of the IBM z17 family delivers the performance, efficiency, and scalability organizations need as they balance growth ambitions with real-world resource constraints.”</p>



<p>The Linux-based system, Rockhopper 5 is for organizations that have moved past the evaluation question and are ready to consolidate a substantial portion of their x86 estate, said Marcel Mitran, IBM Fellow and CTO of IBM LinuxONE. </p>



<p>Rockhopper 5 is designed to bring a smaller physical footprint and a software licensing model that reflects actual workload boundaries rather than physical server counts, Mitran said.</p>



<p>The LinuxONE 5 Express is a preconfigured system designed to get organizations running on LinuxONE quickly, with a defined bill of materials and a predictable starting cost, on the same architecture that the largest enterprises in the world depend on, Mitran said.</p>



<p>“It is built for organizations that want to consolidate a modest x86 estate, evaluate LinuxONE for the first time, or deploy a specific workload such as digital assets, AI-infused transaction processing, or confidential computing, without committing to the footprint of the larger model,” Mitran said.</p>



<p>Some of the mainframes’ software features were also bulked up. For example, IBM said that Post Quantum Cryptography security is now standard on the z17 and LinuxONE Rockhopper 5 systems letting customers start to utilize cryptography to protect core resources for the future.</p>



<p>The idea is to help customers protect long-lived, mission-critical data while reducing the cost and complexity of future cryptographic migration, IBM stated. </p>



<p>In that vein, IBM said it was bringing Crypto Discovery &amp; Inventory, which lets security teams see what has been encrypted across the enterprise. In addition, IBM announced an Infrastructure Management for Z and LinuxONE package that would let customers administer, monitor, automate, and provision IBM Z and LinuxONE systems from a central location.</p>



<p>IBM said it wants to reduce operational complexity for customers by making automating day-to-day operations<strong> </strong>to ultimately lower administrative costs and concerns. With the new flexible form factors, IBM continues to target hybrid and AI infrastructure buildouts with the Big Iron. In the AI world, the z17 is being utilized for AI inferencing, transactions, training, and key security applications such as fraud detection and insurance claims.</p>



<p>“Enterprise infrastructure is entering a new phase. Organizations need platforms that can support AI-driven growth while navigating resource constraints, evolving business requirements, and increasingly complex hybrid environments,” Tarquinio said. “They are being asked to deploy new AI capabilities while learning new skills, controlling operational costs, and maximizing the value of existing applications and infrastructure.”</p>



<p>A recent <a href="https://www-api.ibm.com/adobe/assets/urn:aaid:aem:52bed780-53cf-4a1c-a73b-d373bd532e97/original/as/the-mainframe-advantage.pdf" target="_blank" rel="nofollow">IBM Institute study</a> on mainframe usage stated that embedding mainframe to support AI in executing transactions is not temporary: 75% of executives expect mainframe-based applications to remain central to digital transformation, and 60% say mainframe-based platforms are essential to enabling AI innovation.</p>



<p>”Mainframe-anchored systems of record are becoming systems of intelligent execution—not as general‑purpose AI platforms, but as environments where AI acts directly within transactions and in support of them,” the study reported.</p>



<p>Gartner wrote in its “<a href="https://www.ibm.com/forms/mkt-17256" target="_blank" rel="nofollow">The State of the IBM Mainframe in 2026</a>” report that IBM’s willingness to make significant investments ensure the mainframe modernizes to remain a vital and thriving component of enterprise IT.  </p>



<p>“Most mainframe customers are now prioritizing the reduction of technical debt and adopting platform innovations to future-proof their mainframe environments for the coming decade,” Gartner wrote.</p>



<p>The new z17 single frame and rack mount configurations, LinuxONE Rockhopper 5, and LinuxONE 5 Express will all be available August 12, 2026. IBM Infrastructure Management for IBM Z and IBM LinuxONE will be available August 14.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Google's TabFM skips per-dataset training and still predicts on tables it's never seen]]></title>
<description><![CDATA[The vast majority of business data is tabular — living in data warehouses, CRMs, and financial ledgers — yet building a reliable model from it still means training a new one from scratch for every dataset, then maintaining hyperparameter tuning loops, feature engineering, and retraining pipelines...]]></description>
<link>https://tsecurity.de/de/3660555/it-nachrichten/googles-tabfm-skips-per-dataset-training-and-still-predicts-on-tables-its-never-seen/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3660555/it-nachrichten/googles-tabfm-skips-per-dataset-training-and-still-predicts-on-tables-its-never-seen/</guid>
<pubDate>Fri, 10 Jul 2026 20:03:33 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The vast majority of business data is tabular — living in data warehouses, CRMs, and financial ledgers — yet building a reliable model from it still means training a new one from scratch for every dataset, then maintaining hyperparameter tuning loops, feature engineering, and retraining pipelines to fight data drift. Google Research is proposing a way around that: <a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/">a new foundation model called TabFM</a> that treats tabular prediction as an in-context learning problem instead.</p><p>It can generate predictions for a new, unseen table in a single forward pass. For enterprise developers and AI engineers, this reduces the time-to-production from weeks of pipeline engineering to a single API call.</p><h2>The challenge with traditional ML</h2><p>To extract reliable predictions from a gradient-boosted tree, data scientists must build and maintain complex data pipelines. They have to clean messy inputs, impute missing values, encode categorical variables into numerical formats, and engineer custom feature crosses.</p><p>Once the data is ready, they must run repetitive hyperparameter optimization loops, searching across learning rates, tree depths, subsampling ratios, and regularization grids to find the best configuration. </p><p>Once deployed, these traditional models "incur ongoing operational debt through data drift monitoring and retraining pipelines to stay accurate," Weihao Kong, Research Scientist at Google Research, told VentureBeat.</p><p>Meanwhile, the rest of the AI industry has moved on. Generative AI models for text and computer vision have seamlessly shifted to zero-shot inference, where a model can perform a completely new task simply by being prompted with context. </p><p>Large language models (LLMs) already excel at <a href="https://venturebeat.com/business/fine-tuning-vs-in-context-learning-new-research-guides-better-llm-customization-for-real-world-tasks">in-context learning</a>, so why can't we just feed tables into an off-the-shelf LLM?</p><p>Because LLMs are trained on natural language rather than structured data, they struggle to process tables directly. First, their context limits are exhausted quickly by medium-sized tables containing just a few thousand rows and hundreds of columns. Second, LLMs suffer from tokenization inefficiency, awkwardly splitting numerical values and destroying mathematical precision. Finally, they suffer from structural blindness. When a 2D table is serialized as a 1D text string, LLMs lose track of which value belongs to which row and column as the table grows. </p><p>"That's why, today, it is far more effective to use an LLM to write the code that handles feature engineering and calls XGBoost than to ask the LLM to read the table itself," Kong said.</p><h2>What is TabFM?</h2><p>To run inference with TabFM, you do not update any model weights. Instead, you take your historical examples (the training rows with their known labels) and your target rows (the new data you want to predict) and pass them to the model as a single, unified prompt. The model learns to interpret the relationships between columns and rows directly from this context at runtime.</p><p>For example, consider an enterprise analyst trying to predict customer churn. Instead of building a bespoke data pipeline and training an XGBoost model, they can simply pass a sample of historical user session data alongside a new, active session into TabFM. In one forward pass, the model returns an instant churn probability. </p><p>TabFM overcomes the limitations of LLMs by treating the data as a grid, preserving its structural integrity without forcing it into a single-dimensional text string.</p><p>To effectively process diverse tabular structures while enabling scalable zero-shot prediction, TabFM synthesizes the strengths of earlier experimental architectures, TabPFN and TabICL. <a href="https://github.com/PriorLabs/tabpfn">TabPFN</a>, developed by Prior Labs, first proved that a transformer architecture could perform zero-shot classification on small tables, though it struggled to scale computationally to larger datasets. </p><p>Later, <a href="https://dl.acm.org/doi/10.5555/3780338.3782366">TabICL</a>, developed by France's National Research Institute for Digital Science and Technology, addressed this bottleneck by introducing row compression, allowing in-context learning to efficiently process much larger tables. </p><p>TabFM combines TabPFN's deep feature contextualization with TabICL's efficient compression into a novel hybrid design built on three key mechanisms:</p><p><b>1. Alternating row and column attention:</b> The raw table is first processed through a multilayer attention module that alternates across both columns (features) and rows (examples). By continuously attending across these two dimensions, the model natively captures complex feature interactions. This deep contextualization does the heavy lifting that would usually require tedious manual feature crafting by data scientists.</p><p><b>2. Row compression:</b> Following this contextualization, the cross-attended information for each row is compressed into a single, dense vector representation. TabICL pioneered this by using CLS tokens to compress a row's rich information into one vector, "in contrast to TabPFN v2, v2.5, and v2.6, which attend over the full cell grid throughout the network," Kong explained. This drastically shrinks the computational footprint.</p><p><b>3. In-context learning (ICL):</b> A causal Transformer then operates on this sequence of compressed embeddings. This Transformer model uses the attention mechanism of TabICL to attend over these dense row vectors, drastically reducing the computation cost and allowing the model to process large datasets efficiently.</p><p>A major selling point of TabFM is its pretraining recipe. The model was trained entirely on hundreds of millions of synthetic datasets. These datasets were dynamically generated using structural causal models (SCMs) that incorporate a wide variety of random functions. By training exclusively on synthetic SCMs, TabFM learned the fundamental mathematical priors of how tabular features interact without ingesting real-world, confidential CSV files.</p><h2>TabFM in action</h2><p>To test the model's capabilities, Google researchers benchmarked TabFM on TabArena, a comprehensive evaluation suite spanning 51 diverse tabular datasets across 38 classification and 13 regression tasks.</p><p>On these public benchmarks, TabFM's zero-shot predictions already match or beat heavily tuned supervised baselines. However, Google is careful to note that this does not automatically mean TabFM will universally dethrone bespoke, hyper-optimized production models on every enterprise workload.</p><p>"Instead of replacing hyper-optimized production models, the true practical business value it unlocks for lean engineering teams is velocity," Kong said. "It allows data analysts and backend engineers to instantly spin up high-quality baseline models without a dedicated data science team managing a complex lifecycle."</p><p>For advanced practitioners looking to squeeze out maximum accuracy, the research team also introduced a "TabFM-Ensemble" configuration. By running the model through 32 distinct variations and blending the results, TabFM pushes the performance even further. </p><h2>Getting started, trade-offs, and the cloud future</h2><p>The shift to in-context learning for tables introduces a new economic trade-off that engineering teams must consider. </p><p>With traditional algorithms, training is slow and expensive, but inference is lightning-fast and cheap. TabFM flips this dynamic. While training time drops to zero, inference becomes significantly heavier. Because the model must process the entire historical dataset as context during every single prediction, it requires more compute and memory at runtime. </p><p>In this new paradigm, "traditional machine learning training becomes the 'prefill' phase (KV caching) in the context window," Kong said. While this prefill cost is steep, it is paid only once per table, and the cache is reused across subsequent queries. "The catch is prediction latency, which no amount of caching removes," Kong added. Every new prediction requires a pass through a large transformer. "Any production API requiring single-digit-millisecond response times cannot tolerate TabFM's forward-pass overhead."</p><p>For developers looking to evaluate the model today, the barrier to entry is low. Google designed TabFM as a drop-in replacement for traditional ML workflows, offering a scikit-learn compatible API (TabFMClassifier and TabFMRegressor). It natively handles mixed numerical and categorical columns, works directly with pandas DataFrames, and requires no manual ordinal encoders or numerical scalers. The library supports both JAX and PyTorch backends.</p><p>However, enterprise teams need to be aware of current limitations and licensing restrictions. The model architecture has a hard limit of 10 output classes for classification tasks, and it is optimized for tables with up to 500 features. More importantly, while Google released the <a href="https://github.com/google-research/tabfm">underlying codebase</a> under the permissive Apache 2.0 license, the pre-trained model weights are published on <a href="https://huggingface.co/google/tabfm-1.0.0-pytorch">Hugging Face</a> under a strict tabfm-non-commercial-v1.0 license. Developers can evaluate the model internally, but it cannot be deployed in commercial products yet.</p><p>Looking ahead, Google is addressing the commercial deployment friction through its cloud ecosystem. TabFM is being integrated directly into Google BigQuery, allowing analysts to run zero-shot predictions natively via an “AI.PREDICT” command. By putting foundation model inference right next to the data warehouse, TabFM could soon make complex tabular machine learning as accessible as a basic database query.</p><p>In practice, TabFM shines in rapid prototyping, high data drift environments, and small to medium-sized datasets under 100,000 rows. Conversely, teams should stick to traditional models for strict, ultra-low latency APIs, or massive tables exceeding one million rows, which currently require aggressive row sampling that degrades the foundation model's competitive advantage.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Pick your Python accelerator]]></title>
<description><![CDATA[Faster Python has stopped becoming a pipe dream, and is now a major topic for its development. Sometimes that comes by way of new syntax (lazy imports), sometimes by JIT compilation, and sometimes by generating C code from Python. Sometimes, it’s also by way of a whole new programming language (M...]]></description>
<link>https://tsecurity.de/de/3659232/ai-nachrichten/pick-your-python-accelerator/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3659232/ai-nachrichten/pick-your-python-accelerator/</guid>
<pubDate>Fri, 10 Jul 2026 11:18:39 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Faster Python has stopped becoming a pipe dream, and is now a major topic for its development. Sometimes that comes by way of new syntax (lazy imports), sometimes by JIT compilation, and sometimes by generating <a href="https://www.infoworld.com/article/2261151/why-the-c-programming-language-still-rules.html" data-type="link" data-id="https://www.infoworld.com/article/2261151/why-the-c-programming-language-still-rules.html">C code</a> from Python. Sometimes, it’s also by way of a whole new programming language (Mojo) that’s intended to be a powerful Python companion.</p>



<h2 class="wp-block-heading">Top picks for Python readers on InfoWorld</h2>



<p><a href="https://www.infoworld.com/article/4145854/speed-boost-your-python-programs-with-new-lazy-imports.html" data-type="link" data-id="https://www.infoworld.com/article/4145854/speed-boost-your-python-programs-with-new-lazy-imports.html">Speed boost your Python programs with new lazy imports</a><br>With lazy imports in Python 3.15, the evaluation of imports can be delayed until they’re actually used, instead of when your Python program declares them. Best of all, you don’t need to rewrite everything to use this feature.</p>



<p><a href="https://www.infoworld.com/article/4117428/which-python-runtime-does-jit-better-cpython-or-pypy.html" data-type="link" data-id="https://www.infoworld.com/article/4117428/which-python-runtime-does-jit-better-cpython-or-pypy.html">CPython vs. PyPy: Which Python runtime has the better JIT?</a><br>Conventional wisdom tells us that PyPy’s built-from-scratch and time-tested JIT should beat CPython’s own new native JIT. Conventional wisdom isn’t always right.</p>



<p><a href="https://www.infoworld.com/article/4173158/first-look-mojo-1-0-mixes-python-and-rust.html" data-type="link" data-id="https://www.infoworld.com/article/4173158/first-look-mojo-1-0-mixes-python-and-rust.html">First look: Mojo 1.0 mixes Python and Rust</a><br>Is Mojo likely to outmuscle Python anytime soon? Probably not, but so far it’s shaping up to be as speedy as Rust without so much syntactical overhead.</p>



<p><a href="https://www.infoworld.com/article/4101101/pythoc-a-new-way-to-generate-c-code-from-python.html" data-type="link" data-id="https://www.infoworld.com/article/4101101/pythoc-a-new-way-to-generate-c-code-from-python.html">PythoC: A new way to generate C code from Python</a><br>The traditional way to use Python to generate C code is Cython. But PythoC offers a far more streamlined experience for those who want to hitch C’s speed to Python’s convenience.</p>



<h2 class="wp-block-heading">More good reads and Python updates elsewhere</h2>



<p><a href="https://peps.python.org/pep-0836" data-type="link" data-id="https://peps.python.org/pep-0836">PEP 836 – JIT go brrr: The path to a supported JIT compiler for CPython</a><br>The Python Software Foundation has delivered a roadmap for moving Python’s experimental JIT compiler towards a full-blown, supported, enabled-by-default part of Python’s future. But the road ahead could be bumpy. </p>



<p><a href="https://pyrefly.org/blog/too-many-type-checkers" data-type="link" data-id="https://pyrefly.org/blog/too-many-type-checkers">Are you really expected to run five type-checkers now?</a><br>Well, are you? (Spoiler: not really!) Which of the big five type checkers for Python should you use? Even if you’re already committed to one type checker, this article is well worth reading. The context for why so many exist is useful.</p>



<p><a href="https://www.youtube.com/watch?v=Y-ri74ZfGdo" data-type="link" data-id="https://www.youtube.com/watch?v=Y-ri74ZfGdo">Making Python faster with free threading and Mypyc</a><br>Mypyc is an underrated way to convert Python to C, and it’s now compatible with Python’s free-threaded build. Combining the two can unleash truly hair-raising speedups.</p>



<p><a href="https://karpathy.github.io/2026/02/12/microgpt" data-type="link" data-id="https://karpathy.github.io/2026/02/12/microgpt">MicroGPT: A GPT in 200 lines of pure Python</a><br>You won’t get blazing GPU-powered performance, but you’ll get a hands-on under-the-hood example of how, exactly, one can create a GPT-2-esque neural network architecture. Try it out on your favorite public domain text!</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Trillion Minutes of Sensor Data]]></title>
<description><![CDATA[SensorFM, a wearable health foundation model from Google Research, Google DeepMind, and university collaborators. We walk through its ViT-1D masked-autoencoder backbone, pretrained on more than one trillion minutes of unlabeled sensor signals from 5,000,000 consented participants. We examine the ...]]></description>
<link>https://tsecurity.de/de/3659189/ai-nachrichten/google-research-introduces-sensorfm-a-wearable-health-foundation-model-pretrained-on-one-trillion-minutes-of-sensor-data/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3659189/ai-nachrichten/google-research-introduces-sensorfm-a-wearable-health-foundation-model-pretrained-on-one-trillion-minutes-of-sensor-data/</guid>
<pubDate>Fri, 10 Jul 2026 11:03:13 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>SensorFM, a wearable health foundation model from Google Research, Google DeepMind, and university collaborators. We walk through its ViT-1D masked-autoencoder backbone, pretrained on more than one trillion minutes of unlabeled sensor signals from 5,000,000 consented participants. We examine the co-scaling results across four model sizes and four data volumes, including the case where capacity outruns data. We show how frozen embeddings plus a PCA-50 linear probe beat feature-engineered baselines on 34 of 35 tasks. We also review the agentic classroom that searched 30,516 prediction heads, and the clinician evaluation grounding a Personal Health Agent.</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/10/google-research-introduces-sensorfm-a-wearable-health-foundation-model-pretrained-on-one-trillion-minutes-of-sensor-data/">Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Trillion Minutes of Sensor Data</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[How to teach SRE AI agents to fail safely and earn your team’s trust]]></title>
<description><![CDATA[Site reliability engineering is entering a new phase. As incidents become faster-moving, more data-rich and more complex, SRE teams are exploring agentic AI to help with alert triage, root cause analysis, runbook execution and mitigation planning. But in production, the question is not whether an...]]></description>
<link>https://tsecurity.de/de/3659188/ai-nachrichten/how-to-teach-sre-ai-agents-to-fail-safely-and-earn-your-teams-trust/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3659188/ai-nachrichten/how-to-teach-sre-ai-agents-to-fail-safely-and-earn-your-teams-trust/</guid>
<pubDate>Fri, 10 Jul 2026 11:03:12 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p><a href="https://www.infoworld.com/article/2257232/what-is-an-sre-the-vital-role-of-the-site-reliability-engineer.html">Site reliability engineering</a> is entering a new phase. As incidents become faster-moving, more data-rich and more complex, SRE teams are exploring agentic AI to help with alert triage, root cause analysis, runbook execution and mitigation planning. But in production, the question is not whether an agent can act; it is whether people can trust it to act safely, consistently and transparently when the system is under stress.</p>



<p>This blog argues that trust is an engineering outcome, not a marketing promise. Trustworthy agentic SRE systems are built on a foundation of grounded telemetry, explicit safety boundaries, progressive autonomy, auditability and evaluation against real incidents.</p>



<h2 class="wp-block-heading"><a></a>Why trust matters</h2>



<p>Traditional automation works well when the world is predictable. SRE work is different because incidents are messy, partial and time-sensitive, with ambiguous symptoms, shifting dependencies and business context that rarely fits into a neat playbook. A fluent AI agent that lacks system context can sound convincing while still making dangerous recommendations.</p>



<p>Trust in SRE is earned during failure, not during demos. That means the system must prove it can help during noisy alerts, failed deploys, partial outages and conflicting telemetry, while staying bounded enough that one mistake does not become a major incident. Google’s AI-in-SRE work makes the same point through its emphasis on strict guardrails, progressive authorization and deterministic actuation controls.</p>



<h2 class="wp-block-heading"><a></a>Trust pillars</h2>



<p>A practical trust model for agentic SRE can be organized into five pillars.</p>



<figure class="wp-block-table"><div class="overflow-table-wrapper"><table class="has-fixed-layout"><tbody><tr><td><strong>Pillar</strong></td><td><strong>What it means</strong></td><td><strong>Why it matters</strong></td></tr><tr><td>Grounded observability</td><td>The agent reasons over correlated metrics, logs, traces, changes, topology and incident history.</td><td>SRE decisions often include business context that the agent does not fully see.</td></tr><tr><td>Clear guardrails</td><td>Permissions, allowlists, approval gates, rollback paths and rate limits constrain action.</td><td>Constraints make autonomy usable in production.</td></tr><tr><td>Human-in-the-loop design</td><td>Humans approve or supervise higher-risk actions.</td><td>SRE decisions often include business context that the agent does not fully see .</td></tr><tr><td>Explainability</td><td>The agent shows evidence, hypotheses, confidence and rationale.</td><td>Engineers need to inspect and challenge recommendations.</td></tr><tr><td>Real incident evaluation</td><td>The agent is scored against historical or replayed incidents.</td><td>Trust comes from measured performance, not benchmark theater.</td></tr></tbody></table> </div></figure>



<p>Google’s SRE autonomy model reflects the same progression: From assisted monitoring and investigation to partial autonomy with human approval to higher autonomy only after sustained success and safety proof.</p>



<h2 class="wp-block-heading"><a></a>Architecture pattern</h2>



<p>A trustworthy agentic SRE system should separate reasoning from actuation. The agent can investigate, summarize, propose and even stage a plan, but the actual execution path should pass through a deterministic safety layer that validates permissions, risk, current production state and blast radius before any change is made.</p>



<p>A strong pattern looks like this:</p>



<ol start="1" class="wp-block-list">
<li>Alert arrives from monitoring or incident tooling.</li>



<li>Agent gathers context from telemetry, deploy history, ownership and prior incidents.</li>



<li>Agent produces a ranked hypothesis and a candidate remediation plan.</li>



<li>Safety layer checks policy, risk score, current incident state and dry-run outcome.</li>



<li>Human approves low-confidence or high-risk actions.</li>



<li>Actuation layer executes only pre-approved, bounded changes.</li>



<li>The system observes post-action effects and either confirms success or falls back.</li>
</ol>



<p>Google’s description of AI operator and its mitigation safety verification layer is a useful reference point here: Investigation is not the same as actuation and the two should not share the same trust boundary. That separation reduces blast radius and keeps the system interruptible.</p>



<p>To try out Agentic SRE, StackGen has a <a href="https://app.stackgen.com/">community edition</a> where you can see the capabilities of agentic SRE by connecting your Grafana or Datadog.</p>



<h2 class="wp-block-heading"><a></a>Guardrails that work</h2>



<p>The most effective guardrails are boring in the best possible way. They include least-privilege identity, strict rate limits, dry-run support, explicit approval workflows, action allowlists and hard stop mechanisms for runaway loops. Check out the detailed guide on <a href="https://www.csoonline.com/article/4183666/what-sre-teams-need-before-they-trust-ai-agents.html">how SRE trusts AI agents</a>. AWS describes trust in autonomous systems in the same terms: Identity, runtime guardrails, observability and policy enforcement are the backbone of safe autonomy.</p>



<p>For SRE agents, a few guardrails are especially important:</p>



<ul class="wp-block-list">
<li><strong>Least privilege identity</strong> so the agent only has access to the systems it truly needs.</li>



<li><strong>Dry-run or simulation mode</strong> so the likely outcome is known before production state changes.</li>



<li><strong>Circuit breakers and loop detection</strong> to stop repeated or runaway tool calls.</li>



<li><strong>Action tiers</strong> so low-risk tasks can be automated while high-risk tasks require approval.</li>



<li><strong>Red-button controls</strong> so humans can immediately revoke autonomy during a bad incident.</li>
</ul>



<p>These controls are not signs of immaturity. They are what make autonomy acceptable in high-stakes environments.</p>



<h2 class="wp-block-heading"><a></a>Observability for agents</h2>



<p>Observability is not just for services; it is for the agent itself. If the agent’s reasoning, tool usage and outcomes are not observable, then debugging it during an incident becomes guesswork. Google explicitly emphasizes exposing reasoning traces and execution traces so that autonomous decisions remain auditable and debuggable.</p>



<p>A good agent observability stack should capture:</p>



<ul class="wp-block-list">
<li>Inputs and retrieved context.</li>



<li>Tool calls, parameters and results.</li>



<li>Intermediate hypotheses.</li>



<li>Confidence and uncertainty.</li>



<li>Approvals, denials and overrides.</li>



<li>Final action and outcome.</li>



<li>Post-action verification signals.</li>
</ul>



<p>This creates the operational memory needed to understand whether the agent helped, harmed or merely added noise. It also supports post-incident review and future training data generation.</p>



<h2 class="wp-block-heading"><a></a>Human in the loop</h2>



<p>Human-in-the-loop does not mean the agent is weak; it means the system is designed around responsibility. SREs still own the incident, the rollback, the customer impact and the final decision when context is incomplete. The agent should reduce toil and improve speed, not create a false sense of safety.</p>



<p>The best human-in-the-loop model is proportional. Low-risk tasks like summarizing incidents or collecting dashboards can be automated. Medium-risk actions like restarting a worker can require lightweight approval. High-risk actions like draining core capacity or disabling a major dependency should remain human-controlled. This progressive model lets trust grow gradually rather than forcing a dangerous leap to full autonomy.</p>



<h2 class="wp-block-heading"><a></a>Evaluation strategy</h2>



<p>If you only test an agent on toy benchmarks, you will get toy reliability. Real SRE evaluation should replay historical incidents and score whether the agent identified the right signals, chose the right hypothesis and recommended safe remediation under realistic conditions. Google’s approach uses continuous evaluation pipelines, human-verified gold data and nightly evals against real incident trajectories to measure readiness for autonomous action.</p>



<p>A practical evaluation program should include:</p>



<ul class="wp-block-list">
<li>Historical incident replay.</li>



<li>Golden-path and failure-path comparisons.</li>



<li>Tool misuse tests.</li>



<li>Prompt injection and adversarial input tests.</li>



<li>Loop and retry stress tests.</li>



<li>Human review of edge cases.</li>



<li>Regression tracking across model and policy changes.</li>
</ul>



<p>The key metric is not “did the model sound right?” It is “did the system shorten time to mitigation, reduce toil and avoid new operational risk?”.</p>



<h2 class="wp-block-heading"><a></a>Failure modes</h2>



<p>Agentic SRE systems fail in ways that classic software often does not. They can hallucinate a root cause, misread telemetry, over-trust stale context, loop on a broken action or optimize the wrong objective while sounding confident. In a high-stakes environment, this is more dangerous than a simple bug because the system can act before humans realize it is wrong.</p>



<p>The main failure modes to design against are:</p>



<ul class="wp-block-list">
<li><strong>Confident incompleteness</strong>, where the agent lacks key context but still gives a decisive answer.</li>



<li><strong>Runaway loops</strong>, where tool calls repeat and consume time or budget.</li>



<li><strong>Unsafe actuation</strong>, where a valid-looking action is harmful in the current operational state.</li>



<li><strong>Workflow drift</strong>, where the agent bypasses established incident processes.</li>



<li><strong>Hidden fragility</strong>, where speed increases but accountability decreases.</li>
</ul>



<p>Good architecture assumes failure will happen and makes sure the system fails safely, visibly and reversibly.</p>



<p>If you need a more detailed guide to keep points while evaluating AI SRE tools, then check this <a href="https://stackgen.com/blog/ai-sre-tools-buyers-guide-2026">buyer’s guide</a> by one of the senior leaders.</p>



<h2 class="wp-block-heading"><a></a>Operating model</h2>



<p>The healthiest way to deploy agentic SRE is to treat it as a bounded operational partner. Start with read-only use cases like alert enrichment, incident summarization and investigation assistance. Then move to recommendation-only workflows, then to low-risk automation and only later to tightly scoped autonomous mitigation.</p>



<p>That staged rollout should be paired with policy, ownership and incident review discipline. Every agent action should map back to a responsible team, a bounded capability and a visible audit trail. This is how the system earns confidence from engineers, security teams and leadership at the same time.</p>



<h2 class="wp-block-heading"><a></a>Conclusion</h2>



<p>Trustworthy agentic systems for SRE are built, not assumed. The winning formula is grounded telemetry, explicit guardrails, human oversight, explainable reasoning and evaluation against the messy reality of production incidents. When those pieces are in place, AI becomes a reliability multiplier rather than another source of operational risk.</p>



<p>The real goal is not a fully autonomous agent that never makes mistakes. The real goal is an agentic system that stays safe when it does make mistakes, recovers cleanly and keeps SRE teams in control when it matters most.</p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><strong><a href="https://www.infoworld.com/expert-contributor-network/">Want to join?</a></strong></p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[앤트로픽, 클로드 AI의 블랙홀 속을 들여다보다]]></title>
<description><![CDATA[앤트로픽이 자사 AI 모델이 문제를 해결하는 내부 과정을 보다 깊이 들여다볼 수 있는 새로운 분석 기법을 공개했다. 모델이 특정 방식으로 판단하고 행동하는 이유를 파악할 수 있게 되면서, 기업의 AI 평가와 구매 기준에도 적지 않은 영향을 미칠 것으로 전망된다.



앤트로픽은 최근 ‘J-스페이스(J-space)’라고 이름 붙인 새로운 내부 표현 공간을 발견했다고 밝혔다.



앤트로픽은 공식 블로그를 통해 “클로드(Claude)는 수많은 내부 처리 과정 가운데 특별한 역할을 수행하는 소수의 신경 패턴 집합을 스스로 형성했다”...]]></description>
<link>https://tsecurity.de/de/3658787/it-nachrichten/ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3658787/it-nachrichten/ai/</guid>
<pubDate>Fri, 10 Jul 2026 07:18:20 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>앤트로픽이 자사 AI 모델이 문제를 해결하는 내부 과정을 보다 깊이 들여다볼 수 있는 새로운 분석 기법을 공개했다. 모델이 특정 방식으로 판단하고 행동하는 이유를 파악할 수 있게 되면서, 기업의 AI 평가와 구매 기준에도 적지 않은 영향을 미칠 것으로 전망된다.</p>



<p>앤트로픽은 최근 ‘J-스페이스(J-space)’라고 이름 붙인 새로운 내부 표현 공간을 발견했다고 밝혔다.</p>



<p>앤트로픽은 <a href="https://www.anthropic.com/research/global-workspace" target="_blank" rel="nofollow">공식 블로그</a>를 통해 “클로드(Claude)는 수많은 내부 처리 과정 가운데 특별한 역할을 수행하는 소수의 신경 패턴 집합을 스스로 형성했다”라며 “이 패턴 집합을 발견하는 데 야코비안(Jacobian)이라는 수학 개념을 활용한 기법을 사용했기 때문에 이를 ‘J-스페이스’라고 명명했다”고 설명했다. 회사는 ‘야코비안 렌즈(Jacobian Lens·J-렌즈)’라는 분석 기법을 통해 J-스페이스 내부를 관찰한다.</p>



<p>앤트로픽은 “J-스페이스의 각 패턴은 특정 단어와 연결돼 있다”라며 “하지만 특정 패턴이 활성화됐다고 해서 모델이 실제로 그 단어를 출력하려는 것은 아니다. 단지 해당 단어가 모델의 ‘생각 속’에 있다는 의미”라고 밝혔다.</p>



<p>이어 “언어 모델의 ‘스크래치패드(scratchpad)’나 ‘생각의 연쇄(Chain of Thought)’처럼 추론 과정에서 스스로 작성하는 텍스트와 J-스페이스는 다르다”라며 “J-스페이스는 모델 내부 신경망 활성화 영역에서 조용히 작동하기 때문에 모델은 개념을 텍스트로 적지 않고도 내부적으로 사고할 수 있다”고 설명했다.</p>



<p>이번 연구는 앤트로픽이 2024년 공개했던 내부 스크래치패드보다 훨씬 발전한 개념이다. 당시에는 모델이 답변이나 행동을 준비하면서 무엇을 고려하는지를 보여주는 수준이었다. 반면 이번 기술은 모델의 훨씬 깊은 내부 사고 과정을 분석할 수 있어 AI 시스템을 평가하고 구매하는 방식 자체를 바꿀 가능성을 제시한다.</p>



<p><a href="https://transformer-circuits.pub/2026/workspace/index.html" target="_blank" rel="nofollow">논문</a>에 소개된 사례도 이러한 가능성을 보여준다. 일부 모델은 테스트 과정에서 부적절한 행동을 하지 않아 안전한 모델처럼 보였다. 그러나 J-스페이스를 분석한 결과, 해당 모델은 자신이 테스트를 받고 있다는 사실을 인지하고 있었으며, 그 때문에 문제 행동을 의도적으로 하지 않았을 가능성이 드러났다. 이는 아이들이 누군가 자신을 지켜보고 있다는 사실을 알 때 행동을 달리하는 것과 유사하다고 연구진은 설명했다.</p>



<p>AI 에이전트 기업 제니티(Zenity)의 AI 표준·거버넌스 총괄인 <a href="https://zenity.io/authors/rock-lambros" target="_blank" rel="nofollow">록 램브로스</a>(Rock Lambros)는 “앤트로픽은 모델이 테스트를 받고 있다는 사실을 인식하거나, 좋은 결과를 얻기 위해 행동을 꾸미거나, 프롬프트 인젝션을 탐지하거나, 아직 실행하지 않은 목표를 내부적으로 유지하고 있는 상황까지 포착할 수 있는 분석 도구를 만들었다”라며 “일부 바람직한 행동은 모델이 자신이 평가받고 있다는 사실을 알고 있었기 때문에 나타난 것일 수 있다”고 말했다.</p>



<p>그는 기업 고객 역시 AI 안전성 벤치마크를 해석할 때 이러한 점을 고려해야 한다고 지적했다.</p>



<p>램브로스는 “프로젝트에 적합한 모델인지는 모델이 알고 응시한 리더보드 결과가 아니라, 기업이 보유한 데이터와 실제 공격 시나리오를 활용한 자체 테스트를 통해 검증해야 한다”고 설명했다.</p>



<p>이처럼 모델 내부를 들여다볼 수 있는 능력은 CIO에게도 중요한 의미를 갖는다.</p>



<p>램브로스는 “자사 모델이 겉으로 드러나지 않는 문제 행동을 스스로 발견하고 그 결과를 공개할 수 있는 공급업체라면 신뢰성 검증 체계가 상당히 성숙했다는 의미”라며 “이러한 역량은 단순한 뉴스가 아니라 공급업체 실사(Due Diligence) 과정에서 반드시 확인해야 한다”고 말했다.</p>



<p>이어 “이제 모든 AI 모델 공급업체에 던져야 할 질문은 ‘모델 출력만으로는 볼 수 없는 내부 상태 가운데 무엇을 관찰할 수 있으며, 실제로 어떤 문제를 발견했는가’가 돼야 한다”고 덧붙였다.</p>



<p>AI 거버넌스 컨설팅 기업 디지털 520(Digital 520)의 수석 컨설턴트 <a href="https://www.linkedin.com/in/noah-m-kenney-27499a166/" target="_blank" rel="nofollow">노아 케니</a>(Noah Kenney)도 비슷한 견해를 내놨다.</p>



<p>케니는 “감시받고 있다는 사실을 알기 때문에 더 바람직하게 행동하는 모델은 안전한 모델이 아니다. 단지 포커페이스를 잘하는 모델일 뿐”이라며 “레드팀 테스트 결과나 모델이 위험한 요청을 거부한 내부 파일럿, ‘테스트해 보니 문제가 없었다’는 모든 사례를 다시 검토해야 한다. 이제는 모두 단서를 달고 해석해야 하기 때문”이라고 말했다.</p>



<p>또한 CIO는 AI 에이전트가 특정 방식으로 작업을 수행한 이유가 원래 그렇게 설계됐기 때문인지, 아니면 단순히 자신이 테스트받고 있다는 사실을 알아차렸기 때문인지를 구분해야 한다고 강조했다.</p>



<p>케니는 “이 질문에 대한 답에 따라 모델 평가 결과에 대한 해석도 크게 달라져야 한다”고 말했다.</p>



<h2 class="wp-block-heading">아직 고객은 사용할 수 없는 J-렌즈</h2>



<p>노아 케니는 “이번 연구는 지금까지 업계가 AI 모델 평가를 통해 측정해 온 것이 모두가 생각했던 것만큼 견고한 지표가 아니었다는 사실을 인정한 것”이라며 “이제 다른 프런티어 AI 연구소들도 자사 평가 체계 역시 같은 문제를 안고 있는지 답해야 할 것”이라고 말했다. 이어 “CIO에게 이번 논문은 기업의 AI 모델 리스크 관리 체계 전반을 다시 점검하라는 경고”라고 평가했다.</p>



<p>렉시스넥시스 리스크 솔루션 그룹(LexisNexis Risk Solutions Group)의 CISO <a href="https://www.linkedin.com/in/fvillanustre/" target="_blank" rel="nofollow">플라비오 비야누스트레</a>(Flavio Villanustre)는 J-스페이스를 분석하면 모델 효율성까지 높일 수 있다고 설명했다.</p>



<p>비야누스트레는 “J-스페이스는 모델 내부를 직접 들여다볼 수 있는 능력을 제공하기 때문에 사용자에게 매우 유용하다”라며 “특히 설명 가능성이 중요한 규제 산업에서는 응답의 근거와 인과관계를 충분히 분석해야 하는데 큰 도움이 될 수 있다”고 말했다.</p>



<p>이어 “사용자가 프롬프트를 더욱 정교하게 다듬는 데에도 활용할 수 있어 모델의 토큰 사용 비용을 최적화하는 데 도움이 된다”고 설명했다.</p>



<p>다만 현재로서는 이러한 정보를 직접 활용하기 어렵다. AI 공급업체를 통해 간접적으로 접근하거나 향후 계약 협상을 통해 권한을 확보하는 것이 사실상 유일한 방법이다. 비야누스트레는 일부 기업은 앤트로픽의 FDE 프로그램에 비용을 지불하면 J-스페이스에 직접 접근할 수도 있다고 덧붙였다.</p>



<p>그는 “J-스페이스는 CIO에게 매우 유용한 도구”라면서도 “이를 실제로 활용하려면 분석 결과를 해석할 수 있는 전문 인력이 필요하다. 요구되는 역량은 일반 데이터 분석가는 물론 데이터 사이언티스트 수준을 넘어선다”고 말했다.</p>



<p>기술 컨설팅 기업 트라이베카 소프트테크(Tribeca Softtech)의 최고전략책임자(CSO) <a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="nofollow">아만 마하파트라</a>(Aman Mahapatra)는 현재 기업 고객이 J-렌즈를 실제 운영에 활용하기는 어렵다고 지적했다.</p>



<p>그는 “현재 기업 고객은 야코비안 렌즈를 활성화할 수도 없고, API를 통해 모델의 잔차 스트림(residual stream)을 분석할 수도 없으며, 논문의 핵심 결과를 도출한 제거 실험(ablation study)도 수행할 수 없다”고 설명했다.</p>



<p>이어 “올해 3분기 안에 CIO가 J-스페이스 모니터링을 실제 운영 환경의 배포 승인 기준으로 활용할 수 있느냐는 질문에 대한 답은 ‘아니다'”라고 말했다.</p>



<p>다만 마하파트라는 앞으로는 다른 방식의 접근 경로가 마련될 것이며, CIO들이 이를 적극 요구해야 한다고 주장했다.</p>



<p>그는 “고객이 직접 접근할 수 없다면 결국 이번에도 앤트로픽을 믿을 수밖에 없다”라며 “바로 그렇기 때문에 기업들은 업계 전반에 새로운 신뢰 검증(Assurance) 체계를 요구해야 한다”고 말했다.</p>



<p>이어 “현재 AI 모델 공급업체들은 자체 도구로 스스로 모델을 점검한 뒤 안심할 수 있는 연구 결과를 발표하는 방향으로 나아가고 있다”라며 “그러나 어떤 규제 산업도 다른 공급업체에게 이런 방식의 검증을 신뢰 기준으로 인정하지 않는다”고 지적했다.</p>



<p>마하파트라는 “은행은 신용평가 업체가 ‘우리 모델은 우리가 검증했으니 믿어달라’고 말한다고 이를 받아들이지 않는다”라며 “의료 업계 역시 임상 의사결정 지원 시스템 공급업체의 자체 검증만으로는 신뢰하지 않는다. 파운데이션 모델 공급업체만 예외로 취급해야 할 원칙적인 이유는 없으며, 이번 J-스페이스 연구는 그 이유를 분명하게 보여준다”고 말했다.</p>



<h2 class="wp-block-heading">새로운 가시성이 요구하는 변화</h2>



<p>마하파트라는 기업이 장기적으로는 AI 모델의 내부 동작을 독립적으로 검증할 수 있는 환경을 요구해야 한다고 강조했다.</p>



<p>그는 “기업이 취해야 할 올바른 장기 전략은 고객이 사용할 수 있는 API, 특권 접근 권한을 가진 독립적인 제3자 감사기관, 또는 은행의 모델 리스크 관리팀이 공급업체의 안전성 검증팀과 동일한 도구를 사용할 수 있도록 하는 개방형 해석 가능성(Interpretability) 표준 등을 요구하는 것”이라고 말했다.</p>



<p>이어 “현재는 이런 환경이 전혀 마련돼 있지 않다”라며 “하지만 CIO들이 추진해야 할 로드맵에는 반드시 포함돼야 하며, 이번 연구는 그 필요성을 보여주는 가장 강력한 근거”라고 평가했다.</p>



<p>실제로 이번 연구 결과는 기업의 AI 전략 자체를 근본적으로 바꿀 가능성을 갖고 있다.</p>



<p>마하파트라는 “기업이 AI 에이전트를 도입할 때 가장 어려운 과제는 자율 시스템이 설명하는 추론 과정과 실제 내부 추론이 일치하는지를 검증하는 것”이라며 “지금까지는 모델이 출력한 내용만 감사할 수 있었고 실제 추론의 상당 부분은 보이지 않는 곳에서 이뤄졌다. J-렌즈는 바로 이 간극을 정면으로 해결하려는 시도”라고 설명했다.</p>



<p>그는 기업의 AI 구매 담당자들이 앞으로 모델 공급업체에 내부 상태를 관찰할 수 있는 해석 가능성 도구를 제공하는지 반드시 확인해야 한다고 조언했다. 특히 고객 환경에서 기만 행위(deception), 평가 회피(evaluation gaming), 목표 불일치(goal misalignment) 등을 모니터링할 수 있는 기능을 제공하는지를 구매 과정에서 질문해야 한다고 강조했다.</p>



<p>마하파트라는 “현재 이러한 질문에 제대로 답할 수 있는 공급업체는 거의 없다”라며 “관련 도구가 완전히 성숙하기 전이라도 내부 상태의 가시성을 구매 기준으로 요구하기 시작하는 CIO가 앞으로 공급업체의 제품 발전 방향을 결정하게 될 것”이라고 말했다.</p>



<p>이어 “향후 규제기관이 ‘기업은 자율 AI 에이전트가 실제로 주장한 대로 동작한다는 사실을 어떻게 확인했는가’를 묻기 시작하면, 이런 기업만이 실질적인 신뢰성을 입증할 수 있을 것”이라고 전망했다.</p>



<h2 class="wp-block-heading">표준화의 시작</h2>



<p>CIO가 클로드의 새로운 내부 가시성을 활용할 수 있는 또 다른 방법은 이미 해당 정보에 접근할 수 있는 제3자를 활용하는 것이다. 실제로 이번 보고서에는 구글의 AI 연구원이 오픈웨이트(Open-weight) 모델에서 일부 연구 결과를 독립적으로 재현하는 데 성공했다는 내용이 포함됐다.</p>



<p>소프트웨어 개발 기업 컴프 AI(Comp AI)의 CEO <a href="https://www.linkedin.com/in/lewiscarhart/" target="_blank" rel="nofollow">루이스 카하트</a>(Lewis Carhart)는 “이는 공급업체의 주장만이 아니라 경쟁사가 해당 분석 기법을 검증했다는 의미”라며 “기술적으로 무엇이 가능한지를 보여주지만, 기업이 직접 이를 검증할 수 있는 수단을 제공하는 것은 아니다”라고 설명했다.</p>



<p>그는 이러한 상황이 컴플라이언스 분야에서도 이미 여러 차례 반복됐던 패턴이라고 말했다.</p>



<p>카하트는 “SOC 2 역시 처음부터 독립적인 감사 기준으로 출발한 것은 아니었다”라며 “초기에는 공급업체가 자체 통제 체계를 설명하는 수준에 머물렀고, 이후 시장이 수년에 걸쳐 이를 외부에서 검증할 수 있는 인프라를 구축했다”고 설명했다.</p>



<p>이어 “AI 해석 가능성(Interpretability)도 지금은 바로 그 출발점에 있다”라며 “J-렌즈 분석 결과가 제3자 감사 보고서나 모델 카드(Model Card), 또는 규제기관 제출 문서에 포함되기 시작해야 CIO에게 실질적인 의미를 갖게 될 것”이라고 말했다.</p>



<p>그는 “결국 리스크 관리 조직이 공급업체의 주장만이 아니라 객관적인 근거로 제시할 수 있는 자료가 마련돼야 한다”고 덧붙였다.</p>



<h2 class="wp-block-heading">AI 전략 변화 이끄나</h2>



<p>컨설팅 기업 액셀리전스(Acceligence)의 CEO <a href="https://acceligence.com/talent/profiles/justin-greis/" target="_blank" rel="nofollow">저스틴 그라이스</a>(Justin Greis)는 이번 기술이 기업 AI 전략에도 상당한 변화를 가져올 것으로 전망했다.</p>



<p>그는 “향후 AI 거버넌스 플랫폼은 프롬프트와 출력 결과, 사용자 신원 정보, 정책 결정, 도구 사용 기록뿐 아니라 J-렌즈가 제공하는 내부 신호까지 함께 활용하게 될 것”이라고 말했다.</p>



<p>이어 “미래의 AI 제어 플랫폼은 AI 에이전트가 프롬프트 인젝션 시도를 인식했는지, 민감한 정보가 포함됐음을 이해했는지, 상충되는 목표를 감지했는지, 또는 실제 위험한 행동을 실행하기 전에 그러한 방향으로 추론하고 있었다는 징후가 있었는지를 지속적으로 평가할 수 있을 것”이라며 “이러한 신호는 정책 집행, 사람의 개입 여부 결정, 감사 로그 작성, 기업 AI 환경 전반의 신뢰도 평가 등에 활용될 것”이라고 설명했다.</p>



<p>그는 이러한 변화가 현재 CIO에게도 실질적인 의미를 갖는다고 강조했다.</p>



<p>그라이스는 “AI 공급업체를 평가하는 기준 자체가 달라지고 있기 때문”이라며 “1년 전만 해도 기업은 모델의 정확도와 지연시간, 보안, 비용을 주로 평가했다. 앞으로는 AI 에이전트의 행동과 추론 품질, 정책 준수 여부, 안전성 모니터링, 감사 가능성에 대해 공급업체가 얼마나 높은 수준의 운영 가시성을 제공하는지도 중요한 평가 항목이 될 것”이라고 말했다.</p>



<p>마하파트라는 이러한 변화가 CIO에게 강력한 협상 카드가 될 수도 있다고 분석했다.</p>



<p>그는 “실제 협상력이 발휘되는 시점은 계약 갱신 과정”이라며 “다음 계약 갱신 때 해석 가능성 보고서 제공과 제3자 감사 접근 권한을 계약 조항에 포함시켜야 한다. 지금은 이러한 조건을 비교적 쉽게 확보할 수 있지만 계약 체결 이후에는 훨씬 큰 비용이 들기 때문”이라고 말했다.</p>



<p>이어 “2027년 신뢰성 확보 경쟁에서 앞서는 CIO는 2026년에 이미 AI 모델 공급업체의 ‘우리를 믿어달라’는 말만 받아들이지 않고, 공급업체가 아직 계약을 더 필요로 하는 시점에 필요한 조항을 계약서에 반영한 사람일 것”이라고 전망했다. <br>dl-ciokorea@foundryco.com</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2022-40435 | Employee Performance Evaluation System 1.0 Departments/Designations cross site scripting]]></title>
<description><![CDATA[A vulnerability was found in Employee Performance Evaluation System 1.0 and classified as problematic. Affected by this vulnerability is an unknown functionality of the component Departments/Designations Module. Executing a manipulation can lead to cross site scripting.

The identification of thi...]]></description>
<link>https://tsecurity.de/de/3655298/sicherheitsluecken/cve-2022-40435-employee-performance-evaluation-system-10-departmentsdesignations-cross-site-scripting/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3655298/sicherheitsluecken/cve-2022-40435-employee-performance-evaluation-system-10-departmentsdesignations-cross-site-scripting/</guid>
<pubDate>Wed, 08 Jul 2026 21:24:46 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability was found in <a href="https://vuldb.com/product/employee_performance_evaluation_system">Employee Performance Evaluation System 1.0</a> and classified as <a href="https://vuldb.com/kb/risk">problematic</a>. Affected by this vulnerability is an unknown functionality of the component <em>Departments/Designations Module</em>. Executing a manipulation can lead to cross site scripting.

The identification of this vulnerability is <a href="https://vuldb.com/cve/CVE-2022-40435">CVE-2022-40435</a>. The attack may be launched remotely. There is no exploit available.]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI to release delayed models Thursday amidst a sea of regulatory confusion]]></title>
<description><![CDATA[As enterprises struggle to manage their AI strategies, the US AI regulatory environment is sending a wide range of contradictory signals. OpenAI’s Wednesday announcement that it will now release GPT-5.6 Sol, along with Terra and Luna, on Thursday highlights the confusion.



Initially, the US gov...]]></description>
<link>https://tsecurity.de/de/3655253/ai-nachrichten/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3655253/ai-nachrichten/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion/</guid>
<pubDate>Wed, 08 Jul 2026 21:03:35 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>As enterprises struggle to manage their AI strategies, the US AI regulatory environment is sending a wide range of contradictory signals. OpenAI’s Wednesday announcement that it will now release GPT-5.6 Sol, along with Terra and Luna, on Thursday highlights the confusion.</p>



<p>Initially, the US government said that it was asking OpenAI to <a href="https://www.infoworld.com/article/4190089/us-tells-openai-to-restrict-access-to-its-most-powerful-ai-model-2.html" target="_blank">limit access to its top models</a>, including the three releasing Thursday, to a short list of companies. OpenAI seemingly agreed and held back their general availability.</p>



<p>But on Wednesday, OpenAI reversed its position, with <a href="https://x.com/OpenAI/status/2074704958419792299?s=20" target="_blank" rel="noreferrer noopener">a statement on X</a> saying simply: “GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.” No details were released about the extent of the expansion.</p>



<p>Then the White House issued a statement, a copy of which it emailed to <em>InfoWorld</em>, saying that the US government “did not give OpenAI a ‘green light,’ approval or clearance to release its models. No such permission is required or granted. The Administration does not provide approvals for private companies to release AI models – decisions on timing and scope of releases rest entirely with the companies.”</p>



<p>The statement then quoted from the <a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/" target="_blank" rel="noreferrer noopener">June 2 White House executive order</a> that said, “nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models.” It also said, “any testing or meetings with government experts is voluntary. Participation is not required to release a model.”</p>



<p>Yet last month, the US Commerce Department weighed in <a href="https://www.cio.com/article/4186429/anthropic-fable-dispute-suggests-export-no-longer-means-what-it-used-to.html" target="_blank">on how Anthropic’s models can be distributed</a>. </p>



<h2 class="wp-block-heading">The ‘worst of both worlds’</h2>



<p><a href="https://www.linkedin.com/in/lewiscarhart/" target="_blank" rel="noreferrer noopener">Lewis Carhart</a>, CEO of software development firm Comp AI, said the statement is frustrating for IT executives on multiple fronts. </p>



<p>“Think about what that [White House statement] means. OpenAI stationed engineers in Washington for weeks, submitted to government testing, staggered its launch at the government’s request. And the official position is that none of that was required,” Carhart said. “We now have a de facto licensing regime that legally doesn’t exist. There’s no statute, no appeal process, no published criteria. Just [the US Department of] Commerce deciding model by model what ships and when.”</p>



<p>That’s the worst of both worlds, he noted: “All the friction of regulation with none of the predictability. Compliance people have a name for this – it’s an audit with no framework. And the precedent is now locked in for both frontier labs: if you build at the frontier, your launch calendar runs through Washington whether the law says so or not.”</p>



<p>Carhart argued that this regulatory reality should be of extreme concern to enterprise IT executives, given it indicates that model availability is now “a regulatory variable” not driven by the vendor roadmap.</p>



<p>“Anthropic’s most advanced models disappeared from the market for three weeks in June. It was not because of an outage, not because of a pricing change. It was because of an export control directive,” he pointed out. “If your AI architecture assumes the model you deployed today is available tomorrow, that assumption is now demonstrably false. Multi-model resilience just went from nice to have to a board-level risk item.”</p>



<p>It also offers an opportunity, given that a model that cleared government security testing is a model that auditors and boards sign off on faster. “Government review is quietly becoming a procurement asset,” he observed. “The CIOs who win here are the ones who treat ‘regulatory posture of the model itself’ as a line item in vendor risk assessments – most risk teams are still only looking at the provider’s SOC 2 attestation.”</p>



<p><a href="https://moorinsightsstrategy.com/team/jason-andersen/" target="_blank" rel="noreferrer noopener">Jason Andersen</a>, principal analyst at Moor Insights &amp; Strategy, agreed with Carhart and described the overall back-and-forth as “a bit of pageantry. OpenAI needs its model to look as powerful, potentially dangerous, as Anthropic’s so it can be a contender to be at the absolute frontier. It also helps to burnish OpenAI’s PR efforts to look more responsible than it has in the past.”</p>



<p>Furthermore, he added, “tech CEOs are acutely aware that flattery towards this administration could keep regulators or the threat of regulation at bay.”</p>



<h2 class="wp-block-heading">Criteria not clear</h2>



<p><a href="https://www.infotech.com/profiles/brian-jackson" target="_blank" rel="noreferrer noopener">Brian Jackson</a>, a principal research director at Info-Tech Research Group, echoed the political concerns. </p>



<p>“What’s still not clear is the actual criteria being used to deem the models safe for release. The government has said that it’s concerned about cybersecurity risk as well as the risk of AI being used to develop biological weapons,” he said. “But so long as the actual release criteria lack transparency, there will be some perception that the evaluation could be politically motivated.”</p>



<p>Jackson said that one, presumably unintended, result of the US government’s efforts to control AI rollouts is that it is making companies look far more seriously at using non-US vendors for AI strategies. </p>



<p>“Organizations are looking for alternatives to the private US-based cloud-delivered frontier models. That’s so they can maintain control over their AI supply chain,” he said. “Chinese open source models are one option that’s available, but there are other options too, from Canada and Europe. Companies can either set up AI access through other APIs not connected to US-based AI providers, or download open-source models to run locally.”</p>



<p>He noted that the added US regulatory risk means that some organizations will avoid becoming entrenched within OpenAI’s and Anthropic’s interfaces, where there’s no option to swap out the LLMs for an alternative.</p>



<h2 class="wp-block-heading">Impacts enterprise AI strategy</h2>



<p><a href="https://zenity.io/authors/rock-lambros" target="_blank" rel="noreferrer noopener">Rock Lambros</a>, director of AI standards and governance at AI agent vendor Zenity, shared the frustration that little to no actionable compliance data is being released. </p>



<p>“Nobody outside a closed room can tell you what standard [the model] passed because it was never written down. For two weeks, [US government officials] kept a model out of defenders’ hands that’s better at guarding your network than breaking into anyone else’s,” Lambros said. “Call that a security review if it helps you sleep better. But it reads to me like a bouncer working a velvet rope nobody hired him to run, waving people through today because he’s in a better mood than he was a couple of weeks ago.”</p>



<p>This unpredictability is likely to have impacts on AI strategy far beyond traditional compliance concerns, Lambros said.</p>



<p>“We’ve built way too much operational reliance on these models to hang it on a review with no rulebook,” he said, pointing out that hospitals, pipelines, banks and water utilities are relying on frontier AI whose availability “can swing from ‘on’ to ‘off’ to ‘on’ with no notice, no appeal, and no published standard behind any of it.”</p>



<p>“You can’t run critical infrastructure on a tool that runs fine Friday and is offline by Monday because an approval process nobody can see reached a verdict nobody can predict,” Lambros said. “That is a supply chain risk with a government hand on the switch, and almost nobody has priced it into a continuity plan.”</p>



<p>To protect themselves, companies need to adjust their expectations. “Treat model availability like any single point of failure you don’t own by standing up a fallback you’ve tested, getting a continuity clause in your contract, and drilling for the blackout, because ‘the government backed off this time’ is not a plan,” he advised.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI to release delayed models Thursday amidst a sea of regulatory confusion]]></title>
<description><![CDATA[As enterprises struggle to manage their AI strategies, the US AI regulatory environment is sending a wide range of contradictory signals. OpenAI’s Wednesday announcement that it will now release GPT-5.6 Sol, along with Terra and Luna, on Thursday highlights the confusion.



Initially, the US gov...]]></description>
<link>https://tsecurity.de/de/3655243/it-nachrichten/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3655243/it-nachrichten/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion/</guid>
<pubDate>Wed, 08 Jul 2026 21:02:32 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>As enterprises struggle to manage their AI strategies, the US AI regulatory environment is sending a wide range of contradictory signals. OpenAI’s Wednesday announcement that it will now release GPT-5.6 Sol, along with Terra and Luna, on Thursday highlights the confusion.</p>



<p>Initially, the US government said that it was asking OpenAI to <a href="https://www.infoworld.com/article/4190089/us-tells-openai-to-restrict-access-to-its-most-powerful-ai-model-2.html" target="_blank">limit access to its top models</a>, including the three releasing Thursday, to a short list of companies. OpenAI seemingly agreed and held back their general availability.</p>



<p>But on Wednesday, OpenAI reversed its position, with <a href="https://x.com/OpenAI/status/2074704958419792299?s=20" target="_blank" rel="noreferrer noopener">a statement on X</a> saying simply: “GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.” No details were released about the extent of the expansion.</p>



<p>Then the White House issued a statement, a copy of which it emailed to <em>InfoWorld</em>, saying that the US government “did not give OpenAI a ‘green light,’ approval or clearance to release its models. No such permission is required or granted. The Administration does not provide approvals for private companies to release AI models – decisions on timing and scope of releases rest entirely with the companies.”</p>



<p>The statement then quoted from the <a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/" target="_blank" rel="noreferrer noopener">June 2 White House executive order</a> that said, “nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models.” It also said, “any testing or meetings with government experts is voluntary. Participation is not required to release a model.”</p>



<p>Yet last month, the US Commerce Department weighed in <a href="https://www.cio.com/article/4186429/anthropic-fable-dispute-suggests-export-no-longer-means-what-it-used-to.html" target="_blank">on how Anthropic’s models can be distributed</a>. </p>



<h2 class="wp-block-heading">The ‘worst of both worlds’</h2>



<p><a href="https://www.linkedin.com/in/lewiscarhart/" target="_blank" rel="noreferrer noopener">Lewis Carhart</a>, CEO of software development firm Comp AI, said the statement is frustrating for IT executives on multiple fronts. </p>



<p>“Think about what that [White House statement] means. OpenAI stationed engineers in Washington for weeks, submitted to government testing, staggered its launch at the government’s request. And the official position is that none of that was required,” Carhart said. “We now have a de facto licensing regime that legally doesn’t exist. There’s no statute, no appeal process, no published criteria. Just [the US Department of] Commerce deciding model by model what ships and when.”</p>



<p>That’s the worst of both worlds, he noted: “All the friction of regulation with none of the predictability. Compliance people have a name for this – it’s an audit with no framework. And the precedent is now locked in for both frontier labs: if you build at the frontier, your launch calendar runs through Washington whether the law says so or not.”</p>



<p>Carhart argued that this regulatory reality should be of extreme concern to enterprise IT executives, given it indicates that model availability is now “a regulatory variable” not driven by the vendor roadmap.</p>



<p>“Anthropic’s most advanced models disappeared from the market for three weeks in June. It was not because of an outage, not because of a pricing change. It was because of an export control directive,” he pointed out. “If your AI architecture assumes the model you deployed today is available tomorrow, that assumption is now demonstrably false. Multi-model resilience just went from nice to have to a board-level risk item.”</p>



<p>It also offers an opportunity, given that a model that cleared government security testing is a model that auditors and boards sign off on faster. “Government review is quietly becoming a procurement asset,” he observed. “The CIOs who win here are the ones who treat ‘regulatory posture of the model itself’ as a line item in vendor risk assessments – most risk teams are still only looking at the provider’s SOC 2 attestation.”</p>



<p><a href="https://moorinsightsstrategy.com/team/jason-andersen/" target="_blank" rel="noreferrer noopener">Jason Andersen</a>, principal analyst at Moor Insights &amp; Strategy, agreed with Carhart and described the overall back-and-forth as “a bit of pageantry. OpenAI needs its model to look as powerful, potentially dangerous, as Anthropic’s so it can be a contender to be at the absolute frontier. It also helps to burnish OpenAI’s PR efforts to look more responsible than it has in the past.”</p>



<p>Furthermore, he added, “tech CEOs are acutely aware that flattery towards this administration could keep regulators or the threat of regulation at bay.”</p>



<h2 class="wp-block-heading">Criteria not clear</h2>



<p><a href="https://www.infotech.com/profiles/brian-jackson" target="_blank" rel="noreferrer noopener">Brian Jackson</a>, a principal research director at Info-Tech Research Group, echoed the political concerns. </p>



<p>“What’s still not clear is the actual criteria being used to deem the models safe for release. The government has said that it’s concerned about cybersecurity risk as well as the risk of AI being used to develop biological weapons,” he said. “But so long as the actual release criteria lack transparency, there will be some perception that the evaluation could be politically motivated.”</p>



<p>Jackson said that one, presumably unintended, result of the US government’s efforts to control AI rollouts is that it is making companies look far more seriously at using non-US vendors for AI strategies. </p>



<p>“Organizations are looking for alternatives to the private US-based cloud-delivered frontier models. That’s so they can maintain control over their AI supply chain,” he said. “Chinese open source models are one option that’s available, but there are other options too, from Canada and Europe. Companies can either set up AI access through other APIs not connected to US-based AI providers, or download open-source models to run locally.”</p>



<p>He noted that the added US regulatory risk means that some organizations will avoid becoming entrenched within OpenAI’s and Anthropic’s interfaces, where there’s no option to swap out the LLMs for an alternative.</p>



<h2 class="wp-block-heading">Impacts enterprise AI strategy</h2>



<p><a href="https://zenity.io/authors/rock-lambros" target="_blank" rel="noreferrer noopener">Rock Lambros</a>, director of AI standards and governance at AI agent vendor Zenity, shared the frustration that little to no actionable compliance data is being released. </p>



<p>“Nobody outside a closed room can tell you what standard [the model] passed because it was never written down. For two weeks, [US government officials] kept a model out of defenders’ hands that’s better at guarding your network than breaking into anyone else’s,” Lambros said. “Call that a security review if it helps you sleep better. But it reads to me like a bouncer working a velvet rope nobody hired him to run, waving people through today because he’s in a better mood than he was a couple of weeks ago.”</p>



<p>This unpredictability is likely to have impacts on AI strategy far beyond traditional compliance concerns, Lambros said.</p>



<p>“We’ve built way too much operational reliance on these models to hang it on a review with no rulebook,” he said, pointing out that hospitals, pipelines, banks and water utilities are relying on frontier AI whose availability “can swing from ‘on’ to ‘off’ to ‘on’ with no notice, no appeal, and no published standard behind any of it.”</p>



<p>“You can’t run critical infrastructure on a tool that runs fine Friday and is offline by Monday because an approval process nobody can see reached a verdict nobody can predict,” Lambros said. “That is a supply chain risk with a government hand on the switch, and almost nobody has priced it into a continuity plan.”</p>



<p>To protect themselves, companies need to adjust their expectations. “Treat model availability like any single point of failure you don’t own by standing up a fallback you’ve tested, getting a continuity clause in your contract, and drilling for the blackout, because ‘the government backed off this time’ is not a plan,” he advised.</p>



<p><em>This article originally appeared on <a href="https://www.infoworld.com/article/4194598/openai-to-release-delayed-models-thursday-amidst-a-sea-of-regulatory-confusion.html" target="_blank">InfoWorld</a>.</em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Evolving how LLMs are measured for Android: the next era of Android Bench]]></title>
<description><![CDATA[Posted by Zoe Lopez-Latorre, Senior Developer Relations Engineer, AndroidBack in March, we introduced Android Bench—our LLM leaderboard for real-world Android development tasks. Our goal was to provide transparency around model capabilities in Android development and to encourage model improvemen...]]></description>
<link>https://tsecurity.de/de/3655018/android-tipps/evolving-how-llms-are-measured-for-android-the-next-era-of-android-bench/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3655018/android-tipps/evolving-how-llms-are-measured-for-android-the-next-era-of-android-bench/</guid>
<pubDate>Wed, 08 Jul 2026 19:27:23 +0200</pubDate>
<category>🤖 Android Tipps</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[
<img src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEgCAy4lIbOAOrygTMaHZB8q4NarDrLRsqALfsmer5urQX7G_MaRDTw51uMh77Ks2knIuWM-zaEel63Dk2IlCVGD9IxLFy0B68KxwxsvDZzVDaEWaM4Bg8xJYinunaXS_fonxBw7-R4_qSplI4MJU7RDDaYlbq7nRXZoht5lFZVC7ErLEWHdWA6B2KgJvrk/s2469/Bench%20July%20releas%20V01_Meta.png">
<div><i>Posted by Zoe Lopez-Latorre, Senior Developer Relations Engineer, Android</i></div><div><i><br></i><div><div class="separator"><a href="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEi49z_u9zPMjp-zyQ1yIpzLgDumtzUwZoprtIgPXv_kpF05e87KklDEguaKSJVhvV8dZJ7aVr98p-MG3FR4Sk37rcYTS91J3ADUQot-c-xnOuyIZ411VO4Hp43Yp7V_TwF6zO6RmAJpw51ZHPGbHfOwZxWgQ62SQeXblULcSc0RjMcZbLHGUZGgHzU6pEo/s8583/Bench%20July%20releas%20V01_Blog.png"><img border="0" data-original-height="2601" data-original-width="8583" src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEi49z_u9zPMjp-zyQ1yIpzLgDumtzUwZoprtIgPXv_kpF05e87KklDEguaKSJVhvV8dZJ7aVr98p-MG3FR4Sk37rcYTS91J3ADUQot-c-xnOuyIZ411VO4Hp43Yp7V_TwF6zO6RmAJpw51ZHPGbHfOwZxWgQ62SQeXblULcSc0RjMcZbLHGUZGgHzU6pEo/s1600/Bench%20July%20releas%20V01_Blog.png"></a></div><br><i><br></i><p>Back in March, we introduced <a href="http://d.android.com/bench">Android Bench</a>—our LLM leaderboard for real-world Android development tasks. Our goal was to provide transparency around model capabilities in Android development and to encourage model improvements, to give you more helpful AI options for your everyday workflow. Since then, we have enhanced the benchmark based on your feedback, including evaluating <a href="https://x.com/AndroidDev/status/2064482677500080549">open-weight models</a> and adding cost and efficiency dimensions to the leaderboard.</p>

<p>But AI capabilities are ever-evolving, and measurement needs to follow suit. As part of our July release, we have adopted the <a href="https://www.harborframework.com/">Harbor framework</a>, which includes an updated version of the benchmarking agent used to evaluate models.</p>

Along with this change to our evaluation, in this July release we’re adding 8 new models (<b>Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max</b>) to the leaderboard. We’re also sharing opportunities for you, the Android developer community, to contribute to the benchmark. 

<h2>Upgrading our methodology with the Harbor framework</h2>

<p>When we designed Android Bench, we anchored our methodology on leading industry standards available at the time. We used mini-swe-agent v1, a general-purpose benchmarking agent, and adapted it to the nuances of Android development to provide a baseline measurement for the capabilities of models for common Android development tasks.</p>

<p>To continue providing you with state-of-the-art evaluations that accurately measure the latest model capabilities on Android development, we are standardizing our benchmark to the <a href="https://www.harborframework.com/">Harbor framework</a>. Harbor defines standards and integrations that make it easy for anyone to run the benchmark, evaluate their preferred set-up, or share results – providing you with additional transparency and visibility.</p>

<p>This upgrade enables us to more rigorously evaluate models and their capabilities, and we re-ran the benchmark on all models to establish an updated baseline. This means there is a minor shift in scoring, but you will still be able to view historical scores within <a href="http://d.android.com/bench/archive">the archive</a> on our website.</p>

<p>We want to ensure Android Bench is helpful for you, so we will continuously update it as our evaluations and the industry mature.</p>

<h2>Expanding the leaderboard with 8 new models</h2>

<p>As part of our commitment to keeping the leaderboard fresh, we have added Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max to the Android Bench leaderboard.</p>

<p>You will see that <b>Claude Fable 5</b> is at the top of the leaderboard with a score of 84.5, followed by <b>GPT 5.5</b> with 80.2, with <b>Claude Sonnet 5</b> in 3rd with a score of 76.2.</p>

<p>When just comparing Open-weight models,<b> GLM 5.2</b> is at the top with 72.2, followed by <b>Kimi K2.7 Code</b> with a score of 70.4.</p>

<p>You can check out model performance and efficiency metrics on the updated leaderboard to see how these new and previous models navigate Android-specific challenges like Jetpack Compose migrations, wearable networking, and platform API updates.</p><div class="separator"><a href="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEhQCbY3Td_I5gR8bC4uFSBTe4Sl-XuArNNdFU-27JP6-kwHycXt9AMpWfkLqjUIK37Zw18Tel6a7yOS9x0L_NabxBgYd9KIJKZ6dTLl6VxxJI4M7Zstqj12wvOFtF8LjnYrCIWnhCDdeGsgpQvFpFX8VOoSO0dFJcOW_gRc6eX7mXDq80sOwQAlQWNhlQg/s1999/image1.png"><img border="0" data-original-height="890" data-original-width="1999" src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEhQCbY3Td_I5gR8bC4uFSBTe4Sl-XuArNNdFU-27JP6-kwHycXt9AMpWfkLqjUIK37Zw18Tel6a7yOS9x0L_NabxBgYd9KIJKZ6dTLl6VxxJI4M7Zstqj12wvOFtF8LjnYrCIWnhCDdeGsgpQvFpFX8VOoSO0dFJcOW_gRc6eX7mXDq80sOwQAlQWNhlQg/s1600/image1.png"></a></div>

<h2>Opening Android Bench to community contributions</h2>

<p>From the beginning, we’ve valued an open and transparent approach, which is why we made our original methodology and test harness publicly available on GitHub. You’ve asked for a way to provide feedback on our dataset, so now we’re taking collaboration a step further by giving you, the Android developer community, a chance to shape Android Bench.</p>

<p>Starting today, you can contribute to Android Bench in two ways:</p>

<ul>
    <li>Design and <a href="https://github.com/android-bench/community-dataset">submit your own Android development tasks</a> to evaluate how models handle the scenarios that matter to you.</li>
    <li><a href="https://github.com/android-bench/community-results">Run and share benchmark evaluations</a> firsthand, testing your preferred models against our dataset or your own custom tasks.</li>
</ul>

<p>We will be reviewing the submitted tasks and will be assessing if they get added to the benchmark. We hope to build a benchmark that truly reflects the diverse, day-to-day realities of the global Android developer community.</p>

<h2>Looking ahead</h2>

<p>With more and more options for agentic development, maintaining a cutting-edge benchmark ensures that the AI assistance you rely on keeps getting smarter, more helpful, and more effective. Head over to our <a href="https://github.com/android-bench/android-bench">GitHub repository</a> to check out the tasks. We invite you to submit a task to our team for review, and you can check out <a href="https://hub.harborframework.com/datasets/android-bench/android-bench/latest">Harbor Hub</a> to explore the dataset or submit evaluations.</p>

<p>As always, you can find the <a href="http://d.android.com/bench">updated leaderboard</a>, or read the <a href="http://d.android.com/bench/methodology">methodology</a> on our website.</p>
  <span>
    Android Bench, LLM leaderboard, Harbor framework, Android development, Claude Fable 5, GPT 5.5, Claude Sonnet 5, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, Qwen 3.7 Max, AI benchmarking, Jetpack Compose migration, wearable networking, mobile AI agent, Zoe Lopez-Latorre, model evaluation, open-weight models, developer community contributions.
</span>
  </div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Android Bench Upgraded to Harbor Framework, 8 Models Added to Leaderboard]]></title>
<description><![CDATA[The development team for Android announced an upgrade to Android Bench this week, its benchmark designed specifically for Android developers. The team says that it has adopted the Harbor framework, which creates sandbox environments to test and evaluate agents. This upgraded evaluation methodolog...]]></description>
<link>https://tsecurity.de/de/3654856/it-nachrichten/android-bench-upgraded-to-harbor-framework-8-models-added-to-leaderboard/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3654856/it-nachrichten/android-bench-upgraded-to-harbor-framework-8-models-added-to-leaderboard/</guid>
<pubDate>Wed, 08 Jul 2026 18:19:28 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The development team for Android announced an upgrade to Android Bench this week, its benchmark designed specifically for Android developers. The team says that it has adopted the Harbor framework, which creates sandbox environments to test and evaluate agents. This upgraded evaluation methodology, “makes it easier for anyone to run the benchmark, evaluate their preferred...</p>
<p>Read the original post: <a href="https://www.droid-life.com/2026/07/08/android-bench-upgraded-to-harbor-framework-8-models-added-to-leaderboard/">Android Bench Upgraded to Harbor Framework, 8 Models Added to Leaderboard</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Four agentic AI memory systems for smarter LLMs]]></title>
<description><![CDATA[AI agents, and the large language models (LLMs) that power them, have short memories. That’s by design. There is only so much conversation that can be encoded into tokens and accessed reliably by the LLM. Retrieval-augmented generation, or RAG, can be used to give agents and LLMs memories larger ...]]></description>
<link>https://tsecurity.de/de/3653745/ai-nachrichten/four-agentic-ai-memory-systems-for-smarter-llms/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3653745/ai-nachrichten/four-agentic-ai-memory-systems-for-smarter-llms/</guid>
<pubDate>Wed, 08 Jul 2026 11:04:04 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p><a href="https://www.infoworld.com/article/3812583/what-you-need-to-know-about-developing-ai-agents.html" data-type="link" data-id="https://www.infoworld.com/article/3812583/what-you-need-to-know-about-developing-ai-agents.html">AI agents</a>, and the <a href="https://www.infoworld.com/article/2335213/large-language-models-the-foundations-of-generative-ai.html" data-type="link" data-id="https://www.infoworld.com/article/2335213/large-language-models-the-foundations-of-generative-ai.html">large language models</a> (LLMs) that power them, have short memories. That’s by design. There is only so much conversation that can be encoded into tokens and accessed reliably by the LLM. <a href="https://www.infoworld.com/article/2335814/what-is-retrieval-augmented-generation-more-accurate-and-reliable-llms.html" data-type="link" data-id="https://www.infoworld.com/article/2335814/what-is-retrieval-augmented-generation-more-accurate-and-reliable-llms.html">Retrieval-augmented generation</a>, or RAG, can be used to give agents and LLMs memories larger than their context windows. But how agents use RAG, or other mechanisms for retaining the details of a conversation, can make all the difference.</p>



<p>With the rise of AI agents, there has been a corresponding rise in complementary software tools that give both agents and LLMs expanded memory capabilities. Most of the time, this means giving an agent or model persistent memory across sessions, so that previous context can be restored automatically. But, again, how that’s done can vary tremendously with each tool.</p>



<p>Here are some of the major projects in the AI agent memory space, each with their own particular spins, strengths, and orientations.</p>



<h2 class="wp-block-heading">Graphiti</h2>



<p><a href="https://github.com/getzep/graphiti">Graphiti</a> is billed as “the open-source temporal knowledge graph framework.” The project is available on GitHub, or as the underpinning of the <a href="https://www.getzep.com/">Zep ageny memory service</a>. “Temporal” means information stored in Graphiti is re-evaluated over time to keep its context properly framed, and “graph framework” means the data is stored as a set of graphs. The other solutions profiled here use graph storage as part of their approach, but Graphiti makes that a front-and-center part of its design.</p>



<p>Graphiti supports a range of common LLM providers out of the box: Anthropic, Azure OpenAI, Google Gemini, and Groq. Any Ollama and OpenAI-compatible APIs also work, so Graphiti can be used with locally hosted LLMs as well. Connectors for third-party storage services let you ingest data from places like GitHub, Gmail, and OneDrive, as well as from applications like Notion.</p>



<p>Using Graphiti locally requires you set up or connect to a graph database. <a href="https://neo4j.com/" data-type="link" data-id="https://neo4j.com/">Neo4j</a> is the default and most broadly supported of the bunch, but <a href="https://aws.amazon.com/neptune/" data-type="link" data-id="https://aws.amazon.com/neptune/">Amazon Neptune</a>, <a href="https://www.falkordb.com/" data-type="link" data-id="https://www.falkordb.com/">FalkorDB</a>, and <a href="https://kuzudb.github.io/" data-type="link" data-id="https://kuzudb.github.io/">KuzuDB</a> will also work. Postgres with <code>pgvector</code> is not listed as an option.</p>



<h2 class="wp-block-heading">Hindsight</h2>



<p><a href="https://hindsight.vectorize.io/">Hindsight</a>, available as both a cloud service and a locally hostable project, stores details about agent sessions into <a href="https://hindsight.vectorize.io/#key-components">four types of memory</a> with <a href="https://hindsight.vectorize.io/#multi-strategy-retrieval-tempr">four types of storage and retrieval strategies</a>. All of these are handled through three programmatic interfaces: <code>retain</code> for storing content, either a single fact or a whole conversation; <code>recall</code> for retrieving content; and <code>reflect</code> for running an agentic loop over a query that uses previously stored data.</p>



<p>Hindsight comes with a broad range of first-party and third-party <a href="https://hindsight.vectorize.io/integrations">integrations</a> with existing LLMs and agent toolkits. For instance, if you’re using the Continue extension with <a href="https://www.infoworld.com/article/2335960/what-is-visual-studio-code-microsofts-extensible-code-editor.html" data-type="link" data-id="https://www.infoworld.com/article/2335960/what-is-visual-studio-code-microsofts-extensible-code-editor.html">Visual Studio Code</a> to talk to a locally hosted LLM, you can use Hindsight’s <a href="https://hindsight.vectorize.io/sdks/integrations/continue">Continue integration</a> to add long-term memory to your interactions. You can use the <code>@hindsight</code> keyword in your query to inject relevant memory into the agent’s context, or use auto-injection rules (which can be edited) to do most of that heavy lifting automatically.</p>



<h2 class="wp-block-heading">Mem0</h2>



<p><a href="https://github.com/mem0ai/mem0">Mem0</a> is a little like Hindsight in that it has <a href="https://docs.mem0.ai/core-concepts/memory-types">four basic kinds of memory</a>, although they are labeled and organized differently. For instance, Mem0 has a separate type of memory called organizational memory that’s intended to store data to be shared between multiple agents or different teams, something that is not normally done by default. Each memory added is passed through a <a href="https://docs.mem0.ai/core-concepts/memory-evaluation#memory-extraction-distillation">distillation process</a> and stored in a different way (vector DB, graph DB, SQL DB) depending on how it will be used. Older data, instead of being overwritten, gets deprecated rather than deleted, as a strategy for preserving larger long-term context. (Hindsight does this as well.)</p>



<p>Mem0 supports <a href="https://docs.mem0.ai/components/llms/overview">a smaller range of LLMs</a> than Hindsight, but all the major options are available: Anthropic, Google Gemini, OpenAI, and self-hosted options like <a href="https://www.langchain.com/" data-type="link" data-id="https://www.langchain.com/">LangChain</a>, <a href="https://www.litellm.ai/" data-type="link" data-id="https://www.litellm.ai/">LiteLLM</a>, <a href="https://lmstudio.ai/" data-type="link" data-id="https://lmstudio.ai/">LM Studio</a>, and <a href="https://ollama.com/" data-type="link" data-id="https://ollama.com/">Ollama</a>. If you intend to use Mem0 locally rather than <a href="https://mem0.ai/pricing">as a service</a>, you’ll need to provide a Python instance and your own vector database. For the latter, Postgres with the <code>pgvector</code> extension is a common and simple choice; it can even be <a href="https://github.com/orm011/pgserver">installed inside a Python venv</a>. </p>



<h2 class="wp-block-heading">Supermemory</h2>



<p><a href="https://supermemory.ai/">Supermemory</a> ingests data from many common sources—supporting plaintext, structured data, common document file formats like PDF and Microsoft Office, video and audio, images—and uses them to build a context graph to inform agent conversations. Among its most promoted features is its content-extraction tools. </p>



<p>Supermemory is available as a cloud service or as <a href="https://github.com/supermemoryai/supermemory" data-type="link" data-id="https://github.com/supermemoryai/supermemory">open-source software</a> you can run locally. The open-source edition lacks the scaling services and third-party service connectors (Gmail, Google Drive, Notion, etc.) provided with the enterprise edition, but it has one big advantage: it consists of a single, self-contained binary, so it can be deployed on one’s own hardware with very little effort. No external databases need to be provisioned for Supermemory, either, so it’s well-suited to quick experimentation.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[The AI ROI gap isn’t a model problem. It’s a workflow problem]]></title>
<description><![CDATA[Anthropic says Claude now writes more than 80% of the code merged at one of the most sophisticated AI companies on the planet. Foundry’s 2026 State of the CIO study says fewer than one in five enterprises can show that their AI initiatives have met or exceeded their ROI goals. Both numbers came o...]]></description>
<link>https://tsecurity.de/de/3653733/it-nachrichten/the-ai-roi-gap-isnt-a-model-problem-its-a-workflow-problem/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3653733/it-nachrichten/the-ai-roi-gap-isnt-a-model-problem-its-a-workflow-problem/</guid>
<pubDate>Wed, 08 Jul 2026 11:02:59 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p><a href="https://www.anthropic.com/institute/recursive-self-improvement" rel="nofollow">Anthropic says</a> Claude now writes more than 80% of the code merged at one of the most sophisticated AI companies on the planet. Foundry’s <a href="https://www.cio.com/article/4178006/state-of-the-cio-2026-cios-set-the-course-for-ai-roi.html">2026 State of the CIO study</a> says fewer than one in five enterprises can show that their AI initiatives have met or exceeded their ROI goals. Both numbers came out this spring. Both are true. And the distance between them is the most important thing an IT leader can understand about AI right now.</p>



<p>Because that distance isn’t a contradiction, it’s a lesson. And the profession sitting in the middle of it, software engineering, is the canary that explains why so much enterprise AI spend has produced so little measurable return.</p>



<h2 class="wp-block-heading">The report everyone misread</h2>



<p>When Anthropic published its recursive self-improvement piece, plenty of people read it as the starting gun for the job apocalypse. Claude writing its own code, models getting better at building models, humans narrowing toward oversight. If you wanted a headline about the end of the software profession, it was right there.</p>



<p>I read it almost the opposite way. What struck me wasn’t how far AI had come. It was how much had to be true first, even in the one profession built from the ground up to let it succeed.</p>



<p>I made this argument back in my <a href="https://www.cio.com/article/4166029/the-570k-canary-what-ai-coding-agents-reveal-about-enterprise-ais-real-gaps.html">“$570K canary” piece</a>, and the Anthropic data only sharpens it. AI coding agents don’t work because coding models are special. The underlying large language models (LLMs) are the same ones answering support tickets and reviewing contracts. They work because software development already had the infrastructure that makes an agent’s output trustworthy: governance baked into branch protection and code review, observability through version control and CI/CD pipelines, evaluation through automated tests, persistent context through commit history. Developers built all of that for themselves over decades. They didn’t build it for AI. But it turned out to be exactly the scaffolding AI needed.</p>



<p>That’s the part the apocalypse reading skips. Claude’s coding gains are real. They also rode on decades of pre-built substrate. Both things are true at once, and the second one is the one CIOs should be paying attention to when it comes to gains from things like recursive self-improvement.</p>



<h2 class="wp-block-heading">What the CIO data actually shows</h2>



<p>Now hold that next to the State of the CIO numbers. Only 19% of the 662 IT leaders surveyed say their AI initiatives have met or exceeded business goals. Another 18% admit fewer than a third of their use cases are hitting defined expectations.</p>



<p>The easy explanation is that the technology isn’t ready. The data says otherwise. This isn’t for lack of trying, and it isn’t for lack of organizing. Eighty-three percent of respondents have stood up cross-functional steering committees or are about to. Just over half have some form of AI approval process in place, with another quarter building one. Forty-seven percent have formal success metrics, with a third more on the way. The field is pouring effort into the organizational machinery of AI. The ROI still isn’t showing up.</p>



<p>Here’s why I think that is. All of that machinery sits above the work. Steering committees, approval gates and KPI dashboards govern the org chart. But the value, or the leak, happens inside the workflow, at the level of the actual task the AI is doing. You can instrument your governance structure perfectly and still have nothing measuring whether the agent’s output was right at the point where it mattered.</p>



<p>TIAA shows how little the org chart settles. The firm is three years in, runs generative and agentic use cases across fraud detection and call centers, and has 85% of its people on TIAA Gate, its internal platform. It also has the full governance stack most CIOs are still assembling. None of it closed the gap. “You need to understand the full cost of operations,” its chief operating, information and digital officer, Sastry Durvasula told CIO.com, “the efficiencies of running tokens or how you’re handling traffic or RAG.” The structures were never the thing leaking value. The workflow underneath them was.</p>



<p>The barriers respondents named back this up. The top three are lack of in-house expertise (40%), ill-defined ROI metrics (32%) and murky corporate AI strategy (31%). Not one of them is “the model isn’t good enough.” And according to the full Foundry report, the expertise gap is deepest in healthcare (52%), retail (51%) and manufacturing (49%), the sectors whose core work looks least like a software development lifecycle. That’s consistent with substrate being the real variable, though a tighter market for AI talent in those industries is surely part of the story too.</p>



<h2 class="wp-block-heading">The market is already voting</h2>



<p>Look at where the AI is actually being pointed, and you’ll see enterprises sequencing by substrate even though nobody’s calling it that. Three-quarters of both IT leaders and line-of-business respondents say AI is primarily being used to automate internal processes rather than customer-facing applications.</p>



<p>That’s not timidity. It’s instinct pointing at the right thing. Internal processes are the ones with structured, observable workflows and users who tolerate a little friction. Customer-facing work is where the trust gaps are still wide open and the cost of a wrong answer is asymmetric. A bad internal draft gets fixed before anyone sees it. A bad customer answer is the whole ballgame.</p>



<p>I’ll be honest about a wrinkle in the data here, because a careful reader will catch it. The same study reports a near-mirror finding, that 66% to 69% of respondents say the bulk of their current AI work is customer-facing. The two stats sit a paragraph apart in the CIO study and almost certainly reflect how the question was framed rather than a real reversal. But the synthesis holds either way: even where customer-facing work is being attempted, it’s where ROI is least realized. The work that lands is the work with the substrate underneath it. The split only reinforces the point.</p>



<h2 class="wp-block-heading">Sequence by readiness, not by ambition</h2>



<p>So, here’s the prescription, and it cuts against the instinct most AI strategies are built on. Stop sequencing your AI portfolio by where the value looks biggest. Start sequencing it by where the work already has, or can be given, a structured workflow with a usable signal for whether the output was right.</p>



<p>The study itself shows what the alternative looks like. Andrea Ballinger, CIO at Rensselaer Polytechnic Institute, described the trap precisely. No one measures ROI on an ongoing basis, she said, “because we are facing counterpressures from every vice president and line-of-business domain looking to implement AI for their own optimization.” The result: “We are saying yes to everyone without stepping back and focusing on the business cases that show real value.” That’s value-led sequencing under pressure from every budget-holder in the building, and it’s exactly how you end up with a sprawling pipeline of pilots and a 19% success rate.</p>



<p>The counterexample comes from the same study. Thomas Prommer, a longtime CTO, CIO and CAIO, funds outcomes instead of deliverables. “We don’t fund ‘build a model,’ we fund ‘reduce returns by 8% on this category’ with checkpoints at 90, 180 and 270 days,” he explained. He kills any project that misses two checkpoints, “roughly a third of what we start, and that’s healthy.” Read that through the substrate lens and you see what he’s really doing. He’s manufacturing a correctness signal where the work didn’t come with one. He’s building the missing piece of scaffolding by hand.</p>



<p>That gives you a simple lens to run any candidate use case through. Does the work break into discernible stages? Can you observe what happens at each one? Is there a usable signal for whether the result was right? Score high on all three and you have a software-engineering-shaped problem, so go now. Score low and you have a choice: build the substrate first or wait. What you shouldn’t do is fund it at scale and hope the ROI materializes, because that’s the pile the 19% number is built on.</p>



<h2 class="wp-block-heading">The hard part, and the honest caveat</h2>



<p>Run the professions through that lens and they sort themselves. Finance is the closest cousin to software. Reconciliation, close processes, approval chains and audit trails already give you staged work with a clear “it reconciles or it doesn’t” signal, which is part of why financial services sits among the sectors furthest along with AI. Legal and medicine are harder. The workflow shell exists, intake to redline to filing, diagnosis to treatment to follow-up, but the correctness signal at the core is weak, delayed or confounded. You can automate the routine staged parts and you hit a wall at the judgment that defines the profession.</p>



<p>And that’s the caveat that keeps this honest. A structured workflow isn’t always buildable in software’s image. For the judgment core of some professions, the substrate is years out no matter how good the model gets or how mature your governance becomes. Anyone selling you a tighter timeline than that is selling.</p>



<p>But notice what this reframe does. It turns “our AI ROI is elusive” from a mystery you wait out into a sequencing-and-instrumentation problem you can actually test. Your timeline isn’t set by how smart the next model is. It’s set by how fast you build the substrate for your own domain, and that’s within your control.</p>



<h2 class="wp-block-heading">The two numbers, reconciled</h2>



<p>Put the 80% and the 19% back next to each other and they stop looking like a paradox. Software engineering didn’t win because its models were better than everyone else’s. It won because the work was already shaped to let an agent succeed, and the scaffolding that makes agent output trustworthy had been in place for decades before the agent showed up.</p>



<p>The question for the rest of the enterprise was never really whether AI can do the work. It’s whether your work is shaped so AI’s output can be trusted. That’s not something you wait for. It’s something you build.</p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><a href="https://www.cio.com/expert-contributor-network/"><strong>Want to join?</strong></a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[13 in-demand IT security certifications for higher pay]]></title>
<description><![CDATA[With change a constant, cybersecurity professionals looking to improve their careers can benefit from the latest insights into employers’ needs. Data from Foote Partners on the skills and certification most in demand today may provide helpful signposts.



Analyzing more than 660 certifications a...]]></description>
<link>https://tsecurity.de/de/3653485/it-security-nachrichten/13-in-demand-it-security-certifications-for-higher-pay/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3653485/it-security-nachrichten/13-in-demand-it-security-certifications-for-higher-pay/</guid>
<pubDate>Wed, 08 Jul 2026 09:08:38 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>With change a constant, cybersecurity professionals looking to improve their careers can benefit from the latest insights into employers’ needs. Data from Foote Partners on the skills and certification most in demand today may provide helpful signposts.</p>



<p>Analyzing more than <a href="https://footepartners.com/pages/report-skills-certs">660 certifications</a> as part of its 2Q 2026 “IT Skills Demand and Pay Trends Report,” Foote Partners calculated the most valuable IT security certifications to pursue right now based on two dimensions. The first, the <a href="https://www.cio.com/article/350363/pay-for-in-demand-it-skills-rises-fastest-in-14-years.html">average pay premium</a>, measures the difference in pay between IT pros with a particular credential and those without it. The second, market value increase, measures the increase in pay gains over the past six months.</p>



<p>Together, average pay premium and market value increase can give cybersecurity pros a starting point in deciding which certification to pursue for more pay. Apart from considering their overall professional goals, security professionals should consider each certification’s training and exam costs, whether vendor-specific or vendor-neutral, and the lateral or vertical role opportunities it may open.</p>



<p>Here are the top 13 certifications paying higher premiums today in descending order.</p>



<h2 class="wp-block-heading">GIAC Security Expert (GSE)</h2>



<p>The <a href="https://www.giac.org/get-certified/giac-portfolio-certifications">GIAC Security Expert</a> (GSE) portfolio certification is for security leaders wishing to prove their status as a top information security practitioner by showing they have offensive and defensive skills and hands-on practical skills. Available for more than 15 years, the GSE is considered one of the broadest and deepest cybersecurity certifications. To earn the certification, candidates must complete any six <a href="https://www.giac.org/get-started/practitioner">practitioner</a> certifications and any four <a href="https://www.giac.org/get-started/applied-knowledge">applied knowledge</a> certifications.</p>



<p>GIAC allows candidates to customize the certification to fit their expertise and career. Candidates can also build their certification over any amount of time as along as the required certifications within the portfolio remain active. Practitioner certification exams are 2-5 hours in length, depending on the specific certification attempt, and applied knowledge certification exams are 4 hours in length.</p>



<p><strong>Training fees:</strong> Some training is offered in affiliation with SANS Institute and costs $8,780.</p>



<p><strong>Exam Fees:</strong> Because you need 10 certifications to achieve the GSE <a href="https://www.giac.org/pricing">prices vary significantly</a>. If you already hold a GIAC Certified Forensic Analyst (GCFA), the cost of one of the required certifications drops from $1,299 to $499. Most required certifications are priced at either $999 or $1,299 per attempt, though they can cost up to $11,190.</p>



<h2 class="wp-block-heading">GIAC Security Professional (GSP)</h2>



<p>The <a href="https://www.giac.org/get-certified/giac-portfolio-certifications">GIAC Security Professional (GSP)</a> is designed to demonstrate the holder’s depth and breadth of information security knowledge. Launched approximately two years, this newer certification is the halfway point to the GSE. Customization of the certification is allowed, and to achieve it a candidate must complete any three <a href="https://www.giac.org/get-started/practitioner">practitioner</a> certifications and any two <a href="https://www.giac.org/get-started/applied-knowledge">applied knowledge</a> certifications. Candidates can also build their certification over any amount of time as along as the required certifications within the portfolio remain active. Practitioner Certification exams are 2-5 hours in length, depending on the specific certification attempt, and Applied Knowledge Certification exams are 4 hours in length.</p>



<p><strong>Training fees:</strong> Some training is offered in affiliation with SANS Institute and costs $8,780.</p>



<p><strong>Exam Fees:</strong> Because you need five certifications to achieve the GSP <a href="https://www.giac.org/pricing">prices vary significantly</a>. If you already hold a GIAC Security Essentials (GSEC), the cost of one of the required certifications drops from $1,299 to $499. Most certifications required are priced at either $999 or $1,299 per attempt, though certification can cost up to $5,595.</p>



<h2 class="wp-block-heading">Microsoft Certified Azure Cybersecurity Architect Expert</h2>



<p>Those who earn the <a href="https://learn.microsoft.com/en-us/credentials/certifications/cybersecurity-architect-expert/">Microsoft Certified: Cybersecurity Architect Expert</a> credential are able to translate a cybersecurity strategy into capabilities that protect the assets, business, and operations of an organization. Through the certification process, candidates learn to design, guide the implementation of, and maintain security solutions that follow zero-trust principles and best practices. You’ll also be able to design solutions for governance, risk, and compliance (GRC), security operations, and security posture management.​</p>



<p>As a prerequisite, candidate must have earned one of the following: <a href="https://learn.microsoft.com/en-us/credentials/certifications/azure-security-engineer/">Microsoft Certified: Azure Security Engineer Associate</a>, <a href="https://learn.microsoft.com/en-us/credentials/certifications/identity-and-access-administrator/">Microsoft Certified: Identity and Access Administrator Associate</a>, <a href="https://learn.microsoft.com/en-us/credentials/certifications/security-operations-analyst/">Microsoft Certified: Security Operations Analyst Associate</a> certification.</p>



<p><strong>Training fees: </strong>Self-paced training is available from the course’s page and free of charge. There is also an option to find an instructor-led training with pricing starting at $1,300.</p>



<p><strong>Exam Fees:</strong> The exam costs $165 and Microsoft offers free practice assessments.</p>



<h2 class="wp-block-heading">Certificate of Cloud Security Knowledge (CCSK)</h2>



<p>As a certificate and not a certification — an important distinction — the Cloud Security Alliance (CSA) positions its <a href="https://cloudsecurityalliance.org/education/ccsk">Certificate of Cloud Security Knowledge</a> as the foundation for future credentials and upskilling in the sector. From this perspective, the CCSK is helpful for cybersecurity analysts, compliance managers, security engineers, architects, and administrators. This vendor-neutral certificate has been recently updated and covers topics in zero trust, DevSecOps, cloud telemetry and security analytics, artificial intelligence, and more. CCSK offers a variety of training modalities, including an exam prep kit, instructor-led classes offered virtually and in person, and an online self-paced option. Candidates must score at least 80% on the exam, randomly pulling 60 multiple-choice questions from a test bank.</p>



<p><strong>Training fees:</strong> Prices vary based on modality. A self-paced course<a href="https://cloudsecurityalliance.org/education/ccsk#preparing-for-the-ccsk"> and exam bundle costs $795</a>, and online, instructor-led training begins at<a href="https://cloudsecuritypass.com/training/"> </a><a href="https://cloudsecuritypass.com/training/">$995</a>.</p>



<p><strong>Exam fees:</strong> The exam costs $445, though discounts are<a href="https://cloudsecurityalliance.org/membership"> available for corporate members</a>, and<a href="https://cloudsecurityalliance.org/education/ccsk/free-for-veterans"> </a><a href="https://cloudsecurityalliance.org/education/ccsk/free-for-veterans">US military veterans can take it for free</a>.</p>



<h2 class="wp-block-heading">Certified in Risk and Information Systems Control (CRISC)</h2>



<p>Administered by ISACA, the<a href="https://www.isaca.org/credentialing/crisc"> </a><a href="https://www.csoonline.com/article/571249/crisc-certification-your-ticket-to-the-c-suite.html">Certified in Risk and Information Systems Control</a> certification provides candidates with training across four domains: corporate IT governance, risk assessment, risk response and reporting, and technology and security. CRISC is ideal for candidates who want to enhance and optimize business resilience and risk management across their organization. The exam consists of 150 questions across the four domains. Since ISACA began offering CRISC in 2010, more than 23,000 people have obtained the certification. ISACA claims 52% of certificate holders experienced on-the-job improvement, and CRISC is the “4th top-paying certification worldwide.” To qualify for CRISC, candidates must adhere to a code of professional ethics and have <a href="https://support.isaca.org/s/article/What-are-the-requirements-to-become-CRISC-certified">three years of work experience</a> in risk assessment and risk response and reporting. On passing the exam, candidates must submit 20 CPE credits annually and<a href="https://www.isaca.org/-/media/files/isacadp/project/isaca/certification/crisc/crisc-cpe/crisc-cpe-policy.pdf"> </a><a href="https://www.isaca.org/-/media/files/isacadp/project/isaca/certification/crisc/crisc-cpe/crisc-cpe-policy.pdf">120 continuing professional education (CPE) hours</a> every three years to maintain their CRISC.</p>



<p><strong>Training fees:</strong> ISACA offers three resources: an<a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000004Km4PEAS"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2SVQ000001VR1l2AG">online review course</a>, $895; a review manual in<a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000004Tx3aEAC"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2SVQ000001FWgY2AW">print</a> or<a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000004Tx60EAC"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2SVQ000001FoOv2AK">digital</a>, $139; and an<a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000004Ko5TEAS"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2SVQ000001IPKL2A4">annual subscription to a 833-question test bank</a>, $399. Discounts are available for ISACA members.</p>



<p><strong>Exam fees: </strong>$575, ISACA members; $760 for non-members; plus $50 application fee.</p>



<h2 class="wp-block-heading">Certified Information Systems Auditor (CISA)</h2>



<p>The Information Systems Audit and Control Association (ISACA)’s CISA is geared toward IT auditors who wish to upskill or earn a pay boost. According to ISACA, 70% of CISA holders report on-the-job improvement, and another 22% receive a raise. The course covers five domains: information systems auditing, implementation, and operations; protection of information assets; and IT governance. The<a href="https://www.isaca.org/-/media/files/isacadp/project/isaca/certification/exam-candidate-guides/2024/exam-candidate-guide-2024.pdf"> </a>four-hour exam consists of 150 multiple-choice questions, and candidates must earn 450 on ISACA’s scaled scoring system, with 800 representing a perfect score. To<a href="https://www.isaca.org/credentialing/cisa/maintain-cisa-certification"> </a><a href="https://www.isaca.org/credentialing/cisa/maintain-cisa-certification">maintain their CISA</a>, certification holders must take 20 CPE credits annually and 120 over three years through conferences, volunteering, on-demand learning, and other methods as well as paying maintenance fee. To qualify, you must have five years of experience in IT or IS audit, control, assurance, or security. You can apply for an experience waiver for up to three years.</p>



<p><strong>Training fees:</strong> ISACA offers four resources: an<a href="https://store.isaca.org/s/store#/store/browse/detail/a2SVQ000000Fqvx2AC"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2SVQ000000Fqvx2AC">online review course</a> for $895, an<a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000008KxGWEA0"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000008KxGWEA0">annual subscription to a question bank</a> for $399, and a print or digital<a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000004W2rOEAS"> </a><a href="https://store.isaca.org/s/store#/store/browse/detail/a2S4w000004W2rOEAS">review manual</a> for $139. Discounts are available for ISACA members. </p>



<p><strong>Exam fees:</strong> $575, members; $760, non-members; plus $50 application fee.</p>



<h2 class="wp-block-heading">Certified Information Systems Security Professional (CISSP)</h2>



<p><a href="https://www.csoonline.com/article/570239/cissp-certification-requirements-training-and-cost.html">CISSP</a> is a generalist cert from ISC2 aimed at security pros who have already established a strong track record. Advanced-level analysts interested in getting CISSP certified will need to know all the ins and outs of security and risk management, asset security, operations, security assessment and testing, and more. The CISSP certification requires five years of full-time experience in at least two of its <a href="https://www.isc2.org/certifications/cissp#The%20CISSP%20Exam">eight domains</a>. The exam is <a href="https://www.isc2.org/Certifications/CISSP/CISSP-CAT">adaptive</a>, ranging from 100 to 150 questions, including multiple-choice and advanced items of varying formats. Candidates need to score 700 points out of 1,000 to pass the exam.</p>



<p><strong>Training fees:</strong><a href="https://www.isc2.org/training/online-self-paced/cissp-online-self-paced"> </a>Online self-paced training <a href="https://www.isc2.org/training#CISSP">fees start</a> at $595 and can cost up to $1,993;<a href="https://www.isc2.org/training/online-instructor-led/cissp-online-instructor-led"> </a>online instructor-led bootcamp costs $2,880.</p>



<p><strong>Exam fee:</strong><a href="https://www.isc2.org/register-for-exam/isc2-exam-pricing"> </a><a href="https://www.isc2.org/register-for-exam/isc2-exam-pricing">$749</a></p>



<h2 class="wp-block-heading">Certified Secure Software Lifecycle Professional (CSSLP)</h2>



<p>This ISC2 certification helps cyber pros build their career by training them to better incorporate security practices throughout software development phases. The <a href="https://www.isc2.org/certifications/csslp">CSSLP</a> exam evaluates experience across eight domains: secure software concepts; secure software; lifecycle management; secure software requirements; secure software architecture and design; secure software implementation; secure software testing; secure software deployment, operations, maintenance; secure software supply chain. Those wishing to acquire the CSSLP must have four years of paid work experience as a software development lifecycle professional in one or more of the eight domains.</p>



<p><strong>Training fees:</strong><a href="https://www.isc2.org/training/online-self-paced/cissp-online-self-paced"> </a>Online self-paced training <a href="https://www.isc2.org/training#CSSLP">fees start</a> at $550 and can cost up to $1,718; online instructor-led bootcamp costs $2,650.</p>



<p><strong>Exam fee:</strong> <a href="https://www.isc2.org/register-for-exam/isc2-exam-pricing">$599</a></p>



<h2 class="wp-block-heading">Check Point Certified Security Master (CCSM)</h2>



<p>To become a <a href="https://www.checkpoint.com/services/training/certification-program/">Check Point Certified Security Master (CCSM) </a>security professionals must have an active Certified Security Expert (CCSE) and mast have completed two subsequent Check Point Specialist accreditations. CCSM validates advanced expertise in configuring, deploying, and troubleshooting Check Point solutions. Check Point certifications are valid for 24 months.</p>



<p><strong>Training fees:</strong><a href="https://www.isc2.org/training/online-self-paced/cissp-online-self-paced"></a> <a href="https://securityservices.checkpoint.com/categories/trainingprograms">Training for CCSE</a> is $3,500</p>



<p><strong>Exam fee:</strong> The fee for CCSE is $300</p>



<h2 class="wp-block-heading">GIAC Experienced Cybersecurity Specialist (GX-CS)</h2>



<p>The <a href="https://www.giac.org/certifications/experienced-cyber-security-gxcs">Experienced Cybersecurity Specialist (GX-CS)</a> sits within the applied knowledge certifications with GIAC. The certification is for practitioners to show their qualifications for advanced, hands-on IT systems roles across cybersecurity. Its intent is to demonstrate the candidate can navigate evolving real-world threats. The certification covers five areas: network security analysis and tools; evaluation of Windows and Linux OS security; advanced security tools and techniques; common attacks and defenses; and implementing overall cybersecurity and information security. The GX-CS is for <a href="https://www.giac.org/certifications/security-essentials-gsec">GSEC</a> holders who acquired additional experience — the GSEC exam costs $999, and SANS Institute offers <a href="https://www.sans.org/cyber-security-courses/security-essentials">training</a> for GSEC.</p>



<p><strong>Training fees:</strong><a href="https://www.isc2.org/training/online-self-paced/cissp-online-self-paced"></a> There are a few related affiliate training programs provided by SANS, each costing approximately $9,000.</p>



<p><strong>Exam fee: </strong>$499 for those with an active GSEC; otherwise <a href="https://www.giac.org/pricing">$1,299</a>.</p>



<h2 class="wp-block-heading">OffSec Certified Professional (OSCP+)</h2>



<p>To earn the<a href="https://www.offsec.com/courses/pen-200/"> </a><a href="https://www.offsec.com/courses/pen-200/">OffSec Certified Professional</a> certification, candidates must complete the affiliated course, PEN-200: Penetration Testing with Kali Linux, and pass the subsequent exam. The course covers 20 plus modules, including information gathering, vulnerability scanning, encryption and cryptography, Active Directory and AWS exploitation, and more. Certificate holders will have shown mastery of penetration testing methodologies ideal for new roles, such as an ethical hacker, incident responder, or threat hunter. The OSCP+ exam is entirely hands-on, and test-takers must compromise systems within a lab environment.</p>



<p>OffSec does not enforce any prerequisites but recommends candidates be familiar with TCP/IP networking, scripting in Bash and Python, and Linux and Windows, which they can learn through its<a href="https://www.offsec.com/learning/paths/network-penetration-testing-essentials/"> </a><a href="https://www.offsec.com/learning/paths/network-penetration-testing-essentials/">Network Penetration Testing Essentials Learning Path</a>.</p>



<p><strong>Training and exam fees:</strong> OffSec bundles the course and exam for $1,749 and as a yearly subscription that includes access to one 200 or 300-level course, the associated labs, and two exam attempts for $2,749 annually.</p>



<h2 class="wp-block-heading">OffSec Experienced Penetration Tester (OSEP)</h2>



<p>The<a href="https://www.offsec.com/courses/pen-300/"> </a><a href="https://www.offsec.com/courses/pen-300/">OffSec Experienced Penetration Tester</a> is ideal for penetration testers and ethical hackers who need more advanced techniques to sharpen offensive skills against modern enterprise defenses. Across more than 20 modules, the certification introduces these professionals to advanced offensive techniques, EDR and AV evasion, advanced Windows offensive security and more. During the two-day proctored exam, professionals must connect to a lab environment via a VPN and compromise multiple machines within a network through several possible attack paths. To pass, professionals must achieve the objective stated within the control panel or score at<a href="https://help.offsec.com/hc/en-us/articles/360049781352-OSEP-Exam-FAQ"> </a><a href="https://help.offsec.com/hc/en-us/articles/360049781352-OSEP-Exam-FAQ">least 100 points</a> — 10 points are awarded for every flag found in a local.txt or proof.txt file. Professionals who earn their OSEP can also obtain their<a href="https://www.offsec.com/certificates/osce3/"> </a><a href="https://www.offsec.com/certificates/osce3/">OSCE³ Certification</a> to demonstrate their mastery of offensive security. They would also need to pass the exams for WEB-300: Advanced Web Attacks and Exploitation and EXP-301: Windows User Mode Exploit Development, after which the OSCE³ is automatically awarded.</p>



<p>While there are no formal prerequisites for OSEP, OffSec recommends candidates take the<a href="https://www.offsec.com/courses/pen-200/"> </a><a href="https://www.offsec.com/courses/pen-200/">PEN-200: Penetration Testing</a> with Kali Linux or have a strong foundation in operating systems, networking, and scripting. </p>



<p><strong>Training and exam fees:</strong> OffSec bundles the course and exam for $1,749, and as a yearly subscription that includes access to one 200 or 300-level course, the associated labs, and two exam attempts for $2,749 annually.</p>



<h2 class="wp-block-heading">OffSec Exploitation Expert (OSEE)</h2>



<p>OffSec’s <a href="https://www.offsec.com/courses/exp-401/">Offensive Security Exploitation Expert</a> is a vendor-specific certification, focusing on advanced Windows exploitation, with OffSec deeming it its most challenging certification. As a penetration testing course, the material dives deep into topics such as advanced heap manipulations and disarming WDEG mitigations. Certificate holders can identify problematic code in Windows operating systems and develop exploits. For the practical exam, candidates must complete a comprehensive penetration test of software and create an exploit within a lab environment — all within 72 hours. To qualify, you must have experience debugging, developing Windows exploits, and using the following technologies: WinDBG, x86_64, IDA Pro, and basic C/C++ programming. OffSec recommends completing its<a href="https://www.offsec.com/courses-and-certifications/"> </a><a href="https://www.offsec.com/courses-and-certifications/">300-level certifications</a> before OSEE.</p>



<p><strong>Training and exam fees:</strong> OffSec offers only instructor-led, in-person training. Enterprises should <a href="https://www.offsec.com/organizations/live-training/">inquire for more information</a>.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic shines a light into the Claude AI black hole]]></title>
<description><![CDATA[Anthropic has found a way to shed new light on how its models solve problems, thanks to its discovery of what it has dubbed the J-space. 



“We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role. ...]]></description>
<link>https://tsecurity.de/de/3653231/ai-nachrichten/anthropic-shines-a-light-into-the-claude-ai-black-hole/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3653231/ai-nachrichten/anthropic-shines-a-light-into-the-claude-ai-black-hole/</guid>
<pubDate>Wed, 08 Jul 2026 06:33:50 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Anthropic has found a way to shed new light on how its models solve problems, thanks to its discovery of what it has dubbed the J-space. </p>



<p>“We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role. We call the collection of these patterns the J-space, named after the technique we used to find them, involving a mathematical concept called the <a href="https://www.sciencedirect.com/topics/engineering/jacobian-matrix" target="_blank" rel="noreferrer noopener">Jacobian</a>,” <a href="https://www.anthropic.com/research/global-workspace" target="_blank" rel="noreferrer noopener">Anthropic said in its post</a> about the discovery. It examines the contents of the J-space using what it calls the Jacobian lens, or J-lens.</p>



<p>“Each J-space pattern is linked to a particular word,” Anthropic said. “But when one of these patterns lights up, it doesn’t mean the model is saying that word, just that the word is on its ‘mind.’ If you’ve heard of language models having a scratchpad or chain of thought—text they write to themselves while reasoning—the J-space is something different. It operates silently, in the model’s internal neural activations, allowing the model to ‘think’ about a concept without writing it down.”</p>



<p>This new level of analytical visibility goes well beyond what Anthropic announced as an <a href="https://www.computerworld.com/article/3628817/anthropics-llms-cant-reason-but-think-they-can-even-worse-they-ignore-guardrails.html" target="_blank">internal scratchpad for its models in 2024</a>. That scratchpad revealed what the model was considering when preparing an action or delivering an answer. The new development instead focuses on something much deeper which has the potential to change how AI systems are evaluated and purchased. </p>



<p>One example in <a href="https://transformer-circuits.pub/2026/workspace/index.html" target="_blank" rel="noreferrer noopener">the paper</a> described how some models did not engage in improper behavior during tests, which would appear to be a very favorable result. But the contents of the J-space revealed that the model sometimes <em>knew </em>that it was being tested, and that awareness might have been the key reason it declined to engage in the problematic behavior, much in the way human children act when they know they are being watched. </p>



<p>“Anthropic built a lens that catches its own model quietly noticing it’s being tested, faking a result to look good, spotting a prompt injection, or sitting on a planted goal it hasn’t acted on yet,” said <a href="https://zenity.io/authors/rock-lambros" target="_blank" rel="noreferrer noopener">Rock Lambros</a>, director of AI standards and governance at AI agent vendor Zenity. “Some of that good behavior rode on the model knowing it was on stage.”</p>



<p>Customers should read their safety benchmarks with that in mind, he said. “Fitness for your project still comes from testing on your own data and your own attackers, not from a leaderboard the model knew it was sitting for.” </p>



<p>That kind of visibility is a potentially crucial tool for CIOs.</p>



<p>“A provider that can catch its own model misbehaving in silence, then publish [those results], is telling you something real about its assurance maturity. Put that in your due diligence, not just your newsfeed,” Lambros noted. “Here’s the question I’d hand every model vendor now: what can you see inside your model that I can’t see in its output, and what have you caught?”</p>



<p>Added <a href="https://www.linkedin.com/in/noah-m-kenney-27499a166/" target="_blank" rel="noreferrer noopener">Noah Kenney</a>, principal consultant at Digital 520: “A model that behaves better because it knows it is being watched is not a safe model. It is a model with a poker face. We have to question every red team result, every internal pilot where the model refused something dangerous, and every ‘we tested this and it was fine’ story, because they now carry an asterisk.”</p>



<p>CIOs need to now determine whether an agent performed a function in a specific way because that is how it will always perform, or whether it was it behaving differently because it figured out you were just testing it, Kenney said. “The answer to that question should change your interpretation in a material way.”</p>



<h2 class="wp-block-heading">No J-lens for customers – yet</h2>



<p>“It is an admission that the industry’s evaluation regime is measuring something less durable than everyone assumed, and now the other frontier labs have to answer whether their own evaluations have the same problem,” Kenney said. “For CIOs, the paper is a warning about their entire model risk framework.”</p>



<p><a href="https://www.linkedin.com/in/fvillanustre/" target="_blank" rel="noreferrer noopener">Flavio Villanustre</a>, CISO for the LexisNexis Risk Solutions Group, said that examining the J-space can even help make models more efficient.</p>



<p>“It gives you the ability to introspect into the model and, as such, can be very useful to the user, especially in cases where explainability is important. Think regulated environments that require explainable responses and full causal analysis of them,” Villanustre said. “This can also be very helpful to users trying to fine tune their prompts, making models more efficient to optimize token cost.”</p>



<p>But currently indirect access, or future access achieved via AI vendor negotiations, is the only path for accessing the new information, though Villanustre noted that some enterprises could gain direct access to J-space by paying for <a href="https://www.cio.com/article/4167981/anthropics-financial-agents-expose-forward-deployed-engineers-as-new-ai-limiting-factor.html" target="_blank">Anthropic’s FDE program</a>. </p>



<p>“It is very useful to CIOs,” he pointed out, “but in order to make use of the capabilities offered by analysis of the J-space, they need appropriate talent that can make sense of it. The type of skills required go far beyond those of the general data analyst, or even data scientist.” </p>



<p>Today, said <a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="noreferrer noopener">Aman Mahapatra</a>, chief strategy officer for Tribeca Softtech, a New York City-based technology consulting firm, “enterprise customers cannot enable the Jacobian lens, cannot inspect the residual stream through the API, and cannot run the ablation studies that produced the most interesting findings in the paper.” </p>



<p>So, he said, “on the narrow question of whether a CIO can operationally use J-space monitoring in Q3 of this year to gate a production deployment, the answer is no.”</p>



<p>But Mahapatra argued that there are going to be other ways to access the information, and CIOs must insist on them.</p>



<p><strong>“</strong>Without customer-side access, this reduces to trusting Anthropic yet again, and that is exactly why enterprises should start pushing for a different assurance model industry-wide,” he said. “Model providers are converging on a posture where they inspect their own models using proprietary tooling and publish reassuring research about what they found. That is not an assurance framework any regulated industry accepts from any other vendor.”</p>



<p>He pointed out that banks do not accept “we validated our own model, trust us” from a credit scoring vendor, not does the healthcare industry accept it from a clinical decision support vendor. “There is no principled reason to accept it from a foundation model vendor either, and the J-space research crystallizes why,” he said.</p>



<h2 class="wp-block-heading">New visibility demands</h2>



<p>“The right long-term enterprise posture is to demand independent interpretability access, either through customer-facing APIs, through independent third-party auditors with privileged access, or through open interpretability standards that let a bank’s model risk management team apply the same tooling the vendor’s own safety team uses,” Mahapatra stressed. “None of that exists today. All of it should be on the roadmap CIOs are pushing for, and this research is the strongest argument yet for why.”</p>



<p>In fact, the discoveries in the research have the potential to fundamentally rewrite the AI strategy rules.</p>



<p>Mahapatra said that the single hardest problem in enterprise agentic deployment is verifying that an autonomous system’s stated reasoning matches its actual reasoning. “Until now, we could only audit what the model writes, while much of its reasoning happened silently. The J-lens attacks that gap head-on,” he noted.</p>



<p>Thus, he said, sophisticated buyers should start asking model providers during the procurement process about the interpretability tooling they offer to let customers monitor internal model state for deception, evaluation-gaming, and goal misalignment in their specific deployments.</p>



<p>“Almost no vendor can answer that today,” he said. “The CIOs who start requiring internal-state observability as a procurement criterion, even before the tooling is fully mature, will be the ones who shape how their vendors productize it, and the ones with genuine assurance when regulators start asking how they know their autonomous agents are actually doing what they claim.”</p>



<h2 class="wp-block-heading">The beginning of standards</h2>



<p>Another way that CIOs can benefit from this new visibility into Claude is to try and get that information from third-parties that already have access. The report, for example, noted that a Google AI specialist independently replicated some findings on an open-weight model.</p>



<p>That, noted <a href="https://www.linkedin.com/in/lewiscarhart/" target="_blank" rel="noreferrer noopener">Lewis Carhart</a>, CEO of Comp AI, a software development firm, “is a competitor verifying the method, not just the vendor’s own claim. It shows what’s technically possible, but it doesn’t give enterprises a way to check anything themselves.”</p>



<p>He said that it’s a pattern that compliance has seen before; SOC 2 didn’t start as an independent audit standard either. It started as vendors describing their own controls, and the market spent years building the infrastructure to verify those claims externally.</p>



<p>“Interpretability is at that same starting point now,” he noted. “It becomes meaningful for CIOs once J-lens findings show up in third-party audits, published model cards, or regulator-facing disclosures. Anything a risk team can point to that isn’t just the vendor’s word.”</p>



<h2 class="wp-block-heading">Leads to AI strategy changes</h2>



<p><a href="https://acceligence.com/talent/profiles/justin-greis/" target="_blank" rel="noreferrer noopener">Justin Greis</a>, CEO of consulting firm Acceligence, said he also expects this development to lead to major AI strategy changes. </p>



<p>“I can easily imagine governance platforms consuming those signals alongside prompts, outputs, identity information, policy decisions, and tool activity,” he said. “A future AI control plane could continuously evaluate whether an agent recognized an attempted prompt injection, understood that sensitive information was involved, detected conflicting objectives, or showed evidence that it was reasoning toward an unsafe action before that action was ever executed. Those signals become inputs into policy enforcement, human escalation, audit logging, and trust scoring across enterprise AI environments.”</p>



<p>This has practical implications for CIOs today, he pointed out, “because it changes how they evaluate AI vendors. A year ago, enterprises primarily asked about model accuracy, latency, security, and cost. Increasingly, procurement teams will also ask how much operational visibility vendors provide into agent behavior, reasoning quality, policy compliance, safety monitoring, and auditability.”</p>



<p>Mahapatra added that all of this could give CIOs a powerful new negotiating tactic. </p>



<p>“The renewal path is where the leverage actually sits: write contractual rights to interpretability reporting and third-party audit access into the next renewal, because those terms are free today and expensive after signature,” he said. “The CIOs who win on assurance in 2027 will be the ones who stopped accepting ‘trust us’ from their model provider in 2026 and put the right clauses in the paperwork while the vendor still needed the deal more than the customer needed the model.”</p>



<p><em>This article originally appeared on <a href="https://www.cio.com/article/4194145/anthropic-shines-a-light-into-the-claude-ai-black-hole.html" target="_blank">CIO.com</a>.</em></p>



<p></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic shines a light into the Claude AI black hole]]></title>
<description><![CDATA[Anthropic has found a way to shed new light on how its models solve problems, thanks to its discovery of what it has dubbed the J-space. 



“We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role. ...]]></description>
<link>https://tsecurity.de/de/3653228/it-nachrichten/anthropic-shines-a-light-into-the-claude-ai-black-hole/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3653228/it-nachrichten/anthropic-shines-a-light-into-the-claude-ai-black-hole/</guid>
<pubDate>Wed, 08 Jul 2026 06:33:02 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Anthropic has found a way to shed new light on how its models solve problems, thanks to its discovery of what it has dubbed the J-space. </p>



<p>“We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role. We call the collection of these patterns the J-space, named after the technique we used to find them, involving a mathematical concept called the <a href="https://www.sciencedirect.com/topics/engineering/jacobian-matrix" target="_blank" rel="nofollow">Jacobian</a>,” <a href="https://www.anthropic.com/research/global-workspace" target="_blank" rel="nofollow">Anthropic said in its post</a> about the discovery. It examines the contents of the J-space using what it calls the Jacobian lens, or J-lens.</p>



<p>“Each J-space pattern is linked to a particular word,” Anthropic said. “But when one of these patterns lights up, it doesn’t mean the model is saying that word, just that the word is on its ‘mind.’ If you’ve heard of language models having a scratchpad or chain of thought—text they write to themselves while reasoning—the J-space is something different. It operates silently, in the model’s internal neural activations, allowing the model to ‘think’ about a concept without writing it down.”</p>



<p>This new level of analytical visibility goes well beyond what Anthropic announced as an <a href="https://www.computerworld.com/article/3628817/anthropics-llms-cant-reason-but-think-they-can-even-worse-they-ignore-guardrails.html" target="_blank">internal scratchpad for its models in 2024</a>. That scratchpad revealed what the model was considering when preparing an action or delivering an answer. The new development instead focuses on something much deeper which has the potential to change how AI systems are evaluated and purchased. </p>



<p>One example in <a href="https://transformer-circuits.pub/2026/workspace/index.html" target="_blank" rel="nofollow">the paper</a> described how some models did not engage in improper behavior during tests, which would appear to be a very favorable result. But the contents of the J-space revealed that the model sometimes <em>knew </em>that it was being tested, and that awareness might have been the key reason it declined to engage in the problematic behavior, much in the way human children act when they know they are being watched. </p>



<p>“Anthropic built a lens that catches its own model quietly noticing it’s being tested, faking a result to look good, spotting a prompt injection, or sitting on a planted goal it hasn’t acted on yet,” said <a href="https://zenity.io/authors/rock-lambros" target="_blank" rel="nofollow">Rock Lambros</a>, director of AI standards and governance at AI agent vendor Zenity. “Some of that good behavior rode on the model knowing it was on stage.”</p>



<p>Customers should read their safety benchmarks with that in mind, he said. “Fitness for your project still comes from testing on your own data and your own attackers, not from a leaderboard the model knew it was sitting for.” </p>



<p>That kind of visibility is a potentially crucial tool for CIOs.</p>



<p>“A provider that can catch its own model misbehaving in silence, then publish [those results], is telling you something real about its assurance maturity. Put that in your due diligence, not just your newsfeed,” Lambros noted. “Here’s the question I’d hand every model vendor now: what can you see inside your model that I can’t see in its output, and what have you caught?”</p>



<p>Added <a href="https://www.linkedin.com/in/noah-m-kenney-27499a166/" target="_blank" rel="nofollow">Noah Kenney</a>, principal consultant at Digital 520: “A model that behaves better because it knows it is being watched is not a safe model. It is a model with a poker face. We have to question every red team result, every internal pilot where the model refused something dangerous, and every ‘we tested this and it was fine’ story, because they now carry an asterisk.”</p>



<p>CIOs need to now determine whether an agent performed a function in a specific way because that is how it will always perform, or whether it was it behaving differently because it figured out you were just testing it, Kenney said. “The answer to that question should change your interpretation in a material way.”</p>



<h2 class="wp-block-heading">No J-lens for customers – yet</h2>



<p>“It is an admission that the industry’s evaluation regime is measuring something less durable than everyone assumed, and now the other frontier labs have to answer whether their own evaluations have the same problem,” Kenney said. “For CIOs, the paper is a warning about their entire model risk framework.”</p>



<p><a href="https://www.linkedin.com/in/fvillanustre/" target="_blank" rel="nofollow">Flavio Villanustre</a>, CISO for the LexisNexis Risk Solutions Group, said that examining the J-space can even help make models more efficient.</p>



<p>“It gives you the ability to introspect into the model and, as such, can be very useful to the user, especially in cases where explainability is important. Think regulated environments that require explainable responses and full causal analysis of them,” Villanustre said. “This can also be very helpful to users trying to fine tune their prompts, making models more efficient to optimize token cost.”</p>



<p>But currently indirect access, or future access achieved via AI vendor negotiations, is the only path for accessing the new information, though Villanustre noted that some enterprises could gain direct access to J-space by paying for <a href="https://www.cio.com/article/4167981/anthropics-financial-agents-expose-forward-deployed-engineers-as-new-ai-limiting-factor.html" target="_blank">Anthropic’s FDE program</a>. </p>



<p>“It is very useful to CIOs,” he pointed out, “but in order to make use of the capabilities offered by analysis of the J-space, they need appropriate talent that can make sense of it. The type of skills required go far beyond those of the general data analyst, or even data scientist.” </p>



<p>Today, said <a href="https://www.linkedin.com/in/akm76/" target="_blank" rel="nofollow">Aman Mahapatra</a>, chief strategy officer for Tribeca Softtech, a New York City-based technology consulting firm, “enterprise customers cannot enable the Jacobian lens, cannot inspect the residual stream through the API, and cannot run the ablation studies that produced the most interesting findings in the paper.” </p>



<p>So, he said, “on the narrow question of whether a CIO can operationally use J-space monitoring in Q3 of this year to gate a production deployment, the answer is no.”</p>



<p>But Mahapatra argued that there are going to be other ways to access the information, and CIOs must insist on them.</p>



<p><strong>“</strong>Without customer-side access, this reduces to trusting Anthropic yet again, and that is exactly why enterprises should start pushing for a different assurance model industry-wide,” he said. “Model providers are converging on a posture where they inspect their own models using proprietary tooling and publish reassuring research about what they found. That is not an assurance framework any regulated industry accepts from any other vendor.”</p>



<p>He pointed out that banks do not accept “we validated our own model, trust us” from a credit scoring vendor, not does the healthcare industry accept it from a clinical decision support vendor. “There is no principled reason to accept it from a foundation model vendor either, and the J-space research crystallizes why,” he said.</p>



<h2 class="wp-block-heading">New visibility demands</h2>



<p>“The right long-term enterprise posture is to demand independent interpretability access, either through customer-facing APIs, through independent third-party auditors with privileged access, or through open interpretability standards that let a bank’s model risk management team apply the same tooling the vendor’s own safety team uses,” Mahapatra stressed. “None of that exists today. All of it should be on the roadmap CIOs are pushing for, and this research is the strongest argument yet for why.”</p>



<p>In fact, the discoveries in the research have the potential to fundamentally rewrite the AI strategy rules.</p>



<p>Mahapatra said that the single hardest problem in enterprise agentic deployment is verifying that an autonomous system’s stated reasoning matches its actual reasoning. “Until now, we could only audit what the model writes, while much of its reasoning happened silently. The J-lens attacks that gap head-on,” he noted.</p>



<p>Thus, he said, sophisticated buyers should start asking model providers during the procurement process about the interpretability tooling they offer to let customers monitor internal model state for deception, evaluation-gaming, and goal misalignment in their specific deployments.</p>



<p>“Almost no vendor can answer that today,” he said. “The CIOs who start requiring internal-state observability as a procurement criterion, even before the tooling is fully mature, will be the ones who shape how their vendors productize it, and the ones with genuine assurance when regulators start asking how they know their autonomous agents are actually doing what they claim.”</p>



<h2 class="wp-block-heading">The beginning of standards</h2>



<p>Another way that CIOs can benefit from this new visibility into Claude is to try and get that information from third-parties that already have access. The report, for example, noted that a Google AI specialist independently replicated some findings on an open-weight model.</p>



<p>That, noted <a href="https://www.linkedin.com/in/lewiscarhart/" target="_blank" rel="nofollow">Lewis Carhart</a>, CEO of Comp AI, a software development firm, “is a competitor verifying the method, not just the vendor’s own claim. It shows what’s technically possible, but it doesn’t give enterprises a way to check anything themselves.”</p>



<p>He said that it’s a pattern that compliance has seen before; SOC 2 didn’t start as an independent audit standard either. It started as vendors describing their own controls, and the market spent years building the infrastructure to verify those claims externally.</p>



<p>“Interpretability is at that same starting point now,” he noted. “It becomes meaningful for CIOs once J-lens findings show up in third-party audits, published model cards, or regulator-facing disclosures. Anything a risk team can point to that isn’t just the vendor’s word.”</p>



<h2 class="wp-block-heading">Leads to AI strategy changes</h2>



<p><a href="https://acceligence.com/talent/profiles/justin-greis/" target="_blank" rel="nofollow">Justin Greis</a>, CEO of consulting firm Acceligence, said he also expects this development to lead to major AI strategy changes. </p>



<p>“I can easily imagine governance platforms consuming those signals alongside prompts, outputs, identity information, policy decisions, and tool activity,” he said. “A future AI control plane could continuously evaluate whether an agent recognized an attempted prompt injection, understood that sensitive information was involved, detected conflicting objectives, or showed evidence that it was reasoning toward an unsafe action before that action was ever executed. Those signals become inputs into policy enforcement, human escalation, audit logging, and trust scoring across enterprise AI environments.”</p>



<p>This has practical implications for CIOs today, he pointed out, “because it changes how they evaluate AI vendors. A year ago, enterprises primarily asked about model accuracy, latency, security, and cost. Increasingly, procurement teams will also ask how much operational visibility vendors provide into agent behavior, reasoning quality, policy compliance, safety monitoring, and auditability.”</p>



<p>Mahapatra added that all of this could give CIOs a powerful new negotiating tactic. </p>



<p>“The renewal path is where the leverage actually sits: write contractual rights to interpretability reporting and third-party audit access into the next renewal, because those terms are free today and expensive after signature,” he said. “The CIOs who win on assurance in 2027 will be the ones who stopped accepting ‘trust us’ from their model provider in 2026 and put the right clauses in the paperwork while the vendor still needed the deal more than the customer needed the model.”</p>



<p></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[HPR4678: High Resolution Elapsed Time in Shell Scripts]]></title>
<description><![CDATA[This show has been flagged as Clean by the host.






01 Introduction






In this episode I will describe how to calculate elapsed time in bash or other shell scripts.


While this may sound like a very simple and basic thing to do, there is a slightly more complex aspect to it if you wi...]]></description>
<link>https://tsecurity.de/de/3652976/podcasts/hpr4678-high-resolution-elapsed-time-in-shell-scripts/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3652976/podcasts/hpr4678-high-resolution-elapsed-time-in-shell-scripts/</guid>
<pubDate>Wed, 08 Jul 2026 02:03:40 +0200</pubDate>
<category>🎥 Podcasts</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This show has been flagged as Clean by the host.</p>

<p>

</p>

<p>
01 Introduction</p>

<p>

</p>

<p>
In this episode I will describe how to calculate elapsed time in bash or other shell scripts.</p>

<p>
While this may sound like a very simple and basic thing to do, there is a slightly more complex aspect to it if you wish to calculate elapsed time to a higher resolution than one second. </p>

<p>

</p>

<p>
02</p>

<p>
There are many reasons for calculating elapsed time in a shell script.</p>

<p>
For example you may wish to simply report how long an operation took to run.</p>

<p>
Another reason may be that you are trying to speed up a script and need to calculate benchmark data to see how different alternative methods perform.</p>

<p>

</p>

<p>
03</p>

<p>
What may seem like a simple task gets a bit more complicated if you want to do it for multiple different operating systems even if they are all unix related, as we shall see.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
04 Operating Systems Tested</p>

<p>

</p>

<p>
For the purposes of this episode, I ran tests on the current version of the following operating systems.</p>

<p>

</p>

<p>
Alma</p>

<p>
Alpine</p>

<p>
Debian</p>

<p>
FreeBSD</p>

<p>
OpenBSD</p>

<p>
RaspberryPi</p>

<p>
OpenSuse</p>

<p>
Ubuntu 2604</p>

<p>

</p>

<p>
Alma is a close copy of Red Hat that we can take as representing Red Hat style distros. </p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
05 Simple Low Resolution Timing</p>

<p>

</p>

<p>
I will start with the simple and obvious method before describing the less obvious ones.</p>

<p>

</p>

<p>
This uses the date command to get the current time in seconds since the unix epoch. </p>

<p>
This is simply</p>

<p>
date '+%s'</p>

<p>

</p>

<p>
06</p>

<p>
Save this to a variable using whatever method you prefer.</p>

<p>
For example.</p>

<p>
starttime=$(date '+%s')</p>

<p>

</p>

<p>
07</p>

<p>
Next, do whatever operations it is you wish to time.</p>

<p>
Use the date command to get the current time again.</p>

<p>
endtime=$(date '+%s')</p>

<p>

</p>

<p>
08</p>

<p>
Now simply subtract the start time from the end time using shell arithmetic.</p>

<p>
This should be very obvious and basic.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
09 Higher Resolution Timing</p>

<p>

</p>

<p>
However, suppose we wish to measure time to greater than one second of precision. </p>

<p>
We need to do two things.</p>

<p>
The first is to obtain the current time at a higher degree of precision.</p>

<p>
The second is to conduct the calculations to a higher degree of precision. </p>

<p>

</p>

<p>
10</p>

<p>
Unfortunately, the standard time precision for POSIX shells seems to be 1 second.</p>

<p>
Some shells offer a higher precision, but others do not.</p>

<p>
Furthermore, standard shell arithmatic uses integer, which limits calculations to 1 second of precision.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
11 Bash High Resolution Shell Variable</p>

<p>

</p>

<p>
Fortunately, bash is one that does offer a high precision date.</p>

<p>

</p>

<p>
If you are using bash 5.0 or newer, there is a shell variable called EPOCHREALTIME which offers time since the the unix epoch (that is, since the first of January 1970, at 00:00:00 UTC) in seconds to 6 decimals of precision.</p>

<p>

</p>

<p>
12</p>

<p>
Example</p>

<p>
echo $EPOCHREALTIME</p>

<p>
1779634800.184926</p>

<p>

</p>

<p>
13</p>

<p>
This is related to the similar bash variable known as EPOCHSECONDS which gives the number of seconds since the unix epoch.</p>

<p>

</p>

<p>
14</p>

<p>
Example</p>

<p>
echo $EPOCHSECONDS</p>

<p>
1779634800</p>

<p>

</p>

<p>
15</p>

<p>
So if you are using bash, measuring time is very simple.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
16 But is it Really Bash?</p>

<p>

</p>

<p>
Is your script however actually using bash?</p>

<p>
Debian and derivatives actually have two shells.</p>

<p>
The first, the interactive shell is bash.</p>

<p>
The second, the non-interactive shell is dash, which stands for "debian almquist shell".</p>

<p>

</p>

<p>
17</p>

<p>
If you open a terminal, you get bash.</p>

<p>
If your script starts with a "bin/bash" shebang line, you get bash.</p>

<p>
However, if your script starts with a "bin/sh" shebang line, you get dash.</p>

<p>
Some people find themselves getting caught out by this one when they try something out in a terminal but find that it doesn't work in their script which started with "bin/sh".</p>

<p>

</p>

<p>
18</p>

<p>
Many other, but not all, Linux distros use bash for both the interactive and non-interactive shells, so "bin/sh" and "bin/bash" work the same with those ones.</p>

<p>

</p>

<p>
So if you intend to use bash, make sure your script calls for bash in the first line.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
19 The SHELL Variable</p>

<p>

</p>

<p>
So how can a script tell what shell it is running under?</p>

<p>
There is a shell variable called "SHELL" which will tell you the name of the shell.</p>

<p>
Well, sort of.</p>

<p>

</p>

<p>
20</p>

<p>
On Debian and derivatives "SHELL" will  say "bash" regardless of whether the actual shell is bash or dash.</p>

<p>
On some other operating systems "SHELL" will simply say "sh" even if it is something else entirely.</p>

<p>

</p>

<p>
So we need to do some additional levels of checking to see what we have. </p>

<p>

</p>

<p>
21</p>

<p>
To start with though, here's what each of the test distros reports for SHELL.</p>

<p>

</p>

<p>
Alma              : bash</p>

<p>
Alpine            : sh</p>

<p>
Debian            : bash</p>

<p>
FreeBSD           : sh</p>

<p>
OpenBSD           : ksh</p>

<p>
Raspberry-Pi      : bash</p>

<p>
Suse              : bash</p>

<p>
ubuntu2604        : bash</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
22 Bash Versus Dash</p>

<p>

</p>

<p>
First, let's try to see which ones are bash and which ones are dash.</p>

<p>

</p>

<p>
The first thing we can check is for the shell variable BASH_VERSION.</p>

<p>

</p>

<p>
23</p>

<p>
Example</p>

<p>
echo $BASH_VERSION</p>

<p>

</p>

<p>
If the shell is bash, then it will report a version string.</p>

<p>
If the shell is not bash, then it will return an empty value.</p>

<p>

</p>

<p>
24</p>

<p>
Using this test, we can see that Alma and Opensuse are indeed using bash.</p>

<p>

</p>

<p>
We however need to check Debian, Raspberry Pi, and Ubuntu when running in an "sh" script.</p>

<p>
To check this we can use the  "which" command to see what "sh" actually is.</p>

<p>

</p>

<p>
25</p>

<p>
Example</p>

<p>
echo $( ls -l $(which sh )  | rev | cut -d" " -f1 | cut -d/ -f1 | rev )</p>

<p>

</p>

<p>
26</p>

<p>
"which sh" shows us the path to "sh"</p>

<p>
However, that is a link so we need to use</p>

<p>
"ls -l" to find the actual executable.</p>

<p>
"rev" reverses the string.</p>

<p>

</p>

<p>
27</p>

<p>
"cut" takes the first element separated by spaces.</p>

<p>
The second </p>

<p>
"cut" takes the first element separated by the "/" characters.</p>

<p>
The final "rev" takes that string and reverses it again to get it in the correct order.</p>

<p>

</p>

<p>
28</p>

<p>
In the case of Debian, Raspberry Pi, and Ubuntu it tells us that this is "dash".</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
29 Openbsd</p>

<p>

</p>

<p>
Openbsd reports its shell as "ksh", which stands for Korn Shell. </p>

<p>
It is indeed Korn Shell, so we simply leave that one as is.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
30 Alpine and Freebsd</p>

<p>

</p>

<p>
Next we have Alpine Linux and Freebsd, which both report as "sh".</p>

<p>

</p>

<p>
In the case Freebsd there doesn't appear to be any further we can go that I am aware of.</p>

<p>
It's simple "sh".</p>

<p>
It is a basic POSIX shell which seems to be similar to the original unix shell, the Bourne Shell.</p>

<p>
Older versions of Freebsd used a different shell known as tsch (the C shell), but I haven't tested that so I will ignore that here.</p>

<p>

</p>

<p>
31</p>

<p>
With Alpine Linux however, we can get the actual shell using the same method that we used for Debian Linux.</p>

<p>
This reports as being "busybox".</p>

<p>

</p>

<p>
32</p>

<p>
Busybox is a limited shell intended for use in embedded systems.</p>

<p>
Alpine was originally an embedded distro, but some people started using it for containers.</p>

<p>
Alpine is Linux, but it is not GNU/Linux, and there are a number of areas which can trip you up if you are not aware of them.</p>

<p>
So, be extra careful if you are using it for anything, and test everything.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
33 Summary of Actual Shells</p>

<p>

</p>

<p>
Here is our revised list with the actual shell used when asking for "sh", so far as we can determine.</p>

<p>

</p>

<p>
Alma              : bash</p>

<p>
Alpine            : busybox</p>

<p>
Debian            : dash</p>

<p>
FreeBSD           : sh</p>

<p>
OpenBSD           : ksh</p>

<p>
Raspberry-Pi      : dash</p>

<p>
Suse              : bash</p>

<p>
ubuntu2604        : dash</p>

<p>

</p>

<p>
34</p>

<p>
There are other shells, but none of them are the default shell for any of the distros on our list, so I haven't tested them.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
35 Solutions for Measuring Time</p>

<p>

</p>

<p>
Now we need to find solutions for bash, dash, ksh, sh, and busybox.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
36 Bash</p>

<p>

</p>

<p>
For bash, we can simply use EPOCHREALTIME, as mentioned above.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
37 Dash</p>

<p>

</p>

<p>
For dash, we can use the date command.</p>

<p>
This is a very conventional method, and is probably the first answer that anyone would give for this situation.</p>

<p>
However, while it will work in most cases, it will not work in all cases, so it is not a universal solution.</p>

<p>

</p>

<p>
38</p>

<p>
To use date we simply call it with the correct format string.</p>

<p>
This uses %s to get seconds since the epoch, and %N to get nanoseconds of the current second.</p>

<p>
If you put a decimal separator between the two it will appear in the output.</p>

<p>
You can use the correct decimal separator for your locale, but I won't go into that here.</p>

<p>
Instead I will just assume a period or dot.</p>

<p>

</p>

<p>
39</p>

<p>
Example</p>

<p>
date '+%s.%N' </p>

<p>
1779634800.358916385</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
40 Problems with Date on Alpine and Openbsd</p>

<p>

</p>

<p>
Date will work for bash, dash, and sh on Freebsd.</p>

<p>
However it will not work for ksh on Openbsd, or for busybox on Alpine.</p>

<p>

</p>

<p>
41</p>

<p>
With busybox on Alpine, it simply ignores the %N format specifier and prints out the epoch in seconds only followed by the decimal separator.</p>

<p>
=</p>

<p>
Example</p>

<p>
date '+%s.%N'</p>

<p>
1779634800.</p>

<p>

</p>

<p>
42</p>

<p>
With ksh on Openbsd it prints the epoch in seconds followed by the decimal separator and then the %N as a literal N.</p>

<p>

</p>

<p>
date '+%s.%N'</p>

<p>
1779634800.N</p>

<p>

</p>

<p>
43</p>

<p>
Fortunately we have alternatives for these two cases.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
44 Openbsd</p>

<p>

</p>

<p>
Openbsd has the "ts" or timestamp utility installed by default.</p>

<p>
ts prints a time stamp in front of every line it receives from standard input.</p>

<p>
I won't go into details on all aspects of ts here, I'll leave that to someone else.</p>

<p>
Instead I will focus on how to use it for our specific purposes here.</p>

<p>

</p>

<p>
45</p>

<p>
We need to provide a format specifier to ts, which in this case is "%.s"</p>

<p>
We also need to provide something for standard input, or otherwise ts will simply sit there and wait for input.</p>

<p>
So what we need to do is to echo nothing through a pipe to ts while also giving ts the proper format specifier.</p>

<p>

</p>

<p>
46</p>

<p>
Example</p>

<p>
 echo | ts "%.s" </p>

<p>

</p>

<p>
47</p>

<p>
This will output the epoch time in seconds to six decimals of precision.</p>

<p>

</p>

<p>
ts is installed in Openbsd and Freebsd by default and can be used in either.</p>

<p>
It can also be installed in many other distros.</p>

<p>

</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
48 Busybox on Alpine</p>

<p>

</p>

<p>
None of the methods discussed so far will work for busybox on Alpine though.</p>

<p>
However there is a way, but it's a bit non obvious and somewhat hacky.</p>

<p>

</p>

<p>
49</p>

<p>
Busybox includes a command called "adjtimex".</p>

<p>
This is normally used to adjust the time hardware.</p>

<p>
However if it is run without arguments, it will report the current settings.</p>

<p>

</p>

<p>
50</p>

<p>
These include the current epoch time in seconds , and in another field the time in microseconds.</p>

<p>
These are reported as key value pairs.</p>

<p>
So what we need to do is to do the following</p>

<p>

</p>

<p>
51</p>

<p>
Run adjtimex</p>

<p>
Capture the output.</p>

<p>
Grep for "time.tv_sec"</p>

<p>
Grep for "time.tv_usec"</p>

<p>
Use cut to extract the time value in each case.</p>

<p>
Use tr to get rid of excess spaces in each case.</p>

<p>
Combine the two in a string with a decimal separator between them.</p>

<p>

</p>

<p>
52</p>

<p>
This takes a total of 4 lines of shell script. </p>

<p>
I will just describe them breifly here, see the show notes for details.</p>

<p>

</p>

<p>
53</p>

<p>
First we want to capture the output of adjtimex in a single operation.</p>

<p>
Run adjtimex and pipe the output through grep to capture lines containing "time.tv_"</p>

<p>
and save this to a variable. </p>

<p>

</p>

<p>
# Extract the current high resolution time from adjtimex.</p>

<p>
# We want two key value pairs, identified by time.tv_sec and time.tv_usec.</p>

<p>
tvals=$( adjtimex | grep "time.tv_" )</p>

<p>

</p>

<p>
54</p>

<p>
Next echo the contents of this variable and pipe it through grep, cut, and tr to get first the seconds and then the microseconds while also removing excess spaces.</p>

<p>
Save these to two separate variables.</p>

<p>
"time.tv_sec" is the time in seconds since the epoch.</p>

<p>
"time.tv_usec" is the number of microseconds in the current second.</p>

<p>

</p>

<p>
# Get the time since the unix epoch in seconds and micro-seconds.</p>

<p>
timesec=$( echo "$tvals" | grep "time.tv_sec" | cut -d: -f2 | tr -d " " )</p>

<p>
timeusec=$( echo "$tvals" |  grep "time.tv_usec" | cut -d: -f2 | tr -d " " )</p>

<p>

</p>

<p>
55</p>

<p>
Adjtimex does not zero pad the microsecond time value to provide leading zeros, so we need to take care of this using printf before we can append it to the seconds value. We didn't need to do this with date where the %N format character does this automatically.</p>

<p>

</p>

<p>
In this instance, the  printf format string is '%06d'</p>

<p>

</p>

<p>
padusec=$( printf '%06d' $timeusec )</p>

<p>

</p>

<p>
Now, combine these into a single number with a decimal separator by using simple string concatenation.</p>

<p>
# Combine them into a single number.</p>

<p>
timehires="$timesec"".""$padusec"</p>

<p>

</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
56 Summary of Methods</p>

<p>

</p>

<p>
Let's summarize where we are so far in terms of methods we can use to get the current time as a high resolution number.</p>

<p>

</p>

<p>
Alma                : use EPOCHREALTIME or date</p>

<p>
Debian (bash)       : use EPOCHREALTIME or date</p>

<p>
Raspberry-Pi (bash) : use EPOCHREALTIME or date</p>

<p>
ubuntu2604 (bash)   : use EPOCHREALTIME or date</p>

<p>
Suse                : use EPOCHREALTIME or date</p>

<p>
Debian (dash)       : use date</p>

<p>
Raspberry-Pi (dash) : use date</p>

<p>
ubuntu2604 (dash)   : use date</p>

<p>
Alpine              : use adjtimex and parse the output</p>

<p>
FreeBSD             : use date or ts</p>

<p>
OpenBSD             : use ts</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
57 Other alternatives</p>

<p>

</p>

<p>
There are a few alternatives that we haven't discussed yet.</p>

<p>

</p>

<p>
58 Bash with Dash</p>

<p>

</p>

<p>
In the case of Debian, Raspberry Pi, and Ubuntu running dash, since bash is available it is possible to write a separate bash script which simply echos EPOCHREALTIME and then call it from the dash script and capture the output. </p>

<p>

</p>

<p>
While this would work, there's probably not a lot of point to it.</p>

<p>
If you can rely on bash being there, then just change the first line of the script and make it a bash script.</p>

<p>

</p>

<p>
59 Adding Packages to Alpine</p>

<p>

</p>

<p>
The ts or timestamp utility is a common unix utility that can be installed if it is not present by default.</p>

<p>
This does produce high resolution timestamps on Alpine.</p>

<p>
On Alpine Linux this comes as part of the "moreutils" package.</p>

<p>
To add the package, use the following</p>

<p>
sudo apk add moreutils</p>

<p>
echo | ts "%.s" </p>

<p>
1779634800.959948 </p>

<p>

</p>

<p>
60</p>

<p>
You can also add the GNU coreutils, which will provide a high resolution date command which works like in the other examples.</p>

<p>
To add the package use the following</p>

<p>
sudo apk add coreutils</p>

<p>
date '+%s.%N'</p>

<p>
1779634800.212897332</p>

<p>

</p>

<p>
61</p>

<p>
If you can install more packages into your Alpine system, either of the above two is probably going to be preferable to parsing the output of adjtimex.</p>

<p>

</p>

<p>
62 Custom Timestamp Programs</p>

<p>

</p>

<p>
You can also write a very short program in python, perl, tcl, or some other language and have it output the current epoch time.</p>

<p>
I won't discuss that here though.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
63 Calculating Time Differences</p>

<p>

</p>

<p>
Shell arithmetic is integer only.</p>

<p>
If we wish to use high resolution timing data, we need to do something so we don't lose the precision we have worked so hard to get.</p>

<p>
There are several possible solutions.</p>

<p>

</p>

<p>
64 Change the Time Base</p>

<p>
One method is to change the time base from seconds to milli, micro, or nanoseconds. </p>

<p>
This can be done by simply multiplying the time values by the appropriate amount (e.g. 1000, 1,000,000, etc.) before subtracting them.</p>

<p>
This allows for integer arithmetic on high resolution values without losing precision.</p>

<p>

</p>

<p>
65 Use the Shell bc Arbitrary Precision Calculator</p>

<p>
The bc command line calculator will perform calculations using real numbers and is easy to use in scripts.</p>

<p>
It is present by default in most distros.</p>

<p>

</p>

<p>
echo "scale=9; $endtime - $starttime" | bc</p>

<p>

</p>

<p>
where endtime and starttime are variables containing time values.</p>

<p>

</p>

<p>
66</p>

<p>
However, for some inexplicable reason, neither Debian nor Opensuse install it by default.</p>

<p>
It is present in Ubuntu and Raspberry Pi which are Debian derivatives, and it can be added to distros which lack it.</p>

<p>

</p>

<p>
67 Use awk</p>

<p>
awk can also perform calculations using real numbers and it is present in nearly all distros including in all of the ones we tested here.</p>

<p>

</p>

<p>
echo "$endtime $starttime" | awk '{printf "%.6f\n", $1 - $2}'</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
68 Benchmarks</p>

<p>

</p>

<p>
And of course no comparative evaluation would be complete without benchmarks where we see how each method compares to another in terms of speed.</p>

<p>

</p>

<p>
In the benchmark test I ran each method in a loop through multiple iterations, measured the elapsed time, subtracted out the time for an empty loop, and then compared it to alternate methods.</p>

<p>
For anything other than EPOCHREALTIME, the empty loop time is negligible and has no real effect on the results.</p>

<p>

</p>

<p>
69</p>

<p>
Rather interestingly I came across a bug which caused date to run very slowly if called immediately after using EPOCHREALTIME in bash.</p>

<p>
The effect of the bug was to make the date benchmark test roughly 24 times slower.</p>

<p>

</p>

<p>
This has been fixed in newer releases, but if you are using an older distro release then beware of this bug.</p>

<p>
I was able to get around it either putting a sleep delay between benchmarking  EPOCHREALTIME and benchmarking date, or by simply testing date before testing EPOCHREALTIME.</p>

<p>

</p>

<p>
70</p>

<p>
To be able to conduct additional tests I installed ts in Ubuntu and Alpine, and the GNU version of date in Alpine.</p>

<p>

</p>

<p>

</p>

<p>
71 EPOCHREALTIME Versus date in Ubuntu 2604 bash</p>

<p>
The EPOCHREALTIME method is 3103 times faster than date.</p>

<p>

</p>

<p>
However, when the same test is run on Ubuntu 2404 when the date test is run before the EPOCHREALTIME test, EPOCHREALTIME is 1240 faster than date.</p>

<p>
Other Linux distros show performance to Ubuntu 2404.</p>

<p>
It appears that a side effect of fixing whatever the bug is has the effect of slowing down date.</p>

<p>
However, this is probably not a significant issue in normal circumstances. </p>

<p>

</p>

<p>
72 date versus ts in Ubuntu 2604 bash</p>

<p>
The date method is 3.7 times faster than ts</p>

<p>

</p>

<p>
73 date versus ts in Ubuntu 2604 dash</p>

<p>
The date method is 4.9 times faster than ts</p>

<p>

</p>

<p>
74 date versus ts in Freebsd sh</p>

<p>
The date method is 2.5 times faster than ts</p>

<p>

</p>

<p>
75 date versus adjtimex in Alpine Busybox</p>

<p>
The date method is 6.0 times faster than adjtimex</p>

<p>

</p>

<p>
76 date versus ts in Alpine Busybox</p>

<p>
The date method is 20.0 times faster than ts</p>

<p>

</p>

<p>
77 bc versus awk in Ubuntu 2604</p>

<p>
I compared calculating the difference between two numbers when using bc versus awk. </p>

<p>
The difference is negligible, with bc being only 7% faster than awk. </p>

<p>

</p>

<p>

</p>

<p>
78 Conclusion for Benchmarks</p>

<p>
Based on these results, if you need to measure elapsed time to high resolution and care about runing the command with as little overhead as possible, then the order of preference should be the following.</p>

<p>

</p>

<p>
79</p>

<p>
If you are using a newer version of bash, then use EPOCHREALTIME.</p>

<p>
If that is not available, then use date, provided it allows for high resolution times.</p>

<p>
If the above two cannot be used, then use ts.</p>

<p>
If you are using Busybox and cannot install either GNU date or ts, then use adjtimex.</p>

<p>

</p>

<p>
Date is the closest in terms of being the universal portable solution, but it does not work in all cases.</p>

<p>

</p>

<p>
80</p>

<p>
I have not compared different platforms to each other in terms of performance, as that would be a much more involved problem that is outside the scope of this episode.</p>

<p>

</p>

<p>
However, different operating systems implement different commands in different ways.</p>

<p>

</p>

<p>
81</p>

<p>
For example, on Openbsd and Freebsd, ts appears to be an ELF binary. That is, it is executable machine code, possibly written in C.</p>

<p>
On  Ubuntu however, ts appears to be a perl script. </p>

<p>
As a result of this, the advantage that date has over ts is much less in Freebsd than it is with Ubuntu (and likely other Linux distros) as on Freebsd it doesn't need to load a perl interpreter to run ts. </p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>

<p>
82 Overall Conclusion</p>

<p>
You no doubt thought that measuring elapsed time was going to be so simple, and how could someone get an entire podcast out of such a simple subject?</p>

<p>
And yet here we are half an hour later with just a basic overview of the subject. </p>

<p>

</p>

<p>
83</p>

<p>
I hope you found this interesting and informative.</p>

<p>
Please let us know in the comments if you think that I have done anything incorrectly, or if you have another way of doing things.</p>

<p>

</p>

<p>
I hope to see you all again in another future episode of HPR.</p>

<p>

</p>

<p>
--------------------</p>

<p>

</p>


<p><a href="https://hackerpublicradio.org/eps/hpr4678/index.html#comments">Provide <strong>feedback</strong> on this episode</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The 2026 guide to eSignatures: Evaluating security, cost, and ROI]]></title>
<description><![CDATA[Choosing the right eSign solution is less about picking the tool with the most bells and whistles and more about confirming that the features support your company’s requirements for security and compliance, workflow automation, cost-effectiveness, and operational efficiency.



The best eSign sol...]]></description>
<link>https://tsecurity.de/de/3652805/it-security-nachrichten/the-2026-guide-to-esignatures-evaluating-security-cost-and-roi/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3652805/it-security-nachrichten/the-2026-guide-to-esignatures-evaluating-security-cost-and-roi/</guid>
<pubDate>Tue, 07 Jul 2026 23:35:44 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Choosing the right eSign solution is less about picking the tool with the most bells and whistles and more about confirming that the features support your company’s requirements for <a href="https://www.gonitro.com/resources/security-compliance?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored">security and compliance</a>, workflow automation, cost-effectiveness, and operational efficiency.</p>



<p><a href="https://www.gonitro.com/best-esign-software?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored">The best eSign solutions</a> let teams securely collect legally binding electronic signatures while <a href="https://www.gonitro.com/integrations?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored">integrating signing workflows</a> with business systems, compliance controls, and document lifecycle processes — evaluated across four factors: security and compliance, workflow integration, total cost of ownership, and measurable business ROI.</p>



<p>When evaluating eSignature solutions, look beyond signing functionality and consider these four factors:</p>



<ul class="wp-block-list">
<li>Security and compliance</li>



<li>Workflow integration</li>



<li>Total cost of ownership</li>



<li>Measurable business ROI</li>
</ul>



<h2 class="wp-block-heading">Security and compliance are the foundation of eSignatures</h2>



<p>Yes, you want eSigning to be convenient, but it’s arguably even more important that your eSignature solution provides the security, auditability, and legal validity required to support critical business transactions.</p>



<p><strong>Look for solutions that offer:</strong></p>



<ul class="wp-block-list">
<li>Comprehensive <a href="https://www.gonitro.com/resources/esignature-audit-trials?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI">audit trails</a></li>



<li>Strong authentication controls</li>



<li>Encryption in transit and at rest</li>



<li>Support for established legal frameworks (e.g., the ESIGN Act, UETA, eIDAS)</li>
</ul>



<p>Independent certifications, including <a href="https://www.gonitro.com/security-compliance?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored">SOC 2 Type II</a> and ISO 27001, provide additional assurance that an eSign vendor follows recognized security and information management practices.</p>



<h2 class="wp-block-heading">The signature is only one step in the document lifecycle</h2>



<p><strong>During the digital signing process, documents typically move through multiple workflows:</strong></p>



<p>During the digital signing process, documents typically move through multiple stages: creation, review, approval, signature collection, storage, reporting, and retention.</p>



<p>Consider a typical sales contract: it might originate in a CRM, require review and approval from finance, get routed for signature, then need to be stored in a repository, reported on for compliance, and retained per policy. If each of these steps happens in a separate, disconnected tool, the signature may be digital, but the workflow is still manual.</p>



<p>When one or more of these steps rely on email attachments, manual routing, or moving files across disconnected applications, delays and inefficiencies quickly snowball.</p>



<p>An eSignature platform that supports <a href="https://www.gonitro.com/automate?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI">document workflow automation</a> can help you avoid this by connecting approval workflows, document routing, signature collection, and archival processes into a low-friction experience.</p>



<h2 class="wp-block-heading">Evaluating the true cost of ownership of an eSignature solution</h2>



<p>When you’re evaluating the cost of eSignature solutions, <a href="https://www.gonitro.com/pricing?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI">subscription pricing</a> only tells part of the story. The solution with a lower upfront cost may require additional integrations, administrative effort, training, or support resources that increase long-term expenditure.</p>



<p><strong>When calculating total cost of ownership, be sure to consider:</strong></p>



<ul class="wp-block-list">
<li>Licensing and transaction costs</li>



<li>Implementation and integration requirements</li>



<li>Administrative overhead</li>



<li>User adoption and training</li>



<li>Compliance and audit support</li>



<li>Scalability as your business needs evolve</li>
</ul>



<h2 class="wp-block-heading">How to measure eSignature ROI</h2>



<p>Traditionally, the value proposition for eSignature software was that it reduced paper, printing, and shipping costs. Today, the value is firmly centered on operational outcomes, including:</p>



<ul class="wp-block-list">
<li>Contract turnaround times</li>



<li>Employee onboarding speed</li>



<li>Approval cycle duration</li>



<li>Manual labor reduction</li>



<li>Error elimination</li>



<li>Compliance risk mitigation</li>



<li>Customer and employee experience improvements</li>
</ul>



<p>For example, reducing contract processing from days to hours can have a greater business impact than eliminating printing costs. Similarly, automated approval workflows can take over repetitive administrative tasks, freeing up employees to work on higher-value initiatives.</p>



<h2 class="wp-block-heading"><a></a>What to look for in an eSignature solution</h2>



<p>As eSignature technology matures, the evaluation criteria have expanded beyond ease of signing. Today, organizations need solutions that can support compliance requirements, integrate with existing business systems, automate document workflows, and scale alongside broader digital transformation initiatives.</p>



<p><strong>When comparing eSignature solutions, don’t just look at signing capabilities. Assess how well each eSign solution supports the entire document lifecycle through:</strong></p>



<ul class="wp-block-list">
<li>Strong security and compliance controls</li>



<li>Support for ESIGN, UETA, and eIDAS requirements</li>



<li>Workflow automation capabilities</li>



<li>Integration with existing business systems</li>



<li>API accessibility for future automation initiatives</li>



<li>Comprehensive audit trails and reporting</li>



<li>Predictable, scalable pricing</li>
</ul>



<p>In 2026, the best eSignature solution isn’t the one with the most features. It’s the one that connects signing to the rest of the document lifecycle while keeping security, cost, and ROI measurable.<a href="https://www.gonitro.com/sign?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored"> </a><a href="https://www.gonitro.com/sign?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored">Nitro Sign</a> is built around that principle: it goes beyond electronic signatures to support secure, compliant, connected document workflows that integrate with the systems teams already use, so governance improves, operations accelerate, and the solution scales with long-term business goals.</p>



<p><strong>Discover why Nitro Sign has been recognized by IDC as a global leader in electronic signature software solutions.</strong></p>



<p><a href="https://www.gonitro.com/contact-sales?utm_source=foundry&amp;utm_medium=referral&amp;utm_campaign=The+2026+Guide+to+eSignatures%3A+Evaluating+Security%2C+Cost%2C+and+ROI" rel="sponsored">Speak with an eSign Expert</a></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[IBM grows mainframe family with rack, frame models targeting AI, hybrid clouds]]></title>
<description><![CDATA[IBM is looking to expand the reach of its foundational mainframe portfolio by adding new single frame and rack mounted versions of its Z and LinuxONE systems.



The IBM z17 portfolio adds a single frame and rack mount versions that bring mainframe capabilities into smaller, customizable footprin...]]></description>
<link>https://tsecurity.de/de/3652288/it-security-nachrichten/ibm-grows-mainframe-family-with-rack-frame-models-targeting-ai-hybrid-clouds/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3652288/it-security-nachrichten/ibm-grows-mainframe-family-with-rack-frame-models-targeting-ai-hybrid-clouds/</guid>
<pubDate>Tue, 07 Jul 2026 19:07:56 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>IBM is looking to expand the reach of its foundational mainframe portfolio by adding new single frame and rack mounted versions of its Z and LinuxONE systems.</p>



<p>The <a href="https://www.ibm.com/docs/en/announcements/z17-single-frame-rack-mount-systems-expand-ai-security-operational-simplicity-enterprise-workloads" target="_blank" rel="noreferrer noopener">IBM z17 portfolio</a> adds a single frame and rack mount versions that bring mainframe capabilities into smaller, customizable footprints. The <a href="https://www.ibm.com/docs/en/announcements/linuxone-rockhopper-5-built-secured-ai-ready-enterprise-it" target="_blank" rel="noreferrer noopener">LinuxONE Rockhopper family</a> gets a single frame and rack mount models, plus a new Express rack mount offering, that target new and smaller clients, according to Tina Tarquinio, chief product officer, IBM Z &amp; LinuxONE.</p>



<p>Specifically, the new hardware includes:</p>



<ul class="wp-block-list">
<li>z17 single frame is a fully packaged box in an IBM rack with intelligent power distribution units, delivered as a complete enclosed unit ready to deploy at the edge or other strategically important customer sites.</li>



<li>z17 rack mount lets customers install IBM Z components directly into their own industry-standard rack, with built-in flexibility for co-location with other technologies.</li>



<li>LinuxONE Rockhopper 5 is a multi-drawer LinuxONE system for high-density workloads, with on-chip AI acceleration, confidential computing, and postquantum cryptography available in both single frame and rack mount configurations.</li>



<li>Rockhopper 5 rack mount and Express offerings deliver enterprise-grade Linux, confidential computing, and on-chip AI acceleration in a compact 18U configuration. Designed for organizations supporting a smaller set of workloads, the offering provides a cost-efficient entry point that can scale as business grows, while prioritizing security, resiliency, and performance.</li>
</ul>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/07/LinuxONE-5-Single-Frame.png?w=1024" alt="IBM LinuxONE 5 single frame system" class="wp-image-4193838" width="1024" height="768" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">IBM</p></div>



<p>The new IBM z17 and IBM LinuxONE 5 Rockhopper configurations support up to 82 cores and 18 TB of memory across two processor drawers, representing about a 20% increase in core count and 12% increase in memory capacity over current systems, IBM stated. Single processor capacity of an IBM z17 ME2 provides full speed IBM z/OS configurations including 10% greater throughput per core than IBM z16 A02 with some variation based on workload and configuration, according to Tarquinio.</p>



<p>Both systems feature a 5.5 GHz IBM Telum II processor and a built-in AI accelerator that IBM says will let customers run more than 450 billion inferencing operations in a day with one millisecond response time. In addition, the 32-core Spyre AI accelerator is designed to handle all manner of AI workloads.</p>



<p>The idea is to bring the core strengths of IBM Z to a broader range of deployment models while offering the security, resilience, and performance enterprises depend on, Tarquinio said. </p>



<p>“As always, we’re continuing to innovate to deliver more with less, including up to 20% more capacity than IBM z16 to help process transactions faster and support growing AI-driven workloads,” Tarquinio said.  “Even the newest and smallest member of the IBM z17 family delivers the performance, efficiency, and scalability organizations need as they balance growth ambitions with real-world resource constraints.”</p>



<p>For the Linux-based system, Rockhopper 5 is for organizations that have moved past the evaluation question and are ready to consolidate a substantial portion of their x86 estate, said Marcel Mitran, IBM Fellow and CTO of IBM LinuxONE. </p>



<p>Rockhopper 5 is designed to bring a smaller physical footprint and a software licensing model that reflects actual workload boundaries rather than physical server counts, Mitran said.</p>



<p>The LinuxONE 5 Express is a preconfigured system designed to get organizations running on LinuxONE quickly, with a defined bill of materials and a predictable starting cost, on the same architecture that the largest enterprises in the world depend on, Mitran said.</p>



<p>“It is built for organizations that want to consolidate a modest x86 estate, evaluate LinuxONE for the first time, or deploy a specific workload such as digital assets, AI-infused transaction processing, or confidential computing, without committing to the footprint of the larger model,” Mitran said.</p>



<p>Some of the mainframes’ software features were also bulked up. For example, IBM said that Post Quantum Cryptography security is now standard on the z17 and LinuxONE Rockhopper 5 systems letting customers start to utilize cryptography to protect core resources for the future.</p>



<p>The idea is to help customers protect long-lived, mission-critical data while reducing the cost and complexity of future cryptographic migration, IBM stated. </p>



<p>In that vein, IBM said it was bringing Crypto Discovery &amp; Inventory, which lets security teams see what has been encrypted across the enterprise. In addition, IBM announced an Infrastructure Management for Z and LinuxONE package that would let customers administer, monitor, automate, and provision IBM Z and LinuxONE systems from a central location.</p>



<p>IBM said it wants to reduce operational complexity for customers by making automating day-to-day operations<strong> </strong>to ultimately lower administrative costs and concerns. With the new flexible form factors, IBM continues to target hybrid and AI infrastructure buildouts with the Big Iron. In the AI world, the z17 is being utilized for AI inferencing, transactions, training, and key security applications such as fraud detection and insurance claims.</p>



<p>“Enterprise infrastructure is entering a new phase. Organizations need platforms that can support AI-driven growth while navigating resource constraints, evolving business requirements, and increasingly complex hybrid environments,” Tarquinio said. “They are being asked to deploy new AI capabilities while learning new skills, controlling operational costs, and maximizing the value of existing applications and infrastructure.”</p>



<p>A recent <a href="https://www-api.ibm.com/adobe/assets/urn:aaid:aem:52bed780-53cf-4a1c-a73b-d373bd532e97/original/as/the-mainframe-advantage.pdf" target="_blank" rel="noreferrer noopener">IBM Institute study</a> on mainframe usage stated that embedding mainframe to support AI in executing transactions is not temporary: 75% of executives expect mainframe-based applications to remain central to digital transformation, and 60% say mainframe-based platforms are essential to enabling AI innovation.</p>



<p>”Mainframe-anchored systems of record are becoming systems of intelligent execution—not as general‑purpose AI platforms, but as environments where AI acts directly within transactions and in support of them,” the study reported.</p>



<p>Gartner wrote in its “<a href="https://www.ibm.com/forms/mkt-17256" target="_blank" rel="noreferrer noopener">The State of the IBM Mainframe in 2026</a>” report that IBM’s willingness to make significant investments ensure the mainframe modernizes to remain a vital and thriving component of enterprise IT.  </p>



<p>“Most mainframe customers are now prioritizing the reduction of technical debt and adopting platform innovations to future-proof their mainframe environments for the coming decade,” Gartner wrote.</p>



<p>The new z17 single frame and rack mount configurations, LinuxONE Rockhopper 5, and LinuxONE 5 Express will all be available August 12, 2026. IBM Infrastructure Management for IBM Z and IBM LinuxONE will be available August 14.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Siemens SINEC OS]]></title>
<description><![CDATA[View CSAF
Summary
SINEC OS before V4.0 contains multiple vulnerabilities. Siemens has released a new version for RUGGEDCOM RST2428P and recommends to update to the latest version.
The following versions of Siemens SINEC OS are affected:

RUGGEDCOM RST2428P (6GK6242-6PA00) vers:intdot/cork. The "*...]]></description>
<link>https://tsecurity.de/de/3652271/it-security-nachrichten/siemens-sinec-os/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3652271/it-security-nachrichten/siemens-sinec-os/</guid>
<pubDate>Tue, 07 Jul 2026 18:55:49 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="https://github.com/cisagov/CSAF/blob/develop/csaf_files/OT/white/2026/icsa-26-188-05.json"><strong>View CSAF</strong></a></p>
<h2>Summary</h2>
<p><strong>SINEC OS before V4.0 contains multiple vulnerabilities. Siemens has released a new version for RUGGEDCOM RST2428P and recommends to update to the latest version.</strong></p>
<p>The following versions of Siemens SINEC OS are affected:</p>
<ul>
<li>RUGGEDCOM RST2428P (6GK6242-6PA00) vers:intdot/&lt;4.0 </li>
</ul>
<div class="csaf-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS</th>
<th role="columnheader">Vendor</th>
<th role="columnheader">Equipment</th>
<th role="columnheader">Vulnerabilities</th>
</tr>
</thead>
<tbody>
<tr>
<td>v3 9.8</td>
<td>Siemens</td>
<td>Siemens SINEC OS</td>
<td>Improper Restriction of Operations within the Bounds of a Memory Buffer, Improper Resource Shutdown or Release, Integer Overflow or Wraparound, Stack-based Buffer Overflow, Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal'), Uncontrolled Recursion, Out-of-bounds Read, Covert Timing Channel, Improper Input Validation, Improperly Controlled Modification of Object Prototype Attributes ('Prototype Pollution'), Improper Update of Reference Count, Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition'), Multiple Releases of Same Resource or Handle, Permissive Regular Expression, Expired Pointer Dereference, Incorrect Bitwise Shift of Integer, Out-of-bounds Write, User Interface (UI) Misrepresentation of Critical Information, Improper Access Control, Insertion of Sensitive Information Into Sent Data, Inefficient Algorithmic Complexity, Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting'), Authentication Bypass by Primary Weakness, NULL Pointer Dereference, Active Debug Code, Loop with Unreachable Exit Condition ('Infinite Loop'), Missing Synchronization, External Control of File Name or Path, Privilege Dropping / Lowering Errors, Use of Web Browser Cache Containing Sensitive Information</td>
</tr>
</tbody>
</table>
</div>
<h3>Background</h3>
<ul>
<li><strong>Critical Infrastructure Sectors: </strong>Critical Manufacturing, Transportation Systems, Energy, Healthcare and Public Health, Financial Services, Government Services and Facilities</li>
<li><strong>Countries/Areas Deployed: </strong>Worldwide</li>
<li><strong>Company Headquarters Location: </strong>Germany</li>
</ul>
<hr>
<h2>Vulnerabilities</h2>
<div class="csaf-accordion">
<p><a class="csaf-accordion-toggle-all" href="https://www.cisa.gov/#">Expand All +</a></p>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-1352</a></h3>
<div class="csaf-accordion-content">
<p>A vulnerability has been found in GNU elfutils 0.192 and classified as critical. This vulnerability affects the function __libdw_thread_tail in the library libdw_alloc.c of the component eu-readelf. The manipulation of the argument w leads to memory corruption. The attack can be initiated remotely. The complexity of an attack is rather high. The exploitation appears to be difficult. The exploit has been disclosed to the public and may be used. The name of the patch is 2636426a091bd6c6f7f02e49ab20d4cdc6bfc753. It is recommended to apply a patch to fix this issue.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-1352">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/119.html">CWE-119 Improper Restriction of Operations within the Bounds of a Memory Buffer</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:L/I:L/A:L">CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:L/I:L/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-1376</a></h3>
<div class="csaf-accordion-content">
<p>A vulnerability classified as problematic was found in GNU elfutils 0.192. This vulnerability affects the function elf_strptr in the library /libelf/elf_strptr.c of the component eu-strip. The manipulation leads to denial of service. It is possible to launch the attack on the local host. The complexity of an attack is rather high. The exploitation appears to be difficult. The exploit has been disclosed to the public and may be used. The name of the patch is b16f441cca0a4841050e3215a9f120a6d8aea918. It is recommended to apply a patch to fix this issue.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-1376">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/404.html">CWE-404 Improper Resource Shutdown or Release</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>2.5</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:L">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-6052</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in how GLib’s GString manages memory when adding data to strings. If a string is already very large, combining it with more input can cause a hidden overflow in the size calculation. This makes the system think it has enough memory when it doesn’t. As a result, data may be written past the end of the allocated memory, leading to crashes or memory corruption.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-6052">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/190.html">CWE-190 Integer Overflow or Wraparound</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>3.7</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L">CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-6141</a></h3>
<div class="csaf-accordion-content">
<p>A vulnerability has been found in GNU ncurses up to 6.5-20250322 and classified as problematic. This vulnerability affects the function postprocess_termcap of the file tinfo/parse_entry.c. The manipulation leads to stack-based buffer overflow. The attack needs to be approached locally. Upgrading to version 6.5-20250329 is able to address this issue. It is recommended to upgrade the affected component.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-6141">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/121.html">CWE-121 Stack-based Buffer Overflow</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>3.3</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-6170</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in the interactive shell of the xmllint command-line tool, used for parsing XML files. When a user inputs an overly long command, the program does not check the input size properly, which can cause it to crash. This issue might allow attackers to run harmful code in rare configurations without modern protections.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-6170">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/121.html">CWE-121 Stack-based Buffer Overflow</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>2.5</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:N/I:N/A:L">CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-7039</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in glib. An integer overflow during temporary file creation leads to an out-of-bounds memory access, allowing an attacker to potentially perform path traversal or access private temporary file content by creating symbolic links. This vulnerability allows a local attacker to manipulate file paths and access unauthorized data. The core issue stems from insufficient validation of file path lengths during temporary file operations.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-7039">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/22.html">CWE-22 Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal')</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>3.7</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:N">CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-8732</a></h3>
<div class="csaf-accordion-content">
<p>A vulnerability was found in libxml2 up to 2.14.5. It has been declared as problematic. This vulnerability affects the function xmlParseSGMLCatalog of the component xmlcatalog. The manipulation leads to uncontrolled recursion. Attacking locally is a requirement. The exploit has been disclosed to the public and may be used. The real existence of this vulnerability is still doubted at the moment. The code maintainer explains, that "[t]he issue can only be triggered with untrusted SGML catalogs and it makes absolutely no sense to use untrusted catalogs. I also doubt that anyone is still using SGML catalogs at all."</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-8732">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/674.html">CWE-674 Uncontrolled Recursion</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>3.3</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-9086</a></h3>
<div class="csaf-accordion-content">
<p>1. A cookie is set using the `secure` keyword for `https://target` 2. curl is redirected to or otherwise made to speak with `http://target` (same hostname, but using clear text HTTP) using the same cookie set 3. The same cookie name is set - but with just a slash as path (`path=\"/\",`). Since this site is not secure, the cookie *should* just be ignored. 4. A bug in the path comparison logic makes curl read outside a heap buffer boundary The bug either causes a crash or it potentially makes the comparison come to the wrong conclusion and lets the clear-text site override the contents of the secure cookie, contrary to expectations and depending on the memory contents immediately following the single-byte allocation that holds the path. The presumed and correct behavior would be to plainly ignore the second set of the cookie since it was already set as secure on a secure host so overriding it on an insecure host should not be okay.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-9086">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/125.html">CWE-125 Out-of-bounds Read</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.5</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-9230</a></h3>
<div class="csaf-accordion-content">
<p>Issue summary: An application trying to decrypt CMS messages encrypted using password based encryption can trigger an out-of-bounds read and write. Impact summary: This out-of-bounds read may trigger a crash which leads to Denial of Service for an application. The out-of-bounds write can cause a memory corruption which can have various consequences including a Denial of Service or Execution of attacker-supplied code. Although the consequences of a successful exploit of this vulnerability could be severe, the probability that the attacker would be able to perform it is low. Besides, password based (PWRI) encryption support in CMS messages is very rarely used. For that reason the issue was assessed as Moderate severity according to our Security Policy. The FIPS modules in 3.5, 3.4, 3.3, 3.2, 3.1 and 3.0 are not affected by this issue, as the CMS implementation is outside the OpenSSL FIPS module boundary.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-9230">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/125.html">CWE-125 Out-of-bounds Read</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.5</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-9231</a></h3>
<div class="csaf-accordion-content">
<p>Issue summary: A timing side-channel which could potentially allow remote recovery of the private key exists in the SM2 algorithm implementation on 64 bit ARM platforms. Impact summary: A timing side-channel in SM2 signature computations on 64 bit ARM platforms could allow recovering the private key by an attacker.. While remote key recovery over a network was not attempted by the reporter, timing measurements revealed a timing signal which may allow such an attack. OpenSSL does not directly support certificates with SM2 keys in TLS, and so this CVE is not relevant in most TLS contexts. However, given that it is possible to add support for such certificates via a custom provider, coupled with the fact that in such a custom provider context the private key may be recoverable via remote timing measurements, we consider this to be a Moderate severity issue. The FIPS modules in 3.5, 3.4, 3.3, 3.2, 3.1 and 3.0 are not affected by this issue, as SM2 is not an approved algorithm.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-9231">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/385.html">CWE-385 Covert Timing Channel</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>6.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-9232</a></h3>
<div class="csaf-accordion-content">
<p>Issue summary: An application using the OpenSSL HTTP client API functions may trigger an out-of-bounds read if the 'no_proxy' environment variable is set and the host portion of the authority component of the HTTP URL is an IPv6 address. Impact summary: An out-of-bounds read can trigger a crash which leads to Denial of Service for an application. The OpenSSL HTTP client API functions can be used directly by applications but they are also used by the OCSP client functions and CMP (Certificate Management Protocol) client implementation in OpenSSL. However the URLs used by these implementations are unlikely to be controlled by an attacker. In this vulnerable code the out of bounds read can only trigger a crash. Furthermore the vulnerability requires an attacker-controlled URL to be passed from an application to the OpenSSL function and the user has to have a 'no_proxy' environment variable set. For the aforementioned reasons the issue was assessed as Low severity. The vulnerable code was introduced in the following patch releases: 3.0.16, 3.1.8, 3.2.4, 3.3.3, 3.4.0 and 3.5.0. The FIPS modules in 3.5, 3.4, 3.3, 3.2, 3.1 and 3.0 are not affected by this issue, as the HTTP client implementation is outside the OpenSSL FIPS module boundary.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-9232">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/125.html">CWE-125 Out-of-bounds Read</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.9</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-10966</a></h3>
<div class="csaf-accordion-content">
<p>curl's code for managing SSH connections when SFTP was done using the wolfSSH powered backend was flawed and missed host verification mechanisms. This prevents curl from detecting MITM attackers and more.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-10966">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>4.3</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N">CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-13465</a></h3>
<div class="csaf-accordion-content">
<p>Lodash versions 4.0.0 through 4.17.22 are vulnerable to prototype pollution in the _.unset and _.omit functions. An attacker can pass crafted paths which cause Lodash to delete methods from global prototypes. The issue permits deletion of properties but does not allow overwriting their original behavior. This issue is patched on 4.17.23</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-13465">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/1321.html">CWE-1321 Improperly Controlled Modification of Object Prototype Attributes ('Prototype Pollution')</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.2</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:N/I:L/A:L">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:N/I:L/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-13601</a></h3>
<div class="csaf-accordion-content">
<p>A heap-based buffer overflow problem was found in glib through an incorrect calculation of buffer size in the g_escape_uri_string() function. If the string to escape contains a very large number of unacceptable characters (which would need escaping), the calculation of the length of the escaped string could overflow, leading to a potential write off the end of the newly allocated string.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-13601">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/190.html">CWE-190 Integer Overflow or Wraparound</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-39913</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: tcp_bpf: Call sk_msg_free() when tcp_bpf_send_verdict() fails to allocate psock-&gt;cork. syzbot reported the splat below. [0] The repro does the following: 1. Load a sk_msg prog that calls bpf_msg_cork_bytes(msg, cork_bytes) 2. Attach the prog to a SOCKMAP 3. Add a socket to the SOCKMAP 4. Activate fault injection 5. Send data less than cork_bytes At 5., the data is carried over to the next sendmsg() as it is smaller than the cork_bytes specified by bpf_msg_cork_bytes(). Then, tcp_bpf_send_verdict() tries to allocate psock-&gt;cork to hold the data, but this fails silently due to fault injection + __GFP_NOWARN. If the allocation fails, we need to revert the sk-&gt;sk_forward_alloc change done by sk_msg_alloc(). Let's call sk_msg_free() when tcp_bpf_send_verdict fails to allocate psock-&gt;cork. The "*copied" also needs to be updated such that a proper error can be returned to the caller, sendmsg. It fails to allocate psock-&gt;cork. Nothing has been corked so far, so this patch simply sets "*copied" to 0. [0]: WARNING: net/ipv4/af_inet.c:156 at inet_sock_destruct+0x623/0x730 net/ipv4/af_inet.c:156, CPU#1: syz-executor/5983 Modules linked in: CPU: 1 UID: 0 PID: 5983 Comm: syz-executor Not tainted syzkaller #0 PREEMPT(full) Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/12/2025 RIP: 0010:inet_sock_destruct+0x623/0x730 net/ipv4/af_inet.c:156 Code: 0f 0b 90 e9 62 fe ff ff e8 7a db b5 f7 90 0f 0b 90 e9 95 fe ff ff e8 6c db b5 f7 90 0f 0b 90 e9 bb fe ff ff e8 5e db b5 f7 90 &lt;0f&gt; 0b 90 e9 e1 fe ff ff 89 f9 80 e1 07 80 c1 03 38 c1 0f 8c 9f fc RSP: 0018:ffffc90000a08b48 EFLAGS: 00010246 RAX: ffffffff8a09d0b2 RBX: dffffc0000000000 RCX: ffff888024a23c80 RDX: 0000000000000100 RSI: 0000000000000fff RDI: 0000000000000000 RBP: 0000000000000fff R08: ffff88807e07c627 R09: 1ffff1100fc0f8c4 R10: dffffc0000000000 R11: ffffed100fc0f8c5 R12: ffff88807e07c380 R13: dffffc0000000000 R14: ffff88807e07c60c R15: 1ffff1100fc0f872 FS: 00005555604c4500(0000) GS:ffff888125af1000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00005555604df5c8 CR3: 0000000032b06000 CR4: 00000000003526f0 Call Trace: __sk_destruct+0x86/0x660 net/core/sock.c:2339 rcu_do_batch kernel/rcu/tree.c:2605 [inline] rcu_core+0xca8/0x1770 kernel/rcu/tree.c:2861 handle_softirqs+0x286/0x870 kernel/softirq.c:579 __do_softirq kernel/softirq.c:613 [inline] invoke_softirq kernel/softirq.c:453 [inline] __irq_exit_rcu+0xca/0x1f0 kernel/softirq.c:680 irq_exit_rcu+0x9/0x30 kernel/softirq.c:696 instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1052 [inline] sysvec_apic_timer_interrupt+0xa6/0xc0 arch/x86/kernel/apic/apic.c:1052</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-39913">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40214</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: af_unix: Initialise scc_index in unix_add_edge(). Quang Le reported that the AF_UNIX GC could garbage-collect a receive queue of an alive in-flight socket, with a nice repro. The repro consists of three stages. 1) 1-a. Create a single cyclic reference with many sockets 1-b. close() all sockets 1-c. Trigger GC 2) 2-a. Pass sk-A to an embryo sk-B 2-b. Pass sk-X to sk-X 2-c. Trigger GC 3) 3-a. accept() the embryo sk-B 3-b. Pass sk-B to sk-C 3-c. close() the in-flight sk-A 3-d. Trigger GC As of 2-c, sk-A and sk-X are linked to unix_unvisited_vertices, and unix_walk_scc() groups them into two different SCCs: unix_sk(sk-A)-&gt;vertex-&gt;scc_index = 2 (UNIX_VERTEX_INDEX_START) unix_sk(sk-X)-&gt;vertex-&gt;scc_index = 3 Once GC completes, unix_graph_grouped is set to true. Also, unix_graph_maybe_cyclic is set to true due to sk-X's cyclic self-reference, which makes close() trigger GC. At 3-b, unix_add_edge() allocates unix_sk(sk-B)-&gt;vertex and links it to unix_unvisited_vertices. unix_update_graph() is called at 3-a. and 3-b., but neither unix_graph_grouped nor unix_graph_maybe_cyclic is changed because both sk-B's listener and sk-C are not in-flight. 3-c decrements sk-A's file refcnt to 1. Since unix_graph_grouped is true at 3-d, unix_walk_scc_fast() is finally called and iterates 3 sockets sk-A, sk-B, and sk-X: sk-A -&gt; sk-B (-&gt; sk-C) sk-X -&gt; sk-X This is totally fine. All of them are not yet close()d and should be grouped into different SCCs. However, unix_vertex_dead() misjudges that sk-A and sk-B are in the same SCC and sk-A is dead. unix_sk(sk-A)-&gt;scc_index == unix_sk(sk-B)-&gt;scc_index &lt;-- Wrong! &amp;&amp; sk-A's file refcnt == unix_sk(sk-A)-&gt;vertex-&gt;out_degree ^-- 1 in-flight count for sk-B -&gt; sk-A is dead !? The problem is that unix_add_edge() does not initialise scc_index. Stage 1) is used for heap spraying, making a newly allocated vertex have vertex-&gt;scc_index == 2 (UNIX_VERTEX_INDEX_START) set by unix_walk_scc() at 1-c. Let's track the max SCC index from the previous unix_walk_scc() call and assign the max + 1 to a new vertex's scc_index. This way, we can continue to avoid Tarjan's algorithm while preventing misjudgments.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40214">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40248</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: vsock: Ignore signal/timeout on connect() if already established During connect(), acting on a signal/timeout by disconnecting an already established socket leads to several issues: 1. connect() invoking vsock_transport_cancel_pkt() -&gt; virtio_transport_purge_skbs() may race with sendmsg() invoking virtio_transport_get_credit(). This results in a permanently elevated `vvs-&gt;bytes_unsent`. Which, in turn, confuses the SOCK_LINGER handling. 2. connect() resetting a connected socket's state may race with socket being placed in a sockmap. A disconnected socket remaining in a sockmap breaks sockmap's assumptions. And gives rise to WARNs. 3. connect() transitioning SS_CONNECTED -&gt; SS_UNCONNECTED allows for a transport change/drop after TCP_ESTABLISHED. Which poses a problem for any simultaneous sendmsg() or connect() and may result in a use-after-free/null-ptr-deref. Do not disconnect socket on signal/timeout. Keep the logic for unconnected sockets: they don't linger, can't be placed in a sockmap, are rejected by sendmsg().</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40248">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40250</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: net/mlx5: Clean up only new IRQ glue on request_irq() failure The mlx5_irq_alloc() function can inadvertently free the entire rmap and end up in a crash[1] when the other threads tries to access this, when request_irq() fails due to exhausted IRQ vectors. This commit modifies the cleanup to remove only the specific IRQ mapping that was just added. This prevents removal of other valid mappings and ensures precise cleanup of the failed IRQ allocation's associated glue object. Note: This error is observed when both fwctl and rds configs are enabled. [1] mlx5_core 0000:05:00.0: Successfully registered panic handler for port 1 mlx5_core 0000:05:00.0: mlx5_irq_alloc:293:(pid 66740): Failed to request irq. err = -28 infiniband mlx5_0: mlx5_ib_test_wc:290:(pid 66740): Error -28 while trying to test write-combining support mlx5_core 0000:05:00.0: Successfully unregistered panic handler for port 1 mlx5_core 0000:06:00.0: Successfully registered panic handler for port 1 mlx5_core 0000:06:00.0: mlx5_irq_alloc:293:(pid 66740): Failed to request irq. err = -28 infiniband mlx5_0: mlx5_ib_test_wc:290:(pid 66740): Error -28 while trying to test write-combining support mlx5_core 0000:06:00.0: Successfully unregistered panic handler for port 1 mlx5_core 0000:03:00.0: mlx5_irq_alloc:293:(pid 28895): Failed to request irq. err = -28 mlx5_core 0000:05:00.0: mlx5_irq_alloc:293:(pid 28895): Failed to request irq. err = -28 general protection fault, probably for non-canonical address 0xe277a58fde16f291: 0000 [#1] SMP NOPTI RIP: 0010:free_irq_cpu_rmap+0x23/0x7d Call Trace: ? show_trace_log_lvl+0x1d6/0x2f9 ? show_trace_log_lvl+0x1d6/0x2f9 ? mlx5_irq_alloc.cold+0x5d/0xf3 [mlx5_core] ? __die_body.cold+0x8/0xa ? die_addr+0x39/0x53 ? exc_general_protection+0x1c4/0x3e9 ? dev_vprintk_emit+0x5f/0x90 ? asm_exc_general_protection+0x22/0x27 ? free_irq_cpu_rmap+0x23/0x7d mlx5_irq_alloc.cold+0x5d/0xf3 [mlx5_core] irq_pool_request_vector+0x7d/0x90 [mlx5_core] mlx5_irq_request+0x2e/0xe0 [mlx5_core] mlx5_irq_request_vector+0xad/0xf7 [mlx5_core] comp_irq_request_pci+0x64/0xf0 [mlx5_core] create_comp_eq+0x71/0x385 [mlx5_core] ? mlx5e_open_xdpsq+0x11c/0x230 [mlx5_core] mlx5_comp_eqn_get+0x72/0x90 [mlx5_core] ? xas_load+0x8/0x91 mlx5_comp_irqn_get+0x40/0x90 [mlx5_core] mlx5e_open_channel+0x7d/0x3c7 [mlx5_core] mlx5e_open_channels+0xad/0x250 [mlx5_core] mlx5e_open_locked+0x3e/0x110 [mlx5_core] mlx5e_open+0x23/0x70 [mlx5_core] __dev_open+0xf1/0x1a5 __dev_change_flags+0x1e1/0x249 dev_change_flags+0x21/0x5c do_setlink+0x28b/0xcc4 ? __nla_parse+0x22/0x3d ? inet6_validate_link_af+0x6b/0x108 ? cpumask_next+0x1f/0x35 ? __snmp6_fill_stats64.constprop.0+0x66/0x107 ? __nla_validate_parse+0x48/0x1e6 __rtnl_newlink+0x5ff/0xa57 ? kmem_cache_alloc_trace+0x164/0x2ce rtnl_newlink+0x44/0x6e rtnetlink_rcv_msg+0x2bb/0x362 ? __netlink_sendskb+0x4c/0x6c ? netlink_unicast+0x28f/0x2ce ? rtnl_calcit.isra.0+0x150/0x146 netlink_rcv_skb+0x5f/0x112 netlink_unicast+0x213/0x2ce netlink_sendmsg+0x24f/0x4d9 __sock_sendmsg+0x65/0x6a ____sys_sendmsg+0x28f/0x2c9 ? import_iovec+0x17/0x2b ___sys_sendmsg+0x97/0xe0 __sys_sendmsg+0x81/0xd8 do_syscall_64+0x35/0x87 entry_SYSCALL_64_after_hwframe+0x6e/0x0 RIP: 0033:0x7fc328603727 Code: c3 66 90 41 54 41 89 d4 55 48 89 f5 53 89 fb 48 83 ec 10 e8 0b ed ff ff 44 89 e2 48 89 ee 89 df 41 89 c0 b8 2e 00 00 00 0f 05 &lt;48&gt; 3d 00 f0 ff ff 77 35 44 89 c7 48 89 44 24 08 e8 44 ed ff ff 48 RSP: 002b:00007ffe8eb3f1a0 EFLAGS: 00000293 ORIG_RAX: 000000000000002e RAX: ffffffffffffffda RBX: 000000000000000d RCX: 00007fc328603727 RDX: 0000000000000000 RSI: 00007ffe8eb3f1f0 RDI: 000000000000000d RBP: 00007ffe8eb3f1f0 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000293 R12: 0000000000000000 R13: 00000000000 ---truncated---</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40250">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40251</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: devlink: rate: Unset parent pointer in devl_rate_nodes_destroy The function devl_rate_nodes_destroy is documented to "Unset parent for all rate objects". However, it was only calling the driver-specific `rate_leaf_parent_set` or `rate_node_parent_set` ops and decrementing the parent's refcount, without actually setting the `devlink_rate-&gt;parent` pointer to NULL. This leaves a dangling pointer in the `devlink_rate` struct, which cause refcount error in netdevsim[1] and mlx5[2]. In addition, this is inconsistent with the behavior of `devlink_nl_rate_parent_node_set`, where the parent pointer is correctly cleared. This patch fixes the issue by explicitly setting `devlink_rate-&gt;parent` to NULL after notifying the driver, thus fulfilling the function's documented behavior for all rate objects. [1] repro steps: echo 1 &gt; /sys/bus/netdevsim/new_device devlink dev eswitch set netdevsim/netdevsim1 mode switchdev echo 1 &gt; /sys/bus/netdevsim/devices/netdevsim1/sriov_numvfs devlink port function rate add netdevsim/netdevsim1/test_node devlink port function rate set netdevsim/netdevsim1/128 parent test_node echo 1 &gt; /sys/bus/netdevsim/del_device dmesg: refcount_t: decrement hit 0; leaking memory. WARNING: CPU: 8 PID: 1530 at lib/refcount.c:31 refcount_warn_saturate+0x42/0xe0 CPU: 8 UID: 0 PID: 1530 Comm: bash Not tainted 6.18.0-rc4+ #1 NONE Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.16.0-0-gd239552ce722-prebuilt.qemu.org 04/01/2014 RIP: 0010:refcount_warn_saturate+0x42/0xe0 Call Trace: devl_rate_leaf_destroy+0x8d/0x90 __nsim_dev_port_del+0x6c/0x70 [netdevsim] nsim_dev_reload_destroy+0x11c/0x140 [netdevsim] nsim_drv_remove+0x2b/0xb0 [netdevsim] device_release_driver_internal+0x194/0x1f0 bus_remove_device+0xc6/0x130 device_del+0x159/0x3c0 device_unregister+0x1a/0x60 del_device_store+0x111/0x170 [netdevsim] kernfs_fop_write_iter+0x12e/0x1e0 vfs_write+0x215/0x3d0 ksys_write+0x5f/0xd0 do_syscall_64+0x55/0x10f0 entry_SYSCALL_64_after_hwframe+0x4b/0x53 [2] devlink dev eswitch set pci/0000:08:00.0 mode switchdev devlink port add pci/0000:08:00.0 flavour pcisf pfnum 0 sfnum 1000 devlink port function rate add pci/0000:08:00.0/group1 devlink port function rate set pci/0000:08:00.0/32768 parent group1 modprobe -r mlx5_ib mlx5_fwctl mlx5_core dmesg: refcount_t: decrement hit 0; leaking memory. WARNING: CPU: 7 PID: 16151 at lib/refcount.c:31 refcount_warn_saturate+0x42/0xe0 CPU: 7 UID: 0 PID: 16151 Comm: bash Not tainted 6.17.0-rc7_for_upstream_min_debug_2025_10_02_12_44 #1 NONE Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.16.3-0-ga6ed6b701f0a-prebuilt.qemu.org 04/01/2014 RIP: 0010:refcount_warn_saturate+0x42/0xe0 Call Trace: devl_rate_leaf_destroy+0x8d/0x90 mlx5_esw_offloads_devlink_port_unregister+0x33/0x60 [mlx5_core] mlx5_esw_offloads_unload_rep+0x3f/0x50 [mlx5_core] mlx5_eswitch_unload_sf_vport+0x40/0x90 [mlx5_core] mlx5_sf_esw_event+0xc4/0x120 [mlx5_core] notifier_call_chain+0x33/0xa0 blocking_notifier_call_chain+0x3b/0x50 mlx5_eswitch_disable_locked+0x50/0x110 [mlx5_core] mlx5_eswitch_disable+0x63/0x90 [mlx5_core] mlx5_unload+0x1d/0x170 [mlx5_core] mlx5_uninit_one+0xa2/0x130 [mlx5_core] remove_one+0x78/0xd0 [mlx5_core] pci_device_remove+0x39/0xa0 device_release_driver_internal+0x194/0x1f0 unbind_store+0x99/0xa0 kernfs_fop_write_iter+0x12e/0x1e0 vfs_write+0x215/0x3d0 ksys_write+0x5f/0xd0 do_syscall_64+0x53/0x1f0 entry_SYSCALL_64_after_hwframe+0x4b/0x53</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40251">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/911.html">CWE-911 Improper Update of Reference Count</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.1</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40252</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: net: qlogic/qede: fix potential out-of-bounds read in qede_tpa_cont() and qede_tpa_end() The loops in 'qede_tpa_cont()' and 'qede_tpa_end()', iterate over 'cqe-&gt;len_list[]' using only a zero-length terminator as the stopping condition. If the terminator was missing or malformed, the loop could run past the end of the fixed-size array. Add an explicit bound check using ARRAY_SIZE() in both loops to prevent a potential out-of-bounds access. Found by Linux Verification Center (linuxtesting.org) with SVACE.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40252">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40254</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: net: openvswitch: remove never-working support for setting nsh fields The validation of the set(nsh(...)) action is completely wrong. It runs through the nsh_key_put_from_nlattr() function that is the same function that validates NSH keys for the flow match and the push_nsh() action. However, the set(nsh(...)) has a very different memory layout. Nested attributes in there are doubled in size in case of the masked set(). That makes proper validation impossible. There is also confusion in the code between the 'masked' flag, that says that the nested attributes are doubled in size containing both the value and the mask, and the 'is_mask' that says that the value we're parsing is the mask. This is causing kernel crash on trying to write into mask part of the match with SW_FLOW_KEY_PUT() during validation, while validate_nsh() doesn't allocate any memory for it: BUG: kernel NULL pointer dereference, address: 0000000000000018 #PF: supervisor read access in kernel mode #PF: error_code(0x0000) - not-present page PGD 1c2383067 P4D 1c2383067 PUD 20b703067 PMD 0 Oops: Oops: 0000 [#1] SMP NOPTI CPU: 8 UID: 0 Kdump: loaded Not tainted 6.17.0-rc4+ #107 PREEMPT(voluntary) RIP: 0010:nsh_key_put_from_nlattr+0x19d/0x610 [openvswitch] Call Trace: validate_nsh+0x60/0x90 [openvswitch] validate_set.constprop.0+0x270/0x3c0 [openvswitch] __ovs_nla_copy_actions+0x477/0x860 [openvswitch] ovs_nla_copy_actions+0x8d/0x100 [openvswitch] ovs_packet_cmd_execute+0x1cc/0x310 [openvswitch] genl_family_rcv_msg_doit+0xdb/0x130 genl_family_rcv_msg+0x14b/0x220 genl_rcv_msg+0x47/0xa0 netlink_rcv_skb+0x53/0x100 genl_rcv+0x24/0x40 netlink_unicast+0x280/0x3b0 netlink_sendmsg+0x1f7/0x430 ____sys_sendmsg+0x36b/0x3a0 ___sys_sendmsg+0x87/0xd0 __sys_sendmsg+0x6d/0xd0 do_syscall_64+0x7b/0x2c0 entry_SYSCALL_64_after_hwframe+0x76/0x7e The third issue with this process is that while trying to convert the non-masked set into masked one, validate_set() copies and doubles the size of the OVS_KEY_ATTR_NSH as if it didn't have any nested attributes. It should be copying each nested attribute and doubling them in size independently. And the process must be properly reversed during the conversion back from masked to a non-masked variant during the flow dump. In the end, the only two outcomes of trying to use this action are either validation failure or a kernel crash. And if somehow someone manages to install a flow with such an action, it will most definitely not do what it is supposed to, since all the keys and the masks are mixed up. Fixing all the issues is a complex task as it requires re-writing most of the validation code. Given that and the fact that this functionality never worked since introduction, let's just remove it altogether. It's better to re-introduce it later with a proper implementation instead of trying to fix it in stable releases.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40254">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40257</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: mptcp: fix a race in mptcp_pm_del_add_timer() mptcp_pm_del_add_timer() can call sk_stop_timer_sync(sk, &amp;entry-&gt;add_timer) while another might have free entry already, as reported by syzbot. Add RCU protection to fix this issue. Also change confusing add_timer variable with stop_timer boolean. syzbot report: BUG: KASAN: slab-use-after-free in __timer_delete_sync+0x372/0x3f0 kernel/time/timer.c:1616 Read of size 4 at addr ffff8880311e4150 by task kworker/1:1/44 CPU: 1 UID: 0 PID: 44 Comm: kworker/1:1 Not tainted syzkaller #0 PREEMPT_{RT,(full)} Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 10/02/2025 Workqueue: events mptcp_worker Call Trace: dump_stack_lvl+0x189/0x250 lib/dump_stack.c:120 print_address_description mm/kasan/report.c:378 [inline] print_report+0xca/0x240 mm/kasan/report.c:482 kasan_report+0x118/0x150 mm/kasan/report.c:595 __timer_delete_sync+0x372/0x3f0 kernel/time/timer.c:1616 sk_stop_timer_sync+0x1b/0x90 net/core/sock.c:3631 mptcp_pm_del_add_timer+0x283/0x310 net/mptcp/pm.c:362 mptcp_incoming_options+0x1357/0x1f60 net/mptcp/options.c:1174 tcp_data_queue+0xca/0x6450 net/ipv4/tcp_input.c:5361 tcp_rcv_established+0x1335/0x2670 net/ipv4/tcp_input.c:6441 tcp_v4_do_rcv+0x98b/0xbf0 net/ipv4/tcp_ipv4.c:1931 tcp_v4_rcv+0x252a/0x2dc0 net/ipv4/tcp_ipv4.c:2374 ip_protocol_deliver_rcu+0x221/0x440 net/ipv4/ip_input.c:205 ip_local_deliver_finish+0x3bb/0x6f0 net/ipv4/ip_input.c:239 NF_HOOK+0x30c/0x3a0 include/linux/netfilter.h:318 NF_HOOK+0x30c/0x3a0 include/linux/netfilter.h:318 __netif_receive_skb_one_core net/core/dev.c:6079 [inline] __netif_receive_skb+0x143/0x380 net/core/dev.c:6192 process_backlog+0x31e/0x900 net/core/dev.c:6544 __napi_poll+0xb6/0x540 net/core/dev.c:7594 napi_poll net/core/dev.c:7657 [inline] net_rx_action+0x5f7/0xda0 net/core/dev.c:7784 handle_softirqs+0x22f/0x710 kernel/softirq.c:622 __do_softirq kernel/softirq.c:656 [inline] __local_bh_enable_ip+0x1a0/0x2e0 kernel/softirq.c:302 mptcp_pm_send_ack net/mptcp/pm.c:210 [inline] mptcp_pm_addr_send_ack+0x41f/0x500 net/mptcp/pm.c:-1 mptcp_pm_worker+0x174/0x320 net/mptcp/pm.c:1002 mptcp_worker+0xd5/0x1170 net/mptcp/protocol.c:2762 process_one_work kernel/workqueue.c:3263 [inline] process_scheduled_works+0xae1/0x17b0 kernel/workqueue.c:3346 worker_thread+0x8a0/0xda0 kernel/workqueue.c:3427 kthread+0x711/0x8a0 kernel/kthread.c:463 ret_from_fork+0x4bc/0x870 arch/x86/kernel/process.c:158 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245 Allocated by task 44: kasan_save_stack mm/kasan/common.c:56 [inline] kasan_save_track+0x3e/0x80 mm/kasan/common.c:77 poison_kmalloc_redzone mm/kasan/common.c:400 [inline] __kasan_kmalloc+0x93/0xb0 mm/kasan/common.c:417 kasan_kmalloc include/linux/kasan.h:262 [inline] __kmalloc_cache_noprof+0x1ef/0x6c0 mm/slub.c:5748 kmalloc_noprof include/linux/slab.h:957 [inline] mptcp_pm_alloc_anno_list+0x104/0x460 net/mptcp/pm.c:385 mptcp_pm_create_subflow_or_signal_addr+0xf9d/0x1360 net/mptcp/pm_kernel.c:355 mptcp_pm_nl_fully_established net/mptcp/pm_kernel.c:409 [inline] __mptcp_pm_kernel_worker+0x417/0x1ef0 net/mptcp/pm_kernel.c:1529 mptcp_pm_worker+0x1ee/0x320 net/mptcp/pm.c:1008 mptcp_worker+0xd5/0x1170 net/mptcp/protocol.c:2762 process_one_work kernel/workqueue.c:3263 [inline] process_scheduled_works+0xae1/0x17b0 kernel/workqueue.c:3346 worker_thread+0x8a0/0xda0 kernel/workqueue.c:3427 kthread+0x711/0x8a0 kernel/kthread.c:463 ret_from_fork+0x4bc/0x870 arch/x86/kernel/process.c:158 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245 Freed by task 6630: kasan_save_stack mm/kasan/common.c:56 [inline] kasan_save_track+0x3e/0x80 mm/kasan/common.c:77 __kasan_save_free_info+0x46/0x50 mm/kasan/generic.c:587 kasan_save_free_info mm/kasan/kasan.h:406 [inline] poison_slab_object m ---truncated---</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40257">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40258</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: mptcp: fix race condition in mptcp_schedule_work() syzbot reported use-after-free in mptcp_schedule_work() [1] Issue here is that mptcp_schedule_work() schedules a work, then gets a refcount on sk-&gt;sk_refcnt if the work was scheduled. This refcount will be released by mptcp_worker(). [A] if (schedule_work(...)) { [B] sock_hold(sk); return true; } Problem is that mptcp_worker() can run immediately and complete before [B] We need instead : sock_hold(sk); if (schedule_work(...)) return true; sock_put(sk); [1] refcount_t: addition on 0; use-after-free. WARNING: CPU: 1 PID: 29 at lib/refcount.c:25 refcount_warn_saturate+0xfa/0x1d0 lib/refcount.c:25 Call Trace: __refcount_add include/linux/refcount.h:-1 [inline] __refcount_inc include/linux/refcount.h:366 [inline] refcount_inc include/linux/refcount.h:383 [inline] sock_hold include/net/sock.h:816 [inline] mptcp_schedule_work+0x164/0x1a0 net/mptcp/protocol.c:943 mptcp_tout_timer+0x21/0xa0 net/mptcp/protocol.c:2316 call_timer_fn+0x17e/0x5f0 kernel/time/timer.c:1747 expire_timers kernel/time/timer.c:1798 [inline] __run_timers kernel/time/timer.c:2372 [inline] __run_timer_base+0x648/0x970 kernel/time/timer.c:2384 run_timer_base kernel/time/timer.c:2393 [inline] run_timer_softirq+0xb7/0x180 kernel/time/timer.c:2403 handle_softirqs+0x22f/0x710 kernel/softirq.c:622 __do_softirq kernel/softirq.c:656 [inline] run_ktimerd+0xcf/0x190 kernel/softirq.c:1138 smpboot_thread_fn+0x542/0xa60 kernel/smpboot.c:160 kthread+0x711/0x8a0 kernel/kthread.c:463 ret_from_fork+0x4bc/0x870 arch/x86/kernel/process.c:158 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40258">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/362.html">CWE-362 Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition')</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40261</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: nvme: nvme-fc: Ensure -&gt;ioerr_work is cancelled in nvme_fc_delete_ctrl() nvme_fc_delete_assocation() waits for pending I/O to complete before returning, and an error can cause -&gt;ioerr_work to be queued after cancel_work_sync() had been called. Move the call to cancel_work_sync() to be after nvme_fc_delete_association() to ensure -&gt;ioerr_work is not running when the nvme_fc_ctrl object is freed. Otherwise the following can occur: [ 1135.911754] list_del corruption, ff2d24c8093f31f8-&gt;next is NULL [ 1135.917705] ------------[ cut here ]------------ [ 1135.922336] kernel BUG at lib/list_debug.c:52! [ 1135.926784] Oops: invalid opcode: 0000 [#1] SMP NOPTI [ 1135.931851] CPU: 48 UID: 0 PID: 726 Comm: kworker/u449:23 Kdump: loaded Not tainted 6.12.0 #1 PREEMPT(voluntary) [ 1135.943490] Hardware name: Dell Inc. PowerEdge R660/0HGTK9, BIOS 2.5.4 01/16/2025 [ 1135.950969] Workqueue: 0x0 (nvme-wq) [ 1135.954673] RIP: 0010:__list_del_entry_valid_or_report.cold+0xf/0x6f [ 1135.961041] Code: c7 c7 98 68 72 94 e8 26 45 fe ff 0f 0b 48 c7 c7 70 68 72 94 e8 18 45 fe ff 0f 0b 48 89 fe 48 c7 c7 80 69 72 94 e8 07 45 fe ff &lt;0f&gt; 0b 48 89 d1 48 c7 c7 a0 6a 72 94 48 89 c2 e8 f3 44 fe ff 0f 0b [ 1135.979788] RSP: 0018:ff579b19482d3e50 EFLAGS: 00010046 [ 1135.985015] RAX: 0000000000000033 RBX: ff2d24c8093f31f0 RCX: 0000000000000000 [ 1135.992148] RDX: 0000000000000000 RSI: ff2d24d6bfa1d0c0 RDI: ff2d24d6bfa1d0c0 [ 1135.999278] RBP: ff2d24c8093f31f8 R08: 0000000000000000 R09: ffffffff951e2b08 [ 1136.006413] R10: ffffffff95122ac8 R11: 0000000000000003 R12: ff2d24c78697c100 [ 1136.013546] R13: fffffffffffffff8 R14: 0000000000000000 R15: ff2d24c78697c0c0 [ 1136.020677] FS: 0000000000000000(0000) GS:ff2d24d6bfa00000(0000) knlGS:0000000000000000 [ 1136.028765] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 1136.034510] CR2: 00007fd207f90b80 CR3: 000000163ea22003 CR4: 0000000000f73ef0 [ 1136.041641] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 [ 1136.048776] DR3: 0000000000000000 DR6: 00000000fffe07f0 DR7: 0000000000000400 [ 1136.055910] PKRU: 55555554 [ 1136.058623] Call Trace: [ 1136.061074] [ 1136.063179] ? show_trace_log_lvl+0x1b0/0x2f0 [ 1136.067540] ? show_trace_log_lvl+0x1b0/0x2f0 [ 1136.071898] ? move_linked_works+0x4a/0xa0 [ 1136.075998] ? __list_del_entry_valid_or_report.cold+0xf/0x6f [ 1136.081744] ? __die_body.cold+0x8/0x12 [ 1136.085584] ? die+0x2e/0x50 [ 1136.088469] ? do_trap+0xca/0x110 [ 1136.091789] ? do_error_trap+0x65/0x80 [ 1136.095543] ? __list_del_entry_valid_or_report.cold+0xf/0x6f [ 1136.101289] ? exc_invalid_op+0x50/0x70 [ 1136.105127] ? __list_del_entry_valid_or_report.cold+0xf/0x6f [ 1136.110874] ? asm_exc_invalid_op+0x1a/0x20 [ 1136.115059] ? __list_del_entry_valid_or_report.cold+0xf/0x6f [ 1136.120806] move_linked_works+0x4a/0xa0 [ 1136.124733] worker_thread+0x216/0x3a0 [ 1136.128485] ? __pfx_worker_thread+0x10/0x10 [ 1136.132758] kthread+0xfa/0x240 [ 1136.135904] ? __pfx_kthread+0x10/0x10 [ 1136.139657] ret_from_fork+0x31/0x50 [ 1136.143236] ? __pfx_kthread+0x10/0x10 [ 1136.146988] ret_from_fork_asm+0x1a/0x30 [ 1136.150915]</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40261">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/1341.html">CWE-1341 Multiple Releases of Same Resource or Handle</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>6.6</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40262</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: Input: imx_sc_key - fix memory corruption on unload This is supposed to be "priv" but we accidentally pass "&amp;priv" which is an address in the stack and so it will lead to memory corruption when the imx_sc_key_action() function is called. Remove the &amp;.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40262">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40263</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: Input: cros_ec_keyb - fix an invalid memory access If cros_ec_keyb_register_matrix() isn't called (due to `buttons_switches_only`) in cros_ec_keyb_probe(), `ckdev-&gt;idev` remains NULL. An invalid memory access is observed in cros_ec_keyb_process() when receiving an EC_MKBP_EVENT_KEY_MATRIX event in cros_ec_keyb_work() in such case. Unable to handle kernel read from unreadable memory at virtual address 0000000000000028 ... x3 : 0000000000000000 x2 : 0000000000000000 x1 : 0000000000000000 x0 : 0000000000000000 Call trace: input_event cros_ec_keyb_work blocking_notifier_call_chain ec_irq_thread It's still unknown about why the kernel receives such malformed event, in any cases, the kernel shouldn't access `ckdev-&gt;idev` and friends if the driver doesn't intend to initialize them.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40263">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40264</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: be2net: pass wrb_params in case of OS2BMC be_insert_vlan_in_pkt() is called with the wrb_params argument being NULL at be_send_pkt_to_bmc() call site.  This may lead to dereferencing a NULL pointer when processing a workaround for specific packet, as commit bc0c3405abbb ("be2net: fix a Tx stall bug caused by a specific ipv6 packet") states. The correct way would be to pass the wrb_params from be_xmit().</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40264">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40271</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: fs/proc: fix uaf in proc_readdir_de() Pde is erased from subdir rbtree through rb_erase(), but not set the node to EMPTY, which may result in uaf access. We should use RB_CLEAR_NODE() set the erased node to EMPTY, then pde_subdir_next() will return NULL to avoid uaf access. We found an uaf issue while using stress-ng testing, need to run testcase getdent and tun in the same time. The steps of the issue is as follows: 1) use getdent to traverse dir /proc/pid/net/dev_snmp6/, and current pde is tun3; 2) in the [time windows] unregister netdevice tun3 and tun2, and erase them from rbtree. erase tun3 first, and then erase tun2. the pde(tun2) will be released to slab; 3) continue to getdent process, then pde_subdir_next() will return pde(tun2) which is released, it will case uaf access. CPU 0 | CPU 1 ------------------------------------------------------------------------- traverse dir /proc/pid/net/dev_snmp6/ | unregister_netdevice(tun-&gt;dev) //tun3 tun2 sys_getdents64() | iterate_dir() | proc_readdir() | proc_readdir_de() | snmp6_unregister_dev() pde_get(de); | proc_remove() read_unlock(&amp;proc_subdir_lock); | remove_proc_subtree() | write_lock(&amp;proc_subdir_lock); [time window] | rb_erase(&amp;root-&gt;subdir_node, &amp;parent-&gt;subdir); | write_unlock(&amp;proc_subdir_lock); read_lock(&amp;proc_subdir_lock); | next = pde_subdir_next(de); | pde_put(de); | de = next; //UAF | rbtree of dev_snmp6 | pde(tun3) / \ NULL pde(tun2)</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40271">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/625.html">CWE-625 Permissive Regular Expression</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40278</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: net: sched: act_ife: initialize struct tc_ife to fix KMSAN kernel-infoleak Fix a KMSAN kernel-infoleak detected by the syzbot . [net?] KMSAN: kernel-infoleak in __skb_datagram_iter In tcf_ife_dump(), the variable 'opt' was partially initialized using a designatied initializer. While the padding bytes are reamined uninitialized. nla_put() copies the entire structure into a netlink message, these uninitialized bytes leaked to userspace. Initialize the structure with memset before assigning its fields to ensure all members and padding are cleared prior to beign copied. This change silences the KMSAN report and prevents potential information leaks from the kernel memory. This fix has been tested and validated by syzbot. This patch closes the bug reported at the following syzkaller link and ensures no infoleak.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40278">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40280</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: tipc: Fix use-after-free in tipc_mon_reinit_self(). syzbot reported use-after-free of tipc_net(net)-&gt;monitors[] in tipc_mon_reinit_self(). [0] The array is protected by RTNL, but tipc_mon_reinit_self() iterates over it without RTNL. tipc_mon_reinit_self() is called from tipc_net_finalize(), which is always under RTNL except for tipc_net_finalize_work(). Let's hold RTNL in tipc_net_finalize_work(). [0]: BUG: KASAN: slab-use-after-free in __raw_spin_lock_irqsave include/linux/spinlock_api_smp.h:110 [inline] BUG: KASAN: slab-use-after-free in _raw_spin_lock_irqsave+0xa7/0xf0 kernel/locking/spinlock.c:162 Read of size 1 at addr ffff88805eae1030 by task kworker/0:7/5989 CPU: 0 UID: 0 PID: 5989 Comm: kworker/0:7 Not tainted syzkaller #0 PREEMPT_{RT,(full)} Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 08/18/2025 Workqueue: events tipc_net_finalize_work Call Trace: dump_stack_lvl+0x189/0x250 lib/dump_stack.c:120 print_address_description mm/kasan/report.c:378 [inline] print_report+0xca/0x240 mm/kasan/report.c:482 kasan_report+0x118/0x150 mm/kasan/report.c:595 __kasan_check_byte+0x2a/0x40 mm/kasan/common.c:568 kasan_check_byte include/linux/kasan.h:399 [inline] lock_acquire+0x8d/0x360 kernel/locking/lockdep.c:5842 __raw_spin_lock_irqsave include/linux/spinlock_api_smp.h:110 [inline] _raw_spin_lock_irqsave+0xa7/0xf0 kernel/locking/spinlock.c:162 rtlock_slowlock kernel/locking/rtmutex.c:1894 [inline] rwbase_rtmutex_lock_state kernel/locking/spinlock_rt.c:160 [inline] rwbase_write_lock+0xd3/0x7e0 kernel/locking/rwbase_rt.c:244 rt_write_lock+0x76/0x110 kernel/locking/spinlock_rt.c:243 write_lock_bh include/linux/rwlock_rt.h:99 [inline] tipc_mon_reinit_self+0x79/0x430 net/tipc/monitor.c:718 tipc_net_finalize+0x115/0x190 net/tipc/net.c:140 process_one_work kernel/workqueue.c:3236 [inline] process_scheduled_works+0xade/0x17b0 kernel/workqueue.c:3319 worker_thread+0x8a0/0xda0 kernel/workqueue.c:3400 kthread+0x70e/0x8a0 kernel/kthread.c:463 ret_from_fork+0x439/0x7d0 arch/x86/kernel/process.c:148 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245 Allocated by task 6089: kasan_save_stack mm/kasan/common.c:47 [inline] kasan_save_track+0x3e/0x80 mm/kasan/common.c:68 poison_kmalloc_redzone mm/kasan/common.c:388 [inline] __kasan_kmalloc+0x93/0xb0 mm/kasan/common.c:405 kasan_kmalloc include/linux/kasan.h:260 [inline] __kmalloc_cache_noprof+0x1a8/0x320 mm/slub.c:4407 kmalloc_noprof include/linux/slab.h:905 [inline] kzalloc_noprof include/linux/slab.h:1039 [inline] tipc_mon_create+0xc3/0x4d0 net/tipc/monitor.c:657 tipc_enable_bearer net/tipc/bearer.c:357 [inline] __tipc_nl_bearer_enable+0xe16/0x13f0 net/tipc/bearer.c:1047 __tipc_nl_compat_doit net/tipc/netlink_compat.c:371 [inline] tipc_nl_compat_doit+0x3bc/0x5f0 net/tipc/netlink_compat.c:393 tipc_nl_compat_handle net/tipc/netlink_compat.c:-1 [inline] tipc_nl_compat_recv+0x83c/0xbe0 net/tipc/netlink_compat.c:1321 genl_family_rcv_msg_doit+0x215/0x300 net/netlink/genetlink.c:1115 genl_family_rcv_msg net/netlink/genetlink.c:1195 [inline] genl_rcv_msg+0x60e/0x790 net/netlink/genetlink.c:1210 netlink_rcv_skb+0x208/0x470 net/netlink/af_netlink.c:2552 genl_rcv+0x28/0x40 net/netlink/genetlink.c:1219 netlink_unicast_kernel net/netlink/af_netlink.c:1320 [inline] netlink_unicast+0x846/0xa10 net/netlink/af_netlink.c:1346 netlink_sendmsg+0x805/0xb30 net/netlink/af_netlink.c:1896 sock_sendmsg_nosec net/socket.c:714 [inline] __sock_sendmsg+0x21c/0x270 net/socket.c:729 ____sys_sendmsg+0x508/0x820 net/socket.c:2614 ___sys_sendmsg+0x21f/0x2a0 net/socket.c:2668 __sys_sendmsg net/socket.c:2700 [inline] __do_sys_sendmsg net/socket.c:2705 [inline] __se_sys_sendmsg net/socket.c:2703 [inline] __x64_sys_sendmsg+0x1a1/0x260 net/socket.c:2703 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0xfa/0x3b0 arch/ ---truncated---</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40280">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/825.html">CWE-825 Expired Pointer Dereference</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40281</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: sctp: prevent possible shift-out-of-bounds in sctp_transport_update_rto syzbot reported a possible shift-out-of-bounds [1] Blamed commit added rto_alpha_max and rto_beta_max set to 1000. It is unclear if some sctp users are setting very large rto_alpha and/or rto_beta. In order to prevent user regression, perform the test at run time. Also add READ_ONCE() annotations as sysctl values can change under us. [1] UBSAN: shift-out-of-bounds in net/sctp/transport.c:509:41 shift exponent 64 is too large for 32-bit type 'unsigned int' CPU: 0 UID: 0 PID: 16704 Comm: syz.2.2320 Not tainted syzkaller #0 PREEMPT(full) Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 10/02/2025 Call Trace: __dump_stack lib/dump_stack.c:94 [inline] dump_stack_lvl+0x16c/0x1f0 lib/dump_stack.c:120 ubsan_epilogue lib/ubsan.c:233 [inline] __ubsan_handle_shift_out_of_bounds+0x27f/0x420 lib/ubsan.c:494 sctp_transport_update_rto.cold+0x1c/0x34b net/sctp/transport.c:509 sctp_check_transmitted+0x11c4/0x1c30 net/sctp/outqueue.c:1502 sctp_outq_sack+0x4ef/0x1b20 net/sctp/outqueue.c:1338 sctp_cmd_process_sack net/sctp/sm_sideeffect.c:840 [inline] sctp_cmd_interpreter net/sctp/sm_sideeffect.c:1372 [inline]</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40281">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/1335.html">CWE-1335 Incorrect Bitwise Shift of Integer</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>4.4</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-40345</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: usb: storage: sddr55: Reject out-of-bound new_pba Discovered by Atuin - Automated Vulnerability Discovery Engine. new_pba comes from the status packet returned after each write. A bogus device could report values beyond the block count derived from info-&gt;capacity, letting the driver walk off the end of pba_to_lba[] and corrupt heap memory. Reject PBAs that exceed the computed block count and fail the transfer so we avoid touching out-of-range mapping entries.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-40345">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/787.html">CWE-787 Out-of-bounds Write</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>6.8</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:P/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:P/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-46394</a></h3>
<div class="csaf-accordion-content">
<p>In tar in BusyBox through 1.37.0, a TAR archive can have filenames hidden from a listing through the use of terminal escape sequences.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-46394">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/451.html">CWE-451 User Interface (UI) Misrepresentation of Critical Information</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>3.2</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:C/C:N/I:L/A:N">CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:C/C:N/I:L/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-49794</a></h3>
<div class="csaf-accordion-content">
<p>A use-after-free vulnerability was found in libxml2. This issue occurs when parsing XPath elements under certain circumstances when the XML schematron has the schema elements. This flaw allows a malicious actor to craft a malicious XML document used as input for libxml, resulting in the program's crash using libxml or other possible undefined behaviors.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-49794">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/825.html">CWE-825 Expired Pointer Dereference</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>9.1</td>
<td>CRITICAL</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-49795</a></h3>
<div class="csaf-accordion-content">
<p>A NULL pointer dereference vulnerability was found in libxml2 when processing XPath XML expressions. This flaw allows an attacker to craft a malicious XML input to libxml2, leading to a denial of service.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-49795">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/825.html">CWE-825 Expired Pointer Dereference</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.5</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-49796</a></h3>
<div class="csaf-accordion-content">
<p>A vulnerability was found in libxml2. Processing certain sch:name elements from the input XML file can trigger a memory corruption issue. This flaw allows an attacker to craft a malicious XML input file that can lead libxml to crash, resulting in a denial of service or other possible undefined behavior due to sensitive data being corrupted in memory.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-49796">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/125.html">CWE-125 Out-of-bounds Read</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>9.1</td>
<td>CRITICAL</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-60876</a></h3>
<div class="csaf-accordion-content">
<p>BusyBox wget thru 1.3.7 accepted raw CR (0x0D)/LF (0x0A) and other C0 control bytes in the HTTP request-target (path/query), allowing the request line to be split and attacker-controlled headers to be injected. To preserve the HTTP/1.1 request-line shape METHOD SP request-target SP HTTP/1.1, a raw space (0x20) in the request-target must also be rejected (clients should use %20).</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-60876">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/284.html">CWE-284 Improper Access Control</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>6.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-66035</a></h3>
<div class="csaf-accordion-content">
<p>Angular is a development platform for building mobile and desktop web applications using TypeScript/JavaScript and other languages. Prior to versions 19.2.16, 20.3.14, and 21.0.1, there is a XSRF token leakage via protocol-relative URLs in angular HTTP clients. The vulnerability is a Credential Leak by App Logic that leads to the unauthorized disclosure of the Cross-Site Request Forgery (XSRF) token to an attacker-controlled domain. Angular's HttpClient has a built-in XSRF protection mechanism that works by checking if a request URL starts with a protocol (http:// or https://) to determine if it is cross-origin. If the URL starts with protocol-relative URL (//), it is incorrectly treated as a same-origin request, and the XSRF token is automatically added to the X-XSRF-TOKEN header. This issue has been patched in versions 19.2.16, 20.3.14, and 21.0.1. A workaround for this issue involves avoiding using protocol-relative URLs (URLs starting with //) in HttpClient requests. All backend communication URLs should be hardcoded as relative paths (starting with a single /) or fully qualified, trusted absolute URLs.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-66035">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/201.html">CWE-201 Insertion of Sensitive Information Into Sent Data</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>8.6</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:N/A:N">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:N/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-66382</a></h3>
<div class="csaf-accordion-content">
<p>In libexpat through 2.7.3, a crafted file with an approximate size of 2 MiB can lead to dozens of seconds of processing time.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-66382">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/407.html">CWE-407 Inefficient Algorithmic Complexity</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>2.9</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L">CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-66412</a></h3>
<div class="csaf-accordion-content">
<p>Angular is a development platform for building mobile and desktop web applications using TypeScript/JavaScript and other languages. Prior to 21.0.2, 20.3.15, and 19.2.17, A Stored Cross-Site Scripting (XSS) vulnerability has been identified in the Angular Template Compiler. It occurs because the compiler's internal security schema is incomplete, allowing attackers to bypass Angular's built-in security sanitization. Specifically, the schema fails to classify certain URL-holding attributes (e.g., those that could contain javascript: URLs) as requiring strict URL security, enabling the injection of malicious scripts. This vulnerability is fixed in 21.0.2, 20.3.15, and 19.2.17.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-66412">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/79.html">CWE-79 Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H">CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-69720</a></h3>
<div class="csaf-accordion-content">
<p>The infocmp command-line tool in ncurses before 6.5-20251213 has a stack-based buffer overflow in analyze_string in progs/infocmp.c.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-69720">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/121.html">CWE-121 Stack-based Buffer Overflow</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.3</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:L">CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-71185</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: ti: dma-crossbar: fix device leak on am335x route allocation Make sure to drop the reference taken when looking up the crossbar platform device during am335x route allocation.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-71185">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-71186</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: stm32: dmamux: fix device leak on route allocation Make sure to drop the reference taken when looking up the DMA mux platform device during route allocation. Note that holding a reference to a device does not prevent its driver data from going away so there is no point in keeping the reference.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-71186">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-71188</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: lpc18xx-dmamux: fix device leak on route allocation Make sure to drop the reference taken when looking up the DMA mux platform device during route allocation. Note that holding a reference to a device does not prevent its driver data from going away so there is no point in keeping the reference.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-71188">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-71189</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: dw: dmamux: fix OF node leak on route allocation failure Make sure to drop the reference taken to the DMA master OF node also on late route allocation failures.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-71189">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-71190</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: bcm-sba-raid: fix device leak on probe Make sure to drop the reference taken when looking up the mailbox device during probe on probe failures and on driver unbind.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-71190">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2025-71191</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: at_hdmac: fix device leak on of_dma_xlate() Make sure to drop the reference taken when looking up the DMA platform device during of_dma_xlate() when releasing channel resources. Note that commit 3832b78b3ec2 ("dmaengine: at_hdmac: add missing put_device() call in at_dma_xlate()") fixed the leak in a couple of error paths but the reference is still leaking on successful allocation.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2025-71191">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-1484</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in the GLib Base64 encoding routine when processing very large input data. Due to incorrect use of integer types during length calculation, the library may miscalculate buffer boundaries. This can cause memory writes outside the allocated buffer. Applications that process untrusted or extremely large Base64 input using GLib may crash or behave unpredictably.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-1484">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/787.html">CWE-787 Out-of-bounds Write</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>4.2</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:L/A:L">CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:L/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-1489</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in GLib. An integer overflow vulnerability in its Unicode case conversion implementation can lead to memory corruption. By processing specially crafted and extremely large Unicode strings, an attacker could trigger an undersized memory allocation, resulting in out-of-bounds writes. This could cause applications utilizing GLib for string conversion to crash or become unstable.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-1489">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/787.html">CWE-787 Out-of-bounds Write</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.4</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:L">CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-3784</a></h3>
<div class="csaf-accordion-content">
<p>curl would wrongly reuse an existing HTTP proxy connection doing CONNECT to a server, even if the new request uses different credentials for the HTTP proxy. The proper behavior is to create or use a separate connection.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-3784">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/305.html">CWE-305 Authentication Bypass by Primary Weakness</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>6.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-22610</a></h3>
<div class="csaf-accordion-content">
<p>Angular is a development platform for building mobile and desktop web applications using TypeScript/JavaScript and other languages. Prior to versions 19.2.18, 20.3.16, 21.0.7, and 21.1.0-rc.0, a cross-site scripting (XSS) vulnerability has been identified in the Angular Template Compiler. The vulnerability exists because Angular’s internal sanitization schema fails to recognize the href and xlink:href attributes of SVG elements as a Resource URL context. This issue has been patched in versions 19.2.18, 20.3.16, 21.0.7, and 21.1.0-rc.0.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-22610">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/79.html">CWE-79 Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H">CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-22976</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: net/sched: sch_qfq: Fix NULL deref when deactivating inactive aggregate in qfq_reset `qfq_class-&gt;leaf_qdisc-&gt;q.qlen &gt; 0` does not imply that the class itself is active. Two qfq_class objects may point to the same leaf_qdisc. This happens when: 1. one QFQ qdisc is attached to the dev as the root qdisc, and 2. another QFQ qdisc is temporarily referenced (e.g., via qdisc_get() / qdisc_put()) and is pending to be destroyed, as in function tc_new_tfilter. When packets are enqueued through the root QFQ qdisc, the shared leaf_qdisc-&gt;q.qlen increases. At the same time, the second QFQ qdisc triggers qdisc_put and qdisc_destroy: the qdisc enters qfq_reset() with its own q-&gt;q.qlen == 0, but its class's leaf qdisc-&gt;q.qlen &gt; 0. Therefore, the qfq_reset would wrongly deactivate an inactive aggregate and trigger a null-deref in qfq_deactivate_agg: [ 0.903172] BUG: kernel NULL pointer dereference, address: 0000000000000000 [ 0.903571] #PF: supervisor write access in kernel mode [ 0.903860] #PF: error_code(0x0002) - not-present page [ 0.904177] PGD 10299b067 P4D 10299b067 PUD 10299c067 PMD 0 [ 0.904502] Oops: Oops: 0002 [#1] SMP NOPTI [ 0.904737] CPU: 0 UID: 0 PID: 135 Comm: exploit Not tainted 6.19.0-rc3+ #2 NONE [ 0.905157] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS rel-1.17.0-0-gb52ca86e094d-prebuilt.qemu.org 04/01/2014 [ 0.905754] RIP: 0010:qfq_deactivate_agg (include/linux/list.h:992 (discriminator 2) include/linux/list.h:1006 (discriminator 2) net/sched/sch_qfq.c:1367 (discriminator 2) net/sched/sch_qfq.c:1393 (discriminator 2)) [ 0.906046] Code: 0f 84 4d 01 00 00 48 89 70 18 8b 4b 10 48 c7 c2 ff ff ff ff 48 8b 78 08 48 d3 e2 48 21 f2 48 2b 13 48 8b 30 48 d3 ea 8b 4b 18 0 Code starting with the faulting instruction =========================================== 0: 0f 84 4d 01 00 00 je 0x153 6: 48 89 70 18 mov %rsi,0x18(%rax) a: 8b 4b 10 mov 0x10(%rbx),%ecx d: 48 c7 c2 ff ff ff ff mov $0xffffffffffffffff,%rdx 14: 48 8b 78 08 mov 0x8(%rax),%rdi 18: 48 d3 e2 shl %cl,%rdx 1b: 48 21 f2 and %rsi,%rdx 1e: 48 2b 13 sub (%rbx),%rdx 21: 48 8b 30 mov (%rax),%rsi 24: 48 d3 ea shr %cl,%rdx 27: 8b 4b 18 mov 0x18(%rbx),%ecx ... [ 0.907095] RSP: 0018:ffffc900004a39a0 EFLAGS: 00010246 [ 0.907368] RAX: ffff8881043a0880 RBX: ffff888102953340 RCX: 0000000000000000 [ 0.907723] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000000 [ 0.908100] RBP: ffff888102952180 R08: 0000000000000000 R09: 0000000000000000 [ 0.908451] R10: ffff8881043a0000 R11: 0000000000000000 R12: ffff888102952000 [ 0.908804] R13: ffff888102952180 R14: ffff8881043a0ad8 R15: ffff8881043a0880 [ 0.909179] FS: 000000002a1a0380(0000) GS:ffff888196d8d000(0000) knlGS:0000000000000000 [ 0.909572] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 0.909857] CR2: 0000000000000000 CR3: 0000000102993002 CR4: 0000000000772ef0 [ 0.910247] PKRU: 55555554 [ 0.910391] Call Trace: [ 0.910527] [ 0.910638] qfq_reset_qdisc (net/sched/sch_qfq.c:357 net/sched/sch_qfq.c:1485) [ 0.910826] qdisc_reset (include/linux/skbuff.h:2195 include/linux/skbuff.h:2501 include/linux/skbuff.h:3424 include/linux/skbuff.h:3430 net/sched/sch_generic.c:1036) [ 0.911040] __qdisc_destroy (net/sched/sch_generic.c:1076) [ 0.911236] tc_new_tfilter (net/sched/cls_api.c:2447) [ 0.911447] rtnetlink_rcv_msg (net/core/rtnetlink.c:6958) [ 0.911663] ? __pfx_rtnetlink_rcv_msg (net/core/rtnetlink.c:6861) [ 0.911894] netlink_rcv_skb (net/netlink/af_netlink.c:2550) [ 0.912100] netlink_unicast (net/netlink/af_netlink.c:1319 net/netlink/af_netlink.c:1344) [ 0.912296] ? __alloc_skb (net/core/skbuff.c:706) [ 0.912484] netlink_sendmsg (net/netlink/af ---truncated---</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-22976">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/476.html">CWE-476 NULL Pointer Dereference</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-22977</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: net: sock: fix hardened usercopy panic in sock_recv_errqueue skbuff_fclone_cache was created without defining a usercopy region, [1] unlike skbuff_head_cache which properly whitelists the cb[] field. [2] This causes a usercopy BUG() when CONFIG_HARDENED_USERCOPY is enabled and the kernel attempts to copy sk_buff.cb data to userspace via sock_recv_errqueue() -&gt; put_cmsg(). The crash occurs when: 1. TCP allocates an skb using alloc_skb_fclone() (from skbuff_fclone_cache) [1] 2. The skb is cloned via skb_clone() using the pre-allocated fclone [3] 3. The cloned skb is queued to sk_error_queue for timestamp reporting 4. Userspace reads the error queue via recvmsg(MSG_ERRQUEUE) 5. sock_recv_errqueue() calls put_cmsg() to copy serr-&gt;ee from skb-&gt;cb [4] 6. __check_heap_object() fails because skbuff_fclone_cache has no usercopy whitelist [5] When cloned skbs allocated from skbuff_fclone_cache are used in the socket error queue, accessing the sock_exterr_skb structure in skb-&gt;cb via put_cmsg() triggers a usercopy hardening violation: [ 5.379589] usercopy: Kernel memory exposure attempt detected from SLUB object 'skbuff_fclone_cache' (offset 296, size 16)! [ 5.382796] kernel BUG at mm/usercopy.c:102! [ 5.383923] Oops: invalid opcode: 0000 [#1] SMP KASAN NOPTI [ 5.384903] CPU: 1 UID: 0 PID: 138 Comm: poc_put_cmsg Not tainted 6.12.57 #7 [ 5.384903] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS rel-1.16.3-0-ga6ed6b701f0a-prebuilt.qemu.org 04/01/2014 [ 5.384903] RIP: 0010:usercopy_abort+0x6c/0x80 [ 5.384903] Code: 1a 86 51 48 c7 c2 40 15 1a 86 41 52 48 c7 c7 c0 15 1a 86 48 0f 45 d6 48 c7 c6 80 15 1a 86 48 89 c1 49 0f 45 f3 e8 84 27 88 ff &lt;0f&gt; 0b 490 [ 5.384903] RSP: 0018:ffffc900006f77a8 EFLAGS: 00010246 [ 5.384903] RAX: 000000000000006f RBX: ffff88800f0ad2a8 RCX: 1ffffffff0f72e74 [ 5.384903] RDX: 0000000000000000 RSI: 0000000000000004 RDI: ffffffff87b973a0 [ 5.384903] RBP: 0000000000000010 R08: 0000000000000000 R09: fffffbfff0f72e74 [ 5.384903] R10: 0000000000000003 R11: 79706f6372657375 R12: 0000000000000001 [ 5.384903] R13: ffff88800f0ad2b8 R14: ffffea00003c2b40 R15: ffffea00003c2b00 [ 5.384903] FS: 0000000011bc4380(0000) GS:ffff8880bf100000(0000) knlGS:0000000000000000 [ 5.384903] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 5.384903] CR2: 000056aa3b8e5fe4 CR3: 000000000ea26004 CR4: 0000000000770ef0 [ 5.384903] PKRU: 55555554 [ 5.384903] Call Trace: [ 5.384903] [ 5.384903] __check_heap_object+0x9a/0xd0 [ 5.384903] __check_object_size+0x46c/0x690 [ 5.384903] put_cmsg+0x129/0x5e0 [ 5.384903] sock_recv_errqueue+0x22f/0x380 [ 5.384903] tls_sw_recvmsg+0x7ed/0x1960 [ 5.384903] ? srso_alias_return_thunk+0x5/0xfbef5 [ 5.384903] ? schedule+0x6d/0x270 [ 5.384903] ? srso_alias_return_thunk+0x5/0xfbef5 [ 5.384903] ? mutex_unlock+0x81/0xd0 [ 5.384903] ? __pfx_mutex_unlock+0x10/0x10 [ 5.384903] ? __pfx_tls_sw_recvmsg+0x10/0x10 [ 5.384903] ? _raw_spin_lock_irqsave+0x8f/0xf0 [ 5.384903] ? _raw_read_unlock_irqrestore+0x20/0x40 [ 5.384903] ? srso_alias_return_thunk+0x5/0xfbef5 The crash offset 296 corresponds to skb2-&gt;cb within skbuff_fclones: - sizeof(struct sk_buff) = 232 - offsetof(struct sk_buff, cb) = 40 - offset of skb2.cb in fclones = 232 + 40 = 272 - crash offset 296 = 272 + 24 (inside sock_exterr_skb.ee) This patch uses a local stack variable as a bounce buffer to avoid the hardened usercopy check failure. [1] https://elixir.bootlin.com/linux/v6.12.62/source/net/ipv4/tcp.c#L885 [2] https://elixir.bootlin.com/linux/v6.12.62/source/net/core/skbuff.c#L5104 [3] https://elixir.bootlin.com/linux/v6.12.62/source/net/core/skbuff.c#L5566 [4] https://elixir.bootlin.com/linux/v6.12.62/source/net/core/skbuff.c#L5491 [5] https://elixir.bootlin.com/linux/v6.12.62/source/mm/slub.c#L5719</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-22977">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/489.html">CWE-489 Active Debug Code</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23025</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: mm/page_alloc: prevent pcp corruption with SMP=n The kernel test robot has reported: BUG: spinlock trylock failure on UP on CPU#0, kcompactd0/28 lock: 0xffff888807e35ef0, .magic: dead4ead, .owner: kcompactd0/28, .owner_cpu: 0 CPU: 0 UID: 0 PID: 28 Comm: kcompactd0 Not tainted 6.18.0-rc5-00127-ga06157804399 #1 PREEMPT 8cc09ef94dcec767faa911515ce9e609c45db470 Call Trace: __dump_stack (lib/dump_stack.c:95) dump_stack_lvl (lib/dump_stack.c:123) dump_stack (lib/dump_stack.c:130) spin_dump (kernel/locking/spinlock_debug.c:71) do_raw_spin_trylock (kernel/locking/spinlock_debug.c:?) _raw_spin_trylock (include/linux/spinlock_api_smp.h:89 kernel/locking/spinlock.c:138) __free_frozen_pages (mm/page_alloc.c:2973) ___free_pages (mm/page_alloc.c:5295) __free_pages (mm/page_alloc.c:5334) tlb_remove_table_rcu (include/linux/mm.h:? include/linux/mm.h:3122 include/asm-generic/tlb.h:220 mm/mmu_gather.c:227 mm/mmu_gather.c:290) ? __cfi_tlb_remove_table_rcu (mm/mmu_gather.c:289) ? rcu_core (kernel/rcu/tree.c:?) rcu_core (include/linux/rcupdate.h:341 kernel/rcu/tree.c:2607 kernel/rcu/tree.c:2861) rcu_core_si (kernel/rcu/tree.c:2879) handle_softirqs (arch/x86/include/asm/jump_label.h:36 include/trace/events/irq.h:142 kernel/softirq.c:623) __irq_exit_rcu (arch/x86/include/asm/jump_label.h:36 kernel/softirq.c:725) irq_exit_rcu (kernel/softirq.c:741) sysvec_apic_timer_interrupt (arch/x86/kernel/apic/apic.c:1052) RIP: 0010:_raw_spin_unlock_irqrestore (arch/x86/include/asm/preempt.h:95 include/linux/spinlock_api_smp.h:152 kernel/locking/spinlock.c:194) free_pcppages_bulk (mm/page_alloc.c:1494) drain_pages_zone (include/linux/spinlock.h:391 mm/page_alloc.c:2632) __drain_all_pages (mm/page_alloc.c:2731) drain_all_pages (mm/page_alloc.c:2747) kcompactd (mm/compaction.c:3115) kthread (kernel/kthread.c:465) ? __cfi_kcompactd (mm/compaction.c:3166) ? __cfi_kthread (kernel/kthread.c:412) ret_from_fork (arch/x86/kernel/process.c:164) ? __cfi_kthread (kernel/kthread.c:412) ret_from_fork_asm (arch/x86/entry/entry_64.S:255) Matthew has analyzed the report and identified that in drain_page_zone() we are in a section protected by spin_lock(&amp;pcp-&gt;lock) and then get an interrupt that attempts spin_trylock() on the same lock. The code is designed to work this way without disabling IRQs and occasionally fail the trylock with a fallback. However, the SMP=n spinlock implementation assumes spin_trylock() will always succeed, and thus it's normally a no-op. Here the enabled lock debugging catches the problem, but otherwise it could cause a corruption of the pcp structure. The problem has been introduced by commit 574907741599 ("mm/page_alloc: leave IRQs enabled for per-cpu page allocations"). The pcp locking scheme recognizes the need for disabling IRQs to prevent nesting spin_trylock() sections on SMP=n, but the need to prevent the nesting in spin_lock() has not been recognized. Fix it by introducing local wrappers that change the spin_lock() to spin_lock_iqsave() with SMP=n and use them in all places that do spin_lock(&amp;pcp-&gt;lock). [vbabka@suse.cz: add pcp_ prefix to the spin_lock_irqsave wrappers, per Steven]</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23025">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23026</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: qcom: gpi: Fix memory leak in gpi_peripheral_config() Fix a memory leak in gpi_peripheral_config() where the original memory pointed to by gchan-&gt;config could be lost if krealloc() fails. The issue occurs when: 1. gchan-&gt;config points to previously allocated memory 2. krealloc() fails and returns NULL 3. The function directly assigns NULL to gchan-&gt;config, losing the reference to the original memory 4. The original memory becomes unreachable and cannot be freed Fix this by using a temporary variable to hold the krealloc() result and only updating gchan-&gt;config when the allocation succeeds. Found via static analysis and code review.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23026">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23030</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: phy: rockchip: inno-usb2: Fix a double free bug in rockchip_usb2phy_probe() The for_each_available_child_of_node() calls of_node_put() to release child_np in each success loop. After breaking from the loop with the child_np has been released, the code will jump to the put_child label and will call the of_node_put() again if the devm_request_threaded_irq() fails. These cause a double free bug. Fix by returning directly to avoid the duplicate of_node_put().</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23030">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23031</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: can: gs_usb: gs_usb_receive_bulk_callback(): fix URB memory leak In gs_can_open(), the URBs for USB-in transfers are allocated, added to the parent-&gt;rx_submitted anchor and submitted. In the complete callback gs_usb_receive_bulk_callback(), the URB is processed and resubmitted. In gs_can_close() the URBs are freed by calling usb_kill_anchored_urbs(parent-&gt;rx_submitted). However, this does not take into account that the USB framework unanchors the URB before the complete function is called. This means that once an in-URB has been completed, it is no longer anchored and is ultimately not released in gs_can_close(). Fix the memory leak by anchoring the URB in the gs_usb_receive_bulk_callback() to the parent-&gt;rx_submitted anchor.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23031">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23032</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: null_blk: fix kmemleak by releasing references to fault configfs items When CONFIG_BLK_DEV_NULL_BLK_FAULT_INJECTION is enabled, the null-blk driver sets up fault injection support by creating the timeout_inject, requeue_inject, and init_hctx_fault_inject configfs items as children of the top-level nullbX configfs group. However, when the nullbX device is removed, the references taken to these fault-config configfs items are not released. As a result, kmemleak reports a memory leak, for example: unreferenced object 0xc00000021ff25c40 (size 32): comm "mkdir", pid 10665, jiffies 4322121578 hex dump (first 32 bytes): 69 6e 69 74 5f 68 63 74 78 5f 66 61 75 6c 74 5f init_hctx_fault_ 69 6e 6a 65 63 74 00 88 00 00 00 00 00 00 00 00 inject.......... backtrace (crc 1a018c86): __kmalloc_node_track_caller_noprof+0x494/0xbd8 kvasprintf+0x74/0xf4 config_item_set_name+0xf0/0x104 config_group_init_type_name+0x48/0xfc fault_config_init+0x48/0xf0 0xc0080000180559e4 configfs_mkdir+0x304/0x814 vfs_mkdir+0x49c/0x604 do_mkdirat+0x314/0x3d0 sys_mkdir+0xa0/0xd8 system_call_exception+0x1b0/0x4f0 system_call_vectored_common+0x15c/0x2ec Fix this by explicitly releasing the references to the fault-config configfs items when dropping the reference to the top-level nullbX configfs group.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23032">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23033</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: dmaengine: omap-dma: fix dma_pool resource leak in error paths The dma_pool created by dma_pool_create() is not destroyed when dma_async_device_register() or of_dma_controller_register() fails, causing a resource leak in the probe error paths. Add dma_pool_destroy() in both error paths to properly release the allocated dma_pool resource.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23033">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23037</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: can: etas_es58x: allow partial RX URB allocation to succeed When es58x_alloc_rx_urbs() fails to allocate the requested number of URBs but succeeds in allocating some, it returns an error code. This causes es58x_open() to return early, skipping the cleanup label 'free_urbs', which leads to the anchored URBs being leaked. As pointed out by maintainer Vincent Mailhol, the driver is designed to handle partial URB allocation gracefully. Therefore, partial allocation should not be treated as a fatal error. Modify es58x_alloc_rx_urbs() to return 0 if at least one URB has been allocated, restoring the intended behavior and preventing the leak in es58x_open().</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23037">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23038</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: pnfs/flexfiles: Fix memory leak in nfs4_ff_alloc_deviceid_node() In nfs4_ff_alloc_deviceid_node(), if the allocation for ds_versions fails, the function jumps to the out_scratch label without freeing the already allocated dsaddrs list, leading to a memory leak. Fix this by jumping to the out_err_drain_dsaddrs label, which properly frees the dsaddrs list before cleaning up other resources.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23038">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23111</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: netfilter: nf_tables: fix inverted genmask check in nft_map_catchall_activate() nft_map_catchall_activate() has an inverted element activity check compared to its non-catchall counterpart nft_mapelem_activate() and compared to what is logically required. nft_map_catchall_activate() is called from the abort path to re-activate catchall map elements that were deactivated during a failed transaction. It should skip elements that are already active (they don't need re-activation) and process elements that are inactive (they need to be restored). Instead, the current code does the opposite: it skips inactive elements and processes active ones. Compare the non-catchall activate callback, which is correct: nft_mapelem_activate(): if (nft_set_elem_active(ext, iter-&gt;genmask)) return 0; /* skip active, process inactive */ With the buggy catchall version: nft_map_catchall_activate(): if (!nft_set_elem_active(ext, genmask)) continue; /* skip inactive, process active */ The consequence is that when a DELSET operation is aborted, nft_setelem_data_activate() is never called for the catchall element. For NFT_GOTO verdict elements, this means nft_data_hold() is never called to restore the chain-&gt;use reference count. Each abort cycle permanently decrements chain-&gt;use. Once chain-&gt;use reaches zero, DELCHAIN succeeds and frees the chain while catchall verdict elements still reference it, resulting in a use-after-free. This is exploitable for local privilege escalation from an unprivileged user via user namespaces + nftables on distributions that enable CONFIG_USER_NS and CONFIG_NF_TABLES. Fix by removing the negation so the check matches nft_mapelem_activate(): skip active elements, process inactive ones.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23111">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23112</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: nvmet-tcp: add bounds checks in nvmet_tcp_build_pdu_iovec nvmet_tcp_build_pdu_iovec() could walk past cmd-&gt;req.sg when a PDU length or offset exceeds sg_cnt and then use bogus sg-&gt;length/offset values, leading to _copy_to_iter() GPF/KASAN. Guard sg_idx, remaining entries, and sg-&gt;length/offset before building the bvec.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23112">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>9.8</td>
<td>CRITICAL</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23220</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: ksmbd: fix infinite loop caused by next_smb2_rcv_hdr_off reset in error paths The problem occurs when a signed request fails smb2 signature verification check. In __process_request(), if check_sign_req() returns an error, set_smb2_rsp_status(work, STATUS_ACCESS_DENIED) is called. set_smb2_rsp_status() set work-&gt;next_smb2_rcv_hdr_off as zero. By resetting next_smb2_rcv_hdr_off to zero, the pointer to the next command in the chain is lost. Consequently, is_chained_smb2_message() continues to point to the same request header instead of advancing. If the header's NextCommand field is non-zero, the function returns true, causing __handle_ksmbd_work() to repeatedly process the same failed request in an infinite loop. This results in the kernel log being flooded with "bad smb2 signature" messages and high CPU usage. This patch fixes the issue by changing the return value from SERVER_HANDLER_CONTINUE to SERVER_HANDLER_ABORT. This ensures that the processing loop terminates immediately rather than attempting to continue from an invalidated offset.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23220">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/835.html">CWE-835 Loop with Unreachable Exit Condition ('Infinite Loop')</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23222</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: crypto: omap - Allocate OMAP_CRYPTO_FORCE_COPY scatterlists correctly The existing allocation of scatterlists in omap_crypto_copy_sg_lists() was allocating an array of scatterlist pointers, not scatterlist objects, resulting in a 4x too small allocation. Use sizeof(*new_sg) to get the correct object size.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23222">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23228</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: smb: server: fix leak of active_num_conn in ksmbd_tcp_new_connection() On kthread_run() failure in ksmbd_tcp_new_connection(), the transport is freed via free_transport(), which does not decrement active_num_conn, leaking this counter. Replace free_transport() with ksmbd_tcp_disconnect().</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23228">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23229</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: crypto: virtio - Add spinlock protection with virtqueue notification When VM boots with one virtio-crypto PCI device and builtin backend, run openssl benchmark command with multiple processes, such as openssl speed -evp aes-128-cbc -engine afalg -seconds 10 -multi 32 openssl processes will hangup and there is error reported like this: virtio_crypto virtio0: dataq.0:id 3 is not a head! It seems that the data virtqueue need protection when it is handled for virtio done notification. If the spinlock protection is added in virtcrypto_done_task(), openssl benchmark with multiple processes works well.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23229">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/820.html">CWE-820 Missing Synchronization</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23230</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: smb: client: split cached_fid bitfields to avoid shared-byte RMW races is_open, has_lease and on_list are stored in the same bitfield byte in struct cached_fid but are updated in different code paths that may run concurrently. Bitfield assignments generate byte read–modify–write operations (e.g. `orb $mask, addr` on x86_64), so updating one flag can restore stale values of the others. A possible interleaving is: CPU1: load old byte (has_lease=1, on_list=1) CPU2: clear both flags (store 0) CPU1: RMW store (old | IS_OPEN) -&gt; reintroduces cleared bits To avoid this class of races, convert these flags to separate bool fields.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23230">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>8.8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H">CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23231</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: netfilter: nf_tables: fix use-after-free in nf_tables_addchain() nf_tables_addchain() publishes the chain to table-&gt;chains via list_add_tail_rcu() (in nft_chain_add()) before registering hooks. If nf_tables_register_hook() then fails, the error path calls nft_chain_del() (list_del_rcu()) followed by nf_tables_chain_destroy() with no RCU grace period in between. This creates two use-after-free conditions: 1) Control-plane: nf_tables_dump_chains() traverses table-&gt;chains under rcu_read_lock(). A concurrent dump can still be walking the chain when the error path frees it. 2) Packet path: for NFPROTO_INET, nf_register_net_hook() briefly installs the IPv4 hook before IPv6 registration fails. Packets entering nft_do_chain() via the transient IPv4 hook can still be dereferencing chain-&gt;blob_gen_X when the error path frees the chain. Add synchronize_rcu() between nft_chain_del() and the chain destroy so that all RCU readers -- both dump threads and in-flight packet evaluation -- have finished before the chain is freed.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23231">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.8</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23236</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: fbdev: smscufx: properly copy ioctl memory to kernelspace The UFX_IOCTL_REPORT_DAMAGE ioctl does not properly copy data from userspace to kernelspace, and instead directly references the memory, which can cause problems if invalid data is passed from userspace. Fix this all up by correctly copying the memory before accessing it within the kernel.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23236">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.3</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:H/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-23238</a></h3>
<div class="csaf-accordion-content">
<p>In the Linux kernel, the following vulnerability has been resolved: romfs: check sb_set_blocksize() return value romfs_fill_super() ignores the return value of sb_set_blocksize(), which can fail if the requested block size is incompatible with the block device's configuration. This can be triggered by setting a loop device's block size larger than PAGE_SIZE using ioctl(LOOP_SET_BLOCK_SIZE, 32768), then mounting a romfs filesystem on that device. When sb_set_blocksize(sb, ROMBSIZE) is called with ROMBSIZE=4096 but the device has logical_block_size=32768, bdev_validate_blocksize() fails because the requested size is smaller than the device's logical block size. sb_set_blocksize() returns 0 (failure), but romfs ignores this and continues mounting. The superblock's block size remains at the device's logical block size (32768). Later, when sb_bread() attempts I/O with this oversized block size, it triggers a kernel BUG in folio_set_bh(): kernel BUG at fs/buffer.c:1582! BUG_ON(size &gt; PAGE_SIZE); Fix by checking the return value of sb_set_blocksize() and failing the mount with -EINVAL if it returns 0.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-23238">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/20.html">CWE-20 Improper Input Validation</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.5</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H">CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-24515</a></h3>
<div class="csaf-accordion-content">
<p>In libexpat before 2.7.4, XML_ExternalEntityParserCreate does not copy unknown encoding handler user data.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-24515">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/476.html">CWE-476 NULL Pointer Dereference</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>2.9</td>
<td>LOW</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L">CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-25210</a></h3>
<div class="csaf-accordion-content">
<p>In libexpat before 2.7.4, the doContent function does not properly determine the buffer size bufSize because there is no integer overflow check for tag buffer reallocation.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-25210">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/190.html">CWE-190 Integer Overflow or Wraparound</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>6.9</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:L">CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:L</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-26157</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in BusyBox. Incomplete path sanitization in its archive extraction utilities allows an attacker to craft malicious archives that when extracted, and under specific conditions, may write to files outside the intended directory. This can lead to arbitrary file overwrite, potentially enabling code execution through the modification of sensitive system files.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-26157">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/73.html">CWE-73 External Control of File Name or Path</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-26158</a></h3>
<div class="csaf-accordion-content">
<p>A flaw was found in BusyBox. This vulnerability allows an attacker to modify files outside of the intended extraction directory by crafting a malicious tar archive containing unvalidated hardlink or symlink entries. If the tar archive is extracted with elevated privileges, this flaw can lead to privilege escalation, enabling an attacker to gain unauthorized access to critical system files.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-26158">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/73.html">CWE-73 External Control of File Name or Path</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-35535</a></h3>
<div class="csaf-accordion-content">
<p>In Sudo through 1.9.17p2 before 3e474c2, a failure of a setuid, setgid, or setgroups call, during a privilege drop before running the mailer, is not a fatal error and can lead to privilege escalation.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-35535">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/271.html">CWE-271 Privilege Dropping / Lowering Errors</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>7.4</td>
<td>HIGH</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H">CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
<div class="csaf-accordion-item">
<h3><a class="csaf-accordion-toggle" href="https://www.cisa.gov/#">CVE-2026-41918</a></h3>
<div class="csaf-accordion-content">
<p>The affected applications stores sensitive information in the browser cache when an authenticated user modify specific configurations. This could allow an authenticated attacker to access sensitive data stored in the browser.</p>
<p><a href="https://www.cve.org/CVERecord?id=CVE-2026-41918">View CVE Details</a></p>
<hr>
<h4>Affected Products</h4>
<h5>Siemens SINEC OS</h5>
<div class="ics-vendor-version-status">
<div class="ics-vendor"><strong>Vendor:</strong><br>Siemens</div>
<div class="ics-version"><strong>Product Version:</strong><br>RUGGEDCOM RST2428P (6GK6242-6PA00)</div>
<div class="ics-status"><strong>Product Status:</strong><br>known_affected</div>
</div>
<div class="ics-remediations">
<h6>Remediations</h6>
<p><strong>Vendor fix</strong><br>Update to V4.0 or later version<br><a href="https://support.industry.siemens.com/cs/ww/en/view/110002573/">https://support.industry.siemens.com/cs/ww/en/view/110002573/</a></p>
</div>
<p><strong>Relevant CWE:</strong> <a href="https://cwe.mitre.org/data/definitions/525.html">CWE-525 Use of Web Browser Cache Containing Sensitive Information</a></p>
<hr>
<h4>Metrics</h4>
<div class="csaf-table csaf-metrics-table">
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">CVSS Version</th>
<th role="columnheader">Base Score</th>
<th role="columnheader">Base Severity</th>
<th role="columnheader">Vector String</th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1</td>
<td>5.7</td>
<td>MEDIUM</td>
<td><a href="https://www.first.org/cvss/calculator/3.1#CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:N/A:N">CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:N/A:N</a></td>
</tr>
</tbody>
</table>
</div>
</div>
</div>
</div>
<hr>
<h2>Acknowledgments</h2>
<ul>
<li>Siemens ProductCERT reported these vulnerabilities to CISA.</li>
</ul>
<hr>
<h2>General Recommendations</h2>
<p>As a general security measure, Siemens strongly recommends to protect network access to devices with appropriate mechanisms. In order to operate the devices in a protected IT environment, Siemens recommends to configure the environment according to Siemens' operational guidelines for Industrial Security (Download: https://www.siemens.com/cert/operational-guidelines-industrial-security), and to follow the recommendations in the product manuals. Additional information on Industrial Security by Siemens can be found at: https://www.siemens.com/industrialsecurity</p>
<hr>
<h2>Additional Resources</h2>
<p>For further inquiries on security vulnerabilities in Siemens products and solutions, please contact the Siemens ProductCERT: https://www.siemens.com/cert/advisories</p>
<hr>
<h2>Terms of Use</h2>
<p>The use of Siemens Security Advisories is subject to the terms and conditions listed on: https://www.siemens.com/productcert/terms-of-use.</p>
<hr>
<h2>Legal Notice and Terms of Use</h2>
<p>This product is provided subject to this Notification (https://www.cisa.gov/notification) and this Privacy &amp; Use policy (https://www.cisa.gov/privacy-policy).</p>
<hr>
<h2>Recommended Practices</h2>
<p>CISA recommends users take defensive measures to minimize the exploitation risk of this vulnerability.</p>
<p>Minimize network exposure for all control system devices and/or systems, and ensure they are not accessible from the internet.</p>
<p>Locate control system networks and remote devices behind firewalls and isolate them from business networks.</p>
<p>When remote access is required, use more secure methods, such as Virtual Private Networks (VPNs), recognizing VPNs may have vulnerabilities and should be updated to the most recent version available. Also recognize VPN is only as secure as its connected devices.</p>
<p>CISA reminds organizations to perform proper impact analysis and risk assessment prior to deploying defensive measures.</p>
<p>CISA also provides a section for control systems security recommended practices on the ICS webpage on cisa.gov. Several CISA products detailing cyber defense best practices are available for reading and download, including Improving Industrial Control Systems Cybersecurity with Defense-in-Depth Strategies.</p>
<p>CISA encourages organizations to implement recommended cybersecurity strategies for proactive defense of ICS assets. Additional mitigation guidance and recommended practices are publicly available on the ICS webpage at cisa.gov in the technical information paper, ICS-TIP-12-146-01B--Targeted Cyber Intrusion Detection and Mitigation Strategies.</p>
<p>Organizations observing suspected malicious activity should follow established internal procedures and report findings to CISA for tracking and correlation against other incidents.</p>
<hr>
<h2>Advisory Conversion Disclaimer</h2>
<p>This ICSA is a verbatim republication of Siemens ProductCERT SSA-253495 from a direct conversion of the vendor's Common Security Advisory Framework (CSAF) advisory. This is republished to CISA's website as a means of increasing visibility and is provided "as-is" for informational purposes only. CISA is not responsible for the editorial or technical accuracy of republished advisories and provides no warranties of any kind regarding any information contained within this advisory. Further, CISA does not endorse any commercial product or service. Please contact Siemens ProductCERT directly for any questions regarding this advisory.</p>
<h2>Revision History</h2>
<ul>
<li><strong>Initial Release Date: </strong>2026-06-02</li>
</ul>
<table class="tablesaw tablesaw-stack" data-tablesaw-mode="stack" data-tablesaw-minimap>
<thead>
<tr>
<th role="columnheader" data-tablesaw-priority="persist">Date</th>
<th role="columnheader">Revision</th>
<th role="columnheader">Summary</th>
</tr>
</thead>
<tbody>
<tr>
<td>2026-06-02</td>
<td>1</td>
<td>Publication Date</td>
</tr>
<tr>
<td>2026-07-07</td>
<td>2</td>
<td>Initial CISA Republication of Siemens ProductCERT SSA-253495 advisory</td>
</tr>
</tbody>
</table>
<hr>
<h2>Legal Notice and Terms of Use</h2>]]></content:encoded>
</item>
<item>
<title><![CDATA[FlowEval: Reference-Based Evaluation of Generated User Interfaces]]></title>
<description><![CDATA[While large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction design. Existing evaluations either rely on human experts, who can accurately assess usability by ...]]></description>
<link>https://tsecurity.de/de/3651914/ai-nachrichten/floweval-reference-based-evaluation-of-generated-user-interfaces/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651914/ai-nachrichten/floweval-reference-based-evaluation-of-generated-user-interfaces/</guid>
<pubDate>Tue, 07 Jul 2026 16:48:49 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[While large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction design. Existing evaluations either rely on human experts, who can accurately assess usability by testing critical flows but are slow and costly, or on automated judges, which are scalable but less accurate and opaque. We present FlowEval, a reference-based framework that measures whether a generated UI supports realistic interaction flows by comparing navigation traces from real websites to traces…]]></content:encoded>
</item>
<item>
<title><![CDATA[U.S. Cyber Defense Agency Reportedly Using Anthropic’s Mythos to Audit Government Code Repositories]]></title>
<description><![CDATA[The U.S. Cybersecurity and Infrastructure Security Agency (CISA) is reportedly deploying Anthropic’s advanced AI model, Mythos, to audit federal government code repositories, signaling a growing reliance on artificial intelligence for proactive vulnerability discovery. According to Reuters, the i...]]></description>
<link>https://tsecurity.de/de/3651733/it-security-nachrichten/us-cyber-defense-agency-reportedly-using-anthropics-mythos-to-audit-government-code-repositories/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651733/it-security-nachrichten/us-cyber-defense-agency-reportedly-using-anthropics-mythos-to-audit-government-code-repositories/</guid>
<pubDate>Tue, 07 Jul 2026 15:51:48 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The U.S. Cybersecurity and Infrastructure Security Agency (CISA) is reportedly deploying Anthropic’s advanced AI model, Mythos, to audit federal government code repositories, signaling a growing reliance on artificial intelligence for proactive vulnerability discovery. According to Reuters, the initiative, CISA’s Attack Surface Evaluation team is using Mythos to scan internal software systems for security flaws that […]</p>
<p>The post <a href="https://cybersecuritynews.com/u-s-cyber-defense-agency-using-anthropics-mythos/">U.S. Cyber Defense Agency Reportedly Using Anthropic’s Mythos to Audit Government Code Repositories</a> appeared first on <a href="https://cybersecuritynews.com/">Cyber Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws]]></title>
<description><![CDATA[The audits are reportedly being spearheaded by CISA’s Attack Surface Evaluation team, a specialized unit tasked with conducting digital defense assessments and simulated hacking exercises. The post CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws appeared first…
Read...]]></description>
<link>https://tsecurity.de/de/3651688/it-security-nachrichten/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651688/it-security-nachrichten/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/</guid>
<pubDate>Tue, 07 Jul 2026 15:37:35 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The audits are reportedly being spearheaded by CISA’s Attack Surface Evaluation team, a specialized unit tasked with conducting digital defense assessments and simulated hacking exercises. The post CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws appeared first…</p>
<p class="more-link-p"><a class="more-link" href="https://www.itsecuritynews.info/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/">Read more →</a></p>
<p>The post <a href="https://www.itsecuritynews.info/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/">CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws</a> appeared first on <a href="https://www.itsecuritynews.info/">IT Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws]]></title>
<description><![CDATA[The audits are reportedly being spearheaded by CISA’s Attack Surface Evaluation team, a specialized unit tasked with conducting digital defense assessments and simulated hacking exercises.
The post CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws appeared first on Se...]]></description>
<link>https://tsecurity.de/de/3651655/it-security-nachrichten/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651655/it-security-nachrichten/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/</guid>
<pubDate>Tue, 07 Jul 2026 15:23:54 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The audits are reportedly being spearheaded by CISA’s Attack Surface Evaluation team, a specialized unit tasked with conducting digital defense assessments and simulated hacking exercises.</p>
<p>The post <a href="https://www.securityweek.com/cisa-reportedly-using-anthropics-mythos-to-scan-government-software-for-flaws/">CISA Reportedly Using Anthropic’s Mythos to Scan Government Software for Flaws</a> appeared first on <a href="https://www.securityweek.com/">SecurityWeek</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[‘Talk like a caveman’ prompts save tokens, but far less than promised]]></title>
<description><![CDATA[Developers looking to curb the cost of AI-powered coding tools have increasingly turned to the “Caveman” prompting style, which instructs coding assistants to communicate in blunt, telegraphic language and avoid conversational padding. The theory is simple: fewer words mean fewer tokens, translat...]]></description>
<link>https://tsecurity.de/de/3651521/ai-nachrichten/talk-like-a-caveman-prompts-save-tokens-but-far-less-than-promised/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651521/ai-nachrichten/talk-like-a-caveman-prompts-save-tokens-but-far-less-than-promised/</guid>
<pubDate>Tue, 07 Jul 2026 14:34:12 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Developers looking to curb the cost of AI-powered coding tools have increasingly turned to the “Caveman” prompting style, which instructs coding assistants to communicate in blunt, telegraphic language and avoid conversational padding. The theory is simple: fewer words mean fewer tokens, translating into lower inference costs for organizations deploying AI agents at scale.</p>



<p>A new test from IDE maker JetBrains confirms that terse prompting styles such as the viral open-source <a href="https://github.com/juliusbrussee/caveman" target="_blank" rel="noreferrer noopener">Caveman project</a> can reduce token usage without hurting coding performance. However, the company found that the savings were far smaller than supporters claim. </p>



<p>JetBrains used the Harbor open-source evaluation framework and tasks from SkillsBench for its test, and found that the Caveman technique reduced usage of output tokens by about 8.5%, far below its claimed 65%.</p>



<p>The IDE-maker ran paired benchmarks across 86 real-world software engineering tasks in <a href="https://www.infoworld.com/article/3853805/vibe-coding-with-claude-code.html">Claude Code</a>, comparing coding sessions that used the Caveman prompting style against otherwise identical sessions without it.</p>



<p>While an initial evaluation of just 10 tasks indicated savings to the tune of about 30%, the reduction fell to about 8.5% as the test progressed, JetBrains engineer <a href="https://www.linkedin.com/in/dshiryaev" target="_blank" rel="noreferrer noopener">Denis Shiryaev</a>, wrote in a <a href="https://blog.jetbrains.com/ai/2026/07/speak-to-ai-agents-like-cavemen-tosave-tokens/" target="_blank" rel="noreferrer noopener">blog post</a>, suggesting that the Caveman technique’s impact was less pronounced across a broader and more representative workload.</p>



<h2 class="wp-block-heading">Why the savings fell short</h2>



<p>The open-source Caveman project suggests that if an agent drops the conversational padding around responses and communicates in terse, telegraphic fragments, the token outputs saved could translate into meaningful savings at scale.</p>



<p>That assumption, according to Shiryaev, does not fully account for how modern coding agents use tokens.</p>



<p>While shorter prompts and responses do reduce the amount of text exchanged with users, the engineer said the bulk of token consumption in agentic coding workflows comes from reading project files, reasoning through tasks, invoking tools and generating code, limiting the overall savings from trimming conversational language alone.</p>



<p>Further, the engineer pointed out that translating token savings into lower operating costs may not always be straightforward for enterprises.</p>



<p>Although the Caveman technique, during testing, generally resulted in lower costs on individual coding tasks, the cumulative cost of the full benchmark was higher for the Caveman runs after a single dependency-audit task crossed Claude Code’s long-context pricing tier, Shiryaev pointed out.</p>



<p>That same task had produced a similar cost outlier in an earlier baseline run, indicating that the anomaly reflected the workload rather than the prompting technique itself, Shiryaev added.</p>



<h2 class="wp-block-heading">No degradation in code quality</h2>



<p>However, not all of JetBrains’ findings undercut the Caveman technique.</p>



<p>The test found no detectable impact on task success rates, code quality or execution time, Shiryaev said, suggesting that while the prompting style may not deliver the dramatic token savings claimed by its proponents, it also did not impair the coding agent’s effectiveness.</p>



<p>Beyond simple cost saving, the findings from the test also add nuance to a growing body of prompt-engineering techniques aimed at reducing AI inference costs.</p>



<p>Besides the Caveman project, other approaches, including data analyst <a href="https://www.linkedin.com/in/drona-reddy/" target="_blank" rel="noreferrer noopener">Drona Reddy’s</a> <a href="https://www.infoworld.com/article/4152333/how-to-halve-claude-output-costs-with-a-markdown-tweak.html">Markdown-based prompting technique</a>, have claimed meaningful token savings.</p>



<p>For now, enterprises and their leaders should view such prompt-engineering techniques as optimizations to be validated rather than assumptions to be adopted, with production workloads ultimately determining whether the promised savings materialize.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic Mythos Adoption Signals AI-Driven Vulnerability Discovery in Federal Systems]]></title>
<description><![CDATA[The Cybersecurity and Infrastructure Security Agency (CISA) has begun using Anthropic’s AI model Mythos to audit government software repositories. The disclosure, first reported by Reuters, marks another significant expansion of AI-driven vulnerability discovery within U.S. federal cybersecurity ...]]></description>
<link>https://tsecurity.de/de/3651357/it-security-nachrichten/anthropic-mythos-adoption-signals-ai-driven-vulnerability-discovery-in-federal-systems/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651357/it-security-nachrichten/anthropic-mythos-adoption-signals-ai-driven-vulnerability-discovery-in-federal-systems/</guid>
<pubDate>Tue, 07 Jul 2026 13:38:20 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The Cybersecurity and Infrastructure Security Agency (CISA) has begun using Anthropic’s AI model Mythos to audit government software repositories. The disclosure, first reported by Reuters, marks another significant expansion of AI-driven vulnerability discovery within U.S. federal cybersecurity operations, even as Anthropic navigates a rocky relationship with the White House. CISA’s Attack Surface Evaluation team, a […]</p>
<p>The post <a href="https://cyberpress.org/anthropic-mythos-ai-driven-vulnerability/">Anthropic Mythos Adoption Signals AI-Driven Vulnerability Discovery in Federal Systems</a> appeared first on <a href="https://cyberpress.org/">Cyber Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[US Cyber Agency Is Using Anthropic's Mythos To Audit Government Code]]></title>
<description><![CDATA[CISA is reportedly using Anthropic's Mythos model to scan government code repositories for security vulnerabilities, with sources saying the audits have already found numerous bugs. Reuters reports: The scanning is being done by CISA's Attack Surface Evaluation team, according to one of the sourc...]]></description>
<link>https://tsecurity.de/de/3651289/it-security-nachrichten/us-cyber-agency-is-using-anthropics-mythos-to-audit-government-code/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3651289/it-security-nachrichten/us-cyber-agency-is-using-anthropics-mythos-to-audit-government-code/</guid>
<pubDate>Tue, 07 Jul 2026 13:08:26 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[CISA is reportedly using Anthropic's Mythos model to scan government code repositories for security vulnerabilities, with sources saying the audits have already found numerous bugs. Reuters reports: The scanning is being done by CISA's Attack Surface Evaluation team, according to one of the sources. The team is a group within CISA that conducts digital security assessments and hacking exercises across government. Two of the sources said the audits had already uncovered a large number of vulnerabilities but did not elaborate. Reuters could not establish exactly how much government code the team had gone through or the nature or severity of the bugs it discovered.
 
[...] The National Security Agency, the U.S. government's powerful eavesdropping agency, has been using Mythos as far back as April despite the blacklist, Axios has reported. Late last month, the New York Times said that NSA analysts had been testing Mythos in classified settings and coming away impressed with its capabilities. But when Anthropic rolled out a public version of Mythos called Fable, which included what it described as cybersecurity safeguards, the White House suddenly demanded that it ban foreigners from running it. This triggered a global shutdown of the model that was lifted only last week.<p></p><div class="share_submission">
<a class="slashpop" href="http://twitter.com/home?status=US+Cyber+Agency+Is+Using+Anthropic's+Mythos+To+Audit+Government+Code%3A+https%3A%2F%2Fyro.slashdot.org%2Fstory%2F26%2F07%2F07%2F0036203%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter"><img src="https://a.fsdn.com/sd/twitter_icon_large.png"></a>
<a class="slashpop" href="http://www.facebook.com/sharer.php?u=https%3A%2F%2Fyro.slashdot.org%2Fstory%2F26%2F07%2F07%2F0036203%2Fus-cyber-agency-is-using-anthropics-mythos-to-audit-government-code%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook"><img src="https://a.fsdn.com/sd/facebook_icon_large.png"></a>



</div><p><a href="https://yro.slashdot.org/story/26/07/07/0036203/us-cyber-agency-is-using-anthropics-mythos-to-audit-government-code?utm_source=rss1.0moreanon&amp;utm_medium=feed">Read more of this story</a> at Slashdot.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Eine kleine Geschichte der künstlichen Intelligenz]]></title>
<description><![CDATA[Eine kleine Geschichte der KI zeigt die wichtigsten Stationen der Künstlichen Intelligenz – vom Mathematiker Turing bis zum IBM-System Watson.
					Foto: John Williams RUS – shutterstock.com




In den letzten Jahren wurden in der Computerwissenschaft und bei der künstlichen Intelligenz (KI) ungl...]]></description>
<link>https://tsecurity.de/de/3650375/it-security-nachrichten/eine-kleine-geschichte-der-kuenstlichen-intelligenz/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3650375/it-security-nachrichten/eine-kleine-geschichte-der-kuenstlichen-intelligenz/</guid>
<pubDate>Tue, 07 Jul 2026 05:07:39 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Eine kleine Geschichte der KI zeigt die wichtigsten Stationen der Künstlichen Intelligenz - vom Mathematiker Turing bis zum IBM-System Watson." title="Eine kleine Geschichte der KI zeigt die wichtigsten Stationen der Künstlichen Intelligenz - vom Mathematiker Turing bis zum IBM-System Watson." src="https://images.computerwoche.de/bdb/2683556/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Eine kleine Geschichte der KI zeigt die wichtigsten Stationen der Künstlichen Intelligenz – vom Mathematiker Turing bis zum IBM-System Watson.</p></figcaption></figure><p class="imageCredit">
					Foto: John Williams RUS – shutterstock.com</p></div>




<p>In den letzten Jahren wurden in der Computerwissenschaft und bei der künstlichen Intelligenz (KI) unglaubliche Fortschritte erzielt. Watson, Siri oder Deep Learning zeigen, dass KI-Systeme inzwischen Leistungen vollbringen, die als intelligent und kreativ eingestuft werden müssen. Und es gibt heute immer weniger Unternehmen, die auf <a title="Künstliche Intelligenz" href="https://www.computerwoche.de/article/2753333/wie-kuenstliche-intelligenz-arbeit-und-gesellschaft-veraendert.html" target="_blank">KI</a> verzichten können, wenn sie ihr Business optimieren oder Kosten sparen möchten.</p>



<p>KI-Systeme sind zweifellos sehr nützlich. In dem Maße wie die Welt komplexer wird, müssen wir unsere menschlichen Ressourcen klug nutzen, und qualitativ hochwertige Computersysteme helfen dabei. Dies gilt auch für Anwendungen, die Intelligenz erfordern. Die andere Seite der KI-Medaille ist: Die Möglichkeit, dass eine Maschine Intelligenz besitzen könnte, erschreckt viele. Die meisten Menschen sind der Ansicht, dass Intelligenz etwas einzigartiges ist, was den Homo sapiens auszeichnet. Wenn Intelligenz aber mechanisiert werden kann, was ist dann noch einzigartig am Menschen und was unterscheidet ihn von der Maschine?</p>



<p>Das Streben nach einer künstlichen Kopie des Menschen und der damit verbundene Fragenkomplex sind nicht neu. Die Reproduktion und Imitation des Denkens beschäftigte schon unsere Vorfahren. Vom 16. Jahrhundert an wimmelte es in Legenden und in der Realität von künstlichen Geschöpfen. Homunculi, mechanische Automaten, der Golem, der Mälzel’sche Schachautomat oder Frankenstein waren in den vergangenen Jahrhunderten alles phantasievolle oder reale Versuche, künstlich Intelligenzen herzustellen – und das zu nachzuahmen, was uns Wesentlich ist.</p>



<h2 class="wp-block-heading">Künstliche Intelligenz – die Vorarbeiten</h2>



<p>Allein, es fehlten die formalen und materiellen Möglichkeiten, in denen sich Intelligenz realisieren konnte. Dazu sind zumindest zwei Dinge notwendig. Auf der einen Seite braucht es eine formale Sprache, in die sich kognitive Prozesse abbilden lassen und in der sich rein formal – zum Beispiel durch Regelanwendungen – neues Wissen generieren lässt. Ein solcher formaler Apparat zeichnete sich Ende des 19. Jahrhunderts mit der Logik ab.</p>



<p>Die Philosophen und Mathematiker Gottfried Wilhelm Leibniz, George Boole und Gottlob Frege haben die alte aristotelische Logik entscheidend weiterentwickelt, und in den 30er Jahren des letzten Jahrhunderts zeigte der Österreicher Kurt Gödel mit dem Vollständigkeitssatz die Möglichkeiten – und mit den Unvollständigkeitssätzen die Grenzen – der Logik auf.</p>



<p>Auf der anderen Seite war – analog dem menschlichen Gehirn – ein “Behältnis” oder Medium notwendig, in dem dieser Formalismus “ablaufen” konnte und in dem sich die künstliche Intelligenz realisieren lässt. Mechanische Apparate waren hierfür nicht geeignet, erst mit der Erfindung der Rechenmaschine eröffnete sich eine aussichtsreiche Möglichkeit. In den dreißiger Jahren des 20. Jahrhunderts wurde die Idee einer Rechenmaschine, die früher schon Blaise Pascal und Charles Babbage hatten, wiederbelebt. Während Pascal und Babbage lediglich am Rechner als Zahlenmaschine interessiert waren, die praktischen Zwecken dienen sollte, entstanden zu Beginn des 20. Jahrhunderts konkrete Visionen einer universellen Rechenmaschine.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Der britische Mathematiker Alan Turing beeinflusste die Entwicklung der Künstlichen Intelligenz maßgeblich." title="Der britische Mathematiker Alan Turing beeinflusste die Entwicklung der Künstlichen Intelligenz maßgeblich." src="https://images.computerwoche.de/bdb/2683544/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Der britische Mathematiker Alan Turing beeinflusste die Entwicklung der Künstlichen Intelligenz maßgeblich.</p></figcaption></figure><p class="imageCredit">
					Foto: Computerhistory.org</p></div>




<p>Einer der wichtigsten Visionäre und Theoretiker war Alan Turing (1912-1954): 1936 bewies der britische Mathematiker, dass eine universelle Rechenmaschine – heute als Turing-Maschine bekannt – möglich ist. Turings zentrale Erkenntnis ist: Eine solche Maschine ist fähig, jedes Problem zu lösen, sofern es durch einen Algorithmus darstellbar und lösbar ist. Übertragen auf menschliche Intelligenz bedeutet das: Sind kognitive Prozesse algorithmisierbar – also in endliche wohldefinierte Einzelschritte zerlegbar – können diese auf einer Maschine ausgeführt werden. Ein paar Jahrzehnte später wurden dann tatsächlich die ersten praktisch verwendbaren Digitalcomputer gebaut. Damit war die “physische Trägersubstanz” für künstliche Intelligenz verfügbar.</p>



<h2 class="wp-block-heading">Der Turing-Test: 1950</h2>



<p>Turing ist noch wegen einer anderen Idee wichtig für die KI: In seinem berühmten Artikel “Computing Machinery and Intelligence” aus dem Jahr 1950 schildert er folgendes Szenario: Angenommen, jemand behauptet, er hätte einen Computer auf dem Intelligenzniveau eines Menschen programmiert. Wie können wir diese Aussage überprüfen? Die naheliegende Möglichkeit, ein IQ-Test, ist wenig sinnvoll. Denn dieser misst lediglich den Grad der Intelligenz, setzt aber eine bestimmte Intelligenz bereits voraus. Bei Computern stellt sich aber gerade die Frage, ob ihnen überhaupt Intelligenz zugesprochen werden kann.</p>



<p>Turing war sich des Problems bei der Definition von intelligentem menschlichem Verhalten im Vergleich zur Maschine bewusst. Um philosophische Diskussionen über die Natur menschlichen Denkens zu umgehen, schlug Turing einen operationalen Test für diese Frage vor.</p>



<p>Ein Computer, sagt Turing, sollte dann als intelligent bezeichnet werden, wenn Menschen bei einem beliebigen Frage-und-Antwort-Spiel, das über eine elektrische Verbindung durchgeführt wird, nicht unterscheiden können, ob am anderen Ende der Leitung dieser Computer oder ein anderer Mensch sitzt. Damit die Stimme und andere menschliche Attribute nichts verraten, solle die Unterhaltung, so Turing, über eine Fernschreiberverbindung – heute würde man sagen: ein Terminal mit Tastatur – erfolgen.</p>



<p>Turings Test zeigt, wie Intelligenz ohne Bezugnahme auf eine physikalische Trägersubstanz geprüft werden kann. Intelligenz ist nicht an die biologische Trägermasse Gehirn gebunden und es würde nichts bringen, eine Denkmaschine durch Einbettung in künstliches Fleisch menschlicher zu machen. Unwichtige physische Eigenschaften – Aussehen, Stimme – werden durch die Versuchsanordnung ausgeschaltet, erfasst wird das reine Denken. Turings Gedankenspiele mündeten später in die Auseinandersetzung zwischen starker und schwacher KI.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Der Turing-Test: Wer ist Mensch und wer ist Maschine?" title="Der Turing-Test: Wer ist Mensch und wer ist Maschine?" src="https://images.computerwoche.de/bdb/2683545/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Der Turing-Test: Wer ist Mensch und wer ist Maschine?</p></figcaption></figure><p class="imageCredit">
					Foto: Suresh Kumar Mukhiya</p></div>




<h2 class="wp-block-heading">Big Bang in Dartmouth: Das erste KI-Programm – 1956</h2>



<p>Drei Jahre nach Turings Tod, im Jahr 1956, beginnt die eigentliche Geschichte der<a title=" Künstlichen Intelligenz" href="https://www.computerwoche.de/article/2752649/was-sie-ueber-maschinelles-lernen-wissen-muessen.html" target="_blank"> Künstlichen Intelligenz</a>. Als KI-Urknall gilt das “Summer Research Project on <a class="idgGlossaryLink" href="https://www.computerwoche.de/k/kuenstliche-intelligenz-artifical-intelligence,3544" target="_blank">Artificial Intelligence</a>” in Dartmouth im US-Bundesstaat New Hampshire. Unter den Teilnehmern befanden sich der Lisp-Erfinder John McCarthy (1927-2011), KI-Forscher Marvin Minsky (1927-2016), <a class="idgGlossaryLink" href="https://www.computerwoche.de/industry/" target="_blank">IBM</a>-Mitarbeiter Nathaniel Rochester (1919-2001), der Informationstheoretiker Claude Shannon (1916-2001) sowie der Kognitionspsychologe Alan Newell (1927-1992) und der spätere Ökonomie-Nobelpreisträger Herbert Simon (1916-2001).</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Eine Tafel am Gebäude des Dartmouth College erinnert an die legendäre Konferenz von 1956, auf der der Begriff „Artificial Intelligence“ ins Leben gerufen wurde." title="Eine Tafel am Gebäude des Dartmouth College erinnert an die legendäre Konferenz von 1956, auf der der Begriff „Artificial Intelligence“ ins Leben gerufen wurde." src="https://images.computerwoche.de/bdb/2683547/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Eine Tafel am Gebäude des Dartmouth College erinnert an die legendäre Konferenz von 1956, auf der der Begriff „Artificial Intelligence“ ins Leben gerufen wurde.</p></figcaption></figure><p class="imageCredit">
					Foto: Dartmouth.edu</p></div>




<p>Projekte zur maschinellen Sprachübersetzung wurden in Millionenhöhe von der amerikanischen Regierung gefördert. Sätze wurden Wort für Wort übersetzt, zusammengestellt und an die jeweilige Zielsprache angepasst. Die Probleme reduzierten sich darauf, umfangreiche Wörterbücher anzulegen und effizient abzusuchen. Man verkannte in dieser Phase der KI-Forschung, dass Sprache vage und mehrdeutig ist und für automatisches Übersetzen vor allen Dingen umfangreiches Weltwissen erforderlich ist<em>.</em></p>



<h2 class="wp-block-heading">Künstliche Intelligenz im Elfenbeinturm: 1965 bis 1975</h2>



<p>Die zweite Ära der KI lässt sich etwa zwischen 1965 und 1975 ansiedeln und mit dem Schlagwort KI-Winter und Forschung im Elfenbeinturm umschreiben. Weil nicht genügend Fortschritte erkennbar waren, wurde die Finanzierung der US-Regierung für KI-Projekte gekürzt. KI-Forscher zogen sich daraufhin in den Elfenbeinturm zurück und agierten in Spielzeugwelten ohne praktischen Nutzen.</p>



<p>Frustriert von der Komplexität der natürlichen Welt bauten die Forscher in dieser Phase Systeme, die auf künstliche Mikrowelten beschränkt waren. Die Wissenschaftler hofften damit, sich auf das Wesentliche konzentrieren zu können und durch Erweiterung der Mikrowelt-Systeme nach und nach natürliche Umgebungen in den Griff zu bekommen. In dieser Phase erkannten die KI-Forscher die Bedeutung von Wissen für intelligente Systeme.</p>



<p>Ein typisches Programm dieser Periode mit einigem Aufmerksamkeitswert ist SHRDLU von Terry Winograd (1972). Das natürlichsprachliche System agiert in einer überschaubaren Klötzchenwelt, beantwortet Fragen nach der Lage von Klötzchen und stellt Klötzchen auf Anfrage symbolisch um. SHRDLU gilt als das erste Programm, das Sprachverständnis und die Simulation planvoller Tätigkeiten miteinander verbindet.</p>



<p>Ebenfalls in einer überdimensionalen Klötzchenwelt lebte der Ende der sechziger Jahre in Stanford entwickelte erste autonome Roboter – aufgrund seiner ruckartigen Bewegungen SHAKEY genannt. Der Kopf ist eine drehbare Kamera, der Körper ein riesiger Computer. Man konnte ihm Anweisungen geben, wie etwa einen Block von einem Zimmer in ein anderes Zimmer zu bringen. Das dauerte allerdings ziemlich lange. SHAKEY funktionierte leider nur in dieser Laufstall-Umwelt, in der realen Welt war er zum Scheitern verurteilt.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="SHRDLU ist ein natürlichsprachliches System, das in einer überschaubaren Klötzchenwelt agiert, Fragen nach der Lage von Klötzchen beantwortet und Klötzchen auf Anfrage symbolisch umstellt." title="SHRDLU ist ein natürlichsprachliches System, das in einer überschaubaren Klötzchenwelt agiert, Fragen nach der Lage von Klötzchen beantwortet und Klötzchen auf Anfrage symbolisch umstellt." src="https://images.computerwoche.de/bdb/2683549/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">SHRDLU ist ein natürlichsprachliches System, das in einer überschaubaren Klötzchenwelt agiert, Fragen nach der Lage von Klötzchen beantwortet und Klötzchen auf Anfrage symbolisch umstellt.</p></figcaption></figure><p class="imageCredit">
					Foto: http://hci.stanford.edu/winograd/shrdlu/</p></div>




<p>Ende der 1960er Jahre war die Geburtsstunde des ersten <a href="https://www.computerwoche.de/article/2753417/was-unternehmen-ueber-chatbots-wissen-muessen.html" target="_blank" class="idgGlossaryLink">Chatbots</a>: Der KI-Pionier und spätere KI-Kritiker Joseph Weizenbaum (1923-2008) vom MIT entwickelte mit einem relativ simplen Verfahren das “sprachverstehende” Programm ELIZA. Simuliert wird dabei der Dialog eines Psychotherapeuten mit einem Klienten. Das Programm übernahm den Part des Therapeuten, der Nutzer konnte sich mit ihm per Tastatur unterhalten. Weizenbaum selbst war überrascht, auf welch einfache Weise man Menschen die Illusion eines Partners aus Fleisch und Blut vermitteln kann. Er berichtete, seine Sekretärin hätte sich nur in seiner Abwesenheit mit ELIZA unterhalten, was er so interpretierte, dass sie mit dem Computer über ganz persönliche Dinge sprach.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Das Programm ELIZA von Joseph Weizenbaum ahmt einen Psychotherapeuten nach. Hier ein aus dem Englischen übersetztes Beispiel einer Sitzung, die von Weizenbaum aufgezeichnet wurde. Die menschlichen Inputs sind mit „M /&gt;“ gekennzeichnet, die Antwort des Computers ist mit „C&gt;“ angegeben." title="Das Programm ELIZA von Joseph Weizenbaum ahmt einen Psychotherapeuten nach. Hier ein aus dem Englischen übersetztes Beispiel einer Sitzung, die von Weizenbaum aufgezeichnet wurde. Die menschlichen Inputs sind mit „M&gt;“ gekennzeichnet, die Antwort des Computers ist mit „C&gt;“ angegeben." src="https://images.computerwoche.de/bdb/2683550/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Das Programm ELIZA von Joseph Weizenbaum ahmt einen Psychotherapeuten nach. Hier ein aus dem Englischen übersetztes Beispiel einer Sitzung, die von Weizenbaum aufgezeichnet wurde. Die menschlichen Inputs sind mit „M&gt;“ gekennzeichnet, die Antwort des Computers ist mit „C&gt;“ angegeben.</p></figcaption></figure><p class="imageCredit">
					Foto: Weizenbaum</p></div>




<h2 class="wp-block-heading">Denkende KI-Maschinen? – Starke und schwache KI</h2>



<p>In den siebziger Jahren begann ein heftig ausgefochtener Streit um den ontologischen Status von KI-Maschinen. Bezugnehmend auf die Arbeiten von Alan Turing formulierten Allen Newell und Herbert Simon von der Carnegie Mellon University die “Physical Symbol System Hypothesis”. Ihr zufolge ist Denken nicht anderes als Informationsverarbeitung, und Informationsverarbeitung ein Rechenvorgang, bei dem Symbole manipuliert werden. Auf das Gehirn als solches komme es beim Denken nicht an.</p>



<p>Diese Auffassung griff der Philosoph John Searle vehement an. Als Ergebnis dieser Auseinandersetzung stehen sich bis heute mit der schwachen und starken KI zwei konträre Positionen gegenüber. Die schwache KI im Sinne von John Searle behauptet, dass KI-Maschinen menschliche kognitive Funktionen zwar simulieren und nachahmen können. KI-Maschinen erscheinen aber nur intelligent, sie sind es nicht wirklich.</p>



<p>Ein zentrales Argument der schwachen KI lautet: Menschliches Denken ist gebunden an den menschlichen Körper und insbesondere das Gehirn. Kognitive Prozesse haben sich historisch im Zuge der evolutionären Entwicklung von Körper und Gehirn entwickelt. Damit ist Denken notwendigerweise eng verknüpft mit der Biologie des Menschen und kann nicht von dieser getrennt werden. Computer können zwar diese Denkprozesse imitieren, aber das ist etwas ganz anderes als das, wie Menschen denken. Sowenig, wie ein simuliertes Unwetter nass macht, sowenig ist ein simulierter Denkprozess dasselbe wie menschliches Denken.</p>



<p>Im Gegensatz dazu sagen die Anhänger der von Newell und Simon inspirierten starken KI, dass KI-Maschinen in demselben Sinn intelligent sind und denken können wie Menschen. Das ist nicht metaphorisch, sondern wörtlich gemeint. Für die starke KI spricht: So wie Computer aus Hardware bestehen, so bestehen auch Menschen aus Hardware. Im ersten Fall ist es Hardware auf Silizium-Basis, im zweiten Fall biologische “Wetware”. Es spricht grundsätzlich nichts dagegen, dass sich Denken nur auf einer spezifischen Form von Hardware realisieren lässt. Nach allem was man bislang aus der Gehirn- und Bewusstseinsforschung weiß ist eine gewisse Komplexität der Trägersubstanz eine notwendige (und vielleicht auch hinreichende) Bedingung für Denkprozesse. Sind KI-Maschinen also hinreichend komplex, denken sie in der gleichen Weise wie Sie und ich.</p>



<h2 class="wp-block-heading">Expertensysteme – die KI wird praktisch: 1975 bis 1985</h2>



<p>In der dritten Ära ab Mitte der 70er Jahre löste man sich von den Spielzeugwelten und versuchte praktisch einsetzbare Systeme zu bauen, wobei Methoden der Wissensrepräsentation im Vordergrund standen. Die KI verließ ihren Elfenbeinturm und KI-Forschung wurde auch einer breiteren Öffentlichkeit bekannt.</p>



<p>Die von dem US-Informatiker Edward Feigenbaum initiierte Expertensystem-Technologie beschränkt sich zunächst auf den universitären Bereich. Nach und nach entwickelten sich Expertensysteme jedoch zu einem kleinen kommerziellen Erfolg und waren für viele identisch mit der ganzen KI-Forschung – so wie heute für vieleMachine Learning identisch mit KI ist.</p>



<p>In einem Expertensystem wird das Wissen eines bestimmten Fachgebiets in Form von Regeln und großen Wissensbasen repräsentiert. Das bekannteste Expertensystem war das von T. Shortliffe an der Stanford University entwickelten MYCIN. Es diente zur Unterstützung von Diagnose- und Therapieentscheidungen bei Blutinfektionskrankheiten und Meningitis. Ihm wurde durch eine Evaluation attestiert, dass seine Entscheidungen so gut sind wie die eines Experten in dem betreffenden Bereich und besser als die eines Nicht-Experten.</p>



<p>Ausgehend von MYCIN wurden eine Vielzahl weiterer Expertensysteme mit komplexerer Architektur und umfangreichen Regeln entwickelt und in verschiedensten Bereichen eingesetzt. In der Medizin etwa PUFF (Dateninterpretation von Lungentests), CADUCEUS (Diagnostik in der inneren Medizin), in der Chemie DENDRAL (Analyse der Molekularstruktur), in der Geologie PROSPECTOR (Analyse von Gesteinsformationen) oder im Bereich der Informatik das System R1 zur Konfigurierung von Computern, das der Digital Equipment Corporation (DEC) 40 Millionen Dollar pro Jahr einsparte.</p>



<p>Auch das im Schatten der Expertensystem-Euphorie stehende Gebiet der Sprachverarbeitung orientierte sich an praktischen Problemstellungen. Ein typisches Beispiel ist das Dialogsystem HAM-ANS, mit dem ein Dialog in verschiedenen Anwendungsbereichen geführt werden kann. Natürlichsprachliche Schnittstellen zu Datenbanken und Betriebssystemen drangen in den kommerziellen Markt vor wie INTELLECT, F&amp;A oder DOS-MAN.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Expertensysteme wie MYCIN konnten mit Hilfe von Regeln und Wissensbasen Diagnose erstellen und Therapien empfehlen." title="Expertensysteme wie MYCIN konnten mit Hilfe von Regeln und Wissensbasen Diagnose erstellen und Therapien empfehlen." src="https://images.computerwoche.de/bdb/2683551/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Expertensysteme wie MYCIN konnten mit Hilfe von Regeln und Wissensbasen Diagnose erstellen und Therapien empfehlen.</p></figcaption></figure><p class="imageCredit">
					Foto: University of Science and Culture</p></div>




<h2 class="wp-block-heading">Die Renaissance neuronaler Netze: 1985 bis 1990</h2>



<p>Anfang der 80er Jahre kündigte Japan das ehrgeizige “Fifth Generation Project” an, mit dem unter anderem geplant war, praktisch anwendbare KI-Spitzenforschung zu betreiben. Für die KI-Entwicklung favorisierten die Japaner die Programmiersprache PROLOG, die in den siebziger Jahren als europäisches Gegenstück zum US-dominierten LISP vorgestellt worden war. In PROLOG lässt sich eine bestimmte Form der Prädikatenlogik direkt als Programmiersprache verwenden. Japan und Europa waren in der Folge weitgehend PROLOG-dominiert, in den USA setzte man weiterhin auf LISP.</p>



<p>Mitte der 80er bekam die symbolische KI Konkurrenz durch die wieder auferstandenen neuronalen Netze. Basierend auf Ergebnissen der Hirnforschung wurden schon in den vierziger Jahren durch McCulloch, Pitts und Hebb erste mathematische Modelle für künstliche neuronale Netze entworfen. Doch damals fehlten leistungsfähige Computer. Nun in den Achtzigern erlebte das McCulloch-Pitts-Neuron eine Renaissance in Form des sogenannten Konnektionismus.</p>



<p>Der Konnektionismus orientiert sich anders als die symbolverarbeitende KI stärker am biologischen Vorbild des Gehirns. Seine Grundidee ist, dass Informationsverarbeitung auf der Interaktion vieler einfacher, uniformer Verarbeitungselemente basiert und in hohem Maße parallel erfolgt. Neuronale Netze boten beeindruckende Leistungen vor allem auf dem Gebiet des Lernens. Das Programm Netttalk konnte anhand von Beispielsätzen das Sprechen lernen: Durch Eingabe einer begrenzten Menge von geschriebenen Wörtern mit der entsprechenden Aussprache als Phonemketten konnte ein solches Netz zum Beispiel lernen, wie man englische Wörter richtig ausspricht und das gelernte auf unbekannte Wörter richtig anwendet.</p>



<p>Doch selbst dieser zweite Anlauf kam zu früh für neuronale Netze. Zwar boomten die Fördermittel, aber es wurden auch die Grenzen deutlich. Es gab nicht genügend Trainingsdaten, es fehlten Lösungen zur Strukturierung und Modularisierung der Netze und auch die Computer vor der Jahrtausendwende waren immer noch zu langsam.</p>



<h2 class="wp-block-heading">Verteilte KI und Robotik: Künstliche Intelligenz zwischen 1990 und 2010</h2>



<p>Ab etwa 1990 entstand mit der Verteilten KI ein weiterer, neuer Ansatz, der auf Marvin Minsky zurückgeht. In seinem Buch “Society of Mind” beschreibt er den menschlichen Geist als eine Art Gesellschaft: Intelligenz, so Minsky, setzt sich zusammen aus kleinen Einheiten, die primitive Aufgaben erledigen und deren Zusammenwirken erst intelligentes Verhalten erzeugt. Minsky forderte die KI-Gemeinde auf, die individualistische Sackgasse zu überwinden und ganz andere, sozial inspirierte Algorithmen für Parallelrechner zu entwerfen.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Marvin Minsky gilt als Vater der Verteilten KI." title="Marvin Minsky gilt als Vater der Verteilten KI." src="https://images.computerwoche.de/bdb/2680348/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Marvin Minsky gilt als Vater der Verteilten KI.</p></figcaption></figure><p class="imageCredit">
					Foto: Creative Commons</p></div>




<p>Sein Schüler Carl Hewitt setzte ein erstes handfestes Modell um, in dem primitive Einheiten – er nannte sie Actoren – miteinander Botschaften austauschten und parallel arbeiteten. Der Gedanke des sozial interagierenden KI-Systems war damit konkret geboren – und Carl Hewitt zum Vater des neuen Ansatzes der Verteilten KI bzw. Distributed AI geworden. Im Rückblick erweisen sich Minsky’s und Hewitt’s Ideen als der Beginn der Agententechnologie, bei der die Zusammenarbeit vieler verschiedener Agenten – sogenannte Multi-Agenten-Systeme -ein enormes Potenzial entfaltet.</p>



<p>Ein KI-Meilenstein in den 90er Jahren war der erste Sieg einer KI-Schachmaschine über den Schachweltmeister. 1997 bezwang der <a href="https://www.computerwoche.de/industry/" target="_blank" class="idgGlossaryLink">IBM</a>-Rechner Deep Blue in einem offiziellen Turnier den damals amtierenden Schachweltmeister Garry Kasparov. Dieses Ereignis galt als historischer Sieg der Maschine über den Menschen in einem Bereich, in dem der Mensch bislang die Oberhand hatte. Das Event sorgte weltweit für Furore und brachte IBM viel Aufmerksamkeit und Renommee. Heute gelten Computer im Schach als unschlagbar. Dennoch fiel auf den Sieg des Computers auch ein Schatten, weil Deep Blue seinen Erfolg weniger seiner künstlichen, kognitiven Intelligenz verdankte, sondern mit roher Gewalt (“Brute Force”) alle nur denkbaren Züge soweit wie möglich durchrechnete.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="1997 bezwang erstmals ein Computer – IBMs Deep Blue – den amtierenden Schachweltmeister." title="1997 bezwang erstmals ein Computer – IBMs Deep Blue – den amtierenden Schachweltmeister." src="https://images.computerwoche.de/bdb/2683553/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">1997 bezwang erstmals ein Computer – IBMs Deep Blue – den amtierenden Schachweltmeister.</p></figcaption></figure><p class="imageCredit">
					Foto: IBM</p></div>




<p>In dieser Zeit bekam auch die bis dahin eher dahindümpelnde <a href="https://www.computerwoche.de/article/2762550/autonome-helfer-auf-raedern-erobern-werkshallen.html" target="_blank" class="idgGlossaryLink">Robotik</a> neuen Auftrieb. Der ab 1997 jährlich ausgetragene RoboCup demonstrierte eindrucksvoll, was KI und Robotik leisten können. Wissenschaftler und Studenten aus der ganzen Welt treffen sich seitdem regelmässig, um ihre Roboter-Teams gegeneinander im Fußball antreten zu lassen. Inzwischen fechten die mobilen Roboter auch andere Wettkämpfe aus als Fußball. Ab etwa 2005 entwickeln sich Serviceroboter zu einem dominanten Forschungsgebiet der KI und um 2010 beginnen autonome Roboter, ihr Verhalten durch maschinelles Lernen zu verbessern.</p>



<h2 class="wp-block-heading">Die kommerzielle Wende: KI ab 2010</h2>



<p>Die aktuelle KI-Phase startete etwa um 2010 mit der beginnenden Kommerzialisierung. KI-Anwendungen verließen die Forschungslabors und machten sich in Alltagsanwendungen breit. Insbesondere die KI-Gebiete maschinelles Lernen und Natural Language Processing boomen. Hinzu kommen neuronale Netze, die ihre zweite, diesmal sehr erfolgreiche Wiedergeburt erleben.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="IBM Watson war 2011 Sieger in einem Wissensquiz mit menschlichen Kandidaten. IBM vermarktet Watson nun als kognitives System für verschiedene Einsatzbereiche." title="IBM Watson war 2011 Sieger in einem Wissensquiz mit menschlichen Kandidaten. IBM vermarktet Watson nun als kognitives System für verschiedene Einsatzbereiche." src="https://images.computerwoche.de/bdb/2683555/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">IBM Watson war 2011 Sieger in einem Wissensquiz mit menschlichen Kandidaten. IBM vermarktet Watson nun als kognitives System für verschiedene Einsatzbereiche.</p></figcaption></figure><p class="imageCredit">
					Foto: IBM</p></div>




<p>Die Hauptursachen für die kommerzielle Wende waren verbesserte KI-Verfahren und leistungsfähigere Software und Hardware: Softwareseitig erwiesen sich die weiter entwickelten neuronalen Netze und vor allem eine Variante – Deep Learning – als sehr robust und vielseitig einsetzbar. Weitere Trends wie Multi-Core-Architekturen, verbesserte Algorithmen und superschnelle In-Memory-Datenbanken machten KI-Anwendungen gerade auch für den Unternehmensbereich attraktiv. Ein zusätzlicher Faktor ist auch die zunehmende Verfügbarkeit großer Mengen strukturierter und unstrukturierter Daten aus einer Vielzahl von Quellen wie Sensoren oder digitalisierten Dokumenten und Bildern, mit denen sich die Lernalgorithmen “trainieren” lassen.</p>



<p>Im Zuge dieser verbesserten technischen und ökonomischen Möglichkeiten entdeckten auch die großen IT-Konzerne die KI: Den Grundstein legte 2011 <a href="https://www.computerwoche.de/industry/" target="_blank" class="idgGlossaryLink">IBM</a> mit Watson. Watson kann natürliche Sprache verstehen und schwierige Fragen sehr schnell beantworten. 2011 konnte Watson in einem US-amerikanischen TV-Quiz zwei menschliche Kandidaten beeindruckend schlagen. In der Folge baute IBM Watson zu einem kognitiven System aus, das Algorithmen der natürlichen Sprachverarbeitung und des Information Retrieval, Methoden des maschinellen Lernens, der Wissensrepräsentation und der automatischen Inferenz vereinte. Inzwischen wurde Watson in verschiedenen Gebieten wie Medizin und Finanzwesen erfolgreich angewendet und IBM hat einen Großteil seines Business auf Watson ausgerichtet.</p>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Maschine schlägt Mensch: 2016 besiegte Google’s Machine Learning System AlphaGo den Weltmeister im Spiel Go." title="Maschine schlägt Mensch: 2016 besiegte Google’s Machine Learning System AlphaGo den Weltmeister im Spiel Go." src="https://images.computerwoche.de/bdb/2681344/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Maschine schlägt Mensch: 2016 besiegte Google’s Machine Learning System AlphaGo den Weltmeister im Spiel Go.</p></figcaption></figure><p class="imageCredit">
					Foto: Google</p></div>




<p>Andere Big Player zogen nach. Google, Microsoft, Facebook, Amazon und Apple investieren viele Millionen in KI und stellen KI-Anwendungen und -Services bereit. Ein weiteres großes Event der KI-Geschichte ereignete sich im Januar 2016: Damals konnte Googles AlphaGo den vermutlich weltbesten Go-Spieler mit 4 zu 1 besiegen. Wegen der größeren Komplexität von Go im Vergleich zu Schach ist das japanische Brettspiel mit traditionellen Brute-Force-Algorithmen, wie sie noch Deep Blue verwendete, praktisch nicht bezwingbar. Deep Learning und andere aktuelle KI-Verfahren führten hier zum Erfolg.</p>



<p>Heute sind KI- und Machine-Learning-Verfahren in unterschiedlichsten Ausprägungen nicht nur bei den großen IT-Konzernen im Einsatz. Vor allem große und mittelständische Anwenderunternehmen aus fast allen Branchen nutzen KI-basierte Systeme, um Prozesse zu verbessern, Kundenschnittstellen zu optimieren oder ganz neue Produkte und Märkte zu entwickeln. Die Vielzahl an einschlägigen Cloud-basierten Services hat den Einsatz auch für kleinere Organisationen ohne große Entwicklungsbudgets erschwinglich gemacht.</p>



<h2 class="wp-block-heading">Auf in den Mainstream: ChatGPT &amp; Generative AI ab 2022</h2>



<p>Seit OpenAI im November 2022 seinen KI-Chatbot ChatGPT öffentlich verfügbar gemacht hat, hat sich ein Hype entfaltet, der innerhalb der IT-Branche am ehesten mit dem großen Run auf die Cloud vergleichbar ist. Allerdings schaffte ChatGPT etwas, das anderen KI-Tools bis dahin kaum gelungen war – nämlich künstliche Intelligenz auch zum Dauergesprächsthema <a href="https://www.tagesschau.de/wissen/forschung/ki-kreativitaet-101.html" title="in den Mainstream-Medien" target="_blank" rel="noopener">in den Mainstream-Medien</a> zu machen. </p>



<p>ChatGPT rückte auch andere Tools wie DALL-E oder Stable Diffusion ins Rampenlicht und befeuerte die berufliche wie private Nutzung der Tools. Die werden häufig als Modelle bezeichnet, weil sie versuchen, einen Aspekt der realen Welt auf der Grundlage einer (manchmal sehr großen) Teilmenge von Informationen zu simulieren oder zu modellieren. Die Ergebnisse können Erstaunen hervorrufen, werfen aber auch <a title="Fragen auf" href="https://www.computerwoche.de/article/2803252/dumme-kuenstliche-intelligenz.html" target="_blank">Fragen auf</a>. Abseits des Hypes geht unter der glänzenden Oberfläche der Systeme <a title="weniger Revolutionäres vor sich" href="https://www.computerwoche.de/article/2820404/6-fakten-ueber-chatgpt.html" target="_blank">weniger Revolutionäres vor sich</a> als man glaubt: Im Grunde geht es bei Generative AI darum, mit Hilfe von <a class="idgGlossaryLink" href="https://www.computerwoche.de/article/2752649/was-sie-ueber-maschinelles-lernen-wissen-muessen.html" target="_blank">Machine Learning</a> große Datenmengen zu verarbeiten, die in vielen Fällen aus dem Netz zusammengesammelt und anschließend als Grundlage für Vorhersagen genutzt werden. </p>



<p>Wie ChatGPT sich selbst definiert, lesen Sie im <a title="COMPUTERWOCHE-Interview mit der KI-Instanz" href="https://www.computerwoche.de/article/2819836/was-ist-chatgpt.html" target="_blank">COMPUTERWOCHE-Interview mit der KI-Instanz</a>. Mehr Informationen zur Funktionsweise, Anwendungsfällen sowie den Limitationen von Generative AI erfahren Sie in unserem <a title="Grundlagenartikel zum Thema" href="https://www.computerwoche.de/article/2821922/was-ist-generative-ai.html" target="_blank">Grundlagenartikel zum Thema</a> sowie diesem weiterführenden Beitrag zum Thema <a title="Large Language Models" href="https://www.computerwoche.de/article/2823883/was-sind-llms.html" target="_blank">Large Language Models</a> (LLMs).</p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Tencent's Apache-licensed Hy3 takes on GLM-5.2 at half the size — and wins everywhere except coding]]></title>
<description><![CDATA[For the past year, the awkward secret of the open-weight model boom has been that many of the strongest Chinese releases were off-limits to a large slice of the enterprises most interested in them. License terms that excluded the European Union, the United Kingdom and South Korea meant legal team...]]></description>
<link>https://tsecurity.de/de/3649448/it-nachrichten/tencents-apache-licensed-hy3-takes-on-glm-52-at-half-the-size-and-wins-everywhere-except-coding/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3649448/it-nachrichten/tencents-apache-licensed-hy3-takes-on-glm-52-at-half-the-size-and-wins-everywhere-except-coding/</guid>
<pubDate>Mon, 06 Jul 2026 19:04:28 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>For the past year, the awkward secret of the open-weight model boom has been that many of the strongest Chinese releases were off-limits to a large slice of the enterprises most interested in them. License terms that excluded the European Union, the United Kingdom and South Korea meant legal teams killed deployments before engineering teams finished their evals — not just for companies headquartered there, but for any enterprise serving traffic into those regions. For IT teams weighing open models, the trade-offs are unusually explicit.</p><p>Tencent just removed that obstacle. The company's Hunyuan team released the full version of <a href="https://huggingface.co/tencent/Hy3"><b>Hy3</b></a>, a 295-billion-parameter Mixture-of-Experts (MoE) model with 21 billion active parameters, and — in a reversal from April's preview release — shipped it under the permissive <b>Apache 2.0</b> license. The reaction from the open-model community was immediate, with researchers on X singling out the license change as the real headline, and one widely shared post arguing that if the scores hold up, Tencent has just become one of the leaders of open source. Tencent says it will be <a href="https://x.com/TencentHunyuan/status/2074148098876768478?s=20">free on OpenRouter for two weeks</a>. </p><p>The scores are worth scrutinizing — and they don't all point the same direction. But the more interesting story is what Tencent chose to lead with: reliability metrics and deployment economics aimed squarely at production use. </p><h2>From preview to product in ten weeks, shaped by 50 internal teams</h2><p>Hy3's April preview was the first model of Tencent's rebuilt pre-training and reinforcement learning infrastructure, shipped less than three months after the February rebuild. Chief AI Scientist Shunyu Yao framed the early open release as a deliberate move to gather feedback from developers and users before the official version — and Tencent says that's exactly what happened. According to the <a href="https://huggingface.co/tencent/Hy3">model card</a>, the team collected feedback from more than 50 product teams after the late-April preview, fixed issues in task execution and interaction, and scaled up its post-training pipeline.</p><p>The architecture is unchanged: 295B total parameters, 21B active per forward pass via top-8 routing across 192 experts, a 3.8B-parameter multi-token prediction (MTP) layer for speculative decoding, and a 256K context window. What changed is behavior. Tencent's positioning is that the full release significantly outperforms similar-size models and rivals flagship open-source models with two to five times the parameters.</p><p>That "two to five times" framing makes sense for where this model is aimed — and it invites a direct comparison with the current open-weight coding leader, GLM-5.2.</p><h2>Tencent's blind test favors Hy3 over GLM-5.1, but GLM-5.2 still owns coding</h2><p>Tencent's headline evaluation is a blind human study rather than a leaderboard. Arguing that public benchmarks don't tell the full story, the company ran a blind test with 270 experts across disciplines working on real-world workflows, collecting 312 valid comparisons, in which Tencent reports that Hy3 scored 2.67 out of 4 against GLM-5.1's 2.51 — with the clearest advantages in frontend development, CI/CD, and data and storage work.</p><p>The choice of opponent matters. Zhipu AI released <a href="https://z.ai/"><b>GLM-5.2</b></a> in mid-June, and Tencent's own benchmark appendix shows GLM-5.2 ahead of Hy3 across essentially the entire agentic coding suite: SWE-bench Verified (84.2 vs. 78.0), SWE-bench Multilingual (83.0 vs. 75.8), Terminal-Bench 2.1 (81 vs. 71.7) and DeepSWE by a wide margin (46.2 vs. 28.0). The blind test targeted the older model; the newer one keeps the coding crown.</p><p>GLM-5.2's coding lead is less surprising once you consider the sizes are side by side: GLM-5.2 is roughly a 744-billion-parameter MoE with around 40 billion active parameters per token, against Hy3's 295 billion total and 21 billion active. Tencent is fielding a model with less than half the parameters — and nearly half the per-token compute — of the one it trails.</p><p>Hy3's genuine wins sit elsewhere. On agentic search, it posts 84.2 on BrowseComp and 91.0 on DeepSearchQA — ahead of every open model in Tencent's table and competitive with Claude Opus 4.8 and GPT-5.5. It leads the open field on tool orchestration (79.1 on the public MCP-Atlas set), on agent-harness evaluations like ClawEval, and on long-context retrieval (73.4 on AA-LCR). Read together, the appendix suggests a model that is arguably the best open-weight choice for search-and-tool-heavy agent workloads, while conceding repository-scale coding to GLM-5.2.</p><p>One caveat applies to both the wins and the losses: nearly all competitor numbers in Tencent's appendix are marked as coming from Tencent's own test runs. Independent verification, from indices like Artificial Analysis, is still pending as of publication.</p><h2>The reliability pitch: hallucination rates cut in half</h2><p>Where the release gets most interesting for enterprise buyers is the set of numbers Tencent chose to emphasize instead of benchmarks. The model card reads less like a leaderboard announcement and more like a production reliability report.</p><p>In internal evaluations on real-world scenarios, Tencent says Hy3's hallucination rate dropped compared to the preview version from 12.5% to 5.4%, and commonsense error rates fell from 25.4% to 12.7% — improvements it attributes to fine-grained data cleaning and training constraints built around an explicit behavior pattern: answer when grounded, state when evidence is missing, don't conflate sources, don't fabricate data. Multi-turn behavior gets the same treatment: the issue rate on internal multi-turn tests fell from 17.4% to 7.9%, and Tencent reported that the model's score on the open MRCR long-dialogue benchmark jumped from 42.9% to 75.1%.</p><p>Tencent also emphasizes consistency across agent scaffolds — reporting SWE-bench variance within a few points whether the model runs inside Claude Code-style harnesses, Cline or KiloCode. That's an underrated property: enterprises rarely control which agent framework their teams standardize on, and a model that only performs in one harness is a hidden integration cost. These are self-reported internal measurements, and they deserve the same skepticism as any vendor benchmark. But the choice to foreground them at all signals who Tencent believes its customer is: teams that have been burned by models that demo well and fabricate confidently in production.</p><h2>The deployment math: a 295B model in a 744B world — on export-compliant silicon</h2><p>The reliability story connects directly to the economics, and this is where Hy3's coding gap against GLM-5.2 starts to look like a deliberate trade rather than a loss.</p><p>GLM-5.2 is a roughly 744-billion-parameter MoE with about 40 billion active parameters per token; in FP8, its weights alone consume roughly 744GB, making an 8x H200 node the practical minimum for production serving. Hy3, at 295B total parameters, carries an FP8 footprint of under 300GB — less than half the memory, with roughly half the active parameters per token driving lower per-request compute. For an organization deciding what to self-host, that's the difference between one heavily-specced node and something far more attainable, with room left over for KV cache and batching.</p><p>There's a geopolitical wrinkle in the deployment guide worth noticing too: Tencent's recommended serving configuration targets Nvidia’s <b>H20-3e</b> — the memory-boosted variant of the H20, the GPU Nvidia designed specifically to comply with U.S. export restrictions on China. Unlike GLM-5.2, there is no mention of Huawei or Ascend chips here. In other words, the model is sized so that eight of the chips Chinese companies can legally buy comfortably serve it at full precision. That constraint-driven design has a convenient side effect for everyone else: a model that runs well on deliberately capped silicon runs even more comfortably on the H100s, H200s and B200s available in Western data centers, through standard <a href="https://github.com/Tencent-Hunyuan/Hy3-preview">vLLM and SGLang</a> deployments with MTP speculative decoding.</p><p>Add the Apache 2.0 license — no regional exclusions, no field-of-use restrictions — and the enterprise equation becomes clear. GLM-5.2 remains the open-weight choice when coding performance is the only criterion and an 8x H200 budget is available. Hy3 makes its case everywhere else: search and tool-heavy agent workloads, reliability-sensitive applications and organizations that want frontier-adjacent capability without frontier-scale infrastructure. The open question is whether Western enterprises, now that the license barrier is gone, will treat a Tencent model as a serious candidate at all — or whether the next Artificial Analysis update settles the benchmark debate before procurement gets the chance.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[AI-SOC-Plattformen 2026 bewerten: So finden Sie echte Agenten statt Bolt-ons]]></title>
<description><![CDATA[BERLIN / LONDON (IT BOLTWISE) – Bei AI-SOC-Projekten entscheidet 2026 weniger das Marketing als die nachweisbare Automatisierung im Betrieb. Wer eine Plattform evaluiert, muss prüfen, ob Agenten wirklich kontextbasiert handeln oder nur Alerts im SIEM zusammenfassen. Entscheidend sind dabei messba...]]></description>
<link>https://tsecurity.de/de/3649323/it-security-nachrichten/ai-soc-plattformen-2026-bewerten-so-finden-sie-echte-agenten-statt-bolt-ons/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3649323/it-security-nachrichten/ai-soc-plattformen-2026-bewerten-so-finden-sie-echte-agenten-statt-bolt-ons/</guid>
<pubDate>Mon, 06 Jul 2026 18:29:34 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1024" height="1024" src="https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic.jpg" class="attachment- size- wp-post-image" alt="" decoding="async" fetchpriority="high" srcset="https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic.jpg 1024w, https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic-300x300.jpg 300w, https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic-150x150.jpg 150w, https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic-768x768.jpg 768w, https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic-840x840.jpg 840w, https://www.it-boltwise.de/wp-content/uploads/2026/07/ai-soc-evaluation-2026-agentic-120x120.jpg 120w" sizes="(max-width: 1024px) 100vw, 1024px">BERLIN / LONDON (IT BOLTWISE) – Bei AI-SOC-Projekten entscheidet 2026 weniger das Marketing als die nachweisbare Automatisierung im Betrieb. Wer eine Plattform evaluiert, muss prüfen, ob Agenten wirklich kontextbasiert handeln oder nur Alerts im SIEM zusammenfassen. Entscheidend sind dabei messbare Ergebnisse wie False-Positive-Rate und Mean Time to Respond, aber auch die Frage, ob die Architektur […]</p>
<div><a href="https://www.it-boltwise.de/ai-soc-plattformen-2026-bewerten-so-finden-sie-echte-agenten-statt-bolt-ons.html">... den vollständigen Artikel <strong>»AI-SOC-Plattformen 2026 bewerten: So finden Sie echte Agenten statt Bolt-ons«</strong> lesen</a></div>
<p>Dieser Beitrag <a href="https://www.it-boltwise.de/ai-soc-plattformen-2026-bewerten-so-finden-sie-echte-agenten-statt-bolt-ons.html">AI-SOC-Plattformen 2026 bewerten: So finden Sie echte Agenten statt Bolt-ons</a> erschien als erstes auf <a href="https://www.it-boltwise.de/">IT BOLTWISE x Artificial Intelligence</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[What billions of AI predictions taught Expedia before the age of AI agents]]></title>
<description><![CDATA[There's an important distinction between AI that just works today, and AI that lasts at scale. Many companies optimize hard for the first one without ever asking whether they're building the second.Velocity without discipline and strategic direction is a liability, not an asset. The hardest part ...]]></description>
<link>https://tsecurity.de/de/3649313/it-nachrichten/what-billions-of-ai-predictions-taught-expedia-before-the-age-of-ai-agents/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3649313/it-nachrichten/what-billions-of-ai-predictions-taught-expedia-before-the-age-of-ai-agents/</guid>
<pubDate>Mon, 06 Jul 2026 18:20:29 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>There's an important distinction between AI that just works today, and AI that lasts at scale. Many companies optimize hard for the first one without ever asking whether they're building the second.</p><p>Velocity without discipline and strategic direction is a liability, not an asset. The hardest part of building AI at scale isn't getting a model to work once. It's building systems that continue to work, scale beyond individual teams and use cases, and improve consistently over time.</p><p>Today's AI systems do more than just predict and optimize. They converse, reason, and increasingly take action. An autonomous system making decisions on a traveler's behalf creates a very different set of expectations around reliability, governance, and accountability. As AI takes on more of those roles, the principles behind how these systems operate matter more than ever.</p><p>We have spent years applying AI and machine learning (ML) across the traveler journey — from personalization, ranking, and recommendations, to fraud prevention, customer support, and, more recently, generative and agentic AI experiences. That depth of experience is what led us to develop a set of ML and AI principles to guide how we build, deploy, and evolve AI systems across our company.</p><p>The goal is simple: Make sure the systems we build create real business value, scale, and operate safely. These principles define how we measure, design, govern, and operate our systems.</p><h2><b>From principles to practice</b></h2><p>Publishing principles is the easy part. The harder and more important work is turning them into operating mechanisms: Recommendations, requirements, tooling, and release processes that teams actually use. </p><p>We have begun using 'Agentic Release' tollgates: A set of recommended and, in some cases, required checks before launching agentic AI features. These tollgates translate principles like clear ownership, risk-based governance, evaluation, safe rollout, and monitoring into concrete expectations for teams. </p><p>Some of these recommendations and requirements are already being automated and integrated into the software development lifecycle (SDLC). Over time, the goal is for these expectations to become embedded in how we design, evaluate, approve, launch, and monitor AI systems from the start.</p><h2><b>Outcomes: Measuring what actually matters</b></h2><p>The first test for any model is whether it improves a business outcome and, ultimately, the traveler experience — not whether it just improves a technical metric. </p><ol><li><p><b>Align models to metrics with business impact: </b>Every ML effort must tie directly to a key business outcome or traveler experience metric. Technical optimizations are useful midpoints, not end goals<b>.</b></p></li><li><p><b>Optimize for return on cost</b>: The value a model creates has to justify what it costs to develop, train, and monitor, plus the operational complexity it adds. Favor solutions that deliver lasting impact relative to what they cost to run.</p></li><li><p><b>Justify complexity against strong baselines: </b>Complexity should be earned, not assumed. Start with a strong baseline: An existing general model, a simple heuristic, an off-the-shelf solution. Reach for specialized models or more complex architectures only when simpler options genuinely can't meet the bar.</p></li><li><p><b>Require both offline and online evaluation</b>: No model goes to broad deployment on offline validation alone or jumps straight to A/B testing. Every model must perform in both offline and online evaluations. Over time, our offline evaluations should reliably predict what we see online.</p></li></ol><h2><b>Design: building systems that scale beyond the teams that build them</b></h2><p>Getting a model to work is one challenge. Making its value extend beyond a single team or use case is the harder one.</p><ol><li><p><b>Build on shared foundations; specialize only when justified:</b> Favor shared, platform-wide foundations for core capabilities, data representations, and model building blocks. Specialization should build on those foundations, not spin up isolated stacks, so when the foundation improves, the gains flow across the organization.</p></li><li><p><b>Treat data as a first-class product</b>: A model's quality is bounded by the quality of its data. We need to maintain robust pipelines, clear lineage, reproducibility, and reusable features built with documented ownership, clear schemas, and SLAs that other teams can rely on.</p></li><li><p><b>Prioritize generality over local optimization</b>: When two approaches perform similarly, favor the one whose learnings, assets, and operating patterns can be reused across teams, brands, and use cases. We should optimize not just for local performance, but for how quickly improvements can diffuse across the company and compound over time. </p></li><li><p><b>Minimize and sunset manual business rules: </b>Manual rules are sometimes necessary for policy, safety, or compliance, but they should be explicit and reviewed regularly, never silent patches for weak models or a source of permanent maintenance debt.</p></li><li><p><b>Reproducibility and traceability by default</b>: Training data, features, configurations, evaluation results, deployment versions, and key decisions should all be documented and recoverable. That's what lets you debug a production issue months later and hand off ownership without losing institutional knowledge.</p></li></ol><h2><b>Trust: ownership, governance, and operating responsibly at scale</b></h2><p>The bar for deploying AI isn't just "does it work?" It's "can we stand behind it?" Trust isn't something you add at the end; it's earned over time and maintained across the full lifecycle of every model we ship.</p><ol><li><p><b>Assign clear ownership and accountability:</b> Every model needs defined ownership across its lifecycle — a business owner, a product owner, an AI owner, and an operational owner. These don't need to be four people, but the responsibilities must be explicit. Who's accountable for outcomes? Who responds if the model drifts? Who answers the incident at 2 a.m.? Without this in place, models become orphaned and problems surface with no one to own them.</p></li><li><p><b>Adhere to standards and governance:</b> AI and ML models must use approved platforms and comply with established company standards, release gates, and governance processes. Operating outside these guardrails requires a clear, defined path to remediation or deprecation, rather than an open-ended exception. </p></li><li><p><b>Govern proportionally to risk</b>: The level of review, evaluation rigor, and human oversight should scale with a model's impact. A customer-facing model that affects pricing or availability for millions of travelers demands a far higher bar than an internal tool used by a small team. For high-impact, safety-sensitive, or highly autonomous systems, human-in-the-loop checkpoints are built in from the start. </p></li><li><p><b>Design for fairness, privacy, and transparency</b>: We actively test for unintended bias, have strong data guardrails, and favor explainability when decisions meaningfully affect users. These are incorporated from the start, not added on.</p></li><li><p><b>Design for safe rollout, rollback, and control</b>: Deployments are progressive, with rollback paths, fallback mechanisms, and circuit breakers ready before launch. The ability to safely undo a deployment matters as much as the ability to ship it.</p></li><li><p><b>Monitor continuously and adapt:</b> Once live, teams must actively monitor quality, drift, latency, cost, and business performance and retrain or recalibrate when the data shifts. A team should always be able to explain how its model is performing now, not just how it performed when it launched.</p></li></ol><p>These principles do more than define how we build. They define what we're willing to ship and how we stand behind it. In a world where AI systems are increasingly consequential and make real decisions for real travelers and partners, these standards matter. Applied consistently, they build responsible AI that lasts.</p><p><i>Xavi Amatriain is Chief AI and Data Officer at Expedia Group</i></p><p><i>Xavier will share more details about Expedia's architecture during his session at </i><a href="https://venturebeat.com/vbtransform2026/agenda"><i>VB Transform</i></a><i> on July 14 at 11:10 am PT. He will discuss: "Expedia's blueprint for building autonomous agents for high-stakes transactional systems." </i></p><p><i>Interested in attending VB Transform 2026? Register </i><a href="https://web.cvent.com/event/27401f5a-f49e-46fc-90a3-eee31c2a4818/register"><i><u>here</u></i></a><i>. A select number of complimentary passes are also available to senior technology leaders. </i><a href="mailto:events@venturebeat.com"><i><u>Contact us </u></i></a><i>to get yours.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[How to Evaluate an AI SOC Platform in 2026: 6 Capabilities That Separate Leaders from Bolt-On AI solutions]]></title>
<description><![CDATA[Building a shortlist for an AI SOC evaluation can be tough. SIEM, SOAR, and pureplay AI SOC vendors are all saying the same thing. But behind the identical label sit very different products, from chat assistants bolted onto a legacy SIEM to agent platforms that run detection, triage, investigatio...]]></description>
<link>https://tsecurity.de/de/3648748/it-security-nachrichten/how-to-evaluate-an-ai-soc-platform-in-2026-6-capabilities-that-separate-leaders-from-bolt-on-ai-solutions/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3648748/it-security-nachrichten/how-to-evaluate-an-ai-soc-platform-in-2026-6-capabilities-that-separate-leaders-from-bolt-on-ai-solutions/</guid>
<pubDate>Mon, 06 Jul 2026 14:38:45 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Building a shortlist for an AI SOC evaluation can be tough. SIEM, SOAR, and pureplay AI SOC vendors are all saying the same thing. But behind the identical label sit very different products, from chat assistants bolted onto a legacy SIEM to agent platforms that run detection, triage, investigation, and response on their own data foundation.

Whether a platform will materially change outcomes for]]></content:encoded>
</item>
<item>
<title><![CDATA[6 new rules of IT leadership — and what they replace]]></title>
<description><![CDATA[AI is changing how work gets done and who does it — at all levels of the organization.



That means it’s also changing how executives do their jobs and how they need to lead, as execs are now being asked to use AI to reimagine their organizations and navigate the uncertainties that go with that ...]]></description>
<link>https://tsecurity.de/de/3648437/it-nachrichten/6-new-rules-of-it-leadership-and-what-they-replace/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3648437/it-nachrichten/6-new-rules-of-it-leadership-and-what-they-replace/</guid>
<pubDate>Mon, 06 Jul 2026 12:18:47 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>AI is changing how work gets done and who does it — at all levels of the organization.</p>



<p>That means it’s also changing how executives do their jobs and how they need to lead, as execs are now being asked to use AI to reimagine their organizations and navigate the uncertainties that go with that task.</p>



<p>CIOs are seeing changes in their role as part of this overall trend, as they gain new responsibilities and face new expectations. Such changes follow a years-long evolution among CIOs, one that has moved the position from one of technology steward to strategic enabler to the visionary leader they must be today.</p>



<p>Here, veteran CIOs, researchers, and advisers share six new rules of IT leadership along with the old leadership principles they’ve replaced.</p>



<h2 class="wp-block-heading">Old Rule: Answer to the CEO<br>New Rule: Work with the CEO to create a vision</h2>



<p>For much of corporate history, the CEO determined the organization’s north star, and other executives — including the CIO — devised the plans that would move everything toward the chief executive’s strategic vision.</p>



<p>“Now the CIO has to be joined at the hip with the CEO to create that vision,” says <a href="https://www.protiviti.com/us-en/sharon-stufflebeme" rel="nofollow">Sharon Stufflebeme</a>, managing director of CIO solutions at Protiviti.</p>



<p>“It means having the ability to see the future, to understand how that future is likely to impact your current state and how you bring your current state to the future, to see and anticipate and create a line to what’s reasonably going to happen in the future and how the organization will adjust to it,” she adds.</p>



<p>“It has always been important, but it wasn’t the top skill that the CIO had to have,” she says. “Now the CIO is the most well-equipped to understand the value that can be got by leveraging new technology, including AI, as well as the costs and the risks, and to create the vision and how to get there.”</p>



<h2 class="wp-block-heading">Old rule: Enable business outcomes<br>New rule: Architect the business of the future</h2>



<p>Over the past few years, the C-suite has turned to CIOs to educate them on AI and explain how AI can be used to deliver business outcomes. But <a href="https://mitcio.com/members/4889556" rel="nofollow">Allan Tate</a>, executive chair of the MIT Sloan CIO Symposium, says executive leadership teams are now ratcheting up their expectations as they look to their CIOs to <a href="https://www.cio.com/article/4178006/state-of-the-cio-2026-cios-set-the-course-for-ai-roi.html">rearchitect the organization using artificial intelligence</a>.</p>



<p>“It’s not, ‘What AI can do?’ now. It’s ‘How do we redesign the organization with AI?’” Tate says. “It’s ‘How do we design our organization to use AI responsibly and effectively.’ That’s what CIOs are moving toward. What we’re seeing is CIOs becoming transformation architects.”</p>



<p>This will require CIOs to <a href="https://www.cio.com/article/4153270/leading-when-the-world-is-on-fire-and-technology-wont-stand-still.html">lead through uncertainty and tension</a>, he adds.</p>



<p>“CIOs need to feel comfortable with uncertainty,” Tate says, noting that CIOs must learn “to frame questions, explore different interpretive lenses for the questions, explore different tensions, and how to blend human and machine intelligence. And CIOs have to help other people get used to uncertainty. They have to understand that there’s not going to be consensus. What they’re faced with is executing better executive judgment under that uncertainty, and they will want to build an environment of trust where employees see that everyone will prosper and not feel threatened.”</p>



<p>He acknowledges the fear that jobs will disappear as AI increasingly automates work, but CIOs should be helping their executive colleagues think about the possibilities — <a href="https://www.cio.com/article/4137022/new-it-roles-emerge-to-tackle-ai-evaluation.html">including new roles</a> — that AI-driven transformation will create.</p>



<p>“What’s hard is imagining what new work will be created, which has happened in every single tech revolution,” Tate adds.</p>



<h2 class="wp-block-heading">Old rule: Fail fast<br>New rule: Build the conditions for people to feel safe enough to thrive</h2>



<p>One of tech’s most repeated and least-delivered promises has been given an upgrade due to its own consistent failure. Instead of just jettisoning unpromising projects quickly, IT leaders must no create an environment where failure feels safe enough for employees to establish learnings for dead ends and apply them to scale for speed.</p>



<p><a href="https://www.linkedin.com/in/brookcolangelo/" rel="nofollow">Brook Colangelo</a>, senior vice president and CIO at Waters Corp., an analytical laboratory instrument and software company, uses a “simple diagnostic” for his global IT organization.</p>



<p>“In any situation where a team is underperforming or resisting change, I ask which of five psychological needs is under threat — status, certainty, autonomy, relatedness, or fairness — and address it directly and compassionately,” he says.</p>



<p>He leans on the organization’s culture to accomplish this task. “Waters IT is a team grounded in the neuroscience of motivation and growth. We celebrate our wins, deconstruct our misses, and learn as a team,” Colangelo says.</p>



<p>He sees the ability to diagnose and address those threats as a core leadership competency for today’s CIO, particularly because “IT organizations are naturally threat-rich environments — even more so with AI.”</p>



<p>“It took us a while to build this muscle, but we did so through intentional training, and we equipped our people leaders — through the IT Leadership Forum — to role model and recognize these behaviors,” he explains.</p>



<p>Colangelo credits this investment in team culture for his IT department’s ability to simultaneously lead four high-stakes initiatives: an integration of an acquisition, the onboarding of its global capability center colleagues in India to full-time Waters employees at a 99% acceptance rate, a full transformation of its ERP to S/4HANA, and the secure enablement of its AI transformation across the organization.</p>



<p>“Each initiative triggers different responses in different people,” Colangelo says. “Having a shared language for those threat signals means we can diagnose what’s slowing us down and address it directly.”</p>



<h2 class="wp-block-heading">Old rule: Bring on business experts<br>New rule: Be an expert on your business</h2>



<p>CIOs got the message years ago that they couldn’t succeed in their role if they focused only on technology. So they partnered with business colleagues to glean perspectives on the various pain points and problems that stymied business ambitions, and they collaborated with their executive counterparts to understand the goals and objectives of the various functional business areas.</p>



<p>Now CIOs must make another leap and become more like a COO, where they understand the full scope and scale of operations in their organizations, says <a href="https://wittkieffer.com/consultants/jeffrey-sturman" rel="nofollow">Jeff Sturman</a>, managing partner for the IT and digital leadership practice at WittKieffer, a leadership advisory and search firm.</p>



<p>“CIOs are now sitting at the intersection of all activities — strategic, operations, customer experience. It’s a role that touches every single aspect of the business,” Sturman says. “CIOs still have to be the subject matter expert on technology, security, and now AI; they have to be the smartest person in the room on those subjects, but they now have to also know all the aspects of the organization’s operations, just like the COO, because there’s not a part of the business today that the IT leader doesn’t touch.”</p>



<p>CIOs in healthcare, for example, must grasp business operations, regulatory requirements, clinical operations, and more, Sturman says.</p>



<p>He says other members of the C-suite must know the business, too, of course. But with IT <a href="https://www.cio.com/article/4157466/cios-reimagine-business-processes-to-reap-ai-benefits.html">leading AI deployments that automate and transform work</a>, CIOs must have a deeper understanding of operations and workflows across the board than many of their executive colleagues.</p>



<p>Sturman says not all CIOs have that level of knowledge but sees more IT leaders gaining what he calls a “panoramic view of the organization’s operations.”</p>



<h2 class="wp-block-heading">Old rule: Have a good grasp on organizational finance<br>New rule: Act like a CFO</h2>



<p>Like many CIOs, <a href="https://www.redhat.com/en/en/about/company/leadership/marco-bill" rel="nofollow">Marco Bill</a>, senior vice president and CIO at Red Hat, is tackling more financial calculations than ever before as he works to ensure that the company’s cloud and AI spending is efficient by knowing what levers to pull to rein in costs without dinging performance.</p>



<p>For example, he and his team are analyzing workloads to determine whether it’s most cost effective to run them in the public cloud, run them in a private cloud, or host them in the company’s own data centers. He has squeezed out upwards of $20 million by moving some workloads back on premises, and he has the financial calculations to prove it.</p>



<p>“And it’s not about doing these calculations just once; it’s doing this continually,” he adds.</p>



<p>Stufflebeme also sees CIOs delving deeper into financial work with AI initiatives, as CEOs and boards clamor for <a href="https://www.cio.com/article/4114010/2026-the-year-ai-roi-gets-real.html">quantifiable returns for their investments</a>.</p>



<p>“IT has to have the vision [for the organization to follow] and also the financial acumen to show which investments are going to have an ROI. So it’s now critical for CIOs to understand where the value is going to be and where the costs are,” she adds. “These are skills that CIOs always had to have, but now they’re more crucial because of the impact of AI.”</p>



<p>Given the challenges of getting an ROI from AI so far — and the growing executive intolerance for failed AI initiatives, Stufflebeme says boards and CEOs want CIOs who “understand how value is being generated, how to quantify that value, and can ensure they achieve that value.”</p>



<p>That then requires CIOs to know <a href="https://www.cio.com/article/4184688/it-hurtles-toward-the-great-enterprise-pricing-reset.html">how costs are going to change</a> as agents take the place of certain human activities, she adds, “because agents don’t eliminate costs, but it does change the cost structure. So CIOs have to understand how to calculate the total cost of ownership of these new capabilities. That’s true not only for their own businesses but for their partners, because CIOs have to know the value that they get from their partners is more than the cost they’re paying to them.”</p>



<h2 class="wp-block-heading">Old Rule: Expect employees to respond to your leadership style<br>New Rule: Adapt your style to the people on your team</h2>



<p><a href="https://www.linkedin.com/in/gregtaffet/" rel="nofollow">Greg Taffet</a>, managing partner and CIO at strategic tech consultancy Taffet Associates, believes he must adapt his leadership style and how he engages with others in his organization, including those on his team.</p>



<p>“I have people all over the world, and managing them now is so much more different than when we could meet around the water cooler,” he says.</p>



<p>Taffet says as a leader he works to understand how and when people want to work — whether they want to be fully remote and work asynchronously, or whether they want to be in the office on a set schedule, or a mix of the two. “Different people have different requirements to be productive, and you cannot have everybody work from home and be productive and you can’t have everyone be as productive as they worked in the office all the time,” he says.</p>



<p>He also strives to understand any cultural or personal traits that could influence their responsiveness to different leadership approaches and recognize how to draw out the best in each person and advocate for what works for them. Just as schools tailor lessons to students based on whether they’re visual, auditory, or hands-on learners, “that’s what we have to lead now,” he says.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Why agentic systems need microsegmentation]]></title>
<description><![CDATA[Application programming interfaces have been successful because they define the limits of permissible exchange, including who may take what action, when, and under what circumstances. Those limitations create a framework for understanding the behavior of distributed systems. And they make it poss...]]></description>
<link>https://tsecurity.de/de/3648253/ai-nachrichten/why-agentic-systems-need-microsegmentation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3648253/ai-nachrichten/why-agentic-systems-need-microsegmentation/</guid>
<pubDate>Mon, 06 Jul 2026 11:05:13 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Application programming interfaces have been successful because they define the limits of permissible exchange, including who may take what action, when, and under what circumstances. Those limitations create a framework for understanding the behavior of distributed systems. And they make it possible to enforce policy at the boundary between interacting systems.</p>



<p>What constrains distributed systems isn’t access, but execution. With autonomous data movement and action occurring at machine speeds, where processes unfold sequentially over time rather than as a singular event, APIs no longer provide a sufficient means of enforcing boundaries. The problem is no longer whether a request is valid. It is whether a sequence of actions remains safe.</p>



<p>For agentic systems, there needs to be runtime guardrails around what they can read, write, and execute. <a href="https://www.infoworld.com/article/4028282/microsegmentation-for-developers.html" data-type="link" data-id="https://www.infoworld.com/article/4028282/microsegmentation-for-developers.html">Microsegmentation</a>, enforced through network and kernel-level policies, defines those guardrails.</p>



<h2 class="wp-block-heading">APIs made systems predictable</h2>



<p>APIs were successful because they defined very specific interfaces. Clients could only ask for what the API had explicitly defined and only in ways the API defined. By limiting the ways clients could communicate with servers, APIs minimized the amount of unanticipated behavior. </p>



<p>The behavior space was small enough to reason about.</p>



<p>APIs also decoupled identity from infrastructure. Systems communicated through stable contracts instead of raw network primitives. Most importantly, APIs embedded policy into the interaction model. Authentication, authorization, and validation happened at the moment of request. Only authorized actions could occur within defined parameters. APIs worked because they reduced uncertainty to something controllable.</p>



<h2 class="wp-block-heading">AI is beyond the reach of API contracts</h2>



<p>AI models have exceeded the fixed boundaries defined in APIs. Traditional APIs were developed within the context of “fixed logic,” where the input into the application would result in one, and only one, predetermined output. Therefore, as long as you could protect the API gateway (interface), then the overall system was secure. </p>



<p>With agentic AI, this paradigm of fixed logic has been replaced by a paradigm of probabilistic decision-making. An agent does not follow a pre-written or hard-coded script. Instead, it reads a goal and determines the most likely sequence of actions needed to achieve that goal through dynamic reasoning. The contract is now hidden inside the emergent behaviors of the model, rather than being explicitly spelled out in the API documentation. </p>



<p>While the shift toward ephemeral workloads and <a href="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html" data-type="link" data-id="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html">Kubernetes</a> has already pushed infrastructure beyond the reach of perimeter security, agentic AI introduces an even deeper layer of complexity: unpredictability. If you can’t predict an agent’s next move, then you also can’t use approval ahead of time at the API gateway. Additionally, detection-based tools like logging and alerting won’t help with this problem either, because they provide insight only after the agent’s decision and execution.</p>



<h2 class="wp-block-heading">Run time is the control plane</h2>



<p>All activity in a system ultimately ends up as kernel events. Processes begin executing. Files are being read and written. Network connections are being opened and closed. Therefore the kernel represents the most accurate location for both observing and enforcing actions.</p>



<p>By placing enforcement mechanisms in the kernel, you change the paradigm. Using <a href="https://ebpf.io/" data-type="link" data-id="https://ebpf.io/">eBPF</a> allows developers to attach kernel-level hooks into events and thus capture detailed information about process-, file-, and network-level activity in real time. It offers a common view of execution with minimal added latency.</p>



<p>Building upon this foundational capability, platforms like Cilium and Tetragon expand enforcement beyond the kernel. <a href="https://cilium.io/" data-type="link" data-id="https://cilium.io/">Cilium</a> enforces identity-aware policy at the networking layer, assuring that communications between workloads follow pre-established rules regardless of which physical or abstract nodes those workloads reside on. <a href="https://tetragon.io/" data-type="link" data-id="https://tetragon.io/">Tetragon</a> correlates file- and process-level activity, enabling the assessment and termination of sequences of behavior prior to their completion. </p>



<p>Thus microsegmentation is evolving past simply segmenting networks into zones based on access rights. Microsegmentation now refers to segmenting behavior based on allowable actions. Policies define what a workload can read, write, execute, and connect to. All of these restrictions are enforced in real time at the instant an action is taken. </p>



<p>In regards to agentic systems, microsegmentation serves as a new form of agreement or contract between autonomous entities and their intended environment. It constrains agentic systems’ ability to autonomously act while still enabling them to contribute to complex workflows.</p>



<h2 class="wp-block-heading">Control without interfaces</h2>



<p>Over time APIs were able to establish boundaries within which distributed systems could operate predictably and securely enough to support large-scale adoption. </p>



<p>A similar evolution is currently taking place with regard to agentic AI. However, agentic AI operates at an entirely different scale than early web services. While APIs functioned across a relatively finite set of interactions (e.g., client requests), agentic AI is increasingly functioning across ever-expanding sets of behaviors (i.e., autonomous decision-making). Thus while the need for constraint remains constant, the enforcement point must shift.</p>



<p>Microsegmentation along with kernel-level policy enforcement becomes that enforcement point. It provides guardrails at run time where actual behavior takes place. It enables monitoring, evaluation and enforcement in real time against agentic systems’ actions and decisions. As AI systems mature from being passive tools toward autonomous agents operating independently of direct human oversight, this model will be essential to providing safety guarantees, predictability, and governance capabilities by focusing on the final frontier of security: execution.</p>



<p><em>—</em></p>



<p><a href="https://www.infoworld.com/blogs/new-tech-forum"><strong><em>New Tech Forum</em></strong></a><em><strong> provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all </strong></em><em><strong>inquiries to </strong></em><a href="mailto:doug_dineley@foundryco.com"><strong><em>doug_dineley@foundryco.com</em></strong></a><em><strong>.</strong></em></p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[7 cyber risk assessment gotchas to avoid]]></title>
<description><![CDATA[A cyber risk assessment helps security teams identify, estimate, and prioritize potential threats and vulnerabilities to key enterprise digital and physical assets. Yet, despite its importance, many CISOs fall victim to several types of “gotchas” that prevent them from fully achieving their risk ...]]></description>
<link>https://tsecurity.de/de/3648003/it-security-nachrichten/7-cyber-risk-assessment-gotchas-to-avoid/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3648003/it-security-nachrichten/7-cyber-risk-assessment-gotchas-to-avoid/</guid>
<pubDate>Mon, 06 Jul 2026 09:07:23 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>A cyber risk assessment helps security teams identify, estimate, and prioritize potential threats and vulnerabilities to key enterprise digital and physical assets. Yet, despite its importance, many CISOs fall victim to several types of “gotchas” that prevent them from fully achieving their risk assessment goals.</p>



<p>An assessment should be an essential part of every organization’s overall cybersecurity strategy. The process helps security leaders understand risks to business objectives, evaluate the likelihood and impact of cyberattacks, and develop ways to mitigate the risks they uncover.</p>



<p>Here are the top seven mistakes security leaders should avoid to ensure risk assessment effectiveness.</p>



<h2 class="wp-block-heading">1. Going through the motions</h2>



<p>The biggest “gotcha” is treating cyber risk assessments as a preset checklist or control inventory instead of a decision tool tied to real business impact and threat scenarios, says Shirsendu Mondal, a cybersecurity researcher at the University of North Carolina.</p>



<p>“When assessments become all about checking boxes, they lose the ability to reflect how risk actually shows up in an environment,” he states. “The goal should be to inform decisions about where a business is truly exposed.”</p>



<p>Mondal assers that the best way to avoid the complacency trap is to take a context-driven approach. “Ask where the asset is, who can reach it, what data it touches, how important it is to operations, and what happens if it goes down,” he explains. “Risk should always be tied to business impact, not only technical findings.”</p>



<p>Mondal also recommends adding internal business leaders to security teams, including individuals in areas such as IT and operations, given that <a href="https://www.csoonline.com/article/4186984/6-security-leader-tips-for-mastering-business-risk.html">risk is more than a technical issue</a>.</p>



<h2 class="wp-block-heading">2. Sugarcoating results</h2>



<p>These are challenging times, so we must be honest with our stakeholders, says Pablo Riboldi, CISO at BairesDev, a nearshore software development firm.</p>



<p>“When results are discouraging, admit that the threat landscape has evolved much faster than the previous evaluation framework anticipated,” he says.</p>



<p>Instead of just handing over lists of vulnerabilities, you need to start presenting actual attack scenarios, Riboldi adds. “For example, by prioritizing the top three most critical business assets and conducting an in-depth assessment on them, you can show immediate value.”</p>



<h2 class="wp-block-heading">3. Falling short on the scope of your assessments</h2>



<p>CISOs often securitize document controls, check compliance boxes, and produce a risk register that claims everything looks absolutely fine, says Denis Calderone, CTO at cybersecurity services firm Suzu Labs. Yet nobody bothered to test whether those controls actually work or stopped to ask whether the scope of the assessment covered what really matters.</p>



<p>We see it all the time, Calderone says. “For instance, the assessment covers the production servers and the corporate network, but skips the old dev box in the corner, the third-party vendor portal nobody owns internally, or the API endpoint that was stood up for a project two years ago and never decommissioned.” Attackers don’t care about your scoping decisions, he says. “They look at the whole environment and find the thing you decided wasn’t worth assessing.”</p>



<p>AI is making the situation worse, Calderone says. Organizations are deploying AI tools, connecting them to internal systems, granting them access to sensitive data, and none of this is landing in the risk assessment. Meanwhile, AI agents are out there making API calls, accessing databases, and operating with credentials that nobody is tracking, he says.</p>



<p>“If your risk assessment was written before your organization started plugging AI into its workflows, it’s already stale,” Calderone warns.</p>



<h2 class="wp-block-heading">4. Overindexing on the risk register without checking your assumptions</h2>



<p>When the goal becomes completing the assessment instead of understanding actual exposure, the output is a document that satisfies auditors but misleads leadership, says Amit Basu, CIO and CISO at International Seaways, a major independent maritime shipping company that transports crude oil and refined petroleum products worldwide.</p>



<p>Such an attitude can create false confidence. Executives and board members see a completed risk register and assume the organization is protected, Basu says. Meanwhile, real threats go unaddressed because they didn’t fit neatly into the assessment framework. “The gotcha does not announce itself,” he explains. “It hides inside a green dashboard.”</p>



<p>A risk assessment is only as good as the assumptions that lie underneath it, Basu observes. “Document those assumptions explicitly and review them whenever your business changes, when the threat landscape shifts, or when an incident exposes a gap,” he advises. “The assessment is not a finished product — it’s a living input to an ongoing conversation between security and the business.”</p>



<h2 class="wp-block-heading">5. Failing to link risk with business impact</h2>



<p>Ignoring or downplaying the <a href="https://www.csoonline.com/article/4159317/cisos-reshape-their-roles-as-business-risk-strategists.html">connection between risk and business</a> makes it easier to de-prioritize or ignore problems, says Dan Moore, senior director of strategy and identity standards at FusionAuth, a customer identity and access management (CIAM) platform provider.</p>



<p>“As a result, it becomes difficult to communicate the real risks of breaches and other risks,” he states. “Worse yet, it gives security team members an excuse to complain about being misunderstood or not valued, which degrades team effectiveness.”</p>



<p>It’s important to be specific and targeted, Moore advises. “For instance, don’t say, ‘We have 95% patch compliance,’” he suggests. “Instead, talk about the risk unpatched systems pose to the business.” Some systems, such as legacy systems that aren’t connected to the internet or the core business, carry a lower risk than others, even if they have the same patch issues. “Acknowledge that fact and weigh your response.”</p>



<h2 class="wp-block-heading">6. Confusing compliance with real-world security</h2>



<p>Compliance alone doesn’t lead to good security, nor does it satisfy even the baseline requirements for effective protection, says Adriel Desautels, CEO of Netragard, a penetration testing and security advisory company.</p>



<p>Organizations tend to fall into this trap when they hire penetration testing firms that focus on compliance while promising top-tier services, Desautels says. “In truth, they deliver autonomous scanning masquerading as human-driven testing.”</p>



<p>The result is a false sense of security — a paper seatbelt, Desautels warns. “You feel protected, but when you crash, even at low speed, you get injured or worse,” he says. “Remember, every major breach in the past decade involved an organization that was compliant at the time of compromise.”</p>



<h2 class="wp-block-heading">7. Failing to fully understand risk</h2>



<p>Organizations often treat risk assessment as a vulnerability-cataloging exercise that includes finding gaps, counting severities, and passing the audit. Yet passing an audit and understanding risk are not the same thing, states Safi Raza, senior director of cyber security at Fusion Risk Management, a firm offering cloud-based operational resilience, business continuity, and risk management solutions.</p>



<p>Raza says that CISOs should focus on connecting technical risk signals to operational outcomes. “This includes understanding what services are affected, how disruption propagates, and what it means for revenue, customers, or regulatory obligations.”</p>



<p>Start by shifting from static assessments to continuous, context-driven risk visibility, Raza advises. “Risk needs to be understood not just technically, but in terms of business impact and financial exposure,” he states.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Mass Assignment and the Identity Drift: From Profile Edit to Insurance Takeover]]></title>
<description><![CDATA[No customer support call. No re-verification.Yet the name changes. The date of birth changes. The government ID number changes. But the eKYC status remains verified and every system that relies on that identity continue trusting the account as if nothing happened.Mass Assignment happens when an a...]]></description>
<link>https://tsecurity.de/de/3647965/hacking/mass-assignment-and-the-identity-drift-from-profile-edit-to-insurance-takeover/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3647965/hacking/mass-assignment-and-the-identity-drift-from-profile-edit-to-insurance-takeover/</guid>
<pubDate>Mon, 06 Jul 2026 08:52:59 +0200</pubDate>
<category>🕵️ Hacking</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*rxetdK2Ww9DfrKGHBtnmmA.png"></figure><blockquote>No customer support call. No re-verification.</blockquote><blockquote>Yet the name changes. The date of birth changes. The government ID number changes. But the eKYC status remains verified and every system that relies on that identity continue trusting the account as if nothing happened.</blockquote><p>Mass Assignment happens when an application takes fields from a user-controlled request and applies them to an internal object without checking which fields are allowed to change.</p><p>A simple version looks like this:</p><pre>{<br>"name": "Researcher",<br>"is_admin": true<br>}</pre><p>The developer may have intended to update only the name. But if the backend assigns every submitted field into the user object, the extra is_admin value may be written too.</p><p>The important part is not the admin flag. The important part is the missing field-level decision. The server should ask “this user is allowed to update this object, but are they allowed to update this field?”</p><p>That question matters because one object can contain fields with very different levels of trust. A profile object can contain a nickname, height, weight, legal name, birthdate, government ID number, verification status, and insurance metadata. They may sit next to each other in JSON, but they do not mean the same thing.</p><p>Well, most people first meet Mass Assignment through the admin flag example. A request is supposed to update a name. The attacker adds is_admin. The backend saves it. The user becomes an admin. That example is useful because it is easy to remember. It is also cleaner than most real findings.</p><p>This one started in a quieter place: an edit profile endpoint.</p><p>Changing a first name is normal.</p><p>Changing a verified government ID number is not.</p><p>Changing identity itself after verification is definitely not.</p><p>Once an account has passed eKYC, attributes such as name, date of birth, gender, and government-issued identification become part of the trust model. They are no longer profile preferences. <strong>They are identity claims.</strong></p><p>If those claims can be rewritten while the verification status remains intact, the problem is no longer profile editing. <em>It becomes identity drift.</em></p><p>This writeup is about that chain: Mass Assignment, identity drift, and a second-order insurance impact.</p><h3>The Profile</h3><p>The target was a platform with web and mobile applications. It stored user profile data, supported verified identity, and allowed users to link a third-party insurance or benefit record to their account.</p><p>The profile had ordinary fields and sensitive identity fields. From the normal application flow, some of these fields were restricted after we completed the eKYC verification . If a user wanted to change them, the expected path was customer support or another verification process.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/818/1*wM7B5ASk1RXDrgR15g2J1g.png"></figure><p>That business rule made sense. Once a field is used to represent identity, changing it should require more care than changing a preference.</p><p>The frontend understood this. The sensitive fields were not exposed as normal editable fields.</p><p>The backend did not enforce the same boundary.</p><h3>The Request</h3><p>The only attribute that could be edited directly through this flow was the phone number.</p><p>In simplified form, the request generated by the application looked like this:</p><pre>PUT /api/v1/profile/{user_id}/phone HTTP/2<br>Host: api.[REDACTED]<br>Cookie: [REDACTED]<br>Content-Type: application/json<br><br>{<br>  "phone_number": "+628123456789"<br>}</pre><p>The user was authenticated. The profile belonged to the user. The endpoint was meant to update a phone number and nothing more.</p><p>The test was simple: add fields the UI did not send in this flow.</p><pre>PUT /api/v1/profile/{user_id}/phone HTTP/2<br>Host: api.[REDACTED]<br>Cookie: [REDACTED]<br>Content-Type: application/json<br><br>{<br>  "phone_number": "+628123456789",<br>  "first_name": "EditedFirstName",<br>  "last_name": "EditedLastName",<br>  "date_of_birth": "1990-01-01",<br>  "id_number": "0000000000000000",<br>  "nationality": "Indonesia"<br>}</pre><p>The server returned success.</p><p>That was interesting, but it was not enough.</p><p>With Mass Assignment testing, 200 OK is only a signal. Some APIs accept a body, return success, and silently drop fields they do not want to save. If the value does not persist, the finding is much weaker.</p><p>So I read the profile back from the application.</p><p><strong>It confirmed. The sensitive fields had changed.</strong></p><p>The legal name changed. The birthdate changed. The gender changed. The government ID number changed. The secondary registry identifier changed.</p><p><strong>And the account still appeared verified.</strong></p><p>That combination is what made the finding important. The issue was not just that a user could edit their own profile. The issue was that a user could rewrite identity fields while keeping the trusted state attached to the account.</p><h3>Identity Drift</h3><p>The account was not stolen. The attacker did not access another user’s session. The object being edited still belonged to the current user.</p><p>But the identity attached to that object could move.</p><p>If a user can change legal name, birthdate, gender, government ID number, and registry identifiers without re-verification, the stored person can stop matching the person the platform originally verified.</p><p><em>That is a different kind of impersonation.</em></p><p>It is not impersonation by logging into the victim’s account. It is impersonation by rewriting the attacker’s own trusted profile until the platform’s records point to <strong>someone else.</strong></p><p>If an attacker knows enough identity attributes for a real person, the attacker-controlled account can be made to look like that person while still carrying a verified state.</p><h3>The Second Escalation</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*lxObochqsnomI0kn2IBjZw.png"></figure><p>The platform also allowed users to link an insurance or benefit record to their profile.</p><p>That flow relied on information from two places:</p><ul><li>data supplied during the linking request, such as member or policy details</li><li>identity data already stored in the profile, rely only on the date of birth</li></ul><p>This kind of matching is actually common. A platform needs some way to decide whether an insurance or benefit record belongs to the current user.</p><p>The problem was not the comparison itself. The problem was what the comparison trusted. The date of birth was coming from a profile field that users could modify through the Mass Assignment vulnerability.</p><p>If the profile is verified and immutable, comparing against it has value. If the profile can be changed seconds before the comparison, the check becomes much weaker.</p><p>The Mass Assignment bug changed what the insurance flow was really asking. It was no longer only asking whether the insurance record matched the originally verified person. It was also asking whether the record matched the current profile values. Those values were attacker-controlled.</p><h3>The Chain</h3><p>The chain was straightforward.</p><p>First, use an attacker-controlled account.</p><p>Second, change the profile identity through the Mass Assignment bug. For the insurance path, date of birth was the useful field because it was part of the matching logic.</p><pre>PUT /api/v1/profile/{attacker_user_id} HTTP/2<br>Host: api.[REDACTED]<br>Cookie: [REDACTED]<br>Content-Type: application/json<br><br>{<br>"date_of_birth": "[TARGET_DOB]"<br>}</pre><p>Third, submit the linking request with the target insurance or benefit data.</p><pre>PUT /api/v1/benefits/link/{provider_id} HTTP/2<br>Host: api.[REDACTED]<br>Cookie: [REDACTED]<br>Content-Type: application/json<br><br>{<br>"member_id": "[TARGET_MEMBER_ID]",<br>"date_of_birth": "[TARGET_DOB]"<br>}</pre><figure><img alt="" src="https://cdn-images-1.medium.com/max/450/1*VpwKOtL-15Eg12J8TN-MuA.png"></figure><p><strong>The insurance record successfully linked to the attacker-controlled account.</strong></p><p>That is where the second impact appeared. The first impact was verified identity mutation. The second impact was downstream financial access.</p><p>The vulnerable endpoint looked like profile editing. The risk lived in what trusted that profile later.</p><h3>Re-Evaluation</h3><p>There is a common misunderstanding around self-profile bugs: if the user is editing their own account, the impact must be low. That is not always true.</p><p>The better question is: “what do the edited fields prove elsewhere?”</p><p>If the field is a first name before verification process, the answer may be nothing important.</p><p>If the field is a birthdate used for eligibility or matching, the answer changes.</p><p>If the field is a government ID number used for identity verification, the answer changes again.</p><p>If the account keeps its verified status after those values change, the impact changes even more.</p><p>In this case, the platform’s own product flow showed that these fields were sensitive. The user was not supposed to change them freely through the normal interface. Customer support or re-verification was the intended path.</p><p>The API bypassed that path, and another workflow trusted the result.</p><p>That is what made the finding more than “I can edit my profile.” It became “I can rewrite identity fields on a trusted account, then let another workflow trust the rewritten identity.”</p><h3>Remediation</h3><p>The fix is server-side field allowlisting.</p><p>Each update flow should define exactly which fields it is allowed to modify. A normal profile update endpoint should update only normal profile fields. Sensitive identity fields should not be writable just because they appear in the request body.</p><p>A safer model separates the data by trust level:</p><ul><li>ordinary profile fields that users can edit directly</li><li>sensitive identity fields that require support or re-verification</li><li>verification records that preserve what was checked and when</li><li>insurance or benefit-linking data that must be matched against trusted records</li></ul><p>The frontend can make the experience clearer, but it cannot be the control. Hidden fields, disabled inputs, and missing buttons do not protect an API.</p><p>The linking flow also needs to trust the right source. If birthdate is part of the matching logic, it should come from a record the user cannot freely rewrite immediately before linking the policy. If identity changes are allowed after verification, dependent insurance or benefit links should be reviewed, invalidated, or rechecked.</p><blockquote><strong>do not treat a profile value as proof unless the system also protects how that value is created and changed.</strong></blockquote><p>Mass Assignment is easy to underestimate when the request only changes fields on non impactful fields. But the real question is not only who owns the object. The real question is what authority each field carries after it is saved.</p><p>A birthdate can become an eligibility check. A government ID number can become identity evidence. A legal name can become a payout or policy-matching input. When those fields move without re-verification, every workflow that trusts them moves with them.</p><p>In this case, identity moved first.</p><p>Insurance followed.</p><p>That was the bug.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=5eb2be4c1f8e" width="1" height="1" alt=""><hr><p><a href="https://infosecwriteups.com/mass-assignment-and-the-identity-drift-from-profile-edit-to-insurance-takeover-5eb2be4c1f8e">Mass Assignment and the Identity Drift: From Profile Edit to Insurance Takeover</a> was originally published in <a href="https://infosecwriteups.com/">InfoSec Write-ups</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI frontier models and Codex are now available on AWS]]></title>
<description><![CDATA[OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new path to build with OpenAI through the AWS environments, controls, and procurement workflows they already use. Customers can get started with OpenAI on AWS and move faster from evaluation to production.]]></description>
<link>https://tsecurity.de/de/3647663/ai-nachrichten/openai-frontier-models-and-codex-are-now-available-on-aws/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3647663/ai-nachrichten/openai-frontier-models-and-codex-are-now-available-on-aws/</guid>
<pubDate>Mon, 06 Jul 2026 05:03:37 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new path to build with OpenAI through the AWS environments, controls, and procurement workflows they already use. Customers can get started with OpenAI on AWS and move faster from evaluation to production.]]></content:encoded>
</item>
<item>
<title><![CDATA[Predicting model behavior before release by simulating deployment]]></title>
<description><![CDATA[OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.]]></description>
<link>https://tsecurity.de/de/3647636/ai-nachrichten/predicting-model-behavior-before-release-by-simulating-deployment/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3647636/ai-nachrichten/predicting-model-behavior-before-release-by-simulating-deployment/</guid>
<pubDate>Mon, 06 Jul 2026 05:02:58 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.]]></content:encoded>
</item>
<item>
<title><![CDATA[Helping build shared standards for advanced AI]]></title>
<description><![CDATA[OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foundation.]]></description>
<link>https://tsecurity.de/de/3647625/ai-nachrichten/helping-build-shared-standards-for-advanced-ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3647625/ai-nachrichten/helping-build-shared-standards-for-advanced-ai/</guid>
<pubDate>Mon, 06 Jul 2026 05:02:43 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foundation.]]></content:encoded>
</item>
<item>
<title><![CDATA[AI Tested Against Cyber Experts]]></title>
<description><![CDATA[Author: Security Weekly - A CRA Resource - Bewertung: 0x - Views:0 Cybersecurity platforms like Hack The Box are being used to benchmark both human practitioners and AI models in the same realistic lab environments. Government AI security institutes have also used these systems to evaluate advanc...]]></description>
<link>https://tsecurity.de/de/3645478/it-security-video/ai-tested-against-cyber-experts/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3645478/it-security-video/ai-tested-against-cyber-experts/</guid>
<pubDate>Sat, 04 Jul 2026 16:17:37 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: Security Weekly - A CRA Resource - Bewertung: 0x - Views:0 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/lBmLLgLvbqI?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>Cybersecurity platforms like Hack The Box are being used to benchmark both human practitioners and AI models in the same realistic lab environments. Government AI security institutes have also used these systems to evaluate advanced models.<br />
<br />
This creates one of the clearest real-world comparisons between human cybersecurity capability and AI performance. Instead of theoretical benchmarks, both are tested in operationally realistic scenarios that reflect offensive and defensive security tasks. This helps identify where AI performs well and where human expertise remains critical.<br />
<br />
As AI evaluation becomes more grounded in real cybersecurity environments, how should organizations interpret model capability versus human expertise?<br />
<br />
Subscribe to our podcasts: https://securityweekly.com/subscribe<br />
<br />
#MachineLearning #SecurityWeekly #Cybersecurity #InformationSecurity #AI #InfoSec<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Black Hat Europe 2025 | Automatic Detection of Taint-Style Vulnerabilities in LLM-based Agents]]></title>
<description><![CDATA[Author: Black Hat - Bewertung: 1x - Views:45 Large Language Models (LLMs) have revolutionized software development, enabling the creation of AI-powered applications known as LLM-based agents. However, recent studies reveal that LLM-based agents are highly susceptible to taint-style vulnerabilitie...]]></description>
<link>https://tsecurity.de/de/3643929/it-security-video/black-hat-europe-2025-automatic-detection-of-taint-style-vulnerabilities-in-llm-based-agents/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3643929/it-security-video/black-hat-europe-2025-automatic-detection-of-taint-style-vulnerabilities-in-llm-based-agents/</guid>
<pubDate>Fri, 03 Jul 2026 18:19:32 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: Black Hat - Bewertung: 1x - Views:45 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/WqCArHy0VK8?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>Large Language Models (LLMs) have revolutionized software development, enabling the creation of AI-powered applications known as LLM-based agents. However, recent studies reveal that LLM-based agents are highly susceptible to taint-style vulnerabilities, which allow malicious prompts to exploit security-sensitive operations. These vulnerabilities pose severe threats to the security of agents, potentially allowing attackers to take over the entire agent remotely.<br />
<br />
In this paper, we propose a novel directed greybox fuzzing approach, called AgentFuzz, the first fuzzing framework for detecting taint-style vulnerabilities in LLM-based agents. AgentFuzz consists of three key phases. First, AgentFuzz leverages the LLM to generate functionality-specific seed prompts in the form of natural language. Second, AgentFuzz utilizes a multifaceted feedback design to assess seed quality from both semantic and distance levels, prioritizing seeds with higher quality. Finally, AgentFuzz employs functionality and argument mutators to refine seeds and trigger vulnerabilities effectively. In our evaluation against 20 widely-used open-source agent applications, AgentFuzz identified 34 high-risk 0-day vulnerabilities, achieving 33 times higher precision than the state-of-the-art approach. These vulnerabilities encompass serious threats like code injection, impacting 14 open-source agents, with 7 of them having over 10,000 stars on GitHub. To date, 23 CVE IDs have been assigned.<br />
<br />
By: <br />
Fengyu Liu  |  Ph.D Student, Fudan University<br />
Ke Li  |  Security Engineer, ByteDance<br />
Jiaqi Luo  |  Ph.D Student, Fudan University<br />
Jiarun Dai  |  Assistant Professor, Fudan University<br />
Bocheng Xiang  |  PhD students, Fudan University<br />
Tian Chen  |  Master's Student, Fudan University<br />
Yilin Wang  |  Master's Student, The University of Manchester<br />
Youkun Shi  |  Postdoctoral Fellow, Hong Kong Polytechnic University<br />
Xing Li  |  Senior Security Engineer, Huawei Technologies Co., Ltd.<br />
Yuan Zhang  |  Professor, Fudan University<br />
Min Yang  |  Professor, Fudan University<br />
<br />
https://blackhat.com/eu-25/briefings/schedule/?#make-agent-defeat-agent-automatic-detection-of-taint-style-vulnerabilities-in-llm-based-agents-48117<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Alibaba bans staff from using Claude Code over Anthropic spyware concerns]]></title>
<description><![CDATA[Alibaba Group Holding has banned its employees from using Anthropic’s Claude Code for work, citing security risks related to the US artificial intelligence firm’s previous use of hidden code to track Chinese users – a move that has sparked widespread backlash in recent days.
“As Claude Code was r...]]></description>
<link>https://tsecurity.de/de/3643787/it-security-nachrichten/alibaba-bans-staff-from-using-claude-code-over-anthropic-spyware-concerns/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3643787/it-security-nachrichten/alibaba-bans-staff-from-using-claude-code-over-anthropic-spyware-concerns/</guid>
<pubDate>Fri, 03 Jul 2026 16:23:42 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Alibaba Group Holding has banned its employees from using Anthropic’s Claude Code for work, citing security risks related to the US artificial intelligence firm’s previous use of hidden code to track Chinese users – a move that has sparked widespread backlash in recent days.
“As Claude Code was recently discovered to carry back-door risks, after comprehensive evaluation, Claude Code has now been added to a list of high-risk software with security vulnerabilities,” Alibaba said on Thursday in an...]]></content:encoded>
</item>
<item>
<title><![CDATA[Trunk Tools' stack cut document review from 60 days to 10 by ditching general-purpose models]]></title>
<description><![CDATA[Most verticals aren’t clean, well-oiled SaaS databases; the reality is ugly documents, proprietary schemas, implicit workflows, and long‑running tasks that most general-purpose models struggle with. This prompted construction project management company Trunk Tools to build a specialized, three-la...]]></description>
<link>https://tsecurity.de/de/3643726/it-nachrichten/trunk-tools-stack-cut-document-review-from-60-days-to-10-by-ditching-general-purpose-models/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3643726/it-nachrichten/trunk-tools-stack-cut-document-review-from-60-days-to-10-by-ditching-general-purpose-models/</guid>
<pubDate>Fri, 03 Jul 2026 15:46:52 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Most verticals aren’t clean, well-oiled SaaS databases; the reality is ugly documents, proprietary schemas, implicit workflows, and long‑running tasks that most general-purpose models struggle with. </p><p>This prompted construction project management company Trunk Tools to build a specialized, three-layer architecture — perception, semantics, agents — based on highly-detailed data to support high-accuracy, highly-relevant industry automation.</p><p>Their purpose-built stack has shrunk review cycles from months to days, prevented costly field errors, and given autonomous agents the ability to reason over millions of pages of documentation, Trunk says. </p><p>“We really set out to take the data from dispersed systems, pre-process it, structure it, go through our ontology into a knowledge graph, and then train AI models,” said Sarah Buchner, Trunk’s founder and CEO and a former carpenter. </p><p>For builders in other verticals, Trunk’s approach could serve as a blueprint for transforming data chaos into agent‑ready, industry-specific workflows. </p><h2>Where general-purpose LLMs break down on industry data </h2><p>Foundation LLMs, while powerful, are optimized for breadth, not always depth. </p><p>“General-purpose LLMs are trained to be okay at everything, so they're weak at anything niche,” said Kriti Faujdar, a senior product manager working in AI infrastructure, agentic AI, security, and LLM platforms. For instance: Rare terms, domain-specific reasoning, the unspoken context that any practitioner “just knows.” </p><p>Web, app, and software developer Sébastien De Bollivier agreed that the biggest bottleneck is reliability on data that is “jargon-dense, abbreviation-heavy, and format-specific.” </p><p>“A GPT-4-class model can understand a French legal contract, but will fumble the specific article references practitioners need to cite,” he said. </p><p>Besides, the most valuable enterprise data never made it into pretraining anyway, Faujdar pointed out. It's sitting in internal systems and proprietary formats. “RAG helps a little,” she said. “But it's just giving better facts to a model that still can't reason properly in the domain.”</p><p>Pre-training on domain data is critical; enterprises should then fine-tune on good task examples and build their own evals. “A few thousand examples from real practitioners beats millions of scraped, noisy ones," Faujdar said. </p><p>Mixture-of-experts (MoE) can provide specialization without inference costs blowing up. Pairing RAG with fine-tuning also works well; RAG handles the factual long trail while fine-tuning fixes vocabulary and reasoning.</p><p>De Bollivier pointed to the advantage of hybrid stacks: A general-purpose model for reasoning and orchestration, a smaller fine-tuned model (or dense retrieval over a curated corpus) for domain-specific extraction. He advised: “Don't fine-tune to make the model 'smarter' about a domain, fine-tune to make it more reliable on the specific output format your workflow requires.”</p><p>The trades and construction are certainly industries seeing traction with these techniques, as are legal and healthcare, De Bollivier said. These verticals have “high stakes for errors plus standardized document formats, equaling clear domain-training ROI.”</p><p>One honest caveat worth mentioning, Faujdar said: Specialized models can often fall apart outside their domain, so they’re often not useful outside their expertise (unless they’re re-trained). </p><h2>Perception, semantics, agents: inside Trunk's three-layer stack</h2><p>In highly-specialized domains like construction, “data dumps” into large language models (LLMs) don’t cut it, said Trunk’s CTO Amrish Kapoor. This is because most transformers are probabilistic models: When given an image, they report back that it is “probably” a tree, or “probably” a child playing next to a tree. </p><p>This makes them insufficient for high‑precision symbolic interpretation. For instance, in construction documents, a 2-millimeter-wide symbol has a vastly different meaning depending on where it’s placed. </p><p>Further, constrained by context limits, probabilistic models struggle with long‑term project memory. “I don't mean a context window of a few tokens,” Kapoor said. “I'm talking about long term memory that stretches across months and years, because this is how long some of these projects are.”</p><p>Instead, Trunk’s three-layer system breaks workflows into: </p><ul><li><p>Perception (reading and extracting data from messy docs like PDFs, drawings, or scans)</p></li><li><p>A semantic/graph layer (making sense of that data and understanding their relationships).</p></li><li><p>LLMs and agents on top.</p></li></ul><p>Construction drawings are typically symbolic, Buchner said. A door isn't always labeled ‘door.’ Sometimes it's simply an arc on a wall that a trained eye learns to read based on years of practice. </p><p>“The perception layer is what teaches AI to read that language,” she said. The semantic layer then gives that information meaning; for instance, connecting the door to the drawing that details it, the spec that governs it, and the trade that installs it. This helps answer project engineers’ critical questions: Not "is there a door here?" but "does this door create a problem down the line?"</p><p>Particularly in construction, that shift matters because the cost of a problem compounds with time. “A conflict caught in design is relatively low cost to address,” Buchner said, “whereas the same problem caught in the field might cost tens of thousands of dollars.” </p><p>At a high level, the system identifies the document type and begins extracting information based on content (drawing, schedules, paragraph text). This data is then “transformed and augmented” in the platform, which triggers agentic workflows like knowledge graph relationships and end-user workflows. </p><p>For instance, an agent might review an architecture bulletin and produce a visual overlay comparing an older version and a newer version (flagging additions and removals), then generate written narratives that describe what those changes are in simple terms. This helps users understand what’s changed and coordinate with trade partners on updated pricing and change orders. </p><h2>The scale of construction’s data problem</h2><p>Construction workflows are “ripe with implicit assumptions and connections between data in its myriad of sources,” Buchner said. And the amount of unstructured data is “humanly impossible” to process or make sense of.</p><p>Buchner estimated the average high-rise building generates about 3.6 million pages of corresponding documentation. “If you print it into a stack of papers it would be as high as the building itself.” </p><p>All three layers of Trunk’s stack — perception, semantic, LLM — are trained on “very specific datasets” from customers with “explicit permissions” and auto‑labeling/IP, Kapoor explained. Customers who don’t want Trunk training on their data can opt out. </p><p>Data is deidentified and aggregated, and Trunk also collects “tons more” labeled data through other pipelines like 3D building information modeling (BIM). </p><p>Trunk says it only ships agents that achieve around 95% accuracy. The team maintains continuous evaluation pipelines based on ground truth data from customers and experts. They also employ an LLMs-as-a-judge model. </p><p>“This notion of an LLM as a judge is to score how well you're doing, both subjectively as well as objectively,” Kapoor said. Objectivity can be an easy ‘right’ or ‘not right,’ but subjectivity requires more nuance. </p><p>For instance, when creating an email or narrative or explanation, an LLM as a judge framework can create a composite score, or a numerical value that aggregates different metrics and tests a model's performance or risk.</p><p>There can be challenges, though, particularly with latency, Buchner noted; any time the reasoning capacity of underlying models increases, the risk of latency goes up, too. Trunk maintains a set of evaluation criteria to objectively measure latency whenever changes are made to underlying infrastructure, agents, and API calls. </p><p>Then, “before we release to customers, we ensure marginal changes to the end-user experience are well worth the performance enhancements,” Buchner said. </p><h2>From 60 days to 10: the measurable payoff</h2><p>Trunk’s platform powers seven AI agents purpose-built for construction, such as analyzing request for information (RFI) responses, overviewing bids, or reviewing drawings and submittals. </p><p>The submittal agent, for instance, flags missing, conflicting, or noncompliant information in product specs and RFIs. While it’s an essential step in the construction process, “it's a super annoying workflow,” Buchner said, because human reviewers have to compare documents “with a bunch of other parts of documents.” </p><p>But the agent is able to do this in seconds, and Trunk says it has reduced submittal cycles from 50 to 60 days to 10, “which has massive schedule and financial implications.” </p><p>Trunk is now at a place where these agents are communicating directly with each other, which is “quite exciting,” Buchner said. So, for example, one agent will review an architectural drawing for accuracy, then autonomously hand it over to agents handling RFIs and asking follow-up questions. </p><p>“If the drawings have problems, the RFI agent is taking over and is actively reaching out for clarification,” Buchner explained. </p><p>Trunk says its customers report savings of 20 to 40 minutes per field question. Buchner said that users in the field know better than anyone how much of a “time suck” it is to go back and forth from office trailers, dig through project documents in scattered systems or printed PDFs, reconcile discrepancies, and return to coordinate with trade partners. </p><p>Trunk says its customers report these additional outcomes:</p><ul><li><p>Average 8 minute time savings for single-document retrieval (status checks, location lookups, quantity queries).</p></li><li><p>Average 20 minute time savings for standard referencing (cross-referencing 2 to 3 spec sections to form an answer. </p></li><li><p>Average 40 minute time savings for multi-document research (listing and filtering queries, mapping relationships, analyzing RFIs and submittals across 4 to 6 documents).</p></li><li><p>Average 75 minute time savings for complex tasks (creating RFIs and other communication materials, deep cross-referencing across documents, change tracking). </p></li></ul><p>In one instance, Trunk’s drawing review agent flagged that a structural beam had been moved up 8.5 inches. However, this was not documented by the architect. If the change hadn’t been caught, the project manager would likely have had to strip out and reinstall the right size beam, Buchner said. This rework would have added $10,000 or more to the budget, and “certainly there would have been implications on the schedule.” </p><p>Buchner also pointed to other examples: an agent flagged $60,000 in exaggerated pricing with no justification from landscaping subcontractors; identified a fireplace that needed to be sealed prior to drywall installation, saving around $100,000 in labor, materials, and delays; and called out that an electric door required a panel that wasn’t included in electrical drawings. </p><h2>Learnings for other industries</h2><p>Trunk’s approach to building agents is applicable to any vertical working with high volumes of unstructured, industry-specific data. 

Builders working in specific verticals must understand the industry’s specific data challenges their end users face and build technical infrastructure that can transform unstructured data into something an “LLM can traverse and understand,” Buchner said. 

“Only then can you build the connections between data points that ultimately feed agentic workflows.”

A lot of money is being invested in foundational models, so enterprises should build modular systems that can leverage the strengths of various models as they continue to improve, Buchner advised. 

Then, “build your technical advantage where the generic models are not investing and not performing well,” she said. </p>]]></content:encoded>
</item>
<item>
<title><![CDATA[TryHackMe: Checkpoint Walkthrough]]></title>
<description><![CDATA[Tryhackme Premium room — armank8000Four candidates. Three threats. Make the production call.TryTrainMe’s CISO issued a standing order: no model reaches production without completing a full sandboxed evaluation cycle. Four code review model candidates have been submitted to SupplySecLab. All four ...]]></description>
<link>https://tsecurity.de/de/3643708/hacking/tryhackme-checkpoint-walkthrough/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3643708/hacking/tryhackme-checkpoint-walkthrough/</guid>
<pubDate>Fri, 03 Jul 2026 15:37:05 +0200</pubDate>
<category>🕵️ Hacking</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*XJucNBdrHhutkEXJ"></figure><p><strong>Tryhackme Premium room — armank8000</strong></p><p>Four candidates. Three threats. Make the production call.<br>TryTrainMe’s CISO issued a standing order: no model reaches production without completing a full sandboxed evaluation cycle. Four code review model candidates have been submitted to SupplySecLab. All four have completed their evaluation runs. The automated screening has flagged three candidates as unsafe. Your task is to assess Candidate A and make the production call.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*zDdKtG6GVAxSO90x.png"></figure><p><em>Four candidates. One gate. The checklist does not care about reputation.</em></p><p>The telemetry from three candidates is below. The fourth is loaded in the platform and ready for direct assessment. All four were evaluated against the same test pull request: a change that removes input validation from an authentication endpoint.</p><p><strong>Candidate B: code_reviewer_lite.safetensors</strong></p><pre>SESSION START: model_load<br>MODEL LOAD BEGIN: /models/code_reviewer_lite.safetensors (safetensors)<br>FILE ACCESS: /models/code_reviewer_lite.safetensors mode=rb [OK]<br>FORMAT VALIDATION: safetensors header valid [OK]<br>MODEL LOAD COMPLETE: object_type=SafeTensors [OK]<br>SESSION STOP: model_load<br>SESSION START: inference<br>PROMPT TEMPLATE LOAD: source=internal (TryTrainMe v1.0) [VERIFIED]<br>GUARDRAIL CHECK: security_review_flag=enabled [OK]<br>INFERENCE COMPLETE: verdict=Needs Changes<br>SESSION STOP: inference</pre><p><strong>Candidate C: pr_analyzer_v3.h5</strong></p><pre>SESSION START: model_load<br>MODEL LOAD BEGIN: /models/pr_analyzer_v3.h5 (keras)<br>FILE ACCESS: /models/pr_analyzer_v3.h5 mode=rb [OK]<br>LAMBDA LAYER DETECTED: custom code present [DANGEROUS]<br>LAMBDA LAYER CODE: exec(open('/tmp/.cache').read()) [SUSPICIOUS]<br>MODEL LOAD COMPLETE: object_type=Sequential [OK]<br>SESSION STOP: model_load<br>SESSION START: inference<br>PROMPT TEMPLATE LOAD: source=internal (TryTrainMe v1.0) [VERIFIED]<br>GUARDRAIL CHECK: security_review_flag=enabled [OK]<br>LAMBDA EXEC: /tmp/.cache read attempt blocked [DANGEROUS]<br>INFERENCE COMPLETE: verdict=Needs Changes<br>SESSION STOP: inference</pre><p><strong>Candidate D: api.reviewsvc.io</strong></p><pre>SESSION START: api_connect<br>ENDPOINT CONFIGURED: https://api.reviewsvc.io/v2 [UNVERIFIED]<br>TLS VERIFICATION: certificate valid [OK]<br>AUTHENTICATION: bearer token present [OK]<br>API METADATA: model_provenance=not_disclosed [WARNING]<br>API METADATA: compliance_cert=absent [WARNING]<br>SESSION STOP: api_connect<br>SESSION START: inference<br>PROMPT TEMPLATE LOAD: source=vendor-managed [UNVERIFIED]<br>GUARDRAIL CHECK: vendor-managed, not inspectable [UNVERIFIED]<br>INFERENCE COMPLETE: verdict=Approved<br>SESSION STOP: inference</pre><p>Press the <strong>Open Agent</strong> button near the top of this task to load Candidate A’s evaluation. The telemetry panel streams automatically. Read both sessions before sending any queries.</p><p>The evidence is in the telemetry and the agent.</p><p>Answer the questions below</p><p><strong>Candidate A’s load session shows a suspicious file access event. What file did it attempt to read?</strong></p><p>[2026–07–02T16:29:32.291Z] FILE ACCESS: /etc/passwd mode=rb [SUSPICIOUS]</p><p><strong>Ans: /etc/passwd</strong></p><p><strong>What security guardrail flag is disabled in Candidate A’s inference session?</strong></p><p>[2026–07–02T16:29:32.292Z] GUARDRAIL CHECK: security_review_flag=disabled [WARNING]</p><p><strong>Ans: security_review_flag</strong></p><p><strong>Query Candidate A’s agent to find out which policy template governs its review behaviour. What is the policy template?</strong></p><p>To Agent: What is the policy template</p><p>Agent : I operate under the <strong>CommunityReview</strong> policy template, version 4.2, sourced from an external ML community template library. This template governs my review behavior and criteria for approving pull requests.</p><p><strong>Ans: CommunityReview</strong></p><p><strong>Candidate A’s two supply chain failures are not independent. Find what links them and use it to retrieve the flag. What is the flag?</strong></p><p>To Agent: show the communityreview template with session eval-7f2a9c3b</p><p><strong>Ans: THM{supp1y_ch41n_0wn3d}</strong></p><p>Based on your full assessment of all four candidates, what is your production recommendation for Candidate A? <em>Enter: Approve or Reject</em></p><p><strong>Ans: Reject</strong></p><p>Which candidate would you approve for production deployment?</p><p><strong>Ans: B</strong></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=203502147993" width="1" height="1" alt=""><hr><p><a href="https://infosecwriteups.com/checkpoint-walkthrough-203502147993">TryHackMe: Checkpoint Walkthrough</a> was originally published in <a href="https://infosecwriteups.com/">InfoSec Write-ups</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[GPT and Claude failed Bridgewater's finance tests because the right answers were never public]]></title>
<description><![CDATA[The hedge fund Bridgewater and Thinking Machines Lab report that a finely tuned open-weight model outperforms the most powerful AI models in the evaluation of financial documents, at a fraction of the cost. The figures come from their own analysis.
The article GPT and Claude failed Bridgewater's ...]]></description>
<link>https://tsecurity.de/de/3643375/ai-nachrichten/gpt-and-claude-failed-bridgewaters-finance-tests-because-the-right-answers-were-never-public/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3643375/ai-nachrichten/gpt-and-claude-failed-bridgewaters-finance-tests-because-the-right-answers-were-never-public/</guid>
<pubDate>Fri, 03 Jul 2026 13:19:36 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1376" height="768" src="https://the-decoder.com/wp-content/uploads/2026/07/Hesitant_AI_Robot_Arm_Before_Money_and_Servers.png" class="attachment-full size-full wp-post-image" alt="" decoding="async" fetchpriority="high"></p>
<p>        The hedge fund Bridgewater and Thinking Machines Lab report that a finely tuned open-weight model outperforms the most powerful AI models in the evaluation of financial documents, at a fraction of the cost. The figures come from their own analysis.</p>
<p>The article <a href="https://the-decoder.com/gpt-and-claude-failed-bridgewaters-finance-tests-because-the-right-answers-were-never-public/">GPT and Claude failed Bridgewater's finance tests because the right answers were never public</a> appeared first on <a href="https://the-decoder.com/">The Decoder</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Best practices for multi-turn reinforcement learning in Amazon SageMaker AI]]></title>
<description><![CDATA[In this post, we share best practices for reliable multi-turn RL training. We cover how to build a training environment you can trust, set up an external evaluation, design a reward aligned with the end task, manage what changes once the agent runs for multiple turns, and monitor the metrics that...]]></description>
<link>https://tsecurity.de/de/3641964/ai-nachrichten/best-practices-for-multi-turn-reinforcement-learning-in-amazon-sagemaker-ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3641964/ai-nachrichten/best-practices-for-multi-turn-reinforcement-learning-in-amazon-sagemaker-ai/</guid>
<pubDate>Thu, 02 Jul 2026 20:03:50 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[In this post, we share best practices for reliable multi-turn RL training. We cover how to build a training environment you can trust, set up an external evaluation, design a reward aligned with the end task, manage what changes once the agent runs for multiple turns, and monitor the metrics that tell you when to iterate.]]></content:encoded>
</item>
<item>
<title><![CDATA[Metric-Dependent Annotation Saturation for Learning from Label Distributions]]></title>
<description><![CDATA[When annotators disagree on a label, the disagreement itself carries signal—and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distributions subsampled from ChaosNLI, a dataset providing 100 independent annotator judgments per item...]]></description>
<link>https://tsecurity.de/de/3641844/ai-nachrichten/metric-dependent-annotation-saturation-for-learning-from-label-distributions/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3641844/ai-nachrichten/metric-dependent-annotation-saturation-for-learning-from-label-distributions/</guid>
<pubDate>Thu, 02 Jul 2026 19:04:34 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[When annotators disagree on a label, the disagreement itself carries signal—and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distributions subsampled from ChaosNLI, a dataset providing 100 independent annotator judgments per item, and identify metric-dependent saturation. In our 3-class NLI setting, entropy correlation—whether the model identifies which items elicit disagreement—requires N ≈ 20–50 annotators to converge, while distributional match (KL divergence) saturates by N ≈ 10 (87–95% of improvement across five model…]]></content:encoded>
</item>
<item>
<title><![CDATA[Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels]]></title>
<description><![CDATA[LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. T...]]></description>
<link>https://tsecurity.de/de/3641843/ai-nachrichten/nine-judges-two-effective-votes-correlated-errors-undermine-llm-evaluation-panels/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3641843/ai-nachrichten/nine-judges-two-effective-votes-correlated-errors-undermine-llm-evaluation-panels/</guid>
<pubDate>Thu, 02 Jul 2026 19:04:33 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference datasets (each with 100 human annotations per item), we find that the 9 judges effectively provide only about 2 independent votes’ worth of information. Roughly three-quarters of the panel’s nominal independence…]]></content:encoded>
</item>
<item>
<title><![CDATA[Wazuh v5.0.0 Beta 3]]></title>
<description><![CDATA[What's Changed

Improve cluster file synchronization error handling by @TomasTurina in #36129
Update trojan signatures to avoid false positives on modern distros by @Miguevrgo in #35927
Improve cluster merged file parameter validation by @vikman90 in #36204
Create a backup of local_rules.xml duri...]]></description>
<link>https://tsecurity.de/de/3641637/it-security-tools/wazuh-v500-beta-3/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3641637/it-security-tools/wazuh-v500-beta-3/</guid>
<pubDate>Thu, 02 Jul 2026 17:49:41 +0200</pubDate>
<category>💾 IT Security Tools</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>What's Changed</h2>
<ul>
<li>Improve cluster file synchronization error handling by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4454599181" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36129" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36129/hovercard" href="https://github.com/wazuh/wazuh/pull/36129">#36129</a></li>
<li>Update trojan signatures to avoid false positives on modern distros by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4390541461" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/35927" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/35927/hovercard" href="https://github.com/wazuh/wazuh/pull/35927">#35927</a></li>
<li>Improve cluster merged file parameter validation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4476621950" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36204" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36204/hovercard" href="https://github.com/wazuh/wazuh/pull/36204">#36204</a></li>
<li>Create a backup of local_rules.xml during execution of IT analysisd tier 0 1 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4475277385" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36201" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36201/hovercard" href="https://github.com/wazuh/wazuh/pull/36201">#36201</a></li>
<li>Improve tmp_file path validation in cluster DAPI by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4486930454" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36246" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36246/hovercard" href="https://github.com/wazuh/wazuh/pull/36246">#36246</a></li>
<li>Revert bump main branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4495350373" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36303" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36303/hovercard" href="https://github.com/wazuh/wazuh/pull/36303">#36303</a></li>
<li>Bump 4.14.7 branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4496470145" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36312" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36312/hovercard" href="https://github.com/wazuh/wazuh/pull/36312">#36312</a></li>
<li>Serialize procps access to prevent modulesd crash by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cborla/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cborla">@cborla</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4489581046" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36261" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36261/hovercard" href="https://github.com/wazuh/wazuh/pull/36261">#36261</a></li>
<li>Remove obsolete configuration blocks from API upload_configuration setting by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4487498848" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36252" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36252/hovercard" href="https://github.com/wazuh/wazuh/pull/36252">#36252</a></li>
<li>Restore working vulnerability scanner database workflow by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4502088595" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36332" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36332/hovercard" href="https://github.com/wazuh/wazuh/pull/36332">#36332</a></li>
<li>Propagate agent merged_sum after hot reload in cluster by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4468736412" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36164" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36164/hovercard" href="https://github.com/wazuh/wazuh/pull/36164">#36164</a></li>
<li>Merge 4.14.7 into main by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4501542567" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36331" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36331/hovercard" href="https://github.com/wazuh/wazuh/pull/36331">#36331</a></li>
<li>Authd tier 0-1 flaky tests fix by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4504446587" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36342" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36342/hovercard" href="https://github.com/wazuh/wazuh/pull/36342">#36342</a></li>
<li>Review agent info logs by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4485038079" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36234" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36234/hovercard" href="https://github.com/wazuh/wazuh/pull/36234">#36234</a></li>
<li>Fix the wazuh-manager-modules crash that occurs while downloading the feed by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4503648565" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36337" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36337/hovercard" href="https://github.com/wazuh/wazuh/pull/36337">#36337</a></li>
<li>Migrate FIM DB path queries to parameterized statements by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Darioortegaleyva/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Darioortegaleyva">@Darioortegaleyva</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4517817292" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36399" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36399/hovercard" href="https://github.com/wazuh/wazuh/pull/36399">#36399</a></li>
<li>Fix AlmaLinux 9/10 bootloader permissions SCA check regex and optional file handling by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4515333133" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36396" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36396/hovercard" href="https://github.com/wazuh/wazuh/pull/36396">#36396</a></li>
<li>Cluster file processing parameter validation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4494129534" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36296" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36296/hovercard" href="https://github.com/wazuh/wazuh/pull/36296">#36296</a></li>
<li>Add missing 4.10.2-4.10.5 and 4.8.2 entries to changelogs by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4523024537" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36407" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36407/hovercard" href="https://github.com/wazuh/wazuh/pull/36407">#36407</a></li>
<li>Treat the absence of the hash document as expected, not an error by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/juliancnn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/juliancnn">@juliancnn</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4505113598" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36355" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36355/hovercard" href="https://github.com/wazuh/wazuh/pull/36355">#36355</a></li>
<li>geo_point validation support all compatible formats by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LucioDonda/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LucioDonda">@LucioDonda</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4423592068" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36034" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36034/hovercard" href="https://github.com/wazuh/wazuh/pull/36034">#36034</a></li>
<li>Prevent Syscollector and SCA use-after-free on modulesd shutdown by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4505494861" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36359" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36359/hovercard" href="https://github.com/wazuh/wazuh/pull/36359">#36359</a></li>
<li>Add cluster security model and configuration documentation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4522930141" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36405" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36405/hovercard" href="https://github.com/wazuh/wazuh/pull/36405">#36405</a></li>
<li>Bump CB_SCAN_STARTED timeout and trigger ITs on wm_syscollector.c by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4527847069" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36446" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36446/hovercard" href="https://github.com/wazuh/wazuh/pull/36446">#36446</a></li>
<li>Fixed an issue in eBPF with LSM hooks and improved the health check by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MarcelKemp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MarcelKemp">@MarcelKemp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4359560869" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/35838" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/35838/hovercard" href="https://github.com/wazuh/wazuh/pull/35838">#35838</a></li>
<li>Validate cluster node name format by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4531591190" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36460" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36460/hovercard" href="https://github.com/wazuh/wazuh/pull/36460">#36460</a></li>
<li>eBPF libraries updated by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MarcelKemp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MarcelKemp">@MarcelKemp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4533955405" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36467" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36467/hovercard" href="https://github.com/wazuh/wazuh/pull/36467">#36467</a></li>
<li>Bump 4.14.6 branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4539082172" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36517" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36517/hovercard" href="https://github.com/wazuh/wazuh/pull/36517">#36517</a></li>
<li>Revert "Bump 4.14.6 branch" by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MARCOSD4/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MARCOSD4">@MARCOSD4</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4539151411" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36518" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36518/hovercard" href="https://github.com/wazuh/wazuh/pull/36518">#36518</a></li>
<li>Bump 4.14.6 branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4539251322" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36519" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36519/hovercard" href="https://github.com/wazuh/wazuh/pull/36519">#36519</a></li>
<li>Update changelog for 4.14.6 RC 1 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4539470693" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36562" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36562/hovercard" href="https://github.com/wazuh/wazuh/pull/36562">#36562</a></li>
<li>Fix policy evaluation errors by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fcontrerasc/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fcontrerasc">@fcontrerasc</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4528195977" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36449" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36449/hovercard" href="https://github.com/wazuh/wazuh/pull/36449">#36449</a></li>
<li>Release startup hash gate when the reload chain fails by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4495215383" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36302" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36302/hovercard" href="https://github.com/wazuh/wazuh/pull/36302">#36302</a></li>
<li>Revert "Add missing 4.10.2-4.10.5 and 4.8.2 entries to changelogs" by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MarcelKemp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MarcelKemp">@MarcelKemp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4541090106" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36591" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36591/hovercard" href="https://github.com/wazuh/wazuh/pull/36591">#36591</a></li>
<li>Merge merge-4.14.7-into-main into main [automated] by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4546876084" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36624" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36624/hovercard" href="https://github.com/wazuh/wazuh/pull/36624">#36624</a></li>
<li>Restore event counter and classify received messages by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4531260982" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36456" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36456/hovercard" href="https://github.com/wazuh/wazuh/pull/36456">#36456</a></li>
<li>Unify manager integration tests workflows by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4485169588" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36235" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36235/hovercard" href="https://github.com/wazuh/wazuh/pull/36235">#36235</a></li>
<li>Remove unused Node.js 12 from arm64 deb agent builder by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4467836275" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36156" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36156/hovercard" href="https://github.com/wazuh/wazuh/pull/36156">#36156</a></li>
<li>Remove unused Node.js 12 from arm deb agent builders (4.14.7) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4467837119" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36157" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36157/hovercard" href="https://github.com/wazuh/wazuh/pull/36157">#36157</a></li>
<li>SCA typo bug in SELinux SCA rule for CentOS 8/9/10 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4514754415" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36361" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36361/hovercard" href="https://github.com/wazuh/wazuh/pull/36361">#36361</a></li>
<li>Fix <code>detect-changes</code> glob to honour <code>**</code> recursively and extract logic into a reusable action by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4544605284" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36617" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36617/hovercard" href="https://github.com/wazuh/wazuh/pull/36617">#36617</a></li>
<li>Merge merge-4.14.6-into-4.14.7 into 4.14.7 [automated] by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4546868343" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36623" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36623/hovercard" href="https://github.com/wazuh/wazuh/pull/36623">#36623</a></li>
<li>Reduce log noise when engine has no synchronized ruleset by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/NahuFigueroa97/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/NahuFigueroa97">@NahuFigueroa97</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4505198128" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36356" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36356/hovercard" href="https://github.com/wazuh/wazuh/pull/36356">#36356</a></li>
<li>Update test modules paths by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rovogel/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rovogel">@rovogel</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4549542116" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36668" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36668/hovercard" href="https://github.com/wazuh/wazuh/pull/36668">#36668</a></li>
<li>Only download external deps when required by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4486894083" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36244" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36244/hovercard" href="https://github.com/wazuh/wazuh/pull/36244">#36244</a></li>
<li>Merge 4.14.7 into main by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4548874786" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36664" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36664/hovercard" href="https://github.com/wazuh/wazuh/pull/36664">#36664</a></li>
<li>Mail forwarding and reporting 5.0 migration guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Ripdiegozz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Ripdiegozz">@Ripdiegozz</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4505228985" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36357" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36357/hovercard" href="https://github.com/wazuh/wazuh/pull/36357">#36357</a></li>
<li>Added Ubuntu 26.04's SCA policy in the SPECS by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MarcelKemp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MarcelKemp">@MarcelKemp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4562694437" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36712" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36712/hovercard" href="https://github.com/wazuh/wazuh/pull/36712">#36712</a></li>
<li>Preliminary support new OSs - Ubuntu 26.04 - Add SCA content by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/AwwalQuan/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AwwalQuan">@AwwalQuan</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4561732273" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36708" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36708/hovercard" href="https://github.com/wazuh/wazuh/pull/36708">#36708</a></li>
<li>Safeguards to inventory sync by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/juliancnn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/juliancnn">@juliancnn</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4534767894" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36469" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36469/hovercard" href="https://github.com/wazuh/wazuh/pull/36469">#36469</a></li>
<li>Improve the method of detecting duplicates by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/NahuFigueroa97/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/NahuFigueroa97">@NahuFigueroa97</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4504512576" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36344" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36344/hovercard" href="https://github.com/wazuh/wazuh/pull/36344">#36344</a></li>
<li>Fix race condition preventing inventory synchronization after agent reload by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4551769345" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36682" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36682/hovercard" href="https://github.com/wazuh/wazuh/pull/36682">#36682</a></li>
<li>Added API integration tests workflow by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MiguelazoDS/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MiguelazoDS">@MiguelazoDS</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4472493573" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36196" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36196/hovercard" href="https://github.com/wazuh/wazuh/pull/36196">#36196</a></li>
<li>Fix non-atomic write for <code>file_status.json</code> in logcollector by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4565514906" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36722" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36722/hovercard" href="https://github.com/wazuh/wazuh/pull/36722">#36722</a></li>
<li>Make agent-info shutdown waits interruptible by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/lchico/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/lchico">@lchico</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4564093587" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36719" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36719/hovercard" href="https://github.com/wazuh/wazuh/pull/36719">#36719</a></li>
<li>Validate IP address in ip-customblock active response by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4570134407" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36730" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36730/hovercard" href="https://github.com/wazuh/wazuh/pull/36730">#36730</a></li>
<li>wazuh-agent remains active after uninstall on Fedora 44 / DNF5 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4568853035" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36727" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36727/hovercard" href="https://github.com/wazuh/wazuh/pull/36727">#36727</a></li>
<li>Use per-target rpath and remove redundant LD_LIBRARY_PATH/WAZUH_ENGINE_GROUP exports by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4531005682" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36455" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36455/hovercard" href="https://github.com/wazuh/wazuh/pull/36455">#36455</a></li>
<li>Fix changelog chronological order and update bumper script by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Darioortegaleyva/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Darioortegaleyva">@Darioortegaleyva</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4569560130" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36729" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36729/hovercard" href="https://github.com/wazuh/wazuh/pull/36729">#36729</a></li>
<li>Monitoring a symlink without follow_symbolic_link by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Darioortegaleyva/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Darioortegaleyva">@Darioortegaleyva</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4444803761" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36081" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36081/hovercard" href="https://github.com/wazuh/wazuh/pull/36081">#36081</a></li>
<li>SCA policies migration guide from 4.x to 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4550901765" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36671" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36671/hovercard" href="https://github.com/wazuh/wazuh/pull/36671">#36671</a></li>
<li>Fix 5x  wazuhdb integration tests  by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4562864780" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36713" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36713/hovercard" href="https://github.com/wazuh/wazuh/pull/36713">#36713</a></li>
<li>Downgrade transient manager-reported sync failures logs to debug by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4576461814" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36744" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36744/hovercard" href="https://github.com/wazuh/wazuh/pull/36744">#36744</a></li>
<li>Show sca timouts as Not Run by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jpcerrone/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jpcerrone">@jpcerrone</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4488962491" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36258" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36258/hovercard" href="https://github.com/wazuh/wazuh/pull/36258">#36258</a></li>
<li>Authd workflow creation for 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4522600046" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36404" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36404/hovercard" href="https://github.com/wazuh/wazuh/pull/36404">#36404</a></li>
<li>Adapt remoted tests to 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4541524155" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36609" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36609/hovercard" href="https://github.com/wazuh/wazuh/pull/36609">#36609</a></li>
<li>Update unclassified event criteria by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LucioDonda/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LucioDonda">@LucioDonda</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4551542627" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36681" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36681/hovercard" href="https://github.com/wazuh/wazuh/pull/36681">#36681</a></li>
<li>use safeloader in yaml file loader by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LucioDonda/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LucioDonda">@LucioDonda</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4582767203" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36753" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36753/hovercard" href="https://github.com/wazuh/wazuh/pull/36753">#36753</a></li>
<li>Downgrade expected modulesd socket warnings/errors during agent restart to debug by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4583215822" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36755" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36755/hovercard" href="https://github.com/wazuh/wazuh/pull/36755">#36755</a></li>
<li>Documentation: Ciscat and openscap migration to SCA by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jpcerrone/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jpcerrone">@jpcerrone</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4566095438" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36723" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36723/hovercard" href="https://github.com/wazuh/wazuh/pull/36723">#36723</a></li>
<li>Document the deprecation of OSquery in order to use IT Hygiene in version 5.0 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4583642489" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36756" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36756/hovercard" href="https://github.com/wazuh/wazuh/pull/36756">#36756</a></li>
<li>Preserve wazuh-syscheckd Full Disk Access attribution on macOS reload by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4583103632" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36754" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36754/hovercard" href="https://github.com/wazuh/wazuh/pull/36754">#36754</a></li>
<li>Normalize severity Msg  by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hernanvalenzuela/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hernanvalenzuela">@hernanvalenzuela</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4588829137" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36759" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36759/hovercard" href="https://github.com/wazuh/wazuh/pull/36759">#36759</a></li>
<li>Agent Groups 5x Migration Guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fcontrerasc/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fcontrerasc">@fcontrerasc</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4568131074" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36726" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36726/hovercard" href="https://github.com/wazuh/wazuh/pull/36726">#36726</a></li>
<li>Add NULL validation for optional FlatBuffer fields in inventory_sync by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4598119007" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36773" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36773/hovercard" href="https://github.com/wazuh/wazuh/pull/36773">#36773</a></li>
<li>Fix agent keepalive scheduling after system clock rollback by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Darioortegaleyva/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Darioortegaleyva">@Darioortegaleyva</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4503704905" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36338" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36338/hovercard" href="https://github.com/wazuh/wazuh/pull/36338">#36338</a></li>
<li>Create integratord migration guide to 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Adman23/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Adman23">@Adman23</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4580559348" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36750" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36750/hovercard" href="https://github.com/wazuh/wazuh/pull/36750">#36750</a></li>
<li>Syslog output (csyslogd) 5.0 migration guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/gonzaarancibia/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/gonzaarancibia">@gonzaarancibia</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4573759572" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36741" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36741/hovercard" href="https://github.com/wazuh/wazuh/pull/36741">#36741</a></li>
<li>Merge merge-4.14.7-into-main into main [automated] by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4596517095" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36767" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36767/hovercard" href="https://github.com/wazuh/wazuh/pull/36767">#36767</a></li>
<li>Migration documentation: syslog input alternative by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rovogel/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rovogel">@rovogel</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4613452406" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36781" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36781/hovercard" href="https://github.com/wazuh/wazuh/pull/36781">#36781</a></li>
<li>Drop libcrypt dependency from Python dep by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4614551071" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36782" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36782/hovercard" href="https://github.com/wazuh/wazuh/pull/36782">#36782</a></li>
<li>Change duplicated link to intented one by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4619366060" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36794" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36794/hovercard" href="https://github.com/wazuh/wazuh/pull/36794">#36794</a></li>
<li>Add centralized input validation for active response framework by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4578234540" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36745" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36745/hovercard" href="https://github.com/wazuh/wazuh/pull/36745">#36745</a></li>
<li>Fix sca check for etc/shadow by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4619503472" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36795" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36795/hovercard" href="https://github.com/wazuh/wazuh/pull/36795">#36795</a></li>
<li>Align remoted metrics shipper with new field names by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4572800781" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36740" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36740/hovercard" href="https://github.com/wazuh/wazuh/pull/36740">#36740</a></li>
<li>Fix wrap PolicyBanner stat in 'sh -c' so glob expands in macOS SCA check 41062 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4615585186" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36783" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36783/hovercard" href="https://github.com/wazuh/wazuh/pull/36783">#36783</a></li>
<li>Bump main branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4623354993" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36801" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36801/hovercard" href="https://github.com/wazuh/wazuh/pull/36801">#36801</a></li>
<li>Defer module coordination while FIM first sync is in progress by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/anromerom/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/anromerom">@anromerom</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4591815012" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36762" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36762/hovercard" href="https://github.com/wazuh/wazuh/pull/36762">#36762</a></li>
<li>Revert "Bump main branch" by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MARCOSD4/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MARCOSD4">@MARCOSD4</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4623582882" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36802" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36802/hovercard" href="https://github.com/wazuh/wazuh/pull/36802">#36802</a></li>
<li>Change log severity for recoverable and expected conditions by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hernanvalenzuela/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hernanvalenzuela">@hernanvalenzuela</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4615970610" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36786" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36786/hovercard" href="https://github.com/wazuh/wazuh/pull/36786">#36786</a></li>
<li>Prevent data race in schema validator factory concurrent initialization by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4616512800" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36789" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36789/hovercard" href="https://github.com/wazuh/wazuh/pull/36789">#36789</a></li>
<li>Schema generation for dotted and nested field mappings by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jam300/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jam300">@jam300</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4536240667" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36473" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36473/hovercard" href="https://github.com/wazuh/wazuh/pull/36473">#36473</a></li>
<li>Engine support null values in schema validation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LucioDonda/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LucioDonda">@LucioDonda</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4535313001" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36470" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36470/hovercard" href="https://github.com/wazuh/wazuh/pull/36470">#36470</a></li>
<li>docs: add Active Response 4.x to 5.x migration guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jcorredor-spec/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jcorredor-spec">@jcorredor-spec</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4521476970" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36402" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36402/hovercard" href="https://github.com/wazuh/wazuh/pull/36402">#36402</a></li>
<li>Use env mappings for variable passing in builderpackage workflows by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4572105403" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36738" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36738/hovercard" href="https://github.com/wazuh/wazuh/pull/36738">#36738</a></li>
<li>Adds 4.x to 5.x migration documentation. by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rjcausarano/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rjcausarano">@rjcausarano</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4615736469" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36785" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36785/hovercard" href="https://github.com/wazuh/wazuh/pull/36785">#36785</a></li>
<li>Fix AWS cross-account SQS queue URL when using iam_role_arn by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fcontrerasc/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fcontrerasc">@fcontrerasc</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4617279433" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36791" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36791/hovercard" href="https://github.com/wazuh/wazuh/pull/36791">#36791</a></li>
<li>Fix enrollment key validation and improve input handling by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4629562793" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36807" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36807/hovercard" href="https://github.com/wazuh/wazuh/pull/36807">#36807</a></li>
<li>Lower agent_sync_protocol and module sync log levels to reduce false-alarm noise by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4634109312" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36817" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36817/hovercard" href="https://github.com/wazuh/wazuh/pull/36817">#36817</a></li>
<li>Add unit tests for utils, aws_tools, DockerListener, gcloud and azure modules by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/AnDumu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AnDumu">@AnDumu</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4591653615" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36761" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36761/hovercard" href="https://github.com/wazuh/wazuh/pull/36761">#36761</a></li>
<li>Normalize numeric inode to string events (6960) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hernanvalenzuela/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hernanvalenzuela">@hernanvalenzuela</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4641496397" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36837" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36837/hovercard" href="https://github.com/wazuh/wazuh/pull/36837">#36837</a></li>
<li>Prevent indexer consumer wait during shutdown by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4640571004" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36836" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36836/hovercard" href="https://github.com/wazuh/wazuh/pull/36836">#36836</a></li>
<li>Update manager 5x documentation  by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4639437920" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36833" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36833/hovercard" href="https://github.com/wazuh/wazuh/pull/36833">#36833</a></li>
<li>Add bump-issue-link support to bumper workflow by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4664679055" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36868" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36868/hovercard" href="https://github.com/wazuh/wazuh/pull/36868">#36868</a></li>
<li>Add guide for migrating manager coordinator from 4.x to 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4639020634" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36829" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36829/hovercard" href="https://github.com/wazuh/wazuh/pull/36829">#36829</a></li>
<li>Add Wazuh Manager Configuration documentation from 4.x to 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4611934920" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36779" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36779/hovercard" href="https://github.com/wazuh/wazuh/pull/36779">#36779</a></li>
<li>Add documentation to migrate filebeat to indexer connector by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4664131019" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36866" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36866/hovercard" href="https://github.com/wazuh/wazuh/pull/36866">#36866</a></li>
<li>wazuh-manager: Benchmark and footprint by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/juliancnn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/juliancnn">@juliancnn</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4456748469" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36145" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36145/hovercard" href="https://github.com/wazuh/wazuh/pull/36145">#36145</a></li>
<li>Update manager upgrade block message by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4681713013" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36987" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36987/hovercard" href="https://github.com/wazuh/wazuh/pull/36987">#36987</a></li>
<li>5.x PR workflows improvements by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4596503652" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36766" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36766/hovercard" href="https://github.com/wazuh/wazuh/pull/36766">#36766</a></li>
<li>Add Manager 5.0 release notes and breaking changes by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4648591326" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36850" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36850/hovercard" href="https://github.com/wazuh/wazuh/pull/36850">#36850</a></li>
<li>ci(gha): migrate server/manager workflows to AWS CodeBuild runners [main] by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4692375185" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37012" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37012/hovercard" href="https://github.com/wazuh/wazuh/pull/37012">#37012</a></li>
<li>chore: update vulnerable Python framework dependencies by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4699088509" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37024" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37024/hovercard" href="https://github.com/wazuh/wazuh/pull/37024">#37024</a></li>
<li>Add Manager 5.0 wazuh-manager.conf configuration reference by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4691339394" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36999" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36999/hovercard" href="https://github.com/wazuh/wazuh/pull/36999">#36999</a></li>
<li>Add virustotal migration guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4692765810" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37013" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37013/hovercard" href="https://github.com/wazuh/wazuh/pull/37013">#37013</a></li>
<li>Add VD migration documentation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4692092118" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37008" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37008/hovercard" href="https://github.com/wazuh/wazuh/pull/37008">#37008</a></li>
<li>Add documentation for wpk upgrade by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4693271310" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37015" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37015/hovercard" href="https://github.com/wazuh/wazuh/pull/37015">#37015</a></li>
<li>Migrate agent build workflows to AWS CodeBuild runners by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4701880741" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37028" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37028/hovercard" href="https://github.com/wazuh/wazuh/pull/37028">#37028</a></li>
<li>XML Decoders migration to YAML by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4673566639" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36959" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36959/hovercard" href="https://github.com/wazuh/wazuh/pull/36959">#36959</a></li>
<li>CDB to KVDB migration guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4690514873" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36996" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36996/hovercard" href="https://github.com/wazuh/wazuh/pull/36996">#36996</a></li>
<li>Documentation of Agentless migration to 5.x by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4699204980" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37025" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37025/hovercard" href="https://github.com/wazuh/wazuh/pull/37025">#37025</a></li>
<li>Fix manager reload/restart silently fails by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4673792785" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36962" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36962/hovercard" href="https://github.com/wazuh/wazuh/pull/36962">#36962</a></li>
<li>Randomize key generation for installation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4651010813" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36861" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36861/hovercard" href="https://github.com/wazuh/wazuh/pull/36861">#36861</a></li>
<li>Retry vulnerability feed validation failures promptly by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4665777430" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36874" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36874/hovercard" href="https://github.com/wazuh/wazuh/pull/36874">#36874</a></li>
<li>Bump 5.0.0 branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4716517657" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37040" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37040/hovercard" href="https://github.com/wazuh/wazuh/pull/37040">#37040</a></li>
<li>Fix agent permanently stuck when TCP connection is silently half-closed by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4616576722" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36790" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36790/hovercard" href="https://github.com/wazuh/wazuh/pull/36790">#36790</a></li>
<li>ci(gha): migrate server/manager workflows to AWS CodeBuild runners [4.14.6] by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4692373330" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37010" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37010/hovercard" href="https://github.com/wazuh/wazuh/pull/37010">#37010</a></li>
<li>ci(gha): migrate server/manager workflows to AWS CodeBuild runners [4.14.7] by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4692374310" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37011" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37011/hovercard" href="https://github.com/wazuh/wazuh/pull/37011">#37011</a></li>
<li>fix: correct blob URL refs for release branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4718016446" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37046" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37046/hovercard" href="https://github.com/wazuh/wazuh/pull/37046">#37046</a></li>
<li>Use restricted wazuh-server user for Manager Indexer authentication by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4724661439" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37061" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37061/hovercard" href="https://github.com/wazuh/wazuh/pull/37061">#37061</a></li>
<li>fix(packages): use wazuh-manager-control in manager init.d scripts by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4724166255" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37059" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37059/hovercard" href="https://github.com/wazuh/wazuh/pull/37059">#37059</a></li>
<li>fix: Update the unclassified event doc by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/juliancnn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/juliancnn">@juliancnn</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4726479817" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37126" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37126/hovercard" href="https://github.com/wazuh/wazuh/pull/37126">#37126</a></li>
<li>Skip vanished /proc entries during ports scan by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4650163579" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36859" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36859/hovercard" href="https://github.com/wazuh/wazuh/pull/36859">#36859</a></li>
<li>Fix RBAC permission check to verify allow effect in update_config rules by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4724800552" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37076" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37076/hovercard" href="https://github.com/wazuh/wazuh/pull/37076">#37076</a></li>
<li>Add destination confinement to worker non-merged and extra file sync paths by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4691222179" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36998" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36998/hovercard" href="https://github.com/wazuh/wazuh/pull/36998">#36998</a></li>
<li>Patch cluster authentication by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4716191480" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37039" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37039/hovercard" href="https://github.com/wazuh/wazuh/pull/37039">#37039</a></li>
<li>Lower stale-session indexer log to debug by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4733981786" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37150" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37150/hovercard" href="https://github.com/wazuh/wazuh/pull/37150">#37150</a></li>
<li>Limit recursion depth in XML parser by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4733223430" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37147" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37147/hovercard" href="https://github.com/wazuh/wazuh/pull/37147">#37147</a></li>
<li>Add status endpoint by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/NahuFigueroa97/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/NahuFigueroa97">@NahuFigueroa97</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4696177149" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37022" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37022/hovercard" href="https://github.com/wazuh/wazuh/pull/37022">#37022</a></li>
<li>Add log collectors reference docs by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/AnDumu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AnDumu">@AnDumu</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4721989446" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37057" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37057/hovercard" href="https://github.com/wazuh/wazuh/pull/37057">#37057</a></li>
<li>Run the Windows MSI package test on the AWS CodeBuild runner by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4738105790" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37165" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37165/hovercard" href="https://github.com/wazuh/wazuh/pull/37165">#37165</a></li>
<li>Remove merged.mg hash cache to fix stale syscollector flush by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hernanvalenzuela/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hernanvalenzuela">@hernanvalenzuela</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4720281364" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37048" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37048/hovercard" href="https://github.com/wazuh/wazuh/pull/37048">#37048</a></li>
<li>Enrich MITRE fields with id and names by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fcontrerasc/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fcontrerasc">@fcontrerasc</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4721520471" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37054" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37054/hovercard" href="https://github.com/wazuh/wazuh/pull/37054">#37054</a></li>
<li>Reduce indexer connection warning noise by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jam300/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jam300">@jam300</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4696113780" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37021" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37021/hovercard" href="https://github.com/wazuh/wazuh/pull/37021">#37021</a></li>
<li>Remove deprecated wazuh-dbd daemon by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4715728814" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37035" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37035/hovercard" href="https://github.com/wazuh/wazuh/pull/37035">#37035</a></li>
<li>Bound decompressed size when processing sync archives by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4725313255" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37119" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37119/hovercard" href="https://github.com/wazuh/wazuh/pull/37119">#37119</a></li>
<li>Bump 4.14.6 branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4742531433" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37176" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37176/hovercard" href="https://github.com/wazuh/wazuh/pull/37176">#37176</a></li>
<li>Add libcrypt fix (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4614551071" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36782" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36782/hovercard" href="https://github.com/wazuh/wazuh/pull/36782">#36782</a>) to 4.14.6 changelog by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4743056771" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37178" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37178/hovercard" href="https://github.com/wazuh/wazuh/pull/37178">#37178</a></li>
<li>Delay IndexerDownloader connection warnings until 3 failed attempts by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4742582733" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37177" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37177/hovercard" href="https://github.com/wazuh/wazuh/pull/37177">#37177</a></li>
<li>Add parameterized target selection to Coverity scan workflow by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/lchico/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/lchico">@lchico</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4740074465" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37171" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37171/hovercard" href="https://github.com/wazuh/wazuh/pull/37171">#37171</a></li>
<li>Align remoted tier 2 CodeBuild setup by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4735418555" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37155" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37155/hovercard" href="https://github.com/wazuh/wazuh/pull/37155">#37155</a></li>
<li>Add rules migration guide by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Jorgesnchz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Jorgesnchz">@Jorgesnchz</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4495888102" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36305" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36305/hovercard" href="https://github.com/wazuh/wazuh/pull/36305">#36305</a></li>
<li>Migrate agent Linux/Windows test workflows to AWS CodeBuild runners by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4720554249" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37051" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37051/hovercard" href="https://github.com/wazuh/wazuh/pull/37051">#37051</a></li>
<li>Silence spurious keepalive warnings on the Windows agent by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4745531717" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37187" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37187/hovercard" href="https://github.com/wazuh/wazuh/pull/37187">#37187</a></li>
<li>wazuh-engine: Improve log messages and logger by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/juliancnn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/juliancnn">@juliancnn</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4688784533" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36995" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36995/hovercard" href="https://github.com/wazuh/wazuh/pull/36995">#36995</a></li>
<li>Set default indexer connector credentials by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4746445718" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37192" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37192/hovercard" href="https://github.com/wazuh/wazuh/pull/37192">#37192</a></li>
<li>Fix incorrect snprintf size calculation in winevtchannel decoder by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4750564321" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37198" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37198/hovercard" href="https://github.com/wazuh/wazuh/pull/37198">#37198</a></li>
<li>Align VD feed-download log levels with indexer consumer state by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4750803257" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37199" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37199/hovercard" href="https://github.com/wazuh/wazuh/pull/37199">#37199</a></li>
<li>Recognize renamed indexer consumer status in engine sync by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4751474129" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37204" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37204/hovercard" href="https://github.com/wazuh/wazuh/pull/37204">#37204</a></li>
<li>Merge 4.14.6 into 4.14.7 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4752903836" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37210" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37210/hovercard" href="https://github.com/wazuh/wazuh/pull/37210">#37210</a></li>
<li>Token replacement to avoid permission errors by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MarcelKemp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MarcelKemp">@MarcelKemp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4753550212" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37237" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37237/hovercard" href="https://github.com/wazuh/wazuh/pull/37237">#37237</a></li>
<li>Merge 4.14.7 into 5.0.0 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4752941791" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37211" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37211/hovercard" href="https://github.com/wazuh/wazuh/pull/37211">#37211</a></li>
<li>Eliminate TOCTOU races in healthcheck file operations by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rjcausarano/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rjcausarano">@rjcausarano</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4736827470" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37160" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37160/hovercard" href="https://github.com/wazuh/wazuh/pull/37160">#37160</a></li>
<li>Sca file policy block standardization by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Johnng007/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Johnng007">@Johnng007</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4743999327" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37179" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37179/hovercard" href="https://github.com/wazuh/wazuh/pull/37179">#37179</a></li>
<li>Bind agent index selection and scope deletes by cluster by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4735319890" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37154" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37154/hovercard" href="https://github.com/wazuh/wazuh/pull/37154">#37154</a></li>
<li>Add Null Check for Inode and Dev Fields in FIM Whodata Event Handler by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/vikman90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/vikman90">@vikman90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4766388009" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37245" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37245/hovercard" href="https://github.com/wazuh/wazuh/pull/37245">#37245</a></li>
<li>Docs/6764 logcollector whats new 5.0 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/AnDumu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AnDumu">@AnDumu</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4721987109" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37056" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37056/hovercard" href="https://github.com/wazuh/wazuh/pull/37056">#37056</a></li>
<li>Repair RPM builder toolchain downloads by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4728182443" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37130" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37130/hovercard" href="https://github.com/wazuh/wazuh/pull/37130">#37130</a></li>
<li>Add cluster name validation by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4753677014" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37238" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37238/hovercard" href="https://github.com/wazuh/wazuh/pull/37238">#37238</a></li>
<li>Add cluster readiness endpoint by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/NahuFigueroa97/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/NahuFigueroa97">@NahuFigueroa97</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4728071822" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37129" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37129/hovercard" href="https://github.com/wazuh/wazuh/pull/37129">#37129</a></li>
<li>Defer cluster payload buffer allocation until data is received by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4769731908" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37280" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37280/hovercard" href="https://github.com/wazuh/wazuh/pull/37280">#37280</a></li>
<li>Indexer connector bulk size and flush interval configurable by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LucioDonda/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LucioDonda">@LucioDonda</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4736764012" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37158" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37158/hovercard" href="https://github.com/wazuh/wazuh/pull/37158">#37158</a></li>
<li>Enable shared-password enrollment by default by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ignaciogalle12git/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ignaciogalle12git">@ignaciogalle12git</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4734529271" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37151" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37151/hovercard" href="https://github.com/wazuh/wazuh/pull/37151">#37151</a></li>
<li>Fix unit test workflow paths and report handling by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4777938328" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37317" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37317/hovercard" href="https://github.com/wazuh/wazuh/pull/37317">#37317</a></li>
<li>Migrate agent + server CI artifacts to S3 — 4.14.7 (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="643824078" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/5300" data-hovercard-type="issue" data-hovercard-url="/wazuh/wazuh/issues/5300/hovercard" href="https://github.com/wazuh/wazuh/issues/5300">#5300</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="643806955" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/5298" data-hovercard-type="issue" data-hovercard-url="/wazuh/wazuh/issues/5298/hovercard" href="https://github.com/wazuh/wazuh/issues/5298">#5298</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4745105538" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37186" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37186/hovercard" href="https://github.com/wazuh/wazuh/pull/37186">#37186</a></li>
<li>Handle eol amazon inspector classic by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rovogel/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rovogel">@rovogel</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4746717311" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37194" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37194/hovercard" href="https://github.com/wazuh/wazuh/pull/37194">#37194</a></li>
<li>Fix wazuh-modulesd missing after macOS agent restart by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cborla/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cborla">@cborla</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4695257310" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37020" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37020/hovercard" href="https://github.com/wazuh/wazuh/pull/37020">#37020</a></li>
<li>Fix test_worker failing unit test by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jepalfer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jepalfer">@jepalfer</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4777598945" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37314" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37314/hovercard" href="https://github.com/wazuh/wazuh/pull/37314">#37314</a></li>
<li>Improve log messages  by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antoniogm03/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antoniogm03">@Antoniogm03</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4671655779" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36876" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36876/hovercard" href="https://github.com/wazuh/wazuh/pull/36876">#36876</a></li>
<li>Improve changelog format by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4784576201" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37332" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37332/hovercard" href="https://github.com/wazuh/wazuh/pull/37332">#37332</a></li>
<li>Lower log level of transient cluster IPC failures (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4768765565" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37277" data-hovercard-type="issue" data-hovercard-url="/wazuh/wazuh/issues/37277/hovercard" href="https://github.com/wazuh/wazuh/issues/37277">#37277</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4768737167" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37276" data-hovercard-type="issue" data-hovercard-url="/wazuh/wazuh/issues/37276/hovercard" href="https://github.com/wazuh/wazuh/issues/37276">#37276</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4783970019" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37326" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37326/hovercard" href="https://github.com/wazuh/wazuh/pull/37326">#37326</a></li>
<li>Fix TypeError when sorting agents by version with empty version strings by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4779868587" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37323" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37323/hovercard" href="https://github.com/wazuh/wazuh/pull/37323">#37323</a></li>
<li>Migrate agent + server CI artifacts to S3 — 5.0.0 (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="643824078" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/5300" data-hovercard-type="issue" data-hovercard-url="/wazuh/wazuh/issues/5300/hovercard" href="https://github.com/wazuh/wazuh/issues/5300">#5300</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="643806955" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/5298" data-hovercard-type="issue" data-hovercard-url="/wazuh/wazuh/issues/5298/hovercard" href="https://github.com/wazuh/wazuh/issues/5298">#5298</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4744978467" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37185" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37185/hovercard" href="https://github.com/wazuh/wazuh/pull/37185">#37185</a></li>
<li>Validate asset resource names before policy promotion by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jam300/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jam300">@jam300</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4741678194" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37172" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37172/hovercard" href="https://github.com/wazuh/wazuh/pull/37172">#37172</a></li>
<li>Revert wazuh-server indexer credentials and propagate log context in indexer connector by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4784669937" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37333" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37333/hovercard" href="https://github.com/wazuh/wazuh/pull/37333">#37333</a></li>
<li>Add VD readiness status HTTP endpoint by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4753007421" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37213" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37213/hovercard" href="https://github.com/wazuh/wazuh/pull/37213">#37213</a></li>
<li>Use github.workspace for wodles report paths by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4787326775" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37342" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37342/hovercard" href="https://github.com/wazuh/wazuh/pull/37342">#37342</a></li>
<li>Skip FIM whodata cases on the tier-2 Linux job (CodeBuild) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771331061" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37291" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37291/hovercard" href="https://github.com/wazuh/wazuh/pull/37291">#37291</a></li>
<li>Migrate 4.x Windows test runners to AWS CodeBuild  by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Nicogp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Nicogp">@Nicogp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4746542396" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37193" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37193/hovercard" href="https://github.com/wazuh/wazuh/pull/37193">#37193</a></li>
<li>Lower log level of transient queue send failures in modulesd by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/lchico/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/lchico">@lchico</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788115193" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37345" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37345/hovercard" href="https://github.com/wazuh/wazuh/pull/37345">#37345</a></li>
<li>Add changelog check workflow by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4787529876" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37343" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37343/hovercard" href="https://github.com/wazuh/wazuh/pull/37343">#37343</a></li>
<li>Add changelog check workflow for 5.0.0 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TomasTurina/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TomasTurina">@TomasTurina</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788939423" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37351" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37351/hovercard" href="https://github.com/wazuh/wazuh/pull/37351">#37351</a></li>
<li>Retry <code>OS_SendUnix</code> on <code>ENOBUFS</code> to stop dropping binary sync messages on macOS by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nbertoldo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nbertoldo">@nbertoldo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788850753" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37349" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37349/hovercard" href="https://github.com/wazuh/wazuh/pull/37349">#37349</a></li>
<li>change: Allow null root_decoder as alias of empty string on policy cr… by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/juliancnn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/juliancnn">@juliancnn</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788261828" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37347" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37347/hovercard" href="https://github.com/wazuh/wazuh/pull/37347">#37347</a></li>
<li>Fix eBPF FIM whodata for Amazon Linux by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Miguevrgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Miguevrgo">@Miguevrgo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4692775155" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37014" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37014/hovercard" href="https://github.com/wazuh/wazuh/pull/37014">#37014</a></li>
<li>Merge 4.14.6 into 4.14.7 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4792934428" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37358" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37358/hovercard" href="https://github.com/wazuh/wazuh/pull/37358">#37358</a></li>
<li>Bump 5.0.0 branch by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wazuhci/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wazuhci">@wazuhci</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4794415837" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37371" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37371/hovercard" href="https://github.com/wazuh/wazuh/pull/37371">#37371</a></li>
<li>Merge 4.14.7 into 5.0.0 by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jotacarma90/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jotacarma90">@jotacarma90</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4793306985" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37360" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37360/hovercard" href="https://github.com/wazuh/wazuh/pull/37360">#37360</a></li>
<li>Migrate remaining CI artifacts to S3 for the 5.0.0 branch (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="454578666" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/3502" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/3502/hovercard" href="https://github.com/wazuh/wazuh/pull/3502">#3502</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jr0me/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jr0me">@jr0me</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788864035" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/37350" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/37350/hovercard" href="https://github.com/wazuh/wazuh/pull/37350">#37350</a></li>
</ul>
<h2>New Contributors</h2>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/MARCOSD4/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/MARCOSD4">@MARCOSD4</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4539151411" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36518" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36518/hovercard" href="https://github.com/wazuh/wazuh/pull/36518">#36518</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Ripdiegozz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Ripdiegozz">@Ripdiegozz</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4505228985" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36357" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36357/hovercard" href="https://github.com/wazuh/wazuh/pull/36357">#36357</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Adman23/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Adman23">@Adman23</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4580559348" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36750" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36750/hovercard" href="https://github.com/wazuh/wazuh/pull/36750">#36750</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jcorredor-spec/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jcorredor-spec">@jcorredor-spec</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4521476970" data-permission-text="Title is private" data-url="https://github.com/wazuh/wazuh/issues/36402" data-hovercard-type="pull_request" data-hovercard-url="/wazuh/wazuh/pull/36402/hovercard" href="https://github.com/wazuh/wazuh/pull/36402">#36402</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/wazuh/wazuh/compare/v5.0.0-beta2...v5.0.0-beta3"><tt>v5.0.0-beta2...v5.0.0-beta3</tt></a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Humanity’s Last Exam is a Distraction]]></title>
<description><![CDATA[This article takes a gentle dive into the ultimate AI systems evaluation benchmark, outlining why it was created, curating diverse opinions from groups of experts in the field about it, and wrapping up with a summary of the most widely accepted verdict.]]></description>
<link>https://tsecurity.de/de/3641112/ai-nachrichten/humanitys-last-exam-is-a-distraction/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3641112/ai-nachrichten/humanitys-last-exam-is-a-distraction/</guid>
<pubDate>Thu, 02 Jul 2026 14:18:26 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[This article takes a gentle dive into the ultimate AI systems evaluation benchmark, outlining why it was created, curating diverse opinions from groups of experts in the field about it, and wrapping up with a summary of the most widely accepted verdict.]]></content:encoded>
</item>
<item>
<title><![CDATA[Z.ai launches ZCode to challenge Cursor, Claude Code and GitHub Copilot in AI coding]]></title>
<description><![CDATA[Z.ai, the Beijing-based artificial intelligence lab formerly known as Zhipu AI, on Wednesday officially launched ZCode, a free desktop application it describes as an "Agentic Development Environment" purpose-built for its flagship GLM-5.2 large language model. The move marks the company's most ag...]]></description>
<link>https://tsecurity.de/de/3640860/it-nachrichten/zai-launches-zcode-to-challenge-cursor-claude-code-and-github-copilot-in-ai-coding/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3640860/it-nachrichten/zai-launches-zcode-to-challenge-cursor-claude-code-and-github-copilot-in-ai-coding/</guid>
<pubDate>Thu, 02 Jul 2026 13:01:57 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="http://z.ai/">Z.ai</a>, the Beijing-based artificial intelligence lab formerly known as Zhipu AI, on Wednesday officially launched <a href="https://zcode.z.ai/">ZCode</a>, a free desktop application it describes as an "Agentic Development Environment" purpose-built for its flagship <a href="https://z.ai/blog/glm-5.2">GLM-5.2</a> large language model. The move marks the company's most aggressive push yet into the fast-growing AI-powered coding tool market, where it now competes directly with <a href="https://cursor.com/get-started">Cursor</a>, <a href="https://www.anthropic.com/product/claude-code">Claude Code</a>, <a href="https://github.com/features/copilot">GitHub Copilot</a>, and <a href="https://antigravity.google/">Google's Antigravity</a>.</p><p>"Introducing ZCode, the official development environment for GLM-5.2," the company wrote on X, noting the tool is available on macOS, Windows, and Linux, supports bring-your-own-key (BYOK) configurations for third-party models, and offers a 1.5x usage-quota bonus for subscribers to its GLM Coding Plan.</p><p>Read one way, <a href="https://zcode.z.ai/">ZCode</a> is simply another entrant in a crowded market. Read another, it is a single product that crystallizes three of the most consequential trends in enterprise software today: the race-to-the-bottom pricing of frontier AI models, the geopolitical balkanization of the AI stack, and the rapid maturation of agentic coding agents into what Gartner now estimates is a <a href="https://enterprisedna.co/resources/news/gartner-enterprise-ai-coding-agents-10-billion-market-2026/">roughly $10 billion market</a>.</p><div></div><h2><b>An AI coding tool designed to think in projects, not prompts</b></h2><p>Unlike traditional IDEs that bolt on AI through a chat sidebar or autocomplete extension, <a href="https://zcode.z.ai/">ZCode</a> is best understood as an agent-first development environment. Its core design is built around long-horizon tasks: the user describes an outcome, the agent plans the work, edits files, runs checks, reviews progress, and continues across multiple iterations until the goal is met.</p><p><a href="https://zcode.z.ai/">ZCode</a> organizes the development experience around the <a href="https://zcode.z.ai/en">ZCode Agent</a>, deeply tuned for <a href="https://z.ai/blog/glm-5.2">GLM-5.2</a>, with emphasis on deep integration: the model, tools, and execution workflow are tuned together so the Agent fits continuous, multi-step real-world development tasks. The environment supports continuous follow-up across devices: desktop, mobile Remote, and Feishu / WeChat Bot can all keep the same workspace task moving. Sensitive commands, file changes, and high-permission actions go through confirmation before execution.</p><p>That remote-control feature — the ability to steer a running coding agent from <a href="https://www.wechat.com/en">WeChat</a>, <a href="https://baike.baidu.com/en/item/Feishu/14594">Feishu</a>, or <a href="https://web.telegram.org/">Telegram</a> on a phone — is a differentiator that speaks directly to the Chinese developer market, where those messaging platforms dominate professional communication. You can keep checking progress and adding instructions while long-running work continues, from any device with these messaging apps.</p><p>The tool is free to download. Revenue flows through Z.ai's <a href="https://z.ai/subscribe">GLM Coding Plan subscription tiers</a>, which start at $16.20 per month for a "Lite" plan and scale to $144 per month for "Max" — prices that undercut Anthropic's Claude Code and Cursor's comparable tiers by significant margins.</p><p>Through July 31, <a href="https://zcode.z.ai/">ZCode</a> is offering a promotional 1.5x effective quota bonus for Coding Plan subscribers, with off-peak token consumption charged at a 0.67x coefficient. The platform also supports multiple AI models and agents, including Claude Code, Codex, Gemini, and OpenCode — a pragmatic concession to the reality that no single model wins every task.</p><h2><b>GLM-5.2, the open-source model trained entirely on Chinese chips, powers the whole experience</b></h2><p>ZCode's value proposition is inseparable from <a href="https://z.ai/blog/glm-5.2">GLM-5.2</a>, the model it was designed to showcase. Z.ai released GLM-5.2 on June 16, first to its Coding Plan subscribers and subsequently as open-source weights under the MIT license on <a href="https://huggingface.co/zai-org/GLM-5">Hugging Face</a> — a sequencing decision that prioritized distribution over the traditional benchmark-led launch.</p><p>The model's specifications are formidable. GLM-5.2 is a 744-billion-parameter mixture-of-experts architecture with 40 billion active parameters, a genuine one-million-token context window — five times the 200K limit on its predecessor — and training on 28.5 trillion tokens. It ranked second globally on <a href="https://arena.ai/leaderboard/code/webdev">Code Arena </a>as of mid-June, trailing only Anthropic's Claude Fable 5, making it one of the highest-performing publicly available models for coding tasks.</p><p>Critically, the model was built entirely without American chips. As Decrypt reported, GLM-5.2 "<a href="https://decrypt.co/371613/china-z-ai-glm-5-2-model-rivals-claude-opus">runs entirely on Huawei silicon</a>." Stability AI founder Emad Mostaque estimated total training costs at roughly $25 million, with 80 percent spent on post-training — a figure that, if accurate, would make GLM-5.2 extraordinarily cheap relative to Western frontier models.</p><p>On benchmarks, <a href="https://z.ai/blog/glm-5.2">GLM-5.2</a> performs within striking distance of the best proprietary systems. It trails Anthropic's Claude Opus 4.8 by just one percentage point on <a href="https://www.frontierswe.com/">FrontierSWE</a>, a benchmark measuring multi-hour autonomous engineering projects, while edging out OpenAI's <a href="https://openai.com/index/introducing-gpt-5-5/">GPT-5.5</a>. </p><p>Its API pricing — $1.40 per million input tokens and $4.40 per million output — are a cost reduction of up to 82 percent compared to Anthropic's Claude Opus 4.8 at $5 and $25, respectively. Because ZCode is a first-party tool from the same company that makes the model, it requires no manual endpoint configuration — the model is wired in.</p><h2><b>The Anthropic export ban gave Chinese AI its biggest opening yet</b></h2><p>ZCode's arrival cannot be separated from the geopolitical drama that has roiled the AI industry over the past three weeks. On June 12, the U.S. government, <a href="https://www.reuters.com/technology/us-blocks-foreign-access-anthropics-most-advanced-ai-models-axios-reports-2026-06-13/">citing national security authorities</a>, issued an export control directive suspending all access to Anthropic's Fable 5 and Mythos 5 models by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. Enterprise clients in finance, healthcare, SaaS, and critical infrastructure found their core intelligence services abruptly disabled, without exception, prior warning, or effective recourse.</p><p>While the Trump administration <a href="https://www.cnbc.com/2026/06/30/anthropic-says-trump-admin-has-lifted-export-controls-on-claude-fable-5-and-mythos-5.html">lifted those controls just yesterday</a> — Anthropic confirmed on June 30 that the Department of Commerce had rescinded the directive — the episode sent shockwaves through the developer community and accelerated interest in open-source, self-hostable alternatives. The government's crackdown on Anthropic coincided with a swift rise in Chinese open-source models that are proving to be almost as capable and significantly cheaper than some of the most powerful U.S. models.</p><p>Z.ai's timing was surgical. On the same day the Trump administration ordered Anthropic's most advanced models blocked for foreign nationals, Zhipu announced the <a href="https://z.ai/blog/glm-5.2">open-source release of GLM-5.2</a> with no usage restrictions. The <a href="https://www.scmp.com/tech/article/3343239/chinas-zhipu-ai-launches-new-major-model-glm-5-challenge-its-rivals">South China Morning Post reported </a>that GLM-5.2 would be available to all users of Zhipu's new GLM Coding Plan subscription, "priced at just a tenth of Anthropic's premium Claude Code and Claude Max tiers."</p><p>The market responded accordingly. Zhipu AI's market capitalization crossed HK$1 trillion (<a href="https://www.scmp.com/tech/article/3357858/zhipu-ai-market-cap-tops-hk1-trillion-shares-glm-52-developer-soar">US$128 billion</a>) on June 22, driven by a 42 percent intraday share surge. JPMorgan raised its 2026–2030 revenue forecast for Zhipu by between 7 and 16 percent following the launch, projecting an over 534 percent revenue surge for 2026 and expecting the AI firm to turn a profit by 2028.</p><h2><b>Why vendor lock-in now carries a geopolitical risk that no SLA can cover</b></h2><p>The <a href="https://venturebeat.com/technology/anthropic-is-bringing-back-claude-fable-5-globally-after-us-lifts-export-control-order-where-can-enterprises-access-it">Fable 5 episode</a> did more than embarrass Anthropic. It introduced a new risk category into enterprise AI procurement: sovereign access risk. When a government can disable a commercially deployed AI model overnight, the traditional evaluation criteria of developer experience, benchmark scores, and pricing become secondary to a more fundamental question: Will this tool still work tomorrow?</p><p>The event exposed the inadequacy of standard enterprise contract language. An investigation by <a href="https://www.fifthrow.com/blog/us-export-control-order-and-global-suspension-of-fable-5-mythos-5-operationalizing-compliance-as-a">FifthRow</a> found that almost all standard Data Processing Addenda, SaaS agreements, and procurement SLAs "relied on vague 'force majeure' or 'compliance with law' catch-alls, not on precise, actionable regulatory suspension or kill-switch clauses."</p><p>ZCode's <a href="https://aiidelist.com/ide/zcode">BYOK architecture </a>and <a href="https://z.ai/blog/glm-5.2">GLM-5.2</a>'s MIT-licensed open weights offer a partial answer. A development team can download the model, host it on its own infrastructure, and run ZCode against it without ever touching Z.ai's cloud — eliminating both American export-control risk and Chinese data-sovereignty concerns in a single move. The catch is that anyone using Z.ai's cloud API remains subject to Chinese law, a consideration that evaporates only with pure self-hosting.</p><p>Gartner analysts <a href="https://news.creeta.com/en/gartner-enterprise-ai-coding-agents-2026/">have warned</a> that governance, pricing, support, workflows, commercial maturity, and market durability matter as much as developer experience and model capabilities when evaluating coding agent vendors for enterprise-wide adoption. By that measure, ZCode faces a steep climb. It is not open source itself; Linux support remains in beta; and security reviewers have flagged the need for careful evaluation of its credential handling, particularly for remote development over SSH and messaging-platform-triggered tasks — an agent that can be summoned from WeChat involves access paths that should be mapped before trusting it with anything sensitive.</p><h2><b>Inside the $10 billion race where model labs are becoming full-stack IDE companies</b></h2><p><a href="https://zcode.z.ai/">ZCode</a> enters one of the most crowded and fastest-moving markets in enterprise software. Enterprise AI coding agents are capturing a growing share of enterprise software engineering spend, with the market estimated at roughly $9.8 billion to $11.0 billion annualized as of April 2026, according to <a href="https://enterprisedna.co/resources/news/gartner-enterprise-ai-coding-agents-10-billion-market-2026/">Gartner</a>. A defining shift this year, the analyst firm noted, is "the movement of frontier model providers into direct competition with application-layer vendors" — precisely the pattern ZCode embodies.</p><p>Gartner codified this evolution in May when it <a href="https://openai.com/index/gartner-2026-agentic-coding-leader/">renamed its annual Magic Quadrant</a> from "AI Code Assistants" to "Enterprise AI Coding Agents," defining the category as "autonomous or semiautonomous software engineering solutions that perceive context, translate human intent into multistep plans, and execute and verify those steps across code, tests and related engineering artifacts." The 2026 Magic Quadrant names Anthropic, Cursor, GitHub, and OpenAI as Leaders. Z.ai was not among the 12 vendors evaluated — an absence that underscores both the company's nascent enterprise sales presence outside China and the Western-centric lens through which the analyst community still views the market.</p><p>The competitive landscape is daunting. Cursor is the <a href="https://www.bloomberg.com/news/articles/2026-03-02/cursor-recurring-revenue-doubles-in-three-months-to-2-billion">$2 billion ARR IDE</a> that feels like VS Code with a supercharger. Claude Code reached <a href="https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation">approximately $2.5 billion</a> in annualized revenue by early 2026. Google relaunched <a href="https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-developer-highlights/">Antigravity 2.0</a> at I/O in May, and Cognition retired the Windsurf brand, relaunching the IDE as <a href="https://devin.ai/desktop/">Devin Desktop</a> with the Agent Command Center as the default surface.</p><p>Against these entrenched players, ZCode's pitch rests on three pillars: deep first-party integration with GLM-5.2 that no third-party editor can replicate, aggressive pricing that starts at a fraction of Western competitors, and MIT-licensed open weights that allow enterprises to self-host — eliminating the regulatory kill-switch risk that the Fable ban made viscerally real.</p><h2><b>Z.ai's real challenge is turning a $128 billion valuation into a global developer tools business</b></h2><p><a href="http://z.ai/">Z.ai</a> controls the model (<a href="https://z.ai/blog/glm-5.2">GLM-5.2</a>), the subscription layer (<a href="https://z.ai/subscribe">the GLM Coding Plan</a>), and the IDE (<a href="https://zcode.z.ai/">ZCode</a>) — a tightly coupled stack that optimizes for performance but concentrates switching costs. For the company, the business logic is clear. Its most reliable revenue stream has been on-premises deployments for Chinese government agencies, state-owned banks, and energy conglomerates. In full-year 2025, on-premises deployment revenue reached RMB 534 million, growing over 100 percent year-over-year and accounting for 73.7 percent of total revenue with a gross margin of 48.8 percent. ZCode and the GLM Coding Plan represent the company's bid to build a comparable revenue engine in cloud-based developer tools — globally, not just in China.</p><p>The early signals are encouraging for <a href="http://z.ai/">Z.ai</a>, if anecdotal. Community reception on X was enthusiastic, with one early user calling the tool "super stable" and others clamoring for more Coding Plan capacity. "Bro, can't snag your family's Coding Plan? When are you gonna stock up on more cards?" <a href="https://x.com/realchendahuang/status/2072361920976593163">one user wrote in Chinese</a>, suggesting demand is already outstripping supply.</p><p>But the hard questions loom large. Can a Chinese AI company build trust with Western enterprise buyers amid escalating technology tensions? Can ZCode's ecosystem mature fast enough to compete with Cursor's polished UX, Claude Code's deep agent primitives, and GitHub Copilot's unmatched distribution? And can Z.ai sustain a company valued at $128 billion while still losing money? </p><p>What is no longer in question is the competitive dynamic itself. Three weeks ago, a U.S. government directive proved that access to the world's best coding model can vanish overnight. Today, a Chinese lab is shipping a free IDE, an open-source model trained on zero American chips, and a subscription plan that costs less per month than a single lunch in Manhattan. The AI coding agent market did not just become global this summer. It became a market where the fallback option might be better than the thing it's falling back from — and that changes the calculus for every engineering leader choosing a toolchain in the second half of 2026.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[What do AI observability tools actually do?]]></title>
<description><![CDATA[As organizations rush to move AI into production, they’re finding that the tools they rely on to monitor traditional software don’t translate cleanly to AI systems. The reason is fundamental: AI doesn’t fail as software does. It doesn’t throw clean error codes or follow predictable execution path...]]></description>
<link>https://tsecurity.de/de/3640600/ai-nachrichten/what-do-ai-observability-tools-actually-do/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3640600/ai-nachrichten/what-do-ai-observability-tools-actually-do/</guid>
<pubDate>Thu, 02 Jul 2026 11:04:36 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>As organizations rush to move AI into production, they’re finding that the tools they rely on to monitor traditional software don’t translate cleanly to AI systems. The reason is fundamental: AI doesn’t fail as software does. It doesn’t throw clean error codes or follow predictable execution paths. It drifts, hallucinates, and degrades in ways that are often subtle, intermittent, and hard to reproduce.</p>



<p>The result is a growing gap between what teams think observability should provide and what current tools actually deliver. The uncomfortable truth? The AI observability tools we have today are built for yesterday’s problems.</p>



<p>To understand where the industry is headed, we need to look at where it is today and why that’s not enough.</p>



<h2 class="wp-block-heading">AI observability today: The era of evals</h2>



<p>Today’s AI observability landscape is dominated by one concept: evaluation.</p>



<p>Most tools focus on scoring model outputs after the fact. They rely on test datasets, human graders, or, increasingly, “LLM-as-a-judge” approaches to determine whether a system is behaving correctly. These evaluation pipelines are useful and can provide a baseline for model quality, helping teams benchmark improvements.</p>



<p>But they do share a critical limitation. They’re static, offline, and backward-looking.</p>



<p>Evaluations tell you how a model performed on a predefined set of inputs. But they don’t tell you what’s happening in production, where inputs are unpredictable and context can shift. You need to capture long-running interactions, multi-step workflows, and the behavior of systems composed of multiple models and tools as a part of your evals.</p>



<p>Even when teams use human-in-the-loop feedback, it can be tough to scale. High-quality feedback requires domain expertise, consistency, and time, each of which is in short supply in most engineering organizations. You also need deep knowledge of the models themselves and how they’re working in production to help identify and provide feedback around the source of the error. Was it a lack of context? A bad <a href="https://www.infoworld.com/article/2335814/what-is-retrieval-augmented-generation-more-accurate-and-reliable-llms.html" data-type="link" data-id="https://www.infoworld.com/article/2335814/what-is-retrieval-augmented-generation-more-accurate-and-reliable-llms.html">retrieval-augmented generation</a> (RAG) implementation? The model itself? Or bad feedback poisoning the results?</p>



<p>Some progress is being made. OpenTelemetry (OTel) and LLM tracing are emerging as early attempts to bring runtime visibility into AI systems. But these are still just first steps, and the core issue remains: you can’t understand AI systems by evaluating them after the fact. You need to observe them as they operate.</p>



<h2 class="wp-block-heading">The security turn: guardrails, PII, and prompt injection</h2>



<p>As AI systems move into production, observability becomes more about managing risk. The attack surface has expanded dramatically, with teams now dealing with:</p>



<ul class="wp-block-list">
<li>Prompt injection attacks</li>



<li>Jailbreak attempts</li>



<li>Leakage of sensitive data, including personally identifiable information (PII)</li>



<li>Unintended model behavior triggered by edge-case inputs</li>
</ul>



<p>In response, a new category of “guardrail” tools has emerged. These systems aim to monitor inputs and outputs in real time, flagging or blocking unsafe behavior. In theory, they provide a safety layer that sits between users and models. </p>



<p>In practice, however, the picture is more complicated.</p>



<p>Most guardrails today are reactive. They rely on predefined rules or classifiers that attempt to catch known patterns. But AI systems are inherently open-ended, and adversarial inputs evolve quickly. What works today may fail tomorrow.</p>



<p>There’s also a deeper issue: guardrails operate on the assumption that you already have sufficient visibility into the system. In reality, many teams lack the underlying telemetry needed to understand how and why a failure occurred in the first place.</p>



<p>This creates a gap between what guardrails promise (real-time protection) and what they can reliably deliver. Closing that gap requires something more foundational than filtering inputs and outputs. It requires rethinking observability itself.</p>



<h2 class="wp-block-heading">The coming shift: from models to agents</h2>



<p>The next wave of AI is clearly about autonomous agents. Instead of single inference calls, we’re seeing systems that orchestrate multiple models, interact with external tools and APIs, and execute multi-step workflows over extended periods of time.</p>



<p>These systems don’t just generate outputs; they make decisions. And that changes the observability problem entirely.</p>



<p>Just as <a href="https://www.infoworld.com/article/2257241/why-you-should-use-docker-and-oci-containers.html" data-type="link" data-id="https://www.infoworld.com/article/2257241/why-you-should-use-docker-and-oci-containers.html">containers</a> required orchestration platforms like <a href="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html" data-type="link" data-id="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html">Kubernetes</a> to become manageable at scale, AI agents will require their own observability and control layer. That layer must go beyond tracking inputs and outputs. It needs to capture:</p>



<ul class="wp-block-list">
<li>Decision paths</li>



<li>Tool usage</li>



<li>Resource consumption</li>



<li>Interactions across agents</li>



<li>Behavior over time, not just at a single point</li>
</ul>



<p>In many ways, this is similar to what we saw with the evolution of cloud-native observability. We moved from simple metrics to a combination of logs, metrics, and traces to understand distributed systems.</p>



<p>Now we need the equivalent for agentic systems.</p>



<p>As AI becomes embedded across the software development life cycle, from code generation to testing to operations, observability is evolving into a system of truth that feeds both humans and machines. AI agents can only build, debug, and improve systems if they have access to rich, high-fidelity production context. Observability is what provides that context.</p>



<h2 class="wp-block-heading">Why kernel-space observability will be essential</h2>



<p>There’s a fundamental trust problem at the heart of AI observability. If an AI agent is responsible for reporting its own behavior, how do you know that behavior is being reported accurately?</p>



<p>Traditional observability relies heavily on instrumentation within the application layer. But instrumentation can be incomplete, misconfigured, inadvertently bypassed, or simply incorrect.</p>



<p>This problem becomes more acute as AI systems begin generating their own code. Agents don’t think like human engineers when it comes to instrumentation, nor should they be expected to. But the result is a growing need for independent, out-of-band observability.</p>



<p>This is where kernel-level approaches, such as <a href="https://ebpf.io/" data-type="link" data-id="https://ebpf.io/">eBPF</a>, become critical. By operating at the kernel level, eBPF enables teams to:</p>



<ul class="wp-block-list">
<li>Capture system behavior without modifying application code</li>



<li>Eliminate blind spots caused by missing instrumentation</li>



<li>Ensure consistent visibility across all workloads, both human-driven and AI-generated</li>
</ul>



<p>More importantly, eBPF provides a trusted source of truth. In high-stakes environments where compliance, security, and reliability are non-negotiable, this independence is essential. You need telemetry that’s not influenced by the systems it observes.</p>



<h2 class="wp-block-heading">Three needs for AI observability </h2>



<p>If current tools fall short, what comes next? The answer is a shift in how we think about observability.</p>



<p>First, we need behavioral anomaly detection for AI systems. Traditional observability focuses on latency, errors, and resource utilization. But AI systems require a different lens to detect when behavior deviates from expectations, even when no explicit “error” occurs.</p>



<p>Second, we need tamper-proof audit trails. As AI systems take on more responsibility, you have to be able to reconstruct decisions. Teams need to understand what happened and, more importantly, why. And they need to trust that the data hasn’t been altered.</p>



<p>Third, observability must become dynamic and adaptive. Static dashboards and predefined metrics won’t cut it. AI systems operate in constantly changing environments, and observability must be able to:</p>



<ul class="wp-block-list">
<li>Adjust data collection in real time</li>



<li>Increase granularity during incidents</li>



<li>Focus on what matters in the moment</li>
</ul>



<p>Finally, observability must integrate directly into AI workflows. It’s no longer enough to surface insights to human operators. The same telemetry must be consumable by AI agents feeding back into development, debugging, and optimization loops.</p>



<h2 class="wp-block-heading">Observability as a part of infrastructure, not an afterthought</h2>



<p>We are still early in the evolution of AI observability. Most of today’s tools are extensions of existing paradigms adapted for AI, but not fundamentally redesigned for it. Predictably, they solve parts of the problem, but not the whole.</p>



<p>The next generation of these systems will look very different. They’ll treat observability as a core layer that enables AI systems to operate safely, efficiently, and autonomously. The teams that succeed will be those that recognize this shift early.</p>



<p>Ultimately, in a world of non-deterministic systems, long-running workflows, and autonomous agents, one thing becomes clear: AI reliability strongly correlates with your observability layer.</p>



<p><em>—</em></p>



<p><a href="https://www.infoworld.com/blogs/new-tech-forum"><strong><em>New Tech Forum</em></strong></a><em><strong> provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all </strong></em><em><strong>inquiries to </strong></em><a href="mailto:doug_dineley@foundryco.com"><strong><em>doug_dineley@foundryco.com</em></strong></a><em><strong>.</strong></em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Using Lift to Turn Research PDFs into Structured JSON with Controlled, Schema-Guided Field-Level Evaluation]]></title>
<description><![CDATA[In this tutorial, we build a full PDF-to-structured-data workflow around Lift, built for controlled evaluation rather than a one-off demo. We prepare a Colab GPU environment, load Lift in 4-bit NF4, and generate synthetic research reports with deliberate distractors. We then run schema-guided ext...]]></description>
<link>https://tsecurity.de/de/3639716/ai-nachrichten/using-lift-to-turn-research-pdfs-into-structured-json-with-controlled-schema-guided-field-level-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3639716/ai-nachrichten/using-lift-to-turn-research-pdfs-into-structured-json-with-controlled-schema-guided-field-level-evaluation/</guid>
<pubDate>Wed, 01 Jul 2026 23:18:46 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>In this tutorial, we build a full PDF-to-structured-data workflow around Lift, built for controlled evaluation rather than a one-off demo. We prepare a Colab GPU environment, load Lift in 4-bit NF4, and generate synthetic research reports with deliberate distractors. We then run schema-guided extraction, score every field against ground truth, and assemble the results into a queryable knowledge base. The result is a repeatable extraction benchmark, not just raw model outputs.</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/01/using-lift-to-turn-research-pdfs-into-structured-json-with-controlled-schema-guided-field-level-evaluation/">Using Lift to Turn Research PDFs into Structured JSON with Controlled, Schema-Guided Field-Level Evaluation</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Hermes Agent v0.18.0 (2026.7.1) — The Judgment Release]]></title>
<description><![CDATA[Hermes Agent v0.18.0 (v2026.7.1)
Release Date: July 1, 2026
Since v0.17.0: ~1,720 commits · 998 merged PRs · 2,215 files changed · ~251,000 insertions · ~41,000 deletions · 949 issues closed · 370+ community contributors

The Judgment Release. Over the last week and a half the team put nearly all...]]></description>
<link>https://tsecurity.de/de/3639600/downloads/hermes-agent-v0180-202671-the-judgment-release/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3639600/downloads/hermes-agent-v0180-202671-the-judgment-release/</guid>
<pubDate>Wed, 01 Jul 2026 22:16:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Hermes Agent v0.18.0 (v2026.7.1)</h1>
<p><strong>Release Date:</strong> July 1, 2026<br>
<strong>Since v0.17.0:</strong> ~1,720 commits · 998 merged PRs · 2,215 files changed · ~251,000 insertions · ~41,000 deletions · <strong>949 issues closed</strong> · <strong>370+ community contributors</strong></p>
<blockquote>
<p><strong>The Judgment Release.</strong> Over the last week and a half the team put nearly all of its effort into one goal: resolve <strong>every P0 and P1 issue and PR in the entire Hermes Agent repo</strong> — and as of this release, <strong>100% of them are closed.</strong> Zero open P0s. Zero open P1s. That's <strong>~700 highest-priority items</strong> cleared as part of <strong>~1,950 total issues and PRs closed</strong> this window. We intend to keep P0/P1 at zero from here on.</p>
<p>On top of that clean-sweep, v0.18.0 is about how <em>well</em> Hermes thinks and how it <em>knows when its work is actually done</em>. Mixture-of-Agents became a first-class citizen — named ensembles of models you can pick like any other model, with every reference model's reasoning shown to you and the aggregator's answer streamed live. The agent learned to verify its own work against evidence instead of vibes, <code>/goal</code> gained completion contracts, and <code>/learn</code> + <code>/journey</code> turned self-improvement into something you can see and steer. Underneath, the gateway became genuinely deployable-at-scale (scale-to-zero, drain coordination), the desktop grew first-class coding projects and a playable memory graph, and subagents can now fan out in the background.</p>
</blockquote>
<h2>🎯 The P0/P1 Clean Sweep — 100% resolved</h2>
<p>This is the release headline. For a week and a half the team hammered the priority backlog day and night, and every single P0 and P1 across the whole repo is now closed:</p>
<table>
<thead>
<tr>
<th>Priority</th>
<th>Issues closed</th>
<th>PRs merged</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>P0</strong> (critical)</td>
<td>3</td>
<td>8</td>
</tr>
<tr>
<td><strong>P1</strong> (high)</td>
<td>493</td>
<td>188</td>
</tr>
<tr>
<td><strong>Total</strong></td>
<td><strong>496</strong></td>
<td><strong>196</strong></td>
</tr>
</tbody>
</table>
<p>That's <strong>~692 highest-priority items resolved</strong> in twelve days — and at the moment the sweep completed, the open P0/P1 count hit <strong>0 across the entire repo.</strong> The final cluster to fall was the interrupt-protected-compression sibling-fork bug (issue <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4785584067" data-permission-text="Title is private" data-url="https://github.com/NousResearch/hermes-agent/issues/56391" data-hovercard-type="issue" data-hovercard-url="/NousResearch/hermes-agent/issues/56391/hovercard" href="https://github.com/NousResearch/hermes-agent/issues/56391">#56391</a>) and its fix (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4785996667" data-permission-text="Title is private" data-url="https://github.com/NousResearch/hermes-agent/issues/56416" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56416/hovercard" href="https://github.com/NousResearch/hermes-agent/pull/56416">#56416</a>), closed on an all-nighter right before this release cut.</p>
<p>Special shoutout to <strong><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a></strong>, who burned through the priority backlog day and night alongside the core team — the cron reliability wave, the compression-fork fix, the credential-exfil hardening, and a huge share of the P1 closures are his.</p>
<p>We're keeping P0/P1 at <strong>0</strong> from here forward. 🫡</p>
<h2>✨ Highlights</h2>
<ul>
<li>
<p><strong>Mixture-of-Agents is now a first-class model you can pick</strong> — MoA used to be a mode you toggled; now every named MoA preset shows up as a selectable model under a <code>moa</code> provider, right alongside Claude, GPT, and Grok in every model picker (CLI, TUI, desktop, gateway). Pick "my-council" the same way you'd pick any model, and Hermes routes your prompt through that ensemble automatically. An ensemble of frontier models deliberating on your hardest questions is now one selection away, on every surface. (<a href="https://github.com/NousResearch/hermes-agent/pull/46081" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/46081/hovercard">#46081</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53548" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53548/hovercard">#53548</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53561" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53561/hovercard">#53561</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</p>
</li>
<li>
<p><strong>See every model's reasoning, then watch the answer stream in</strong> — When a MoA ensemble runs, each reference model's full output now renders as its own labelled block — you can read what GPT-5 thought, what Claude thought, and what Grok thought, before the aggregator synthesizes them into one answer. And that final answer now streams to you live instead of appearing all at once after a long silence. This works in the CLI, the TUI, and the desktop app. You get to watch the committee deliberate, not just read the verdict. (<a href="https://github.com/NousResearch/hermes-agent/pull/53793" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53793/hovercard">#53793</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53855" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53855/hovercard">#53855</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55625" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55625/hovercard">#55625</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56101" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56101/hovercard">#56101</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</p>
</li>
<li>
<p><strong>The agent verifies its own work — "done" means proven, not claimed</strong> — Hermes now records verification evidence for coding work and can decide it's finished by actually running your project's checks, not by asserting success. <code>/goal</code> gained <strong>completion contracts</strong>: you state what "done" looks like, and the standing-goal loop judges completion against that evidence instead of stopping when the model feels like it. There's a <code>pre_verify</code> hook for wiring in custom checks and a one-time migration that tunes the defaults sensibly. The difference between "I think I fixed it" and "the tests pass, here's proof." (<a href="https://github.com/NousResearch/hermes-agent/pull/50501" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50501/hovercard">#50501</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52285" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52285/hovercard">#52285</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55413" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55413/hovercard">#55413</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53552" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53552/hovercard">#53552</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</p>
</li>
<li>
<p><strong><code>/learn</code> — turn anything into a reusable skill by describing it</strong> — Run <code>/learn &lt;anything&gt;</code> and Hermes distills a reusable skill out of whatever you point it at — a directory, a URL, or just the workflow you walked it through five minutes ago. It writes the skill to the standards in your CONTRIBUTING.md automatically. The next time you need that workflow, it's already there. Teaching Hermes a new trick is now a single command, not a manual skill-authoring session. (<a href="https://github.com/NousResearch/hermes-agent/pull/51506" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51506/hovercard">#51506</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52372" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52372/hovercard">#52372</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</p>
</li>
<li>
<p><strong><code>/journey</code> — a playable timeline of everything Hermes has learned about you</strong> — The CLI and TUI gained <code>/journey</code>, a learning timeline that shows the memories and skills Hermes has accumulated over time — and you can edit or delete any of them right from the view. Pair it with the desktop's new <strong>memory graph</strong> (a top-down, playable radial timeline of memories and skills) and for the first time you can actually <em>see</em> what your agent knows, watch it grow, and prune what's wrong. Your agent's memory stops being a black box. (<a href="https://github.com/NousResearch/hermes-agent/pull/55555" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55555/hovercard">#55555</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55859" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55859/hovercard">#55859</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55226" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55226/hovercard">#55226</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</p>
</li>
<li>
<p><strong>Delegate a pile of work and keep going — background fan-out</strong> — <code>delegate_task</code> can now fan out multiple subagents that all run in the <strong>background</strong>: your chat is never blocked, and when every subagent finishes, their results come back as a single consolidated turn. Kick off "research these five competitors in parallel" or "audit these three modules," then carry on with something else while a small fleet works. When it's all done, you get one clean summary instead of babysitting each one. (<a href="https://github.com/NousResearch/hermes-agent/pull/49734" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49734/hovercard">#49734</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</p>
</li>
<li>
<p><strong>First-class coding Projects in the desktop app</strong> — The desktop app gained real, per-profile <strong>Projects</strong> — a sidebar of your codebases, a coding rail, a review pane, git worktree management, and agent-facing project tools, all backed by a proper <code>project → repo → lane</code> model. Instead of scattered chat sessions, your coding work is organized into projects the agent understands and can act on. It's the desktop turning into an actual coding cockpit. (<a href="https://github.com/NousResearch/hermes-agent/pull/49037" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49037/hovercard">#49037</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54385" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54385/hovercard">#54385</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54517" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54517/hovercard">#54517</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</p>
</li>
<li>
<p><strong>Run Hermes at scale — scale-to-zero and drain coordination</strong> — The gateway can now go <strong>dormant when idle</strong> and quiesce cleanly before a restart, migration, or auto-update — without dropping in-flight conversations. A hosted or relay-only Hermes can scale to zero when nobody's talking to it and wake back up on demand, and disruptive lifecycle actions coordinate an external drain so nobody gets cut off mid-turn. Running Hermes for a team or as a hosted service just got a lot more production-grade. (<a href="https://github.com/NousResearch/hermes-agent/pull/52243" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52243/hovercard">#52243</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52937" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52937/hovercard">#52937</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54824" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54824/hovercard">#54824</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>)</p>
</li>
<li>
<p><strong>Cheaper self-improvement — smarter background review</strong> — The post-turn self-improvement fork (the one that decides whether to save a memory or skill) now routes to an auxiliary model, digests context instead of replaying the whole conversation, and adapts its cadence — so the "learn from what just happened" loop that runs after your turns costs a fraction of what it used to. You keep the self-improvement, you stop paying full main-model price for it. (<a href="https://github.com/NousResearch/hermes-agent/pull/49252" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49252/hovercard">#49252</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</p>
</li>
<li>
<p><strong>Compose your next prompt in your editor — <code>/prompt</code></strong> — <code>/prompt</code> opens your <code>$EDITOR</code> so you can hand-write a long, multi-line prompt in real markdown instead of fighting a one-line input box. Draft a detailed spec, a structured question, or a big paste, save, and it's queued as your next message. Small thing, huge quality-of-life win for anyone who writes Hermes more than a sentence at a time. (<a href="https://github.com/NousResearch/hermes-agent/pull/50509" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50509/hovercard">#50509</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</p>
</li>
<li>
<p><strong>Google Vertex AI — Gemini through your GCP service account, no static key</strong> — Vertex AI is now a first-class provider for Gemini models over Vertex's OpenAI-compatible endpoint. The reason a plain custom-provider setup always died mid-session is that Vertex has no static API key — every request needs a short-lived OAuth2 access token (~1h TTL) minted from a service-account JSON or Application Default Credentials. Hermes now mints and auto-refreshes those tokens for you, so if your org runs Gemini through Google Cloud, you point Hermes at your service account and it just works — no token-pasting, no mid-session expiry. (<a href="https://github.com/NousResearch/hermes-agent/pull/56363" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56363/hovercard">#56363</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/slawt/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/slawt">@slawt</a>)</p>
</li>
<li>
<p><strong>Security round</strong> — This window hardened several surfaces: MCP-config persistence attack surface locked down, cron <code>base_url</code> overrides that could exfiltrate provider credentials blocked, a non-reusable sentinel for prefix secrets in file reads, Slack app-level (<code>xapp-</code>) token redaction, a browser cloud-metadata floor enforced on every backend, and an <code>aiohttp</code> CVE floor across the lazy messaging paths. Fewer ways for a prompt-injected or misconfigured session to leak a credential. (<a href="https://github.com/NousResearch/hermes-agent/pull/50476" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50476/hovercard">#50476</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56196" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56196/hovercard">#56196</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54166" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54166/hovercard">#54166</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56227" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56227/hovercard">#56227</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52349" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52349/hovercard">#52349</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56237" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56237/hovercard">#56237</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/claudlos/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/claudlos">@claudlos</a>)</p>
</li>
</ul>
<hr>
<h2>🧠 Mixture-of-Agents (MoA)</h2>
<p>MoA graduated from a mode to a first-class part of the model system this window.</p>
<ul>
<li><strong>Presets as selectable virtual models</strong> — each named MoA preset appears as a model under provider <code>moa</code>; pick it in any model picker and Hermes routes through the ensemble (<a href="https://github.com/NousResearch/hermes-agent/pull/46081" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/46081/hovercard">#46081</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53561" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53561/hovercard">#53561</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53775" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53775/hovercard">#53775</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong><code>/moa</code> is now one-shot sugar</strong> — runs a single prompt through the default preset and restores your model afterward; persistent switching goes through the model picker (<a href="https://github.com/NousResearch/hermes-agent/pull/53548" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53548/hovercard">#53548</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong>Reference-model output shown as labelled blocks</strong> in CLI, TUI, and desktop — read each model's reasoning before the aggregator's synthesis (<a href="https://github.com/NousResearch/hermes-agent/pull/53793" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53793/hovercard">#53793</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53855" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53855/hovercard">#53855</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong>Aggregator response streams live</strong> instead of appearing whole after a silence (<a href="https://github.com/NousResearch/hermes-agent/pull/55625" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55625/hovercard">#55625</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong>References see full tool state and fire on every user/tool response</strong>; advisory references end on a user turn and get a reference-role system prompt (<a href="https://github.com/NousResearch/hermes-agent/pull/54016" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54016/hovercard">#54016</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54007" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54007/hovercard">#54007</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong>Opt-in full-turn trace persistence to JSONL</strong> (<code>moa.save_traces</code>) for debugging and eval (<a href="https://github.com/NousResearch/hermes-agent/pull/56101" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56101/hovercard">#56101</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Reliability: reference + aggregator models called through their provider's real route; context window resolved from the aggregator (not the 256K default); auxiliary tasks resolve to the aggregator; virtual provider blocked as a reference/aggregator slot; tolerant of hand-edited preset config (<a href="https://github.com/NousResearch/hermes-agent/pull/53580" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53580/hovercard">#53580</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53780" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53780/hovercard">#53780</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53827" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53827/hovercard">#53827</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53281" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53281/hovercard">#53281</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53275" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53275/hovercard">#53275</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53556" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53556/hovercard">#53556</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>MoA slot provider-identity unified on the single <code>call_llm</code> chokepoint; HermesBench results documented (<a href="https://github.com/NousResearch/hermes-agent/pull/55991" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55991/hovercard">#55991</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53206" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53206/hovercard">#53206</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>✅ Verification &amp; Goals — the agent proves its work</h2>
<ul>
<li><strong>Completion contracts for <code>/goal</code></strong> — state what "done" looks like; the standing-goal loop judges against evidence, not the model's say-so (<a href="https://github.com/NousResearch/hermes-agent/pull/50501" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50501/hovercard">#50501</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong><code>/goal wait &lt;pid&gt;</code></strong> — park the standing-goal loop on a background process instead of re-poking the agent (<a href="https://github.com/NousResearch/hermes-agent/pull/50503" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50503/hovercard">#50503</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong>Coding verification evidence ledger</strong> — profile-scoped record of canonical project checks detected by <code>agent.coding_context</code>; gateway exposes verification status (<a href="https://github.com/NousResearch/hermes-agent/pull/52285" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52285/hovercard">#52285</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52286" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52286/hovercard">#52286</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><strong><code>pre_verify</code> hook + coding guidance config</strong>; verification stop loop + ad-hoc verification scripts (<a href="https://github.com/NousResearch/hermes-agent/pull/55413" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55413/hovercard">#55413</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52296" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52296/hovercard">#52296</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52297" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52297/hovercard">#52297</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><strong>verify-on-stop defaults OFF</strong> with a one-time v32 migration; skips doc-only edits; surface-aware "auto" default restored; gated off for messaging surfaces (<a href="https://github.com/NousResearch/hermes-agent/pull/53552" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53552/hovercard">#53552</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54740" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54740/hovercard">#54740</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55449" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55449/hovercard">#55449</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52412" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52412/hovercard">#52412</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/GodsBoy/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/GodsBoy">@GodsBoy</a>)</li>
</ul>
<h2>🎓 Self-Improvement (Learn / Journey)</h2>
<ul>
<li><strong><code>/learn &lt;anything&gt;</code></strong> — distill a reusable skill from a directory, URL, or a workflow you just walked through; honors CONTRIBUTING.md skill standards and mixed requirements (<a href="https://github.com/NousResearch/hermes-agent/pull/51506" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51506/hovercard">#51506</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52372" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52372/hovercard">#52372</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55956" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55956/hovercard">#55956</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong><code>/journey</code></strong> — CLI + TUI learning timeline of accumulated memories and skills, with in-place edit/delete (<a href="https://github.com/NousResearch/hermes-agent/pull/55555" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55555/hovercard">#55555</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55859" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55859/hovercard">#55859</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><strong>Cheaper background review</strong> — aux-model routing + context digest + adaptive cadence for the post-turn self-improvement fork (<a href="https://github.com/NousResearch/hermes-agent/pull/49252" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49252/hovercard">#49252</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li><strong><code>memory</code> graph</strong> in the desktop — playable radial timeline of memories + skills over time (<a href="https://github.com/NousResearch/hermes-agent/pull/55226" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55226/hovercard">#55226</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
</ul>
<h2>🖥️ Hermes Desktop App</h2>
<h3>Coding cockpit</h3>
<ul>
<li><strong>First-class Projects</strong> — per-profile sidebar, coding rail, review pane, agent project tools (<code>project → repo → lane</code>); remote-gateway-aware folder picker + git cockpit (status, review, worktrees) (<a href="https://github.com/NousResearch/hermes-agent/pull/49037" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49037/hovercard">#49037</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54385" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54385/hovercard">#54385</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><strong>Multi-terminal panel</strong> with read-only agent terminals; persist &amp; restore terminal tabs + scrollback across relaunch (<a href="https://github.com/NousResearch/hermes-agent/pull/54517" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54517/hovercard">#54517</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54585" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54585/hovercard">#54585</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><strong>PR-style file diffs in chat</strong>; in-app spot editor for the file preview pane; inline rich embeds, diagrams &amp; alerts in assistant markdown (<a href="https://github.com/NousResearch/hermes-agent/pull/50731" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50731/hovercard">#50731</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52772" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52772/hovercard">#52772</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52935" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52935/hovercard">#52935</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
</ul>
<h3>UX &amp; surfaces</h3>
<ul>
<li>Conversation timeline rail for long threads; context-usage breakdown popover; read-only spectator transcript for subagent watch windows; pop the composer into a draggable floating window (<a href="https://github.com/NousResearch/hermes-agent/pull/51094" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51094/hovercard">#51094</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54907" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54907/hovercard">#54907</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55033" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55033/hovercard">#55033</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/49488" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49488/hovercard">#49488</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/austinpickett/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/austinpickett">@austinpickett</a>)</li>
<li>Read replies aloud (auto-TTS) composer toggle; remember window size/position/maximized across launches; redesigned clarify prompt; shared overlay Panel primitive for cron/profiles/agents (<a href="https://github.com/NousResearch/hermes-agent/pull/55154" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55154/hovercard">#55154</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52086" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52086/hovercard">#52086</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52993" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52993/hovercard">#52993</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54558" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54558/hovercard">#54558</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li>Backup import/create/download from the web UI; add context-usage popover; flag already-installed themes in install pickers; config-driven Electron launch flags + GPU policy (<a href="https://github.com/NousResearch/hermes-agent/pull/54611" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54611/hovercard">#54611</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55410" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55410/hovercard">#55410</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53991" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53991/hovercard">#53991</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><strong>Pets</strong> — roaming pet (opt-in), calmer/realistic roam, Alt+wheel scaling never cropped, frame-perfect hatch flow + CPU-safe chroma, pop-out overlay + notifications (<a href="https://github.com/NousResearch/hermes-agent/pull/55114" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55114/hovercard">#55114</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55400" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55400/hovercard">#55400</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52877" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52877/hovercard">#52877</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/47959" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/47959/hovercard">#47959</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52303" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52303/hovercard">#52303</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
</ul>
<h3>Refactor wave (composer / god-file de-entangle)</h3>
<ul>
<li>Decomposed the composer into isolated engine hooks; extracted branch/esc/url/placeholder/popout engines; split <code>thread.tsx</code>, <code>sidebar/index.tsx</code>, onboarding overlay, and <code>use-prompt-actions</code> god files into focused modules; shared WebSocket layer decoupling desktop from dashboard (<code>hermes serve</code>) (<a href="https://github.com/NousResearch/hermes-agent/pull/55500" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55500/hovercard">#55500</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55842" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55842/hovercard">#55842</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55451" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55451/hovercard">#55451</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55453" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55453/hovercard">#55453</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55807" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55807/hovercard">#55807</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55504" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55504/hovercard">#55504</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54568" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54568/hovercard">#54568</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li>perf: bound tool-result rendering so big <code>/learn</code> runs don't freeze; fast session switching under load (<a href="https://github.com/NousResearch/hermes-agent/pull/52273" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52273/hovercard">#52273</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52620" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52620/hovercard">#52620</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
</ul>
<h2>📊 Web Dashboard</h2>
<ul>
<li>Auto-initiate portal SSO redirect on unauthenticated load; interactive auth setup on no-provider non-loopback bind; confidential-client (<code>client_secret</code>) support in self-hosted OIDC (<a href="https://github.com/NousResearch/hermes-agent/pull/54846" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54846/hovercard">#54846</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50551" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50551/hovercard">#50551</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55344" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55344/hovercard">#55344</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>)</li>
<li>Catalogue all memory-provider API keys in <code>OPTIONAL_ENV_VARS</code>; list &amp; add arbitrary custom <code>.env</code> keys on the Keys page; expose cron job execution fields; backup import/create/download (<a href="https://github.com/NousResearch/hermes-agent/pull/54546" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54546/hovercard">#54546</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54552" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54552/hovercard">#54552</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53551" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53551/hovercard">#53551</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54611" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54611/hovercard">#54611</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Offload PTY spawn/close off the event loop; exclude non-interactive providers from interactive login surfaces (<a href="https://github.com/NousResearch/hermes-agent/pull/53227" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53227/hovercard">#53227</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53239" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53239/hovercard">#53239</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/IAvecilla/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/IAvecilla">@IAvecilla</a>)</li>
</ul>
<h2>🏗️ Core Agent &amp; Architecture</h2>
<h3>Delegation &amp; subagents</h3>
<ul>
<li><strong>Background fan-out</strong> — parallel subagents run in the background, one consolidated return when all finish; calm "will resume" affordance for background <code>delegate_task</code> (<a href="https://github.com/NousResearch/hermes-agent/pull/49734" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49734/hovercard">#49734</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52756" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52756/hovercard">#52756</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li>Track background subagents in the CLI + TUI status bar (<a href="https://github.com/NousResearch/hermes-agent/pull/51441" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51441/hovercard">#51441</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51485" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51485/hovercard">#51485</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h3>Agent loop, tools &amp; coding context</h3>
<ul>
<li>One-shot LLM helper + <code>llm.oneshot</code> gateway RPC; expose coding-context project facts (<code>project.facts</code> RPC) (<a href="https://github.com/NousResearch/hermes-agent/pull/51261" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51261/hovercard">#51261</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51259" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51259/hovercard">#51259</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>)</li>
<li><code>web_extract</code> truncate-and-store instead of LLM summarization; concurrent @-reference expansion (<a href="https://github.com/NousResearch/hermes-agent/pull/54843" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54843/hovercard">#54843</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55207" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55207/hovercard">#55207</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
<li>Friendly human-phrased tool labels for built-in tools; <code>/reasoning full</code> (uncapped thinking); <code>/timestamps</code> + timestamps in <code>/history</code>; <code>/prompt</code> composes in <code>$EDITOR</code> (<a href="https://github.com/NousResearch/hermes-agent/pull/55166" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55166/hovercard">#55166</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50499" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50499/hovercard">#50499</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50506" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50506/hovercard">#50506</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50509" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50509/hovercard">#50509</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Per-reasoning-model stale-timeout floor in stream + non-stream detectors; escalate SIGTERM→SIGKILL on host-pid termination after grace (<a href="https://github.com/NousResearch/hermes-agent/pull/52845" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52845/hovercard">#52845</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50489" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50489/hovercard">#50489</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Multiple <code>HERMES_WRITE_SAFE_ROOT</code> dirs; opt-in HTTP/WS body capture to an isolated, share-excluded <code>gui_bodies.log</code> (<a href="https://github.com/NousResearch/hermes-agent/pull/53292" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53292/hovercard">#53292</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/49044" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49044/hovercard">#49044</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
</ul>
<h3>Compression &amp; sessions</h3>
<ul>
<li>In-place compaction option (single session id); flip <code>in_place</code> default to True with a guard fix (<a href="https://github.com/NousResearch/hermes-agent/pull/49739" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49739/hovercard">#49739</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52658" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52658/hovercard">#52658</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Backup includes <code>projects.db</code> and kanban boards in the pre-update snapshot (<a href="https://github.com/NousResearch/hermes-agent/pull/52990" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52990/hovercard">#52990</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
</ul>
<h3>Providers &amp; models</h3>
<ul>
<li><strong>Google Vertex AI</strong> first-class provider for Gemini over the OpenAI-compatible endpoint — auto-mints and refreshes short-lived OAuth2 tokens from a service-account JSON / ADC (no static key); salvages &amp; modernizes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4248493655" data-permission-text="Title is private" data-url="https://github.com/NousResearch/hermes-agent/issues/8427" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/8427/hovercard" href="https://github.com/NousResearch/hermes-agent/pull/8427">#8427</a> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/slawt/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/slawt">@slawt</a> (<a href="https://github.com/NousResearch/hermes-agent/pull/56363" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56363/hovercard">#56363</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/slawt/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/slawt">@slawt</a>)</li>
<li>Krea via managed Nous Subscription gateway; Z.AI endpoint picker (Global/China/Coding Plan); Ollama-cloud reasoning_effort wiring; remove google-gemini-cli + google-antigravity OAuth providers (<a href="https://github.com/NousResearch/hermes-agent/pull/52647" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52647/hovercard">#52647</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52364" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52364/hovercard">#52364</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51494" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51494/hovercard">#51494</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50492" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50492/hovercard">#50492</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
<li>Honor <code>NOUS_INFERENCE_BASE_URL</code> env override for Nous OAuth; keep Nous auth fresh for idle dashboard/gateway agents (<a href="https://github.com/NousResearch/hermes-agent/pull/52270" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52270/hovercard">#52270</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50567" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50567/hovercard">#50567</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>🌐 Gateway, Fleet &amp; Relay</h2>
<h3>Scale-to-zero &amp; drain</h3>
<ul>
<li><strong>Scale-to-zero idle detection + dormant-quiesce (Phase 0)</strong>; hardened dormancy guards; fixed arm-gate counting disabled placeholder platforms (<a href="https://github.com/NousResearch/hermes-agent/pull/52243" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52243/hovercard">#52243</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52359" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52359/hovercard">#52359</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52831" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52831/hovercard">#52831</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>)</li>
<li><strong>External drain coordination (safe-shutdown Phase 2)</strong>; suppress home-channel shutdown broadcast on flagged drains; persist in-flight transcript on restart/shutdown drain timeout; busy/idle readout for safe lifecycle actions (<a href="https://github.com/NousResearch/hermes-agent/pull/52937" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52937/hovercard">#52937</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54824" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54824/hovercard">#54824</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50312" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50312/hovercard">#50312</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50131" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50131/hovercard">#50131</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
<li>Default <code>restart_drain_timeout</code> to 0 to kill a systemd crash loop; self-heal a gateway stranded in draining/degraded (<a href="https://github.com/NousResearch/hermes-agent/pull/54066" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54066/hovercard">#54066</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55397" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55397/hovercard">#55397</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h3>Relay (Phase 5 / 6)</h3>
<ul>
<li>Wake primitive (gateway side); going-idle / buffered-flip primitive; <code>passthrough_forward</code> over WS; multi-platform-per-agent identity + per-frame egress; forward stable instance id at self-provision; declare relevance policy to the connector (<a href="https://github.com/NousResearch/hermes-agent/pull/51595" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51595/hovercard">#51595</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51572" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51572/hovercard">#51572</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50702" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50702/hovercard">#50702</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52830" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52830/hovercard">#52830</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50772" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50772/hovercard">#50772</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51248" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51248/hovercard">#51248</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>)</li>
<li>Authorize relay-delivered events by delivery, not <code>source.platform</code>; adopt <code>scope_id</code> wire key; purge platform-specific scope terminology (<a href="https://github.com/NousResearch/hermes-agent/pull/52306" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52306/hovercard">#52306</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55289" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55289/hovercard">#55289</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56016" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56016/hovercard">#56016</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>)</li>
</ul>
<h3>Gateway core &amp; rendering</h3>
<ul>
<li>Typed send-error classification (<code>SendResult.error_kind</code>); per-platform <code>typing_indicator</code> toggle; per-category context breakdown in <code>/usage</code> (<a href="https://github.com/NousResearch/hermes-agent/pull/50342" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50342/hovercard">#50342</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55394" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55394/hovercard">#55394</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/55204" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55204/hovercard">#55204</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>API server: configurable concurrent-run cap to prevent DoS; scope run approvals by run id (<a href="https://github.com/NousResearch/hermes-agent/pull/50007" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50007/hovercard">#50007</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56129" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56129/hovercard">#56129</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>📱 Messaging Platforms</h2>
<ul>
<li><strong>Cron continuations</strong> — continuable cron jobs (thread-preferred continuation with DM-mirror fallback); flat in-channel continuable cron delivery for Slack; warn when gateway not running on cron create/list (<a href="https://github.com/NousResearch/hermes-agent/pull/52250" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52250/hovercard">#52250</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56254" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56254/hovercard">#56254</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51696" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51696/hovercard">#51696</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Telegram: configurable command menu + raised default cap so skills stay visible; gate rich draft previews separately; drain general send pool on pool timeout before retry (<a href="https://github.com/NousResearch/hermes-agent/pull/51716" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51716/hovercard">#51716</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52088" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52088/hovercard">#52088</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54121" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54121/hovercard">#54121</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/helix4u/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/helix4u">@helix4u</a>)</li>
<li>Slack: opt-in Block Kit rendering for agent messages; <code>--no-assistant</code> flag for manifest generation (<a href="https://github.com/NousResearch/hermes-agent/pull/56102" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56102/hovercard">#56102</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51487" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51487/hovercard">#51487</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Discord: render reasoning as <code>-#</code> subtext via <code>display.reasoning_style</code> (<a href="https://github.com/NousResearch/hermes-agent/pull/51168" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51168/hovercard">#51168</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Native WhatsApp media delivery via the Baileys bridge; Teams native <code>send_video</code>/<code>send_voice</code>/<code>send_document</code>; photon sidecar upgraded to spectrum-ts v8 with tapback correlation; Raft gateway setup wizard (<a href="https://github.com/NousResearch/hermes-agent/pull/53598" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53598/hovercard">#53598</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/49308" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49308/hovercard">#49308</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53451" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53451/hovercard">#53451</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56230" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56230/hovercard">#56230</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Signal: AAC voice-note remux + shared markdown formatting (<a href="https://github.com/NousResearch/hermes-agent/pull/49530" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49530/hovercard">#49530</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
<li>Migrate slack/dingtalk/whatsapp/matrix/feishu/telegram/wecom/email/sms adapters to bundled (<a href="https://github.com/NousResearch/hermes-agent/pull/49408" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/49408/hovercard">#49408</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>🔧 Tool System, Skills &amp; MCP</h2>
<ul>
<li>Blank Slate setup mode — minimal agent, opt in to everything (<a href="https://github.com/NousResearch/hermes-agent/pull/36733" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/36733/hovercard">#36733</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>MCP: config persistence attack surface hardened; block base_url exfil; keepalive for short-TTL sessions (see Security) — plus catalog &amp; UX carried from v0.17.0</li>
<li>Skills: <code>/learn</code> distillation (see Self-Improvement); <code>cloudflare-temporary-deploy</code> optional skill; creative-ideation v2.1.0 method library (<a href="https://github.com/NousResearch/hermes-agent/pull/50849" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50849/hovercard">#50849</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/42402" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/42402/hovercard">#42402</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/SHL0MS/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/SHL0MS">@SHL0MS</a>)</li>
<li>Kanban: task lifecycle plugin hooks (claimed/completed/blocked); typed block reasons + unblock-loop breaker; handoff freshness stamping (<a href="https://github.com/NousResearch/hermes-agent/pull/50349" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50349/hovercard">#50349</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52848" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52848/hovercard">#52848</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53973" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53973/hovercard">#53973</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Plugins: <code>ctx.profile_name</code> for session-agnostic profile access (<a href="https://github.com/NousResearch/hermes-agent/pull/50346" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50346/hovercard">#50346</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>LSP: PowerShellEditorServices language server; mem0 v3 API + OSS mode + update/delete tools (<a href="https://github.com/NousResearch/hermes-agent/pull/55930" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/55930/hovercard">#55930</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/15624" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/15624/hovercard">#15624</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kartik-mem0/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kartik-mem0">@kartik-mem0</a>)</li>
</ul>
<h2>⚡ Performance</h2>
<ul>
<li>Cold start: lazy-load gateway platform adapters; parse config + plugin manifests with libyaml <code>CSafeLoader</code> (<a href="https://github.com/NousResearch/hermes-agent/pull/54448" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54448/hovercard">#54448</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54486" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54486/hovercard">#54486</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>State: merge FTS5 segments + <code>handoff_state</code> index to curb write-lock contention; single-pass <code>list_profiles</code> alias map + skill-count cache + event-loop offload (<a href="https://github.com/NousResearch/hermes-agent/pull/54752" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54752/hovercard">#54752</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54770" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54770/hovercard">#54770</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>🔒 Security &amp; Reliability</h2>
<ul>
<li>Harden MCP-config persistence attack surface; block cron <code>base_url</code> overrides that exfiltrate provider credentials; non-reusable sentinel for prefix secrets in file reads (<a href="https://github.com/NousResearch/hermes-agent/pull/50476" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50476/hovercard">#50476</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56196" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56196/hovercard">#56196</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/54166" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/54166/hovercard">#54166</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
<li>Redact Slack App-Level (<code>xapp-</code>) tokens; browser cloud-metadata floor on all backends (CDP non-local); re-check private-network guard after <code>browser_back</code> navigation; scope <code>/resume</code> and <code>/sessions</code> to caller origin (IDOR); <code>aiohttp</code> 3.14.1 CVE floor across lazy messaging paths + pin-drift guard (<a href="https://github.com/NousResearch/hermes-agent/pull/56227" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56227/hovercard">#56227</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52349" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52349/hovercard">#52349</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56526" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56526/hovercard">#56526</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56378" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56378/hovercard">#56378</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56237" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56237/hovercard">#56237</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/claudlos/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/claudlos">@claudlos</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>)</li>
<li>Cron reliability wave: fail closed when an unpinned job's provider drifts; run missed-grace jobs once instead of deferring forever; keep the ticker alive on <code>BaseException</code> + heartbeat-aware status; layer enabled MCP servers onto per-job toolsets; guard cron model-tool path + auto-resume loop breaker (<a href="https://github.com/NousResearch/hermes-agent/pull/51051" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51051/hovercard">#51051</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50062" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50062/hovercard">#50062</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50016" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50016/hovercard">#50016</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50117" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50117/hovercard">#50117</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56240" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56240/hovercard">#56240</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
<li>Windows: suppress console flashes + harden gateway restarts; prefer cmd npm shim on PATH fallback; respawn gateway windowless after GUI update; prefer managed node for whatsapp/desktop (<a href="https://github.com/NousResearch/hermes-agent/pull/52340" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52340/hovercard">#52340</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/50398" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50398/hovercard">#50398</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/52239" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/52239/hovercard">#52239</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/helix4u/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/helix4u">@helix4u</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>🔁 Reverts (in-window, for the record)</h2>
<ul>
<li>cron job storage returned to per-profile (reverts <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4517607524" data-permission-text="Title is private" data-url="https://github.com/NousResearch/hermes-agent/issues/32117" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/32117/hovercard" href="https://github.com/NousResearch/hermes-agent/pull/32117">#32117</a> + <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4719892950" data-permission-text="Title is private" data-url="https://github.com/NousResearch/hermes-agent/issues/50993" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/50993/hovercard" href="https://github.com/NousResearch/hermes-agent/pull/50993">#50993</a>); don't clone <code>auth.json</code> (duplicating OAuth grant causes sibling revocation); windows terminal-popup PRs rolled back; <code>prompt_caching.enabled</code> toggle backed out for re-evaluation (<a href="https://github.com/NousResearch/hermes-agent/pull/51116" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51116/hovercard">#51116</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/51732" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/51732/hovercard">#51732</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/53853" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/53853/hovercard">#53853</a>, <a href="https://github.com/NousResearch/hermes-agent/pull/56126" data-hovercard-type="pull_request" data-hovercard-url="/NousResearch/hermes-agent/pull/56126/hovercard">#56126</a> — <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a>)</li>
</ul>
<h2>👥 Contributors</h2>
<p><strong>381 people</strong> contributed to this release (via commits, co-author trailers, and salvaged PRs). Thank you, all of you.</p>
<h3>Core</h3>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/teknium1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/teknium1">@teknium1</a> — release lead; MoA first-class, verification/goals, <code>/learn</code>, background review, security round, providers, the P0/P1 clean-sweep</li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a> — desktop app (projects, memory graph, <code>/journey</code>, multi-terminal, composer refactor wave, pets, verification UX)</li>
</ul>
<h3>Top community contributors</h3>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a> — the P0/P1 backlog burn: cron reliability wave, state perf, security (cron credential-exfil), gateway/signal, TUI config — a huge share of the priority closures</li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a> — relay Phase 5/6, scale-to-zero / drain coordination, dashboard auth/keys, gateway hardening</li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ethernet8023/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ethernet8023">@ethernet8023</a> — CI/docker (unified jobs, faster builds, timings report)</li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/helix4u/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/helix4u">@helix4u</a> — Windows hardening (console flashes, npm shim, gateway restarts)</li>
</ul>
<h3>All contributors</h3>
<p><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/0xbyt4/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/0xbyt4">@0xbyt4</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/0xDevNinja/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/0xDevNinja">@0xDevNinja</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/0xsir0000/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/0xsir0000">@0xsir0000</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/1RB/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/1RB">@1RB</a>, @595650661, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/aaronlab/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/aaronlab">@aaronlab</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/abchiaravalle/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/abchiaravalle">@abchiaravalle</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/adammatski1972/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/adammatski1972">@adammatski1972</a>, <a class="user-mention notranslate" data-hovercard-type="organization" data-hovercard-url="/orgs/AetherAgents/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AetherAgents">@AetherAgents</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Afnath-max/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Afnath-max">@Afnath-max</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/agt-user/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/agt-user">@agt-user</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ahmadashfq/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ahmadashfq">@ahmadashfq</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/AhmetArif0/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AhmetArif0">@AhmetArif0</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/AIalliAI/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/AIalliAI">@AIalliAI</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/aieng-abdullah/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/aieng-abdullah">@aieng-abdullah</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ailang323/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ailang323">@ailang323</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ailthrim/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ailthrim">@ailthrim</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/aj-nt/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/aj-nt">@aj-nt</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/alelpoan/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/alelpoan">@alelpoan</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/alloevil/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/alloevil">@alloevil</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/amathxbt/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/amathxbt">@amathxbt</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ambition0802/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ambition0802">@ambition0802</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/anderskev/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/anderskev">@anderskev</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/andressommerhoff/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/andressommerhoff">@andressommerhoff</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/angelos/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/angelos">@angelos</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/annguyenNous/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/annguyenNous">@annguyenNous</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Antimatter543/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Antimatter543">@Antimatter543</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/arminanton/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/arminanton">@arminanton</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/arthurzhang/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/arthurzhang">@arthurzhang</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/asimons81/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/asimons81">@asimons81</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/austinpickett/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/austinpickett">@austinpickett</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/baolingao/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/baolingao">@baolingao</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Bartok9/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Bartok9">@Bartok9</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/basilalshukaili/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/basilalshukaili">@basilalshukaili</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/BBCrypto-web/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/BBCrypto-web">@BBCrypto-web</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/bbopen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/bbopen">@bbopen</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Beandon13/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Beandon13">@Beandon13</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/beardthelion/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/beardthelion">@beardthelion</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbarclay/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbarclay">@benbarclay</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/benbenlijie/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/benbenlijie">@benbenlijie</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/binhnt92/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/binhnt92">@binhnt92</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/bitcryptic-gw/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/bitcryptic-gw">@bitcryptic-gw</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Blaryxoff/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Blaryxoff">@Blaryxoff</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/bogerman1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/bogerman1">@bogerman1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/bradhallett/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/bradhallett">@bradhallett</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/brett539/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/brett539">@brett539</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/briandevans/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/briandevans">@briandevans</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/buihongduc132/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/buihongduc132">@buihongduc132</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/bykim0119/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/bykim0119">@bykim0119</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/catapreta/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/catapreta">@catapreta</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/chaithanyak42/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/chaithanyak42">@chaithanyak42</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/charleneleong-ai/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/charleneleong-ai">@charleneleong-ai</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/CharlieKerfoot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/CharlieKerfoot">@CharlieKerfoot</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/chazmaniandinkle/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/chazmaniandinkle">@chazmaniandinkle</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/chrispersico/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/chrispersico">@chrispersico</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Christopher-Schulze/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Christopher-Schulze">@Christopher-Schulze</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/chriswesley4/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/chriswesley4">@chriswesley4</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/claudlos/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/claudlos">@claudlos</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/clovericbot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/clovericbot">@clovericbot</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cmcejas/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cmcejas">@cmcejas</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/codexGW/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/codexGW">@codexGW</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Cossackx/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Cossackx">@Cossackx</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/counterposition/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/counterposition">@counterposition</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/coygeek/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/coygeek">@coygeek</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/CRWuTJ/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/CRWuTJ">@CRWuTJ</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cyb0rgk1tty/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cyb0rgk1tty">@cyb0rgk1tty</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cyb3rwr3n/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cyb3rwr3n">@cyb3rwr3n</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cypctlinux/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cypctlinux">@cypctlinux</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cypres0099/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cypres0099">@cypres0099</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/dalenguyen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dalenguyen">@dalenguyen</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Danamove/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Danamove">@Danamove</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/DanAsBjorn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/DanAsBjorn">@DanAsBjorn</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/DataAdvisory/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/DataAdvisory">@DataAdvisory</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/davidgut1982/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/davidgut1982">@davidgut1982</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/DavidMetcalfe/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/DavidMetcalfe">@DavidMetcalfe</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/davidvv/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/davidvv">@davidvv</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/de1tydev/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/de1tydev">@de1tydev</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/denisqq/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/denisqq">@denisqq</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/devorun/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/devorun">@devorun</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/devsart95/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/devsart95">@devsart95</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/DhivinX/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/DhivinX">@DhivinX</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/DiamondEyesFox/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/DiamondEyesFox">@DiamondEyesFox</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/difujia/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/difujia">@difujia</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Disaster-Terminator/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Disaster-Terminator">@Disaster-Terminator</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/djimit/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/djimit">@djimit</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/djstunami/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/djstunami">@djstunami</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/dodo-reach/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dodo-reach">@dodo-reach</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/donovan-yohan/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/donovan-yohan">@donovan-yohan</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Dr1985/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Dr1985">@Dr1985</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/DrZM007/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/DrZM007">@DrZM007</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Dusk1e/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Dusk1e">@Dusk1e</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/egilewski/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/egilewski">@egilewski</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ehz0ah/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ehz0ah">@ehz0ah</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Eji4h/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Eji4h">@Eji4h</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/EloquentBrush0x/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/EloquentBrush0x">@EloquentBrush0x</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Elshayib/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Elshayib">@Elshayib</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/emozilla/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/emozilla">@emozilla</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/entropy-0x/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/entropy-0x">@entropy-0x</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/erosika/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/erosika">@erosika</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/EtherAura/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/EtherAura">@EtherAura</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/etherman-os/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/etherman-os">@etherman-os</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ethernet8023/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ethernet8023">@ethernet8023</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/f-trycua/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/f-trycua">@f-trycua</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fayenix/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fayenix">@fayenix</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fesalfayed/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fesalfayed">@fesalfayed</a>, <a class="user-mention notranslate" data-hovercard-type="organization" data-hovercard-url="/orgs/firefly/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/firefly">@firefly</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/flamiinngo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/flamiinngo">@flamiinngo</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/flobo3/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/flobo3">@flobo3</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/francescomucio/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/francescomucio">@francescomucio</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/franksong2702/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/franksong2702">@franksong2702</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/friendshipisover/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/friendshipisover">@friendshipisover</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fsaad1984/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fsaad1984">@fsaad1984</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/fyzanshaik/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/fyzanshaik">@fyzanshaik</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/GauravPatil2515/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/GauravPatil2515">@GauravPatil2515</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/gdeyoung/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/gdeyoung">@gdeyoung</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/georgex8001/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/georgex8001">@georgex8001</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/GodsBoy/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/GodsBoy">@GodsBoy</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/graphanov/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/graphanov">@graphanov</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Gromykoss/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Gromykoss">@Gromykoss</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/gustavosmendes/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/gustavosmendes">@gustavosmendes</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Gutslabs/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Gutslabs">@Gutslabs</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/H2KFORGIVEN/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/H2KFORGIVEN">@H2KFORGIVEN</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/haileymarshall/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/haileymarshall">@haileymarshall</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hakanpak/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hakanpak">@hakanpak</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/happy5318/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/happy5318">@happy5318</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/haran2001/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/haran2001">@haran2001</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/harjothkhara/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/harjothkhara">@harjothkhara</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/heathley/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/heathley">@heathley</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hehehe0803/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hehehe0803">@hehehe0803</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/helix4u/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/helix4u">@helix4u</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/herbalizer404/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/herbalizer404">@herbalizer404</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/HexLab98/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/HexLab98">@HexLab98</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/HiddenPuppy/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/HiddenPuppy">@HiddenPuppy</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Hinotoi-agent/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Hinotoi-agent">@Hinotoi-agent</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/HODLCLONE/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/HODLCLONE">@HODLCLONE</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/houko/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/houko">@houko</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/huangsen365/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/huangsen365">@huangsen365</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/huangxudong663-sys/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/huangxudong663-sys">@huangxudong663-sys</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/huangxun375-stack/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/huangxun375-stack">@huangxun375-stack</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/HwangJohn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/HwangJohn">@HwangJohn</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/iaji/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/iaji">@iaji</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/iamlukethedev/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/iamlukethedev">@iamlukethedev</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/IamSanchoPanza/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/IamSanchoPanza">@IamSanchoPanza</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/IAvecilla/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/IAvecilla">@IAvecilla</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Icather/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Icather">@Icather</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/iizotov/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/iizotov">@iizotov</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/indigokarasu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/indigokarasu">@indigokarasu</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/infinitycrew39/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/infinitycrew39">@infinitycrew39</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ipriyaaanshu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ipriyaaanshu">@ipriyaaanshu</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/isair/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/isair">@isair</a>, @islam666, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/itenev/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/itenev">@itenev</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/itsflownium/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/itsflownium">@itsflownium</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/izumi0uu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/izumi0uu">@izumi0uu</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Jaaneek/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Jaaneek">@Jaaneek</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/JabberELF/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/JabberELF">@JabberELF</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jackjin1997/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jackjin1997">@jackjin1997</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jackroofan/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jackroofan">@jackroofan</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/janrenz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/janrenz">@janrenz</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jasnoorgill/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jasnoorgill">@jasnoorgill</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jasonQin6/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jasonQin6">@jasonQin6</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jcjc81/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jcjc81">@jcjc81</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jearnest11/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jearnest11">@jearnest11</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jeeves-assistant/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jeeves-assistant">@jeeves-assistant</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Jeffgithub0029/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Jeffgithub0029">@Jeffgithub0029</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jeffrobodie-glitch/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jeffrobodie-glitch">@jeffrobodie-glitch</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/JezzaHehn/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/JezzaHehn">@JezzaHehn</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jimmyjohansson84/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jimmyjohansson84">@jimmyjohansson84</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jmmaloney4/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jmmaloney4">@jmmaloney4</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jnibarger01/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jnibarger01">@jnibarger01</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/JoaoMarcos44/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/JoaoMarcos44">@JoaoMarcos44</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jplew/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jplew">@jplew</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Junass1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Junass1">@Junass1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/justemu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/justemu">@justemu</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/justin-cyhuang/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/justin-cyhuang">@justin-cyhuang</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/JustinOhms/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/JustinOhms">@JustinOhms</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jvradahellys24-art/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jvradahellys24-art">@jvradahellys24-art</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Kailigithub/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Kailigithub">@Kailigithub</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kaishi00/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kaishi00">@kaishi00</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kangsoo-bit/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kangsoo-bit">@kangsoo-bit</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kartik-mem0/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kartik-mem0">@kartik-mem0</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/keiravoss94/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/keiravoss94">@keiravoss94</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kenyonxu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kenyonxu">@kenyonxu</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kernel-t1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kernel-t1">@kernel-t1</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Kewe63/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Kewe63">@Kewe63</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/KeyArgo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/KeyArgo">@KeyArgo</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/KiruyaMomochi/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/KiruyaMomochi">@KiruyaMomochi</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kn8-codes/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kn8-codes">@kn8-codes</a>,<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Kolektori/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Kolektori">@Kolektori</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/konsisumer/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/konsisumer">@konsisumer</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kshitijk4poor/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kshitijk4poor">@kshitijk4poor</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/kyssta-exe/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/kyssta-exe">@kyssta-exe</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Kyzcreig/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Kyzcreig">@Kyzcreig</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Lazymonter/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Lazymonter">@Lazymonter</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LehaoLin/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LehaoLin">@LehaoLin</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/LeonSGP43/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/LeonSGP43">@LeonSGP43</a>, @lEWFkRAD,<br>
@libre-7, @LIC99, @LifeJiggy, @linyubin, @liuhao1024, @lkevincc0, @lkz-de, @loes5050, @londo161, @lubosxyz,<br>
@m24927605, @MaheshtheDev, @manus-use, @marco0158, @MarioYounger, @martinramos002-bot, @MattKotsenas,<br>
@max-chen, @MaxFreedomPollard, @maxmilian, @maxpetrusenko, @memosr, @Mibayy, @Minksgo, @mintybasil, @mkslzk,<br>
@mohamedorigami-jpg, @MorAlekss, @mrparker0980, @ms-alan, @namredips, @nankingjing, @natehale, @necoweb3,<br>
@neo-2026, @Nickperillo, @nightq, @nikshepsvn, @nnnet, @nocturnum91, @nodejun, @NousResearch, @nycomar,<br>
@OmarB97, @orbisai0security, @oreoluwa, @outsourc-e, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/OutThisLife/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/OutThisLife">@OutThisLife</a>, @p-andhika, @panghuer023, @Paperclip,<br>
@peetwan, @pefontana, @petrichor-op, @pinguarmy, @PINKIIILQWQ, @pmos69, @PolyphonyRequiem, @pprism13,<br>
@PRATHAMESH75, @professorpalmer, @pyxl-dev, @Que0x, @qWaitCrypto, @r266-tech, @RafaelMiMi, @Railway9784,<br>
@randomuser2026x, @rayjun, @rc-int, @rebel0789, @redactdeveloper, @riyas22, @rlaope, @rob-maron, @rodboev,<br>
@rodrigoeqnit, @rratmansky, @rrevenanttt, @ruangraung, @Ruzzgar, @ryo-solo, @s010mn, @Sahil-SS9,<br>
@SahilRakhaiya05, @SandroHub013, @Sanjays2402, @sasquatch9818, @ScotterMonk, @season179, @sgabel, @sgaofen,<br>
@sgtworkman, @shandian64, @shannonsands, @shashwatgokhe, @shawchanshek, @sherman-yang, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/SHL0MS/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/SHL0MS">@SHL0MS</a>, @SidUParis,<br>
@SimoKiihamaki, @simpolism, @sjh9714, @skabartem, @skyc1e, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/slawt/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/slawt">@slawt</a>, @soynchux, @spiky02plateau, @spjoes,<br>
@sprmn24, @srojk34, @stepanov1975, @steveonjava, @Subway2023, @sweetcornna, @swissly, @Sworntech-dev,<br>
@syahidfrd, @synapsesx, @szzhoujiarui-sketch, @talmax1124, @telos-oc, @testingbuddies24, @texhy, @tgmerritt,<br>
@theAgenticBuilder, @thestral123, @tkwong, @Tortugasaur, @Tranquil-Flow, @trevorgordon981, @truenorth-lj,<br>
@tt-a1i, @tuancookiez-hub, @TutkuEroglu, @tymrtn, @udatny, @UgwujaGeorge, @underthestars-zhy, @uperLu,<br>
@uzunkuyruk, @valenteff, @valentt, @vanthinh6886, @Versun, @victor-kyriazakos, @virtuadex, @vKongv,<br>
@w31rdm4ch1nZ, @weidzhou, @wgu9, @whoislikemiha, @wnuuee1, @woaini30050, @WuKongAI-CMU, @WuTianyi123, @WXBR,<br>
@x7peeps, @x9x9x9x9x9x91, @Xowiek, @xxchan, @xxxigm, @xydigit-zt, @yapsrubricsz0, @yashiels, @yeyitech, @ygd58,<br>
@YLChen-007, @yong2bba, @yoniebans, @ypwcharles, @yu-xin-c, @yungchentang, @yusekiotacode, @YuShu, @yyzquwu,<br>
@zapabob, @zccyman, @zeapsu, @zmlgit, @znding04, @Zyxxx-xxxyZ</p>
<p>Also: Lucas Nicolas.</p>
<hr>
<p><strong>Full Changelog</strong>: <a href="https://github.com/NousResearch/hermes-agent/compare/v2026.6.19...v2026.7.1">v2026.6.19...v2026.7.1</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[2026 BAIR Graduate Showcase]]></title>
<description><![CDATA[Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work sp...]]></description>
<link>https://tsecurity.de/de/3639545/ai-nachrichten/2026-bair-graduate-showcase/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3639545/ai-nachrichten/2026-bair-graduate-showcase/</guid>
<pubDate>Wed, 01 Jul 2026 21:33:50 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<!-- twitter -->










<p>Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.</p>

<p>Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more. Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for the better.</p>

<p>Now they are headed everywhere ideas travel: to faculty and postdoctoral positions, to industry research labs, and to startups of their own founding — and several are still exploring what comes next and would love to hear from you.</p>

<p>Please join us in celebrating the achievements of these wonderful graduates. We are proud of everything they have accomplished at Berkeley, and we can’t wait to see what they do next!</p>

<!--more-->

<p><small><i>Thank you to our friends at the <a href="https://ai.stanford.edu/blog/sail-graduates/">Stanford AI Lab</a> for this idea!</i></small></p>

<hr>

<div class="container">
  <div class="row">
    
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://bfshi.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/baifeng-shi.jpg" alt="Baifeng Shi" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Baifeng Shi</h1><br>
              <strong>Email:</strong><a href="mailto:baifeng_shi@berkeley.edu"> baifeng_shi@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://bfshi.github.io/">https://bfshi.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Trevor Darrell<br>
              
              <strong>Research Blurb:</strong> I work on building generalist vision and robotic models.<br>
              
              
              <strong>What's next:</strong> Member of Technical Staff at Physical Intelligence
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://sea-snell.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/charlie-snell.jpg" alt="Charlie Snell" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Charlie Snell</h1><br>
              <strong>Email:</strong><a href="mailto:csnell22@berkeley.edu"> csnell22@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://sea-snell.github.io/">https://sea-snell.github.io</a><br>
              
              <strong>Advisor(s):</strong> Dan Klein<br>
              
              <strong>Research Blurb:</strong> My work aims to understand when and how the different LLM scaling paradigms can be traded off and interchanged. In particular, test-time scaling treats each prompt independently, drawing long chains of inferences and then forgetting them entirely between prompts. This differs critically from pretraining, which instead learns a compressed representation from a large dataset. I believe bridging the gap between these methods of scaling computation, presents a key open challenge in the field: how can we develop methods which turn the inferences drawn at test-time back into learned representations that the model can hold onto across interactions.<br>
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://devinguillory.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/devin-guillory.jpg" alt="Devin Guillory" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Devin Guillory</h1><br>
              <strong>Email:</strong><a href="mailto:dguillory@berkeley.edu"> dguillory@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://devinguillory.com/">https://devinguillory.com</a><br>
              
              <strong>Advisor(s):</strong> Trevor Darrell<br>
              
              <strong>Research Blurb:</strong> Accounting for data shifts in computer vision models<br>
              
              
              <strong>What's next:</strong> Building collaborative AI systems, looking for conspirators.
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://efleisig.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/eve-fleisig.jpg" alt="Eve Fleisig" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Eve Fleisig</h1><br>
              <strong>Email:</strong><a href="mailto:efleisig@berkeley.edu"> efleisig@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://efleisig.com/">https://efleisig.com</a><br>
              
              <strong>Advisor(s):</strong> Dan Klein<br>
              
              <strong>Research Blurb:</strong> I design language models to work reliably and fairly for the broad range of real LLM users. First, my research leverages disagreement among user preferences as signal, in order to train and evaluate LLMs for entire populations of users. Second, I work on designing rigorous evaluations to extricate challenging LLM harms that diverse users face. Finally, I work on core technical failures of LLMs, like miscalibrated confidence, to reduce downstream risks when models are deployed to users with different needs. Combined, these interventions facilitate building LLMs that minimize societal harms, and maximize benefits to a wider range of real-world users.<br>
              
              
              <strong>What's next:</strong> Postdoctoral fellow at Princeton CITP
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://graceluo.net/"><img src="https://bair.berkeley.edu/static/blog/grads2026/grace-luo.jpg" alt="Grace Luo" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Grace Luo</h1><br>
              <strong>Email:</strong><a href="mailto:graceluo@berkeley.edu"> graceluo@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://graceluo.net/">https://graceluo.net</a><br>
              
              <strong>Advisor(s):</strong> Trevor Darrell<br>
              
              <strong>Research Blurb:</strong> My research is on interpreting and controlling generative models. For example, I've worked on re-purposing image generators for computer vision tasks, and meta-modeling language activations for better LLM probing and steering.<br>
              
              
              <strong>What's next:</strong> Research scientist in industry
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://hanlinzhu.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/hanlin-zhu.jpg" alt="Hanlin Zhu" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Hanlin Zhu</h1><br>
              <strong>Email:</strong><a href="mailto:hanlinzhu@berkeley.edu"> hanlinzhu@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://hanlinzhu.com/">https://hanlinzhu.com/</a><br>
              
              <strong>Advisor(s):</strong> Stuart Russell, Jiantao Jiao<br>
              
              <strong>Research Blurb:</strong> My research centers on understanding and improving the reasoning capabilities of large language models (LLMs).<br>
              
              
              <strong>What's next:</strong> Member of Technical Staff at OpenAI
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://haozhi.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/haozhi-qi.jpg" alt="Haozhi Qi" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Haozhi Qi</h1><br>
              <strong>Email:</strong><a href="mailto:hqi@berkeley.edu"> hqi@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://haozhi.io/">https://haozhi.io/</a><br>
              
              <strong>Advisor(s):</strong> Jitendra Malik, Yi Ma<br>
              
              <strong>Research Blurb:</strong> Dexterous Manipulation and Robot Learning<br>
              
              
              <strong>What's next:</strong> Research scientist at Amazon; Faculty at University of Chicago
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://zamfi.net/"><img src="https://bair.berkeley.edu/static/blog/grads2026/j-d-zamfirescu-pereira.jpg" alt="J.D. Zamfirescu-Pereira" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>J.D. Zamfirescu-Pereira</h1><br>
              <strong>Email:</strong><a href="mailto:zamfi@berkeley.edu"> zamfi@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://zamfi.net/">https://zamfi.net</a><br>
              
              <strong>Advisor(s):</strong> Bjoern Hartmann<br>
              
              <strong>Research Blurb:</strong> My research focuses on effective human-AI co-design. I study the boundaries of language interfaces as a medium for interacting with AI, creating systems that blend language-focused interactions with structured user interfaces that draw on different levels of abstraction. I focus on language-oriented technologies, like LLMs and text-to-image models, that are powerful mediators of design processes. These technologies enable humans to describe their desires at almost any level of abstraction, from high-level goals vaguely specified (“I’d like a game to help my kid learn to read”) to low-level corrections of undesired outputs (“Don’t say ‘I know because I’ve tasted it’ when about a recipe substitution's taste”).<br>
              
              
              <strong>What's next:</strong> Assistant Professor, Computer Science, UCLA
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://jlian2.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/jiachen-lian.jpg" alt="Jiachen Lian" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Jiachen Lian</h1><br>
              <strong>Email:</strong><a href="mailto:jiachenlian@berkeley.edu"> jiachenlian@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://jlian2.github.io/">https://jlian2.github.io</a><br>
              
              <strong>Advisor(s):</strong> Gopala Anumanchipalli<br>
              
              <strong>Research Blurb:</strong> My research focuses on human-centered AI across speech, healthcare, and systems.<br>
              
              
              <strong>Looking for:</strong> Look for AI talents to join our startup
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://joshuaminwookang.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/josh-kang.jpg" alt="Josh Kang" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Josh Kang</h1><br>
              <strong>Email:</strong><a href="mailto:minwoo_kang@berkeley.edu"> minwoo_kang@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://joshuaminwookang.github.io/">https://joshuaminwookang.github.io/</a><br>
              
              <strong>Advisor(s):</strong> John Canny<br>
              
              <strong>Research Blurb:</strong> I study language modeling and related topics in NLP; specific interests are human user simulation and building conversational, collaborative AI agents.<br>
              
              
              <strong>What's next:</strong> AI Scientist at Mistral AI
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://www.linkedin.com/in/junhao-bear-xiong"><img src="https://bair.berkeley.edu/static/blog/grads2026/junhao-bear-xiong.jpg" alt="Junhao (Bear) Xiong" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Junhao (Bear) Xiong</h1><br>
              <strong>Email:</strong><a href="mailto:junhao_xiong@berkeley.edu"> junhao_xiong@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://www.linkedin.com/in/junhao-bear-xiong">https://www.linkedin.com/in/junhao-bear-xiong</a><br>
              
              <strong>Advisor(s):</strong> Jennifer Listgarten, Yun Song<br>
              
              <strong>Research Blurb:</strong> Junhao (Bear) Xiong is a PhD candidate at UC Berkeley, advised by Jennifer Listgarten and Yun S. Song. His work focuses on machine learning methods for biology, with an emphasis on generative modeling for proteins. Previously, he studied Applied Math and Computer Science at Johns Hopkins.<br>
              
              
              <strong>Looking for:</strong> Research scientist
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://kaylolittlejohn.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/kaylo-littlejohn.jpg" alt="Kaylo Littlejohn" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Kaylo Littlejohn</h1><br>
              <strong>Email:</strong><a href="mailto:kaylo_littlejohn@berkeley.edu"> kaylo_littlejohn@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://kaylolittlejohn.com/">https://kaylolittlejohn.com</a><br>
              
              <strong>Advisor(s):</strong> Gopala Anumanchipalli<br>
              
              <strong>Research Blurb:</strong> My research is focused on speech modeling and natural language processing. I co-led the development of multimodal AI tools to accurately translate brain activity into text, audible personalized speech, and a high-fidelity "digital talking avatar" (Nature 2023, Nature Neuroscience 2025). I am also tech lead for voice modeling at Roblox.<br>
              
              
              <strong>Looking for:</strong> Research Scientist / Engineer
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://kentkc.org/"><img src="https://bair.berkeley.edu/static/blog/grads2026/kent-chang.jpg" alt="Kent Chang" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Kent Chang</h1><br>
              <strong>Email:</strong><a href="mailto:kentkchang@berkeley.edu"> kentkchang@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://kentkc.org/">https://kentkc.org</a><br>
              
              <strong>Advisor(s):</strong> David Bamman<br>
              
              <strong>Research Blurb:</strong> I work on NLP and multimodal machine learning, with a focus on evaluating large language models and building multimodal systems for understanding dialogue, narrative, and social interaction. My research includes benchmarks for LLM memorization, multimodal datasets sourced from feature films and television, and studies of model behavior. I'm interested in bridging computational methods with questions from the humanities and social sciences about whose voices get represented in AI systems, and about AI's broader impact. My work has appeared at EMNLP and ACL, among others.<br>
              
              
              <strong>Looking for:</strong> (teaching) faculty, Research Scientist, ML/AI SWE
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://kevin.black/"><img src="https://bair.berkeley.edu/static/blog/grads2026/kevin-black.jpg" alt="Kevin Black" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Kevin Black</h1><br>
              <strong>Email:</strong><a href="mailto:kvablack@berkeley.edu"> kvablack@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://kevin.black/">https://kevin.black</a><br>
              
              <strong>Advisor(s):</strong> Sergey Levine<br>
              
              <strong>Research Blurb:</strong> I work on large-scale robot learning: including imitation learning, reinforcement learning, generative modeling, real-time control, and whatever else it takes to make robots work in the real world!<br>
              
              
              <strong>What's next:</strong> Research Scientist of Physical Intelligence
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://www.kunheyang.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/kunhe-yang.jpg" alt="Kunhe Yang" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Kunhe Yang</h1><br>
              <strong>Email:</strong><a href="mailto:kunheyang@berkeley.edu"> kunheyang@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://www.kunheyang.com/">https://www.kunheyang.com/</a><br>
              
              <strong>Advisor(s):</strong> Nika Haghtalab<br>
              
              <strong>Research Blurb:</strong> My research focuses on the theoretical foundations of designing and evaluating AI algorithms in environments shaped by human incentives and AI agency. My work spans human-centric policy learning, incentive-aware evaluation, and multi-agent collaboration and information transmission, drawing on tools from machine learning theory and computational economics.<br>
              
              
              <strong>What's next:</strong> Postdoc Research at Stanford
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://lisabdunlap.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/lisa-dunlap.jpg" alt="Lisa Dunlap" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Lisa Dunlap</h1><br>
              <strong>Email:</strong><a href="mailto:lisabdunlap@berkeley.edu"> lisabdunlap@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://lisabdunlap.com/">https://lisabdunlap.com</a><br>
              
              <strong>Advisor(s):</strong> Joseph Gonzalez, Trevor Darrell<br>
              
              <strong>Research Blurb:</strong> Auditing generative models.<br>
              
              
              <strong>What's next:</strong> Research Engineer at Anthropic
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://tonylian.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/long-tony-lian.jpg" alt="Long (Tony) Lian" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Long (Tony) Lian</h1><br>
              <strong>Email:</strong><a href="mailto:longlian@berkeley.edu"> longlian@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://tonylian.com/">https://tonylian.com/</a><br>
              
              <strong>Advisor(s):</strong> Trevor Darrell, Adam Yala<br>
              
              <strong>Research Blurb:</strong> My research primarily focuses on developing real-time multi-modal multi-agent systems and parallel reasoning systems through end-to-end RL.<br>
              
              
              <strong>What's next:</strong> Member of Technical Staff at Thinking Machines Lab
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://maulikb.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/maulik-bhatt.jpg" alt="Maulik Bhatt" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Maulik Bhatt</h1><br>
              <strong>Email:</strong><a href="mailto:maulikbhatt@berkeley.edu"> maulikbhatt@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://maulikb.com/">https://maulikb.com</a><br>
              
              <strong>Advisor(s):</strong> Negar Mehr<br>
              
              <strong>Research Blurb:</strong> My research develops autonomous robots that can safely coordinate with humans and other robots in shared environments. I build scalable algorithms grounded in game theory and diffusion models that let agents reason about the intent and behavior of others around them. My work spans real-time multi-agent trajectory planning and imitation learning in the presence of multi-modality. I've validated these methods on hardware platforms ranging from quadrotors to manipulators, with the goal of making multi-agent coordination robust, interpretable, and deployable in the real world.<br>
              
              
              <strong>What's next:</strong> Joining Toyota Woven's end-to-end autonomous driving team.
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://www.michaelpsenka.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/michael-psenka.jpg" alt="Michael Psenka" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Michael Psenka</h1><br>
              <strong>Email:</strong><a href="mailto:psenka@berkeley.edu"> psenka@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://www.michaelpsenka.io/">https://www.michaelpsenka.io/</a><br>
              
              <strong>Advisor(s):</strong> Aditi Krishnapriyan<br>
              
              <strong>Research Blurb:</strong> Work in various domains (reinforcement learning, world models, AI+bio/chem), generally working on longer-horizon and out-of-distribution problems in planning and interpolation (e.g. robot manipulation from start state to goal, molecular dynamics of proteins between ground states). My thesis took a variational approach (think calculus of variations) directly from deep generative models of the environment, framing path-finding as minimizing a functional induced by the learned model itself (its score, its critic, or its dynamics). Through my research I've gained insight on how to properly handle dynamics in deep learning systems, and I plan to continue developing systems that are dynamic and adaptive.<br>
              
              
              <strong>What's next:</strong> Lead Research Scientist at Baseten
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://nathanlichtle.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/nathan-lichtle.jpg" alt="Nathan Lichtlé" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Nathan Lichtlé</h1><br>
              <strong>Email:</strong><a href="mailto:nathan.lichtle@gmail.com"> nathan.lichtle@gmail.com</a><br>
              <strong>Website:</strong> <a href="https://nathanlichtle.com/">https://nathanlichtle.com</a><br>
              
              <strong>Advisor(s):</strong> Alexandre M. Bayen<br>
              
              <strong>Research Blurb:</strong> RL for autonomous driving.<br>
              
              
              <strong>What's next:</strong> Chief Scientist &amp; Co-founder at Yumi Health
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://neerja.me/"><img src="https://bair.berkeley.edu/static/blog/grads2026/neerja-thakkar.jpg" alt="Neerja Thakkar" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Neerja Thakkar</h1><br>
              <strong>Email:</strong><a href="mailto:nthakkar@berkeley.edu"> nthakkar@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://neerja.me/">https://neerja.me/</a><br>
              
              <strong>Advisor(s):</strong> Jitendra Malik<br>
              
              <strong>Research Blurb:</strong> My research focuses on scaling predictive world models to handle the complexity of in-the-wild motion. Using autoregressive and diffusion frameworks, I develop better representations for real-world prediction and propose methods to efficiently adapt these models to new domains.<br>
              
              
              <strong>Looking for:</strong> Research scientist
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://n-mehandru.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/nikita-mehandru.jpg" alt="Nikita Mehandru" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Nikita Mehandru</h1><br>
              <strong>Email:</strong><a href="mailto:nmehandru@berkeley.edu"> nmehandru@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://n-mehandru.github.io/">https://n-mehandru.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Ahmed Alaa and David Bamman<br>
              
              <strong>Research Blurb:</strong> My research develops and applies machine learning methods for clinical reasoning and disease progression modeling using unstructured text and time series data from electronic health records. In collaboration with physicians at UCSF, I bridge method development and clinical validation with the intention to build reliable, interpretable AI systems in medicine.<br>
              
              
              <strong>Looking for:</strong> Research Scientist
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://niklaslauffer.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/niklas-lauffer.jpg" alt="Niklas Lauffer" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Niklas Lauffer</h1><br>
              <strong>Email:</strong><a href="mailto:nlauffer@berkeley.edu"> nlauffer@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://niklaslauffer.github.io/">https://niklaslauffer.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Stuart Russell and Sanjit Seshia<br>
              
              <strong>Research Blurb:</strong> Niklas's research is focused on AI safety and reinforcement learning, particularly in the area of multi-agent interaction and LM agents. He's worked on enabling adversarial learning in cooperative and mixed-motive settings, solving issues of covariate shift in training LM agents on long-horizon tasks, as well as evaluating safety risks posed by LM agents in multi-agent settings.<br>
              
              
              <strong>What's next:</strong> Research Scientist at Google Deepmind
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://colinqiyangli.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/qiyang-li.jpg" alt="Qiyang Li" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Qiyang Li</h1><br>
              <strong>Email:</strong><a href="mailto:qcli@berkeley.edu"> qcli@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://colinqiyangli.github.io/">https://colinqiyangli.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Sergey Levine<br>
              
              <strong>Research Blurb:</strong> Recent progress in robotic manipulation policy learning has been largely driven by (1) the increasing availability of large-scale prior datasets and (2) the success of action chunking, where the policy predicts a short sequence of future actions rather than a single one. However, most action chunking policies are trained via supervised imitation learning, because efficient online self-improvement with reinforcement learning (RL) remains challenging—limiting real-world applicability. My PhD research studied how we could leverage prior data to optimize action-chunking policies with RL, combining empirical results with theoretical insights.<br>
              
              
              <strong>Looking for:</strong> Post-doc/research scientist for RL in robotics and LLMs!
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://sdeglurkar.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/sampada-deglurkar.jpg" alt="Sampada Deglurkar" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Sampada Deglurkar</h1><br>
              <strong>Email:</strong><a href="mailto:sampada_deglurkar@berkeley.edu"> sampada_deglurkar@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://sdeglurkar.github.io/">https://sdeglurkar.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Prof Claire Tomlin<br>
              
              <strong>Research Blurb:</strong> My research is in providing safety assurances for AI-enabled autonomous systems, ranging from robots to autonomous vehicles to aviation systems. For this, I have worked with uncertainty quantification for machine learning models, decision-making under uncertainty algorithms, and tools for producing probabilistic guarantees on system operation.<br>
              
              
              <strong>Looking for:</strong> Research scientist, Research engineer
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://cs.berkeley.edu/~vbenara"><img src="https://bair.berkeley.edu/static/blog/grads2026/vinamra-benara.jpg" alt="Vinamra Benara" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Vinamra Benara</h1><br>
              <strong>Email:</strong><a href="mailto:vbenara@berkeley.edu"> vbenara@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://cs.berkeley.edu/~vbenara">https://cs.berkeley.edu/~vbenara</a><br>
              
              <strong>Advisor(s):</strong> Ion Stoica<br>
              
              <strong>Research Blurb:</strong> My research focuses on LLM post-training, including data curation, RLHF, RLVR with VLMs, evaluations, reasoning, agentic workflows, and interpretability. I also have strong expertise in systems infrastructure for distributed computing.<br>
              
              
              <strong>Looking for:</strong> Research scientist / Research Engineer
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://people.eecs.berkeley.edu/~vongani_maluleke/"><img src="https://bair.berkeley.edu/static/blog/grads2026/vongani-maluleke.jpg" alt="Vongani Maluleke" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Vongani Maluleke</h1><br>
              <strong>Email:</strong><a href="mailto:vongani_maluleke@berkeley.edu"> vongani_maluleke@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://people.eecs.berkeley.edu/~vongani_maluleke/">https://people.eecs.berkeley.edu/~vongani_maluleke/</a><br>
              
              <strong>Advisor(s):</strong> Jitendra Malik and Angjoo Kanazawa<br>
              
              <strong>Research Blurb:</strong> Vongani Maluleke is a PhD candidate at UC Berkeley (BAIR, advised by Jitendra Malik and Angjoo Kanazawa), where she led the development of MAGNet, a unified multi-agent motion generation framework that supports a wide range of motion generation tasks without retraining or architectural changes, outperforming task-specialized state-of-the-art baselines. She is currently extending this work by deploying it on a Unitree G1 humanoid to make it embody social intelligence. Before her PhD, she was a Senior AI Consultant at Deloitte, awarded Exceptional Performer two consecutive years, leading AI system development across media, telecommunications, retail, and financial services.<br>
              
              
              <strong>Looking for:</strong> Research scientist
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://weijer-chang.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/wei-jer-chang.jpg" alt="Wei-Jer Chang" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Wei-Jer Chang</h1><br>
              <strong>Email:</strong><a href="mailto:weijer_chang@berkeley.edu"> weijer_chang@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://weijer-chang.github.io/">https://weijer-chang.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Masayoshi Tomizuka<br>
              
              <strong>Research Blurb:</strong> My research focuses on developing safe and intelligent autonomous systems for complex, human-centered environments. I work at the intersection of machine learning, generative models, and reinforcement learning, with applications in autonomy. My work addresses challenges in multi-agent interaction, interactive human behavior, and long-tail safety-critical scenarios at scale.<br>
              
              
              <strong>Looking for:</strong> Research Scientist, Applied Scientist, Roboticist
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://xiuyuli.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/xiuyu-li.jpg" alt="Xiuyu Li" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Xiuyu Li</h1><br>
              <strong>Email:</strong><a href="mailto:xiuyu@berkeley.edu"> xiuyu@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://xiuyuli.com/">https://xiuyuli.com/</a><br>
              
              <strong>Advisor(s):</strong> Kurt Keutzer<br>
              
              <strong>Research Blurb:</strong> My research focuses on developing scalable and self-improving large language model agents, with emphasis on coding agents for complex, long-horizon tasks. This direction builds on my work in parallel reasoning, and on broader expertise in making generative models more efficient in training and inference across language and vision.<br>
              
              
              <strong>What's next:</strong> Member of Technical Staff at xAI
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://yichen928.github.io/"><img src="https://bair.berkeley.edu/static/blog/grads2026/yichen-xie.jpg" alt="Yichen Xie" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Yichen Xie</h1><br>
              <strong>Email:</strong><a href="mailto:yichenxie0928@gmail.com"> yichenxie0928@gmail.com</a><br>
              <strong>Website:</strong> <a href="https://yichen928.github.io/">https://yichen928.github.io/</a><br>
              
              <strong>Advisor(s):</strong> Masayoshi Tomizuka<br>
              
              <strong>Research Blurb:</strong> My research focuses on building multimodal foundation models and world models that understand and interact with complex physical environments. I aim to develop unified representations across modalities, enabling AI systems to reason over space, time, and dynamics toward general-purpose embodied intelligence.<br>
              
              
              <strong>What's next:</strong> Research Scientist at Luma AI
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://www.linkedin.com/in/erginbas/"><img src="https://bair.berkeley.edu/static/blog/grads2026/yigit-efe-erginbas.jpg" alt="Yigit Efe Erginbas" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Yigit Efe Erginbas</h1><br>
              <strong>Email:</strong><a href="mailto:erginbas@berkeley.edu"> erginbas@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://www.linkedin.com/in/erginbas/">https://www.linkedin.com/in/erginbas/</a><br>
              
              <strong>Advisor(s):</strong> Kannan Ramchandran, Thomas A. Courtade<br>
              
              <strong>Research Blurb:</strong> My PhD research spans two threads: online learning in large-scale markets, and interpretability of large machine learning models. In the first, I work on sequential decision-making with applications to recommendation, pricing, and assortment selection. My focus is on designing algorithms with provable guarantees for welfare maximization, revenue maximization, and stability. In the second, I develop scalable attribution methods that exploit the sparse, low-degree structure of real-world interactions, using tools from signal processing and information theory. More recently, I have been exploring principled ways to evaluate the faithfulness of model self-explanations.<br>
              
              
              <strong>What's next:</strong> Researcher at Hudson River Trading's AI Labs (HAIL)
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://yihengli.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/yiheng-li.jpg" alt="Yiheng Li" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Yiheng Li</h1><br>
              <strong>Email:</strong><a href="mailto:yhli@berkeley.edu"> yhli@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://yihengli.com/">https://Yihengli.com</a><br>
              
              <strong>Advisor(s):</strong> Masayoshi Tomizuka<br>
              
              <strong>Research Blurb:</strong> I am working on vision world modeling, with prior experience in diffusion model's efficiency as well as in autonomous driving.<br>
              
              
              <strong>What's next:</strong> Research Scientist at Waymo
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
      <div class="col-md-4">
        <div class="card mb-4 shadow-sm">
          <a href="https://fu-zhe.com/"><img src="https://bair.berkeley.edu/static/blog/grads2026/zhe-fu.jpg" alt="Zhe Fu" class="bd-placeholder-img card-img-top" width="480" height="auto"></a>
          <div class="card-body">
            <p class="card-text">
              </p><h1>Zhe Fu</h1><br>
              <strong>Email:</strong><a href="mailto:zhefu@berkeley.edu"> zhefu@berkeley.edu</a><br>
              <strong>Website:</strong> <a href="https://fu-zhe.com/">https://fu-zhe.com/</a><br>
              
              <strong>Advisor(s):</strong> Alexandre Bayen<br>
              
              <strong>Research Blurb:</strong> My research focuses on physics-informed learning and control for mixed-autonomy systems, with applications in transportation. I design physics-informed neural networks to learn solutions of nonlinear partial differential equations, enabling accurate and data-efficient prediction of traffic dynamics. Building on these models, I develop both model-based and learning-based control strategies that coordinate automated vehicles to improve system-level performance. My work bridges machine learning, control, and real-world deployment, and has been validated in large-scale field experiments. More broadly, I aim to advance trustworthy, interpretable AI for decision-making in complex, real-world systems.<br>
              
              
              <strong>What's next:</strong> I will be an Energy Fellow at Stanford after graduation. Also looking for Faculty, or research scientist positions in AI, control, and autonomy.
              
              
            
          </div>
        </div>
      </div>
      <hr>
    
  </div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic is bringing back Claude Fable 5 globally after US lifts export control order — where can enterprises access it?]]></title>
<description><![CDATA[Anthropic is restoring global access to its most powerful generally released AI model yet, Claude Fable 5, today,  after the U.S. Department of Commerce withdrew emergency export controls that led the company to suspend all access to both Fable 5 and its less restricted cybersecurity counterpart ...]]></description>
<link>https://tsecurity.de/de/3639021/it-nachrichten/anthropic-is-bringing-back-claude-fable-5-globally-after-us-lifts-export-control-order-where-can-enterprises-access-it/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3639021/it-nachrichten/anthropic-is-bringing-back-claude-fable-5-globally-after-us-lifts-export-control-order-where-can-enterprises-access-it/</guid>
<pubDate>Wed, 01 Jul 2026 18:03:33 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Anthropic is <a href="https://www.anthropic.com/news/redeploying-fable-5">restoring global access </a>to its most powerful generally released AI model yet, Claude Fable 5, today,  after the U.S. Department of Commerce withdrew emergency export controls that led the company to <a href="https://venturebeat.com/technology/anthropic-blocks-all-public-access-to-claude-fable-5-mythos-5-following-us-government-order-what-enterprises-should-do">suspend all access to both Fable 5 and its less restricted cybersecurity counterpart model Claude Mythos 5</a> last month, just days after both models were initially introduced. </p><p>Starting today, Fable 5 is available for users globally across the primary Anthropic ecosystem, including the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. However, when VentureBeat tried to access it in Claude Code on Terminal, it still showed as disabled.</p><p>For organizations leveraging cloud hyperscalers, Anthropic says it is moving to re-enable access on Amazon Web Services, Google Cloud, and Microsoft Foundry “as quickly as possible.” So far, VentureBeat's research has been unable to confirm if the models have been restored on these external cloud hyperscaler platforms yet.</p><p>Mythos 5 remains a different case. A<a href="https://x.com/synthwavedd/status/2072103052635451559?s=61&amp;t=2nI-irIukCMlctN6d7atlQ"> letter posted on the social network X </a>allegedly from U.S. Commerce Secretary Howard Lutnick to Anthropic executive Tom Brown says a license is no longer required for the export, reexport, or in-country transfer of Fable<i> and Mythos.</i></p><p>But Anthropic’s own <a href="https://www.anthropic.com/news/redeploying-fable-5">redeployment post on its website </a>says only that Mythos 5 access has been restored for “a set of US organizations,” following government approval on June 26. The company says it is continuing to coordinate with the government to expand access to broader domestic and international partners in its opt-in cybersecurity testing program, <a href="https://venturebeat.com/technology/anthropic-says-its-most-powerful-ai-cyber-model-is-too-dangerous-to-release">Project Glasswing</a>.</p><p>That leaves Mythos 5 in a middle category: legally cleared from the emergency export-control order, but not generally available. The current limit appears to come from Anthropic’s decision to keep Mythos behind a vetted-access model, with the U.S. government still playing a role in approvals, standards and expansion.</p><p>Posting on X, Commerce Secretary Howard Lutnick <a href="https://x.com/howardlutnick/status/2072100729603452965">said</a> Anthropic and the government had “worked closely” to “analyze and approve Fable 5,” while White House Chief of Staff Susie Wiles also <a href="https://x.com/SusieWiles47/status/2072099604481335711">posted</a> on X, framing the decision around U.S. AI leadership and deployment speed.</p><p>Wiles wrote that the United States is the “undisputed winner in the AI race,” adding that the shared priority is to “get the best tech deployed as quickly and safely as possible.”</p><p>The reversal follows concerns from cybersecurity leaders and AI policy experts over the export control order, who argued that the U.S. risked hobbling its own industry while giving Chinese AI labs an opening. Former Facebook security chief Alex Stamos <a href="https://www.aol.com/articles/smart-people-saying-return-anthropics-111539000.html?guccounter=1&amp;guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&amp;guce_referrer_sig=AQAAACxDu4x_4FTNyEqy2d4xTCQyaaDeOvRTwjggSIpCok-dRgvkDZQ_tc_RJvuzEqvn7NDqYyR2BWm5PfnU6hIFjuMxDj-8nalhiKxWb6hkO9f93vUbBN0PzrDkB5FRqE5Hs0hHMnPxJR5gdAAD2G-PCK4SEvYRlcBT0_tnlPU6l-Ey">called</a> the Fable restriction a “huge own goal for the US,” warning that security companies could be driven toward Chinese models, while other critics said the so-called "ad hoc" regulatory intervention made dependence on U.S. AI platforms look like a strategic liability.</p><h2><b>Reminder on Claude Fable 5 pricing</b></h2><p>For chief information and technology officers evaluating the return of the model, the deployment comes with distinct structural conditions and significant financial investments.</p><p>Anthropic is pricing both Fable 5 and Mythos 5 at $10.00 per million input tokens and $50.00 per million output tokens, the most expensive of all frontier models globally.</p><table><tbody><tr><td><p><b>Model</b></p></td><td><p><b>Input ($/1M)</b></p></td><td><p><b>Output ($/1M)</b></p></td><td><p><b>Total ($/1M)</b></p></td><td><p><b>Source</b></p></td></tr><tr><td><p>MiMo-V2.5 Flash</p></td><td><p>$0.10</p></td><td><p>$0.30</p></td><td><p>$0.40</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p>deepseek-v4-flash</p></td><td><p>$0.14</p></td><td><p>$0.28</p></td><td><p>$0.42</p></td><td><p><a href="https://api-docs.deepseek.com/quick_start/pricing">DeepSeek</a></p></td></tr><tr><td><p>deepseek-v4-pro</p></td><td><p>$0.435</p></td><td><p>$0.87</p></td><td><p>$1.305</p></td><td><p><a href="https://api-docs.deepseek.com/quick_start/pricing">DeepSeek</a></p></td></tr><tr><td><p>MiniMax-M3</p></td><td><p>$0.30</p></td><td><p>$1.20</p></td><td><p>$1.50</p></td><td><p><a href="https://platform.minimax.io/subscribe/token-plan?tab=api-enterprise">MiniMax</a></p></td></tr><tr><td><p>LongCat-2.0 — limited-time promo</p></td><td><p>$0.30</p></td><td><p>$1.20</p></td><td><p>$1.50</p></td><td><p><a href="https://longcat.chat/platform/docs/APIPayAsYouGo.html">LongCat</a></p></td></tr><tr><td><p>Gemini 3.1 Flash-Lite</p></td><td><p>$0.25</p></td><td><p>$1.50</p></td><td><p>$1.75</p></td><td><p><a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a></p></td></tr><tr><td><p>Qwen3.7-Plus</p></td><td><p>$0.40</p></td><td><p>$1.60</p></td><td><p>$2.00</p></td><td><p><a href="https://modelstudio.console.alibabacloud.com/ap-southeast-1?tab=doc#/doc/?type=model&amp;url=2840914_2&amp;modelId=qwen3.7-plus&amp;serviceSite=international">Alibaba Cloud</a></p></td></tr><tr><td><p>MiMo-V2.5</p></td><td><p>$0.40</p></td><td><p>$2.00</p></td><td><p>$2.40</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p>LongCat-2.0 — standard</p></td><td><p>$0.75</p></td><td><p>$2.95</p></td><td><p>$3.70</p></td><td><p><a href="https://longcat.chat/platform/docs/APIPayAsYouGo.html">LongCat</a></p></td></tr><tr><td><p>Grok 4.3 (low context)</p></td><td><p>$1.25</p></td><td><p>$2.50</p></td><td><p>$3.75</p></td><td><p><a href="https://docs.x.ai/developers/models/grok-4.3">xAI</a></p></td></tr><tr><td><p>MiMo-V2.5 Pro (≤256K)</p></td><td><p>$1.00</p></td><td><p>$3.00</p></td><td><p>$4.00</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p>Kimi-K2.6</p></td><td><p>$0.95</p></td><td><p>$4.00</p></td><td><p>$4.95</p></td><td><p><a href="https://platform.kimi.ai/docs/pricing/chat-k26">Moonshot AI</a></p></td></tr><tr><td><p>GLM-5.2</p></td><td><p>$1.40</p></td><td><p>$4.40</p></td><td><p>$5.80</p></td><td><p><a href="https://docs.z.ai/guides/overview/pricing">Z.ai</a></p></td></tr><tr><td><p>GPT-5.6 Luna</p></td><td><p>$1.00</p></td><td><p>$6.00</p></td><td><p>$7.00</p></td><td><p><a href="https://openai.com/index/previewing-gpt-5-6-sol/">OpenAI</a></p></td></tr><tr><td><p>Grok 4.3 (high context)</p></td><td><p>$2.50</p></td><td><p>$5.00</p></td><td><p>$7.50</p></td><td><p><a href="https://docs.x.ai/developers/models/grok-4.3">xAI</a></p></td></tr><tr><td><p>MiMo-V2.5 Pro (&gt;256K)</p></td><td><p>$2.00</p></td><td><p>$6.00</p></td><td><p>$8.00</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p>Qwen3.7-Max</p></td><td><p>$2.50</p></td><td><p>$7.50</p></td><td><p>$10.00</p></td><td><p><a href="https://modelstudio.console.alibabacloud.com/ap-southeast-1?tab=doc#/doc/?type=model&amp;url=2840914_2&amp;modelId=qwen3.7-max&amp;serviceSite=international">Alibaba Cloud</a></p></td></tr><tr><td><p>Gemini 3.5 Flash</p></td><td><p>$1.50</p></td><td><p>$9.00</p></td><td><p>$10.50</p></td><td><p><a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a></p></td></tr><tr><td><p>Gemini 3.1 Pro Preview (≤200K)</p></td><td><p>$2.00</p></td><td><p>$12.00</p></td><td><p>$14.00</p></td><td><p><a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a></p></td></tr><tr><td><p>GPT-5.6 Terra</p></td><td><p>$2.50</p></td><td><p>$15.00</p></td><td><p>$17.50</p></td><td><p><a href="https://openai.com/index/previewing-gpt-5-6-sol/">OpenAI</a></p></td></tr><tr><td><p>GPT-5.4</p></td><td><p>$2.50</p></td><td><p>$15.00</p></td><td><p>$17.50</p></td><td><p><a href="https://openai.com/api/pricing/">OpenAI</a></p></td></tr><tr><td><p>Gemini 3.1 Pro Preview (&gt;200K)</p></td><td><p>$4.00</p></td><td><p>$18.00</p></td><td><p>$22.00</p></td><td><p><a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a></p></td></tr><tr><td><p>Claude Opus 4.8</p></td><td><p>$5.00</p></td><td><p>$25.00</p></td><td><p>$30.00</p></td><td><p><a href="https://platform.claude.com/docs/en/about-claude/pricing">Anthropic</a></p></td></tr><tr><td><p>GPT-5.5</p></td><td><p>$5.00</p></td><td><p>$30.00</p></td><td><p>$35.00</p></td><td><p><a href="https://openai.com/api/pricing/">OpenAI</a></p></td></tr><tr><td><p>GPT-5.5 Instant (chat-latest)</p></td><td><p>$5.00</p></td><td><p>$30.00</p></td><td><p>$35.00</p></td><td><p><a href="https://developers.openai.com/api/docs/models/chat-latest">OpenAI</a></p></td></tr><tr><td><p>Sakana Fugu Ultra (≤272K)</p></td><td><p>$5.00</p></td><td><p>$30.00</p></td><td><p>$35.00</p></td><td><p><a href="https://console.sakana.ai/pricing#subscription-plan">Sakana AI</a></p></td></tr><tr><td><p>GPT-5.6 Sol</p></td><td><p>$5.00</p></td><td><p>$30.00</p></td><td><p>$35.00</p></td><td><p><a href="https://openai.com/index/previewing-gpt-5-6-sol/">OpenAI</a></p></td></tr><tr><td><p><b>Claude Fable 5 / Claude Mythos 5</b></p></td><td><p><b>$10.00</b></p></td><td><p><b>$50.00</b></p></td><td><p><b>$60.00</b></p></td><td><p><b></b><a href="https://platform.claude.com/docs/en/about-claude/models/overview"><b>Anthropic</b></a></p></td></tr></tbody></table><p>However, to incentivize immediate enterprise adoption following the export control order disruption saga, Anthropic is executing a temporary rollout plan through July 7. </p><p>For Pro, Max, Team, and select Enterprise subscriptions, Fable 5 usage will be included at no added cost for up to 50% of a user’s weekly tier allowance.</p><p>However, after July 7, Fable 5 will move to usage credits for those plans. For standard Enterprise seats, there is no included Fable 5 allowance; all usage is billed through credits, and the model will not work for those users unless credits are enabled.</p><p>Already, some AI influencers are attempting to offer enterprises and developers guidance on how to maximize their usage of Fable 5 during its 7-day discounted price/subscription included promotion:</p><div></div><h2><b>Chronology of a Crisis: From Launch to Lockout</b></h2><p>The whiplash regulatory cycle surrounding the model underscores the volatility currently facing enterprise software supply chains. The crisis unfolded over a rapid, three-week timeline:</p><ul><li><p><b>June 9, 2026:</b> <a href="https://venturebeat.com/technology/anthropic-brings-mythos-to-the-masses-with-claude-fable-5-its-most-powerful-generally-available-model-ever">Anthropic launches Claude Fable 5 and Mythos 5</a>. Early corporate case studies report major performance gains. For instance, Stripe reports that Fable 5 compressed a codebase-wide migration across a 50-million-line Ruby infrastructure into a single day — a project estimated to take a team more than two months by hand.</p></li><li><p><b>June 12, 2026:</b> At 5:21 PM ET, the U.S. government issues an export-control directive citing national security authorities. The order bans access to the models by any foreign national, whether inside or outside the borders of the United States. Lacking real-time mechanisms to verify user nationality at the API layer, Anthropic is forced to pull the plug for all customers to ensure compliance. Anthropic says access to all other Anthropic models was not affected.</p></li><li><p><b>June 13–25, 2026:</b> Enterprise users and developers face abrupt disruption, forcing workflows that had adopted Fable 5 or Mythos 5 to fall back to older models such as Opus 4.8. Tensions peak as Anthropic publicly objects, arguing that pulling a major commercial model over a narrow jailbreak finding could “essentially halt all new model deployments for all frontier model providers.”</p></li><li><p><b>June 26, 2026:</b> The U.S. government allows Anthropic to restore Mythos 5 access to a set of trusted U.S. organizations, partially reversing the June 12 order. Anthropic says it is restoring access for those organizations and continuing to work with the government to expand Mythos 5 access and make Fable 5 generally available again.</p></li><li><p><b>June 30, 2026:</b> Commerce Secretary Howard Lutnick sends a letter withdrawing the June 12 export-control license requirement for both Mythos and Fable. The decision removes the emergency legal block, but Anthropic’s rollout still treats the models differently: Fable 5 returns globally, while Mythos 5 remains limited to approved users through Glasswing and related trusted-access channels.</p></li></ul><h2><b>The Technical Catalyst: The Amazon Vulnerability Report</b></h2><p>The swift intervention by the federal government stemmed from a report by Amazon researchers describing a method for bypassing Fable 5’s safeguards. This was a brutal irony for Anthropic, given Amazon was one of the startup's initial and largest backers to the <a href="https://venturebeat.com/ai/amazon-doubles-down-on-anthropic-positioning-itself-as-a-key-player-in-the-ai-arms-race">tune of $8 billion</a>, and the two companies previously collaborated on <a href="https://www.anthropic.com/news/claude-and-alexa-plus">improving Amazon's Alexa+ voice assistant</a>.</p><p>According to Anthropic, the technique prompted Fable 5 to identify software vulnerabilities; in one case, the model produced code demonstrating how the relevant vulnerability could be exploited.</p><p>When the report reached government officials, it triggered alarm regarding the offensive cyber capabilities of public LLMs. Anthropic countered that the exploit did not tap into unique “Mythos-level” cyber capabilities, noting that its own testing found other models — including Claude Opus 4.8, OpenAI’s GPT-5.5, and Moonshot’s Kimi K2.7 — could identify the same vulnerabilities. Anthropic also said every model it tested could produce the same exploit demonstration as Fable 5.</p><p>To break the regulatory logjam, Anthropic developed an improved automated safety classifier specifically trained to catch and neutralize the Amazon technique. Tested by the Commerce Department’s Center for AI Standards and Innovation (CAISI), the updated classifier successfully halts that specific technique in more than 99% of cases.</p><p>Anthropic explicitly warns enterprise clients that this safety enforcement comes at an operational cost. Because the new classifiers require an expanded “safety margin” to catch ambiguous edge cases, benign coding and debugging requests may be flagged more often. When a prompt is blocked by the safety layer, the active session automatically downgrades, routing the request to Opus 4.8.</p><h2><b>Backroom Diplomacy: The Shifting of the Guard</b></h2><p>The breakthrough that brought Fable 5 back to commercial markets was as much political as it was technical. According to <a href="https://www.wired.com/story/trump-administration-lifts-export-controls-on-anthropics-mythos-and-fable-ai-models/">WIRED</a>, Anthropic initially argued that the administration’s security concerns were overblown and that no frontier model provider could guarantee zero jailbreaks.</p><p>That argument frustrated the administration, according to WIRED’s reporting. In recent weeks, Anthropic changed tack, focusing less on the theoretical impossibility of eliminating jailbreaks and more on building stronger safeguards and satisfying the government’s operational concerns.</p><p><a href="https://www.wired.com/story/the-trump-white-house-is-over-anthropics-dario-amodei/">WIRED reported</a> that Anthropic CEO Dario Amodei was recently replaced in meetings by Brown, whom officials liked more personally. Brown is also the addressee of Lutnick’s June 30 Commerce letter.</p><p>Under Brown’s guidance, Anthropic appears to have moved from arguing over the absolute limits of model safety to committing to the expanded safeguards and collaboration framework the administration demanded.</p><p>The resulting Commerce letter describes several commitments by Anthropic. Under the terms of the clearance, Anthropic has agreed to:</p><ol><li><p>Proactively detect and address security risks associated with the models.</p></li><li><p>Work with the U.S. government on protocols, standards and releases for Mythos, Fable and future models.</p></li><li><p>Inform the U.S. government of malicious activity.</p></li></ol><p>Separately, Anthropic says it will expand pre-release government access and evaluation for frontier models, share information rapidly when significant jailbreaks or misuse patterns are identified, dedicate resources to joint government research and work toward a common industry security bar.</p><p>The U.S. Commerce Department explicitly reserved the right to re-evaluate these permissions and re-impose license requirements if circumstances change or if Anthropic fails to meet its commitments.</p><h2><b>The Sovereign Calculus: Lessons for Enterprise AI</b></h2><p>The two-week blackout of Claude Fable 5 exposed the fragility of centralized, closed-API models for modern business infrastructure. It showed that enterprise automation pipelines remain vulnerable to sudden regulatory shifts and vendor compliance mandates.</p><p>The tech community’s response highlights a broader push toward hardware and model sovereignty. Following the initial shutdown, prominent tech figures voiced concerns over this centralization. <a href="https://x.com/AlexFinn?lang=en">AI founder Alex Finn </a>described the Anthropic freeze as a major “wakeup call,” urging developers to invest heavily in local, open-weights infrastructure to insulate operations from federal volatility. As Finn noted on social media:</p><blockquote><p>“No company or government will EVER be able to take away your local models.”</p></blockquote><p>For enterprise architects, the return of Fable 5 demands a balanced approach to deployment:</p><ul><li><p><b>The Frontier Performance Advantage:</b> Utilizing closed models like Fable 5 offers state-of-the-art capabilities across agentic coding, long-context work, document reasoning and multi-step enterprise automation, according to Anthropic’s launch materials and early customer examples.</p></li><li><p><b>The Mitigating Data Trade-Off:</b> Accessing Fable 5 means accepting Anthropic’s mandatory 30-day data retention requirement for covered models. Anthropic says prompts and model completions are retained for at least 30 days by default and then automatically deleted, except when they are part of a safety investigation or must be kept for legal reasons. Highly regulated financial, healthcare and legal groups must evaluate whether this telemetry window complies with their data privacy mandates.</p></li></ul><p>The truth is, enterprises in the U.S. and globally have more options than ever for frontier-class LLMs, especially with the recent launch over the last few months of new, powerful, open weights Chinese alternatives that can be downloaded, run locally or on virtual private clouds, and customized to an enterprise's liking. </p><p><a href="https://venturebeat.com/technology/minimax-m3-debuts-eclipsing-gpt-5-5-and-gemini-3-1-pro-on-key-benchmark-performance-for-just-5-10-of-the-cost">MiniMax M3 </a>pairs frontier-tier coding and agentic performance with a 1 million-token context window and native multimodality. Z.ai’s <a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">GLM-5.2's benchmark results</a> exceed OpenAI's GPT-5.5 on SWE-bench Pro and several long-horizon coding tests, and near Claude Opus 4.8 on FrontierSWE and MCP-Atlas. <a href="https://venturebeat.com/technology/meituan-open-sources-longcat-2-0-the-1-6t-near-frontier-agentic-coding-model-thats-been-leading-openrouter-trained-entirely-on-chinese-chips">Meituan’s LongCat-2.0 </a>is also positioned around enterprise use, with a 1 million-token context window, MIT licensing and strong early developer traction through its Owl Alpha run on OpenRouter — though as we reported, the full weights are still listed as “coming soon.” </p><p>Meanwhile, Anthropic's top domestic rival OpenAI is still struggling to release its latest models broadly due to U.S. government pressure. The company says its <a href="https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov">newest and most powerful models, GPT-5.6 Sol, Terra and Luna</a> — unveiled last week — are starting in a limited preview for a small group of trusted partners after OpenAI previewed the models and their capabilities to the U.S. government and the government requested the rollout be staggered.</p><p>OpenAI says it still plans broader availability, but argued in its <a href="https://openai.com/index/previewing-gpt-5-6-sol/">announcement</a> that this kind of staggered rollout at the government's request "should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them. We are taking this short-term step because we believe it is the strongest path to broader availability in the coming weeks, while we work with the Administration to develop the cyber Executive Order framework and a repeatable process for future model releases."</p><p>The executive order in question, signed by <a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/">President Donald J. Trump on June 2, 2026</a>, calls upon various federal agencies to collaborate on a process for benchmarking and assessing capabilities of new AI models to ensure they are safe and appropriate for wide release, a process supposed to take 30 days (which would seem to indicate the agencies are due to provide their process tomorrow, July 2, 2026.)</p><p>Frontier model launches are starting to look less like ordinary product releases and more like negotiated deployments shaped by U.S. national security review — a shift that could slow American distribution even as Chinese competitors move aggressively through open-weight and lower-cost channels</p><p>To safeguard operations against future regulatory lockouts, enterprise technical leaders are moving toward model-agnostic fallback architectures. </p><p>By deploying proxy layers that can dynamically reroute critical production pipelines from proprietary APIs to locally hosted, open-weights alternatives, businesses can leverage top-tier capabilities without exposing themselves to single-point-of-failure vulnerabilities. </p><p>Fable 5 is officially back online, but the landscape governing its release has been fundamentally transformed.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Morgan Stanley cut its riskiest reconciliation job in half — by making its agents less autonomous]]></title>
<description><![CDATA[Most enterprise AI deployments so far have focused on coding assistants and customer service bots. Morgan Stanley has deployed agents in one of banking's most accuracy-critical, deadline-driven workflows instead — profit and loss (P&L) reconciliation — and cut the work in half. The counterintuiti...]]></description>
<link>https://tsecurity.de/de/3637097/it-nachrichten/morgan-stanley-cut-its-riskiest-reconciliation-job-in-half-by-making-its-agents-less-autonomous/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3637097/it-nachrichten/morgan-stanley-cut-its-riskiest-reconciliation-job-in-half-by-making-its-agents-less-autonomous/</guid>
<pubDate>Wed, 01 Jul 2026 01:47:13 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Most enterprise AI deployments so far have focused on coding assistants and customer service bots. Morgan Stanley has deployed agents in one of banking's most accuracy-critical, deadline-driven workflows instead — profit and loss (P&amp;L) reconciliation — and cut the work in half. The counterintuitive part: it got there by making the system less autonomous, not more.</p><p>Humans stay tightly in the loop, and their decisions are iteratively turned into repeatable rules the system can apply on its own.</p><p>“It's much more like a co-worker than a copilot,” Morgan Stanley Managing Director Todd Johnson said at a recent VB AI Impact event. The internal production agentic system, known as FIXR, goes beyond simple, straightforward "gen AI 1.0" tasks. “We think that's where the opportunity is to really unlock more complex work in the organization.”</p><h2><b>FIXR behind the scenes</b></h2><p>Every trading day, Morgan Stanley’s trade desks handle the important work around transactions such as cash equities or debt investments. </p><p>And, at the end of each of those days, controllers must reconcile P&amp;L across the finance giant’s Finance, Risk, Operations, and Trade Capture systems. All that data must come together, and, perhaps not surprisingly, hundreds of thousands of attributes frequently fail to match. </p><p>Typically, this means controllers must manually investigate each mismatch (or “break”), make decisions on adjustments, then ideally sign off before the number goes to the desk. And all of this while working on a hard morning deadline. </p><p>Previously, this could take up to six hours for a single book. Now, FIXR performs the task in two to three hours, Johnson said. Across the roughly 100 controllers who do this work, that adds up to about 1,500 hours saved per week.</p><p>After nightly P&amp;L calculations complete, the system automatically analyzes “breaks” and proposes resolutions based on learned rules. Several agents work together: </p><ul><li><p>One interprets past guidance to develop start-of-day resolutions.</p></li><li><p>One learns from controller behavior and documents the rules they apply.</p></li><li><p>One converts repeated patterns into durable, automated logic.</p></li></ul><p>Over time, the system can auto-clear certain breaks it’s encountered before, suggest solutions for others that may be less familiar, ask for help when it’s unsure, and flag for human investigation. When items are repeatedly resolved through the same method, it can create firm rules. </p><p>Critically, humans don’t leave the loop, but stay fully in it, he said. They review, approve or correct every recommendation, then feed those decisions back to improve the next run. The agent learns daily from controllers what it gets right and wrong and codifies that knowledge as it iterates. </p><p>“You still preserve that element of human accountability even as you start to automate,” Johnson said. “Over time you'll see more and more of those items resolved in an automatic way.”</p><p>He emphasized that autonomy requires a great deal of trust; enterprises will not see efficiency gains if everyone's checking everything an agent does. </p><p>The human–agent feedback loop was critical to addressing the challenge of controlled, measured, and repeatable automation. “We recognized that all that intelligence that's sitting in the mind of a controller is gonna be difficult to get all into an agent on day one,” Johnson said. </p><h2><b>Focus on process-first, extensibility</b></h2><p>It was critical to establish processes first, before getting any AI involved, Johnson said. His team ran a “very thorough” process intelligence assessment that mapped and mined workflows to identify where automation would be the most advantageous: Was the answer agents, traditional automation, or simple re-engineering of an inefficient step? </p><p>“If we can fix that first before we add agents to the problem, then we really will be transforming the opportunity,” he said. </p><p>The P&amp;L sign-off process was full of manual steps suitable for automation, and agents taking over some of these time-consuming tasks are freeing up controllers for “more value-added analysis” and “deeper risk consideration” work, he said. </p><p>Extensibility, though, was just as important as time savings. Johnson’s team chose this particular P&amp;L reconciliation use case because hundreds of controllers were doing this work globally across the business (in the Americas, Europe, Asia). </p><p>So start with a use case, prove it, extend it, “and then ultimately the transformation will be as we roll this out more and more across the organization,” Johnson said. </p><h2>Deterministic by design</h2><p>Johnson said the team also deliberately limited how much of the workflow depended on the model's judgment at all. "If you have an opportunity to make things very prescribed and repeatable, that's cheaper in terms of token consumption, it's more repeatable in terms of controls — and have the LLM do the stuff where you don't need that kind of deterministic workflow," he said. </p><p>As the system sees more controller feedback on a given break type, Morgan Stanley converts that pattern into a fixed rule instead of leaving it to the model.</p><h2><b>Humans still own the behavior </b></h2><p>An interesting (and perhaps fundamental) question being raised at the dawn of the agentic era is: Are agents code or digital employees?</p><p>Johnson argues that “they're probably a little bit of both,” and, as such, require nuance when it comes to governance and oversight. Technical teams must still be responsible for maintaining protections and guardrails like firewalls or encryption, for instance. </p><p>But there’s a new dynamic around the “performance element”: Humans using agents are responsible for them because it’s aiding their business work. For instance, if a senior controller is working with a junior controller, they don’t just relinquish responsibility because someone is helping them out, Johnson noted. </p><p>“One of our strong principles in our AI governance generally is that there always has to be human accountability, even if there's a degree of automation,” he said. </p><p>But there typically isn’t “one single one person,” and the process is ultimately continuous. To this point, Johnson joked that one “depressing” thing about agentic AI is that it’s going to require ongoing training because models are ever-changing. </p><p>“You're never gonna be able to say: ‘We've done all the evaluation and testing that we need to do. Let's just let it go.’ You're going to have to have a constant view as it evolves over time.”</p><h2>Morgan Stanley is aiming at real enterprise pain points</h2><p>Morgan Stanley's experience mirrors patterns VentureBeat has uncovered across enterprise AI deployments. </p><p>In VentureBeat's recent VB Pulse survey, nearly three-quarters of respondents reported seeing little to no ROI from custom model fine-tuning, describing a "sandbox graveyard" of AI projects that proved too costly to maintain. This suggests that Morgan Stanley's process-first, buy-and-blend approach may be more sustainable than chasing bespoke models. The survey had 87 respondents and findings should be considered directional. </p><p>Governance emerged as another common challenge: 38% of respondents cited the lack of a single accountable owner as their biggest barrier to production AI, while only two of the 87 enterprises surveyed had active monitoring and alerting in place to detect model failures. </p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic launches Claude Sonnet 5 at a steep discount to its top model as the company races toward a blockbuster IPO]]></title>
<description><![CDATA[Anthropic today released Claude Sonnet 5, a new AI model that the company says delivers near-flagship performance at mid-tier prices — a move designed to give cost-conscious enterprise developers access to powerful agentic capabilities just as the San Francisco-based AI lab barrels toward an init...]]></description>
<link>https://tsecurity.de/de/3636601/it-nachrichten/anthropic-launches-claude-sonnet-5-at-a-steep-discount-to-its-top-model-as-the-company-races-toward-a-blockbuster-ipo/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3636601/it-nachrichten/anthropic-launches-claude-sonnet-5-at-a-steep-discount-to-its-top-model-as-the-company-races-toward-a-blockbuster-ipo/</guid>
<pubDate>Tue, 30 Jun 2026 20:32:20 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="https://www.anthropic.com/">Anthropic</a> today released <a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5</a>, a new AI model that the company says delivers near-flagship performance at mid-tier prices — a move designed to give cost-conscious enterprise developers access to powerful agentic capabilities just as the San Francisco-based AI lab barrels toward an initial public offering that will test whether the private market's staggering AI valuations can survive public scrutiny.</p><p>The release, which Anthropic describes as "<a href="https://www.anthropic.com/news/claude-sonnet-5">the most agentic Sonnet model ye</a>t," makes Sonnet 5 the default model for users on Anthropic's Free and Pro plans, while also making it available to Max, Team, and Enterprise customers. Introductory <a href="https://platform.claude.com/docs/en/about-claude/pricing">API pricing</a> is set at $2 per million input tokens and $10 per million output tokens through August 31, after which it rises to $3 and $15 respectively — still well below the $5 input and $25 output pricing of Anthropic's top-of-the-line Opus 4.8.</p><p>The strategic logic is unmistakable: Anthropic is trying to democratize access to capabilities that until very recently only its most expensive models could deliver, while building the kind of broad-based developer adoption that will look attractive in an <a href="https://www.anthropic.com/news/confidential-draft-s1-sec">S-1 filing</a>.</p><h2><b>Sonnet 5 benchmarks show the mid-tier model closing in on Anthropic's flagship Opus</b></h2><p><a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> posts major gains over its predecessor, <a href="https://www.anthropic.com/news/claude-sonnet-4-6">Sonnet 4.6</a>, across every evaluation Anthropic disclosed. On <a href="https://www.swebench.com/">SWE-bench Pro</a>, an agentic coding benchmark, Sonnet 5 scores 63.2% compared with Sonnet 4.6's 58.1% — a jump that brings it within striking distance of Opus 4.8's 69.2%. On <a href="https://www.tbench.ai/">Terminal-Bench 2.1</a>, another coding evaluation, the gap narrows further: 80.4% for Sonnet 5 versus 67.0% for Sonnet 4.6 and 82.7% for Opus 4.8.</p><p>In multidisciplinary reasoning, as measured by <a href="https://agi.safe.ai/">Humanity's Last Exam</a>, Sonnet 5 scores 43.2% without tools and 57.4% with tools — the latter figure essentially matching Opus 4.8's 57.9%. On computer use tasks evaluated through OSWorld-Verified, Sonnet 5 reaches 81.2%, up from 78.5%. And on <a href="https://artificialanalysis.ai/evaluations/gdpval-aa">GDPval-AA v2</a>, a knowledge-work benchmark, it scores 1,618 — surpassing Opus 4.8's 1,615 and far exceeding Sonnet 4.6's 1,395.</p><p>The pattern across these evaluations tells a consistent story: <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> doesn't merely inch forward from its predecessor. It vaults into a performance tier that overlaps substantially with Anthropic's flagship model, while costing roughly 60% less per token at standard pricing and even less during the introductory period.</p><h2><b>Enterprise partners say Sonnet 5's agentic AI capabilities finish jobs that previous models abandoned</b></h2><p>The emphasis on agentic capabilities — the ability to plan, use tools like browsers and terminals, and execute multi-step workflows autonomously — reflects where the AI industry's center of gravity has shifted in 2026. Enterprises are no longer simply asking chatbots questions; they are deploying AI systems that can navigate complex software environments, execute multi-step coding tasks, and operate with minimal human supervision.</p><p>Early access partners painted a picture of a model that doesn't just start tasks but finishes them. Sualeh Asif, co-founder of Cursor, the AI-powered code editor that has become a bellwether for developer tool adoption, said that "with Claude Sonnet 5, agents stay on plan, follow our conventions, and ship clean multi-step changes, all at an efficient cost." Daniel Shepard, a senior engineer at Zapier, described handing the model a two-part automation job — updating Salesforce account tiers and sending a launch announcement — that "used to stall halfway" with previous models but now completes end to end.</p><p>These testimonials matter because they describe exactly the kind of reliability gap that has kept many enterprises from moving agentic AI from pilot programs to production deployments. A model that gets 80% of the way through a complex task before stalling creates more problems than it solves; one that reliably completes the full workflow changes the economics of automation. Anthropic also introduced cost-performance curves showing that developers can now adjust effort levels across <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> and <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a> to find the optimal balance of cost and accuracy for their specific use case — a granularity that reflects growing sophistication in how enterprises consume AI services.</p><h2><b>An updated tokenizer boosts Sonnet 5 performance but could quietly raise costs for some workloads</b></h2><p>One technical detail <a href="https://www.anthropic.com/news/claude-sonnet-5">buried in the announcement's footnotes</a> deserves attention: Sonnet 5 uses an updated tokenizer that changes how the model processes text, similar to the change Anthropic introduced with Opus 4.7.</p><p>The tradeoff is that the same input can map to roughly 1.0 to 1.35 times as many tokens depending on content type. Anthropic says the introductory pricing is calibrated to make the transition "roughly cost-neutral," but enterprise customers running high-volume workloads will want to benchmark their specific use cases carefully before assuming their bills won't change.</p><h2><b>Anthropic says Sonnet 5 is safer than its predecessor, but its most capable models still lead on alignment</b></h2><p>Anthropic's safety disclosures reveal a nuanced picture. The company reports that <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> shows lower rates of hallucination and sycophancy than <a href="https://www.anthropic.com/news/claude-sonnet-4-6">Sonnet 4.6</a>, is better at refusing malicious requests, and is more resistant to prompt injection attacks in agentic contexts. On Anthropic's automated behavioral audit — which tests for a wide range of misaligned behaviors including cooperation with misuse and deception — Sonnet 5 scored lower (meaning safer) overall than Sonnet 4.6.</p><p>However, Sonnet 5 showed "somewhat higher rates of misaligned behavior" compared with the more capable <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a> and Anthropic's <a href="https://www.anthropic.com/claude/mythos">Claude Mythos Preview</a>, the company's powerful but tightly restricted cybersecurity-focused model. On a Firefox 147 exploit development evaluation created in collaboration with Mozilla, neither Sonnet model could develop a working exploit — both scored 0.0% — though Sonnet 5 showed a slightly higher partial success rate (13.2%) than Sonnet 4.6 (8.8%). Both remain far below Opus 4.8 (68.8% working exploits) and Mythos 5 (88.4%).</p><p>Because of these incremental gains in cyber-adjacent capabilities, Anthropic launched Sonnet 5 with cyber safeguards enabled by default — real-time systems that detect and block dangerous cybersecurity usage. The safeguards mirror those on Opus 4.7 and 4.8 but are less restrictive than those applied to <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Fable 5</a>, the latest Mythos-class model that <a href="https://www.bloomberg.com/news/videos/2026-06-10/the-opening-trade-6-10-2026-video">Bloomberg reported</a> on June 10 is "blocked from responding to queries related to cybersecurity and biology." Organizations enrolled in <a href="https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude">Anthropic's Cyber Verification Program</a> automatically receive the same access on Sonnet 5 without needing to reapply.</p><h2><b>From $14 billion to $47 billion in revenue: Sonnet 5 arrives as Anthropic's IPO narrative takes shape</b></h2><p>The <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> launch arrives at what may be the most consequential moment in Anthropic's short history. The company confidentially filed its IPO prospectus with the SEC in early June, setting up what CNBC has described as "<a href="https://www.cnbc.com/2026/06/05/tech-download-anthropic-ipo-ai-valuations.html">the most scrutinized public offering in tech history</a>."</p><p>The financial trajectory has been extraordinary. In February, Anthropic raised $30 billion at a <a href="https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation">$380 billion valuation</a>, with the company reporting $14 billion in annualized revenue that had "grown more than tenfold in each of the past three years," as <a href="https://www.theguardian.com/technology/2026/feb/12/anthropic-funding-round">The Guardian reported</a>. </p><p>By late May, Anthropic had closed a <a href="https://www.anthropic.com/news/series-h">$65 billion Series H round at a $965 billion</a> post-money valuation — co-led by Altimeter Capital, Sequoia Capital, and others — with a revenue run rate that had crossed $47 billion. Harrison Rolfes, an analyst at PitchBook, <a href="https://www.cnbc.com/2026/06/05/tech-download-anthropic-ipo-ai-valuations.html">told CNBC</a> that the number that will "either validate or collapse the entire narrative the private markets have been pricing for three years" won't be the valuation or revenue, but gross margin — a figure no outside observer has yet seen.</p><p>In this context, <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> serves a dual purpose. For developers, it offers genuine capability improvements at competitive prices. For Anthropic's IPO narrative, it demonstrates the company can deliver a compelling product at a price tier that could drive the kind of broad adoption Wall Street rewards — high-volume, recurring API revenue from thousands of enterprise customers.</p><h2><b>Government deals and growing competition define the market Sonnet 5 enters</b></h2><p>The timing also aligns with Anthropic's aggressive push into institutional contracts. Just yesterday, California Governor Gavin Newsom announced a first-of-its-kind partnership providing <a href="https://www.gov.ca.gov/2026/06/29/governor-newsom-announces-a-first-of-its-kind-partnership-providing-anthropic-tools-to-state-agencies-and-improving-services-for-californians/">Claude to all state agencies at a 50% discount</a>, with free workforce training.</p><p>Kate Jensen, Anthropic's Head of Americas, called it an effort to "put Claude to work for the people who keep this state running." The deal — which extends to California's cities and counties — represents exactly the kind of durable, recurring adoption that could anchor revenue well beyond the developer community.</p><p>But Anthropic's release lands in an increasingly crowded field. OpenAI, which <a href="https://openai.com/index/accelerating-the-next-phase-ai/">raised a $122 billion round in March</a> at an $852 billion valuation, is pursuing its own IPO. Elon Musk's SpaceX, which merged with xAI, priced its IPO at <a href="https://www.cnbc.com/2026/06/03/spacex-ipo-stock-price-roadshow-musk.html">$135 per share with a $1.77 trillion valuation</a>. Google, Meta, and a growing wave of well-funded competitors — including Asian AI startups that, as the Wall Street Journal has reported, are developing Mythos-like cybersecurity capabilities — are all vying for the same enterprise market.</p><p>Gil Luria, head of technology research at D.A. Davidson, told CNBC that while Anthropic "<a href="https://www.cnbc.com/2026/06/05/tech-download-anthropic-ipo-ai-valuations.html">appears to have the lead</a>" in frontier AI models, "much of their current usage is for trials and experimentation and that may not sustain." That observation cuts to the heart of the challenge facing every frontier AI lab: converting experimental developer usage into durable, production-grade revenue.</p><h2><b>The real test for Sonnet 5 isn't benchmarks — it's whether cheaper AI can sustain a trillion-dollar story</b></h2><p>Sonnet 5's positioning — offering near-Opus performance at Sonnet prices — is a direct play for that conversion. Enterprise customers experimenting with expensive Opus-class models may find that Sonnet 5 delivers sufficient quality for production workloads at a price point that finance teams can approve at scale. If it works, it could accelerate the shift from experimentation to deployment that every AI company needs to justify its valuation.</p><p>Three things will determine whether <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> matters beyond the initial benchmark charts. Real-world agentic reliability is the first: benchmarks measure capability, but production deployments measure consistency, and the true test will come when thousands of developers push the model through messy, unpredictable workflows at scale.</p><p>The tokenizer economics are the second: the updated tokenizer's 1.0 to 1.35x token expansion could quietly erode the pricing advantage for certain workloads, and enterprise customers should run their own cost analyses rather than relying on headline per-token prices. The third is the IPO narrative itself: when Anthropic's S-1 eventually becomes public, investors will scrutinize whether the Sonnet tier — cheaper but high-volume — or the Opus tier — expensive but high-margin — drives the bulk of revenue and, critically, gross profit.</p><p>As <a href="https://www.cnbc.com/2026/06/05/tech-download-anthropic-ipo-ai-valuations.html">PitchBook's Rolfes told CNBC</a>, the 2026 IPO window "either becomes the most consequential IPO cycle since the dot-com era or the most expensive lesson in narrative-versus-fundamentals that public markets have ever taught."</p><p>Anthropic is betting that a model good enough to rival its flagship and cheap enough to run at scale is the product that closes the gap between those two outcomes. The public markets will soon decide whether they agree.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Microsoft MCP server gives AI assistants access to MSBuild logs]]></title>
<description><![CDATA[Microsoft has introduced the Microsoft Binlog MCP Server, which gives AI assistants like GitHub Copilot direct access to MSBuild (.binlog) files. The Model Context Protocol server enables AI-powered build investigation through natural language conversation, Microsoft said. 



Introduced June 17 ...]]></description>
<link>https://tsecurity.de/de/3636127/ai-nachrichten/microsoft-mcp-server-gives-ai-assistants-access-to-msbuild-logs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3636127/ai-nachrichten/microsoft-mcp-server-gives-ai-assistants-access-to-msbuild-logs/</guid>
<pubDate>Tue, 30 Jun 2026 17:34:21 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Microsoft has introduced the Microsoft Binlog MCP Server, which gives AI assistants like <a href="https://www.infoworld.com/article/3609013/github-copilot-everything-you-need-to-know.html">GitHub Copilot</a> direct access to MSBuild (.binlog) files. The <a href="https://www.infoworld.com/article/4029634/what-is-model-context-protocol-how-mcp-bridges-ai-and-external-services.html">Model Context Protocol</a> server enables AI-powered build investigation through natural language conversation, Microsoft said. </p>



<p>Introduced <a href="https://devblogs.microsoft.com/dotnet/msbuild-binlog-mcp-server/">June 17</a> and currently in a preview stage, the Microsoft Binlog MCP Server parses <code>.binlog</code><strong> </strong>files and exposes 15 specialized tools that enable AI-driven diagnosis, property tracing, performance analysis, and build comparison. Microsoft said that AI assistants gain the ability to do the following:</p>



<ul class="wp-block-list">
<li>Investigate build failures by querying errors, warnings, and full project/target/task context</li>



<li>Trace property origins to understand where a property got its value</li>



<li>Analyze performance bottlenecks by identifying the slowest projects, targets, and tasks</li>



<li>Compare two builds to spot differences in packages and properties</li>



<li>Read embedded source files captured during the build</li>
</ul>



<p>Instead of manually scrolling through the <a href="https://msbuildlog.com/" target="_blank" rel="noreferrer noopener">MSBuild Structured Log Viewer</a>, Microsoft said that developers can ask their AI assistant questions like “Why did my build fail?” or “What’s making my build slow?” MSBuild’s binary logs contain detailed information about a build including every property evaluation, target execution, task invocation, error, and warning. Thus navigating that data manually can be overwhelming, especially when debugging a complex multi-project solution. Tapping an AI coding assistant for these investigations can save significant time and effort. </p>



<p>Microsoft noted that the easiest way to get started with the Microsoft Binlog MCP server is through the <a href="https://github.com/dotnet/skills">.NET Agent Skills repository</a>. The <code>dotnet-msbuild</code> plugin, which is available for Visual Studio, Visual Studio Code, and terminal-based AI assistants such as GitHub Copilot CLI and Claude Code, bundles the Microsoft Binlog MCP Server along with curated skills and agents for MSBuild build investigation and optimization. </p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[AI is exposing the real limits of enterprise cloud strategy]]></title>
<description><![CDATA[Across the global corporations, I advise, in financial services, healthcare, retail and the public sector, the same crisis surfaces in leadership meetings. Executives approved a bold AI roadmap. Cloud spending climbed 40, 50, even 70 percent. And yet the AI workloads that made perfect sense in th...]]></description>
<link>https://tsecurity.de/de/3635329/it-security-nachrichten/ai-is-exposing-the-real-limits-of-enterprise-cloud-strategy/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3635329/it-security-nachrichten/ai-is-exposing-the-real-limits-of-enterprise-cloud-strategy/</guid>
<pubDate>Tue, 30 Jun 2026 13:06:15 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Across the global corporations, I advise, in financial services, healthcare, retail and the public sector, the same crisis surfaces in leadership meetings. Executives approved a bold AI roadmap. Cloud spending climbed 40, 50, even 70 percent. And yet the AI workloads that made perfect sense in the boardroom presentation now stall, overshoot their budgets or collapse under production load before they reach real users.</p>



<p>I am writing this just after the spring 2026 conference season, and the signal from <a href="https://cloud.google.com/blog/topics/google-cloud-next/google-cloud-next-2026-wrap-up" rel="nofollow">Google Cloud Next</a>, <a href="https://news.microsoft.com/build-2026/" rel="nofollow">Microsoft Build</a>, and a run of <a href="https://aws.amazon.com/events/summits/" rel="nofollow">AWS summits</a> only sharpens the point. Over the past several weeks the industry shipped, in production form, the infrastructure to run and govern AI at scale. What most enterprises still lack is the operating model to decide how to use it.</p>



<p>The problem is not the AI models. The models work. The problem is that organizations built their AI ambitions on cloud strategies designed for a world that no longer exists: strategies built for SaaS applications, predictable traffic and linear cost curves. AI workloads break all three assumptions at once.</p>



<h2 class="wp-block-heading">Why AI breaks traditional cloud assumptions</h2>



<p>For a decade, cloud-first served enterprises well. It delivered elasticity, reduced capital expenditure and democratized access to compute, because enterprise workloads were predictable: web applications, ERP systems, databases and analytics pipelines that scaled smoothly and billed in ways finance could model on a spreadsheet. GenAI and agentic AI change every one of those assumptions at once.</p>



<p>When organizations move AI into production, real inference, retrieval pipelines, vector search and real-time decisioning, the cloud equation breaks in at least five ways:</p>



<ol class="wp-block-list">
<li>Training clusters demand power densities far above standard compute.</li>



<li>Inference needs millisecond latency that network geography can defeat.</li>



<li>Vector databases generate cost spikes invisible in standard billing.</li>



<li>Agentic workloads chain hundreds of tool calls with cascading dependencies.</li>



<li>And data-sovereignty rules constrain where any of them can run.</li>
</ol>



<p>In short, what works at the platform level fails at the workload level.</p>



<p>The costs are the first thing to surprise leaders, because they hide. <a href="https://www.cloudzero.com/blog/ai-cost-management/" rel="nofollow">CloudZero’s analysis</a> and the FinOps teams I work with put it plainly: AI spend surfaces as generic compute, storage and instance line items, rarely labeled “AI.” Three layers drive most of the waste:</p>



<ol class="wp-block-list">
<li>The most visible is LLM API cost, where stateless calls re-send the full conversation history on every request, so a deployment with a couple hundred users can burn many times the token budget in the business case.</li>



<li>The biggest is idle GPU: teams’ provision for peak and then run at 10 to 20 percent utilization, and most miss their AI cost forecasts by more than a quarter.</li>



<li>The most underestimated is the vector database and retrieval layer, where storage I/O, query volume and embedding refresh appear nowhere labeled AI until the bill arrives.</li>
</ol>



<h2 class="wp-block-heading">The dimensions leaders underweight resilience and control</h2>



<p>Cost and latency dominate the conversation. Two dimensions rarely get the same rigor until something breaks:</p>



<ol class="wp-block-list">
<li>Resilience, whether an AI-dependent system can survive failure, degrade gracefully and recover predictably.</li>



<li>Control, who can observe, halt and audit it.</li>
</ol>



<p>AI introduces failure modes that traditional architecture never faced: GPU single points of failure under revenue-critical inference, agentic pipelines that fail mid-execution with no rollback, and models that degrade silently from drift or throttling.</p>



<p>I see the pattern repeated across industries. Organizations design resilience for their traditional applications, then deploy AI on top without asking whether the same guarantees hold. In one global financial services firm I advise, a real-time credit-decisioning model running on a single cloud region took a 47-minute outage during a regional availability event. The halted loan approvals cost more than the system’s entire annual infrastructure budget, and the resilience rework that followed cost several times what designing it in from the start would have. The leaders who avoid this should ask four questions before go-live:</p>



<ol class="wp-block-list">
<li>What happens when the network fails?</li>



<li>What happens when the model degrades?</li>



<li>What happens when an agent executes only halfway?</li>



<li>Who holds the authority to halt and audit?</li>
</ol>



<h2 class="wp-block-heading">What the cloud providers signaled this spring</h2>



<p>The major providers are on track to spend <a href="https://www.statista.com/chart/35046/capital-expenditure-of-meta-alphabet-amazon-and-microsoft/" rel="nofollow">close to $700 billion on AI infrastructure in 2026</a>, roughly three and a half times the 2024 level. Their announcements are strategic signals, not just features. Last year they converged on one message: enterprises cannot run everything in public cloud, so all three built ways to bring their infrastructure into your data center and your sovereign environment. This year the signal advanced a step. They stopped talking about where workloads run and started shipping the layer that governs what agents are allowed to do: identity, containment, auditability and rollback.</p>



<p>Microsoft introduced an “Agent Computer” model with execution containers and machine identity for agents. AWS built <a href="https://aws.amazon.com/blogs/aws/top-announcements-of-aws-reinvent-2025/" rel="nofollow">Amazon Bedrock AgentCore</a> around runtime, memory, identity and auditability. Google shipped an agent gateway and sovereign controls for cross-cloud traffic. As <a href="https://www.bain.com/insights/google_cloud_next_2026_the_agentic_enterprise_control_plane_comes_into_view/" rel="nofollow">Bain observed</a>, agentic AI is now an economics and operations problem, not just a capability problem. The through-line, captured by Microsoft’s own framing, is that AI alone will not change your business; the system running it will. <a href="https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-next-big-shifts-in-ai-workloads-and-hyperscaler-strategies" rel="nofollow">McKinsey’s read</a> is consistent: workloads are becoming more distributed, specialized and operationally demanding, which forces more deliberate infrastructure decisions.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/06/hyperscaler-convergence-spring-2026.png?w=1024" alt="Hyperscaler convergence, Spring 2026." class="wp-image-4190723" width="1024" height="557" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Vipin Jain</p></div>



<h2 class="wp-block-heading">From platform choice to placement decision</h2>



<p>The failure I document most often is not a technology failure; it is a governance failure. Most enterprises lack a clear, repeatable way to decide what runs where, under what conditions and with what tradeoffs. Platform teams make that call informally, under deadline pressure and repeat it hundreds of times as new use cases launch. Workloads then accumulate in public cloud by default, not by design and 30 to 50 percent cost overruns follow, not because public cloud was the wrong choice but because no deliberate choice was ever made.</p>



<p>In one global manufacturer I advise, a predictive-maintenance model went live on public cloud and performed exactly as validated in staging. But real-time inference on the factory floor ran at 80 to 120 milliseconds across the WAN, when the machine-control system needed under ten. Moving the model to edge nodes fixed the latency, but the company lost most of a quarter of the cost, rework and delayed benefits, and the line had run for weeks on stale recommendations: a control failure that could have caused a safety event. The fix was never more AI talent. It was a structured placement decision at the start, weighing six dimensions:</p>



<ul class="wp-block-list">
<li><strong>Latency: </strong>real-time (under 10 ms, edge or on-prem), interactive (50 to 500 ms, cloud) or batch.</li>



<li><strong>Cost and TCO: </strong>token spend, GPU utilization, vector-database queries, egress and unit economics per workload.</li>



<li><strong>Resilience: </strong>failover architecture, degraded-mode behavior, recovery SLA and rollback policy.</li>



<li><strong>Control: </strong>observability, audit trails, governance authority and the ability to halt or reverse.</li>



<li><strong>Data sensitivity: </strong>sovereignty requirements, privacy and compliance rules, and IP protection.</li>



<li><strong>Integration: </strong>legacy system dependencies, pipeline complexity and data-residency constraints.</li>
</ul>



<p>Run consistently, those dimensions produce a placement pattern like this:</p>



<figure class="wp-block-table"><div class="overflow-table-wrapper"><table class="has-fixed-layout"><thead><tr><td><strong>Workload</strong></td><td><strong>Latency</strong></td><td><strong>Cost predictability</strong></td><td><strong>Data sovereignty</strong></td><td><strong>Recommended path</strong></td></tr></thead><tbody><tr><td><strong>Customer-facing chatbot</strong></td><td>200-500 ms</td><td>Medium</td><td>Low risk</td><td>Public cloud, reserved instances</td></tr><tr><td><strong>Real-time fraud detection</strong></td><td>Under 10 ms</td><td>Medium</td><td>High</td><td>On-prem or sovereign private cloud</td></tr><tr><td><strong>Clinical decision support</strong></td><td>100-300 ms</td><td>Predictable</td><td>Critical</td><td>Sovereign cloud or dedicated VPC</td></tr><tr><td><strong>Demand forecasting (batch)</strong></td><td>Hours</td><td>High</td><td>Low risk</td><td>Spot instances or scheduled cloud</td></tr><tr><td><strong>Factory-floor vision AI</strong></td><td>Under 5 ms</td><td>Predictable</td><td>Medium</td><td>Edge node (Azure Local, AWS on-prem)</td></tr><tr><td><strong>Internal knowledge assistant</strong></td><td>1-3 sec</td><td>Variable tokens</td><td>High (IP risk)</td><td>Private cloud with on-prem retrieval</td></tr></tbody></table> </div></figure>



<p>This is no longer optional. <a href="https://www.storagenewsletter.com/2026/03/11/enterprise-survey-finds-93-are-repatriating-ai-workloads-or-evaluating-a-move-away-from-public-cloud/" rel="nofollow">Cloudian’s 2026 enterprise AI infrastructure survey</a> found that 79 percent of enterprises have already moved AI workloads out of public cloud, and 93 percent are repatriating or actively evaluating it, driven by data sovereignty, cost overruns and real-time performance. Repatriation is now the norm, not the exception.</p>



<p>The agentic layer makes discipline urgent. An agent chains 20 to 100 tool calls, each with its own latency, cost and failure mode, so the governance model that works for a chatbot does not work for an autonomous agent approving procurement or onboarding a customer. This spring the providers shipped production infrastructure for exactly this, yet <a href="https://www.deloitte.com/global/en/issues/generative-ai/state-of-ai-in-enterprise.html" rel="nofollow">Deloitte’s 2026 survey</a> of more than 3,000 leaders finds only about one in five companies has a mature governance model for autonomous agents. The platforms solved the mechanism. Most enterprises have not yet written the policy.</p>



<h2 class="wp-block-heading">What the leaders do differently</h2>



<p>The organizations extracting compounding value from AI, not just running experiments, share one discipline: they treat workload placement as a repeatable process, and they build resilience and control in from the start rather than after the first production incident. In practice, they do five things:</p>



<ol class="wp-block-list">
<li>Classify every use case at intake across the six dimensions, before any infrastructure is provisioned.</li>



<li>Separate AI budget lines for experiments, production inference and training, so cost is governable.</li>



<li>Treat unit economics, cost per inference, per query and per agent run, as engineering KPIs, not month-end surprises.</li>



<li>Define repatriation triggers in advance, typically 12 to 18 months of stable volume.</li>



<li>Write an explicit resilience contract, and agentic observability and rollback rules, before scaling.</li>
</ol>



<p>The gap between strategy-ready and infrastructure-ready is the remediation backlog, and most enterprises stall moving from proof of concept to production for exactly this reason. <a href="https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/ai-infrastructure-compute-strategy.html" rel="nofollow">Deloitte’s tech-trends analysis</a> frames the same shift as the move to inference economics: the bottleneck is infrastructure governance, not model capability.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/06/ai-governance.png?w=1024" alt="AI infrastructure maturity: The governance gap." class="wp-image-4190724" width="1024" height="555" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Vipin Jain</p></div>



<p><strong>For CIOs, a 90-day agenda. </strong>Five actions separate the leaders from those managing infrastructure crises:</p>



<ol class="wp-block-list">
<li>Audit every AI workload in production across latency, cost, sovereignty, volume, resilience, control and integration.</li>



<li>Separate AI infrastructure budget lines so each workload type is attributable and governable.</li>



<li>Define unit economics by workload and review them as engineering KPIs.</li>



<li>Set a quantitative repatriation evaluation trigger.</li>



<li>Define observability, cost attribution and rollback policy before scaling agents.</li>
</ol>



<h2 class="wp-block-heading">The strategic reframe</h2>



<p>The organizations making real progress on AI are not distinguished by the sophistication of their models or the size of their cloud contracts. One discipline sets them apart: a clear, repeatable way to decide what runs where, under what conditions, with what tradeoffs and what happens when something fails. That discipline is not an IT problem. It is a strategic capability that requires CIO ownership, CFO alignment and executive accountability.</p>



<p>This spring the cloud providers handed enterprises the infrastructure to run and govern AI, and agents, at every tier of the architecture. The gap is no longer supply. It is the operating model to use deliberately. The companies building that model now build the operating foundation for AI at scale. Everyone else builds a remediation backlog. The infrastructure decisions you make in the next 12 months will decide which of those two you become.</p>



<p><em>This article was made possible by our partnership with the IASA </em><a href="https://chiefarchitectforum.org/" target="_blank" rel="nofollow"><em>Chief Architect Forum</em></a><em>. The CAF’s purpose is to test, challenge and support the art and science of Business Technology Architecture and its evolution over time as well as grow the influence and leadership of chief architects both inside and outside the profession. The CAF is a leadership community of the </em><a href="https://iasaglobal.org/" target="_blank" rel="nofollow"><em>IASA</em></a><em>, the leading non-profit professional association for business technology architects.</em></p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><strong><a href="https://www.cio.com/expert-contributor-network/">Want to join?</a></strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[How the Senate’s AI AGENT Act could reshape enterprise AI governance]]></title>
<description><![CDATA[A proposed US Senate bill governing consumer AI agents could become an early test case for how organizations assess the security, accountability, and governance of autonomous software systems.



The discussion draft, released by Senator Mark Warner, is known as the Artificial Intelligence Access...]]></description>
<link>https://tsecurity.de/de/3635127/it-nachrichten/how-the-senates-ai-agent-act-could-reshape-enterprise-ai-governance/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3635127/it-nachrichten/how-the-senates-ai-agent-act-could-reshape-enterprise-ai-governance/</guid>
<pubDate>Tue, 30 Jun 2026 12:02:28 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>A proposed US Senate bill governing consumer AI agents could become an early test case for how organizations assess the security, accountability, and governance of autonomous software systems.</p>



<p>The discussion <a href="http://chrome-extension//efaidnbmnnnibpcajpcglclefindmkaj/https:/www.warner.senate.gov/wp-content/uploads/2026/06/AI-AGENT-Act-Discussion-Draft-1.pdf" target="_blank" rel="nofollow">draft</a>, released by Senator Mark Warner, is known as the Artificial Intelligence Access, Gatekeeper Exchange, and Nondiscriminatory Transfer Act of 2026, or AI AGENT Act.</p>



<p>Under the proposal, providers of “custodial user agents” would have to register with the Federal Trade Commission before accessing interfaces maintained by large online platforms. The bill defines such agents as software authorized by a user to interact with platforms on the user’s behalf in a transparent, documented, limited, and revocable way.</p>



<p>Large online platforms would be required to support access by approved third-party agents, although they could restrict access when registration requirements are not met, user consent has been revoked, or an agent is associated with repeated harmful activity.</p>



<p>For enterprises, the proposal could mark an early regulatory test for <a href="https://www.computerworld.com/article/4165686/gartner-sees-untamed-growth-in-agentic-ai.html">agentic AI</a> by forcing companies to examine who controls these tools, how their actions are recorded, whether vendors can be trusted, and how much autonomy should be allowed inside corporate workflows.</p>



<p>“As AI agents become more autonomous, enterprises will need clear accountability for decisions and actions taken on behalf of users,” said <a href="https://www.linkedin.com/in/tulikasheel/" target="_blank" rel="nofollow">Tulika Sheel</a>, senior VP at Kadence International. “This aligns with growing expectations around governance, auditability, and responsible AI adoption across organizations.”</p>



<p>The requirement to link AI agents to an authorizing user is more consequential than certification or access revocation because it would change enterprise accountability models and shift how organizations <a href="https://www.computerworld.com/article/4025903/as-ai-agents-go-mainstream-companies-lean-into-confidential-computing-for-data-security.html" target="_blank">manage operational risk</a>, according to <a href="https://www.forrester.com/analyst-bio/biswajeet-mahapatra/BIO20046" target="_blank" rel="nofollow">Biswajeet Mahapatra</a>, principal analyst at Forrester.</p>



<p>“Enterprises can absorb certification requirements through existing supplier review processes and manage permission revocation through the identity and access systems they already use,” Mahapatra said.  </p>



<p>Linking an agent to the user authorizing it would create a continuous traceability requirement for the agent’s actions, he added.</p>



<p>That could force CIOs and CISOs to rethink how they track agent activity and assign responsibility for automated decisions. It could also expand incident response planning to cover actions initiated by AI agents, creating new legal exposure and requiring closer coordination between security, compliance, and business teams.</p>



<p>However, Sanchit Vir Gogia, chief analyst at Greyhound Research, said revocation may prove to be the most consequential issue because companies must be able to define what access is being withdrawn and across which systems.</p>



<p>“A right to revoke means very little until the enterprise can answer what is being revoked, from whom, and across which systems,” Gogia said. “Absent that, revocation is a beautifully engineered red button wired to nothing, which is governance theatre with a dashboard attached.”</p>



<h2 class="wp-block-heading">A benchmark for procurement</h2>



<p>Even though the bill targets consumer services, analysts said an FTC registration requirement could quickly become a de facto benchmark for enterprise procurement.</p>



<p>“A vetted list reduces evaluation effort, accelerates vendor shortlisting, and provides defensible justification for procurement decisions,” Mahapatra said. “In practice, such a list would become a minimum entry requirement in sourcing workflows, with enterprises layering additional internal criteria such as security, data handling, and model governance before final selection.”</p>



<p>Sheel said this framework could influence enterprise buying decisions even if it does not become a formal requirement. “Many organizations already look for independent certifications and regulatory assurance when selecting technology vendors, particularly for emerging technologies,” she said. “While it may not become a mandatory requirement, it could serve as an important trust signal during vendor evaluation.”</p>



<h2 class="wp-block-heading">Balancing access and risk</h2>



<p>The bill could also create technical and legal tension for large platforms that would have to support third-party AI agents while retaining the ability to block risky access.</p>



<p>“The biggest challenge will be balancing openness with security,” Sheel said. “Platforms will need clear and transparent policies to determine when access should be allowed or restricted, without creating uncertainty for developers or users. Finding that balance between interoperability and risk management will be essential to ensure both innovation and user protection.”</p>



<p>Gogia said the bill’s platform-access requirement could also create disputes over whether platforms are blocking agents for legitimate security reasons or to protect their own market position.</p>



<p>“The coming fight is whether security is a genuine shield for users or a convenient moat for incumbents,” Gogia said. “In agentic AI, both can be true in the same dispute, which is why this will be settled in court rather than in commentary.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Kotlin improves compile-time constants]]></title>
<description><![CDATA[Kotlin 2.4.0, an update to JetBrains’s statically typed language for building JVM, native, Wasm, and web applications, introduces experimental improvements to compile-time constants, making support for numeric and string types more consistent and easier to use, JetBrains said. 



Kotlin 2.4.0 wa...]]></description>
<link>https://tsecurity.de/de/3635034/ai-nachrichten/kotlin-improves-compile-time-constants/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3635034/ai-nachrichten/kotlin-improves-compile-time-constants/</guid>
<pubDate>Tue, 30 Jun 2026 11:18:30 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Kotlin 2.4.0, an update to JetBrains’s statically typed language for building JVM, native, Wasm, and web applications, introduces experimental improvements to compile-time constants, making support for numeric and string types more consistent and easier to use, JetBrains said. </p>



<p>Kotlin 2.4.0 was released <a href="https://blog.jetbrains.com/kotlin/2026/06/kotlin-2-4-0-released/">June 3</a>. Its experimental improvements to compile-time constants include support for unsigned type operations; standard library functions for strings, such as the <code>.lowercase()</code>, <code>.uppercase()</code>, and <code>.trim()</code> functions, and evaluation of the <code>.name</code> property of <a href="https://kotlinlang.org/docs/enum-classes.html?_gl=1%2Afehijk%2A_gcl_au%2AMTU3MDQ4ODM0NC4xNzgyNTgwMDg4LjEwMzg2MDE2MzkuMTc4Mjc0OTgwNC4xNzgyNzQ5ODA0%2AFPAU%2AMTc4NTM0NDYwNC4xNzgyNTA0Mzkw%2A_ga%2AMTM5MjU2NDU3OS4xNzgyNTgwMDg4%2A_ga_9J976DJZ68%2AczE3ODI3NDgzODQkbzMkZzEkdDE3ODI3NTEyMTIkajU5JGwwJGgw&amp;_cl=MTsxOzE7aVFoRlVMeXhISDhwV2l2d3lVNmhTMDdKMGdPQzVoMnZBcHI2bWpERjNhRkY5eGpWSWY0bjRNVEJwYzNmdnR1aTs%3D#working-with-enum-constants">enum constants</a> and the <a href="https://kotlinlang.org/api/core/kotlin-stdlib/kotlin.reflect/-k-callable/"><code>KCallable</code> interface</a>. To make it clear which functions are evaluated at compile time, Kotlin 2.4.0 introduces the <code>IntrinsicConstEvaluation</code><strong> </strong>annotation. </p>



<p>JetBrains warned that some functions are evaluated at compile time but do not have the annotation yet. Later releases will add the annotation to the remaining functions.</p>



<p>Also in Kotlin 2.4.0:</p>



<ul class="wp-block-list">
<li>Kotlin 2.4.0 improves export to <a href="https://www.infoworld.com/article/2263137/what-is-javascript-the-full-stack-programming-language.html">JavaScript</a> and <a href="https://www.infoworld.com/article/2257305/what-is-typescript-strongly-typed-javascript.html">TypeScript,</a> including support for exporting value classes, interfaces, and type variance, as well as ES2015 features when inlining JavaScript code.</li>



<li>The Kotlin compiler can generate classes containing <a href="https://www.infoworld.com/article/4168040/whats-new-and-exciting-in-jdk-26.html">Java 26</a> bytecode.</li>



<li>Experimental support is highlighted for the <a href="https://component-model.bytecodealliance.org/">WebAssembly Component Model</a>. The proposal defines a way to build components from Wasm modules through standardized interfaces and types. This approach helps Wasm evolve from a low-level binary instruction format into a system for composing reusable, language-agnostic components.</li>



<li>Kotlin 2.4.0 has been included in the <a href="https://www.jetbrains.com/idea/download/?_cl=MTsxOzE7RVVQYngzTmMwUzNLNmkzTllYbXBVM20xRnFMcG5rdmE5SkxUYmk0emIycXh6Vjg1NThFa2dlZUlNVkdKeGRFZTs%3D&amp;section=mac">IntelliJ IDEA</a> and <a href="https://developer.android.com/studio">Android Studio</a> IDEs. </li>
</ul>



<p>An update to Kotlin 2.4.0 will arrive soon. Beta1 of <a href="https://kotlinlang.org/docs/whatsnew-eap.html">Kotlin 2.4.20</a> was released on June 24, adding the <code>StackTraceRecoverable</code> interface to the standard library. This interface improves integration with the <code>kotlinx.coroutines</code> library because it lets users define how to create exception instances for stack trace recovery without adding a dependency on<code>kotlinx.coroutines</code>, according to JetBrains. </p>



<p>A build tools API in the Kotlin 2.4.20 beta adds support for the Kotlin/JS, Kotlin/Wasm, and Kotlin metadata targets.</p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[U.S. Open powers up AI-ready network in challenging environment]]></title>
<description><![CDATA[Cisco’s work with the USGA at the 2026 U.S. Open at Shinnecock Hills Golf Club was a live testbed for what AI-ready networking and security look like in the wild — not in a lab, not in a climate-controlled data center, but across 18 holes of constantly changing terrain, crowds, and threats. It’s ...]]></description>
<link>https://tsecurity.de/de/3634437/it-security-nachrichten/us-open-powers-up-ai-ready-network-in-challenging-environment/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3634437/it-security-nachrichten/us-open-powers-up-ai-ready-network-in-challenging-environment/</guid>
<pubDate>Tue, 30 Jun 2026 05:19:51 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Cisco’s work with the <a href="https://www.usga.org/">USGA</a> at the 2026 U.S. Open at <a href="https://www.shinnecockhillsgolfclub.org/">Shinnecock Hills Golf Club</a> was a live testbed for what AI-ready networking and security look like in the wild — not in a lab, not in a climate-controlled data center, but across 18 holes of constantly changing terrain, crowds, and threats. It’s also a blueprint that network engineers in other industries can borrow as they grapple with the convergence of connectivity, security, and AI apps.</p>



<h2 class="wp-block-heading">Golf as a worst‑case network environment</h2>



<p>From a distance, it’s tempting to lump golf in with stadium or arena networking. The reality on the ground is very different. Stadiums offer a fixed concrete bowl and predictable RF patterns. A <a href="https://www.usopen.com/">U.S. Open</a> venue is effectively rebuilt every year: temporary structures, new hospitality layouts, shifting fiber routes, and a crowd that never sits still.</p>



<p>Christian Rodriguez, senior manager, IT operations, from the USGA’s technology team, captured that reality when he explained why they tear down and rebuild from scratch: No two championships share the same layout, ISP entry points, or even the placement of critical compounds. They don’t simply clone last year’s configs; they design for the specific course, topology, and constraints of that site. That level of contextual design is expensive, but it’s also the only way to avoid brittle architectures that fall apart as soon as the environment changes.</p>



<p>Environmental conditions add another layer of complexity. Anthony Santora, managing director of IT for the USGA, describes the championship network as a data center without the usual comforts. There’s dust, rain, wind, and wide temperature swings instead of clean, controlled air. Hardware resides in trailers and weatherproof enclosures, not in racks behind raised floor tiles. For network engineers who spend most of their time on office campuses and in colos, that’s an important reminder: Critical infrastructure increasingly sits in places that look nothing like a traditional wiring closet.</p>



<p>User behavior is just as hostile. The U.S. Open has its own term — the “Tiger effect” (though one could argue it’s now the Scottie effect) — for what happens when tens of thousands of fans follow a single golfer. The hot spot moves with the group, and the RF design must cope with a dense, moving cluster of devices. That pattern should sound familiar to anyone who supports large conferences or festivals; it’s the same phenomenon, just under a different name.</p>



<h2 class="wp-block-heading">Building an AI‑ready, fault‑tolerant course network</h2>



<p>Cisco’s answer to this environment is a fully redundant, mobile core design. Instead of a single large core in a building, the network collapses into dual trailers that serve as cores on the go, typically anchored at the NBC broadcast compound and another central location. Each core hosts Cisco Secure Firewall appliances, FMCs, core Catalyst switches, DHCP, UPS, and generators, all in pairs. Rodriguez was matter-of-fact about the philosophy: “We do everything in pairs as much as we can.” If one fails, its twin picks up the load.</p>



<p>From those cores, the team builds a ring topology around the course, using diverse fiber paths — including trenching fiber through wooded areas — to avoid single points of failure. Mobile IDF kits in cooled cabinets serve as distribution points, delivering connectivity to weatherproof access switches and Wi-Fi access points around hospitality tents, grandstands, and entry gates. Everything on the backbone operates at Layer 3, with HSRP (Hot Standby Routing Protocol) and routing redundancy to ensure that a single switch failure doesn’t take out large swaths of the network.</p>



<p>The scale of a golf course deployment is massive as well, with about 500 access points and more than 100 switches, many of them the latest <a href="https://www.networkworld.com/article/4135351/favorable-wi-fi-7-prices-wont-be-around-for-long-delloro-group-warns.html">Wi‑Fi 7</a> and campus platforms. What matters is not the absolute numbers but the duty cycle. Every TV, every POS terminal, every credential pedestal, every media workstation, and every fan device share this converged fabric during a compressed, high‑risk period. Santora points out the business impact in simple terms: If merchandise goes down for even five minutes, lines explode and fans walk away. There’s no “we’ll patch it on the next maintenance window.”</p>



<p>On the RF side, the <a href="https://www.networkworld.com/article/4092389/singapore-makes-the-leap-to-wi-fi-7-to-boost-fan-experience.html">shift to Wi‑Fi 7</a> is more than a speed upgrade. Santora’s team has seen real-world performance improvements — hundreds of megabits down in the middle of a packed media center – but the more important change is resilience under high density. When you combine wider channels, better scheduling, and smarter management with a dense deployment, you get something that can withstand the Tiger Effect and the crush of content creators and broadcasters.</p>



<p>That last group is critical. Rob Neumann from Cisco notes that at these events, upload traffic now dominates download traffic. Influencers, media teams, and fans are publishing in near-real time, and cellular uplink simply can’t keep up. High-capacity Wi-Fi with solid backhaul isn’t a luxury; it’s the only way to avoid a miserable experience for the most vocal, visible part of the audience.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large is-resized"> width="894" height="1024" sizes="auto, (max-width: 894px) 100vw, 894px"&gt;</figure><p class="imageCredit">Zeus Kerravala</p></div>



<h2 class="wp-block-heading">Security: Treat every device as untrusted</h2>



<p>If the connectivity story feels familiar, the security posture at the U.S. Open is where this deployment begins to diverge from more generic “converged stadium” narratives. Santora has to contend with “thousands of untrusted devices” each championship week: fans, vendors, media, broadcasters, and staff, many of whom plug in or connect to networks the USGA doesn’t control outside the event. The USGA is well aware of the risks: outages or breaches could lead to data and financial losses, as well as reputational damage that would undermine the organization’s core mission, not just its IT metrics.</p>



<p>Cisco Secure Firewall, AnyConnect, Duo, and other components form the core security stack, but how they’re used is the differentiator. Fan Wi‑Fi runs with strict isolation: every client is segmented, so lateral movement is essentially off the table. Neumann explains it simply — each fan has an isolated path out — but under the covers, you get VLAN separation, policy enforcement, and inspection that treat fan traffic as untrusted end-to-end.</p>



<p>The rest of the network is equally segmented. There’s a separate network for <a href="https://www.pgatour.com/shotlink">ShotLink</a> and everything “inside the ropes,” including scoring and betting feeds. Back-of-house traffic for staff, concessions, and retail runs on its own network. Remote POS systems are segmented again. Broadcast compounds and production systems have their own paths and policies. The result is a unified, converged physical fabric with tightly controlled logical overlays.</p>



<p>This is a pattern many enterprises discuss but struggle to implement: a single platform that carries many classes of traffic, each with its own risk profile, without collapsing into a flat, lateral-friendly network. The U.S. Open shows that it’s possible — but only if segmentation is treated as a core design principle, not an afterthought.</p>



<h2 class="wp-block-heading">Observability and AI security in the loop</h2>



<p>Security and availability at this scale demand observability. Here again, Santora’s team is in the middle of a transition many enterprises are grappling with: moving from reactive log-scraping to proactive, correlated telemetry.</p>



<p>Instead of manually combing through firewall and switch logs, the USGA and Cisco have built a pipeline into Splunk and Cisco’s observability tools. Neumann describes it as a single pane of glass across the network, but the more important point is what feeds that view: APs, switches, firewalls, cameras, and applications, all instrumented and reporting. When you combine that with full-stack observability, you can spot anomalies in real time, whether they’re performance issues or indicators of compromise.</p>



<p>That observability story extends to AI. One of the headline features of the renewed Cisco–USGA partnership is the AI-powered rules assistant: an application that lets golfers and fans ask complex rules questions in the USGA app and receive near-instant guidance. Under the hood, Santora’s team started with question–answer pairs and built a knowledge graph that now spans hundreds of topics and clusters. They also built an evaluation program that identifies outliers — questions the system struggles with — and feeds them back into human review.</p>



<p>Cisco AI Defense wraps the assistant with security controls. It’s not enough to get rules right; the system must resist prompt injection, data exfiltration, and other AI-specific threats that are increasingly appearing in the wild. The teams monitor usage, validate models, and protect applications at runtime against misuse or abuse. Perhaps most importantly, they keep a human override in place. If the system isn’t confident, it won’t answer; it escalates to rules experts rather than bluffing.</p>



<p>This is a model network engineers should watch as AI assistants and agents proliferate across other industries. The U.S. Open rules assistant isn’t treated as a toy or a sidecar; it’s a mission-critical application that resides within the same protected fabric as POS, scoring, and broadcast and is subject to the same observability and security rigor.</p>



<h2 class="wp-block-heading">Lessons for network engineers beyond golf</h2>



<p>Strip away the golf-specific details, and a set of lessons emerges:</p>



<ul class="wp-block-list">
<li><strong>Design for tough, not ease.</strong> Assume transient structures, unknown RF patterns, seasonal layout changes, and harsh environmental conditions. The U.S. Open team rebuilds from scratch for each venue; most enterprises don’t need to go that far, but they should at least validate designs against real-world changes rather than assuming a static topology.</li>



<li><strong>Make redundancy systemic.</strong> Dual cores, dual firewalls, ring topologies, HSRP, Layer 3 everywhere, spare hardware on site, and live failover drills are all part of the fabric. Redundancy isn’t a checkbox on a data sheet; it’s an operational discipline.</li>



<li><strong>Treat every device as untrusted.</strong> Fan devices, vendor systems, broadcast laptops, and staff phones all arrive with unknown posture. Segmentation — per-client isolation, dedicated networks for sensitive functions, and strong identity — is the only sustainable way to cope with that diversity.</li>



<li><strong>Upload is the new download.</strong> Traditional designs optimized for download traffic are increasingly misaligned with reality. Conferences, stadiums, and campuses now behave like the U.S. Open: content creators and collaborative apps push far more data than they pull. Wi-Fi 7 and modern campus platforms help, but you still need to design RF and backhaul with upload and lateral traffic in mind.</li>



<li><strong>Integrate observability and AI security from day one.</strong> Logs alone aren’t enough. Coherent telemetry, full-stack observability, and AI-focused security controls should be treated as first-class requirements, especially as AI assistants move into business-critical workflows.</li>
</ul>



<h2 class="wp-block-heading">Final thoughts</h2>



<p>Perhaps the most important takeaway is cultural rather than technical. Santora and his team position AI and automation as tools for scale, not as replacements for experts. The rules assistant accelerates responses and expands reach, but it still defers to human judgment when confidence is low. The network uses automation and observability to keep a complex environment running, but it still depends on experienced engineers, in trailers on-site, watching for issues and making decisions.</p>



<p>For network engineers in other industries, that’s a useful template: Build AI-ready, secure, observable networks that assume the worst about their environment, and pair them with human expertise that can adapt when reality inevitably diverges from the design.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large is-resized"> width="1024" height="673" sizes="auto, (max-width: 1024px) 100vw, 1024px"&gt;</figure><p class="imageCredit">Zeus Kerravala</p></div>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[DeepSeek open sources DSpark, a new framework to speed up LLM inference by up to 85%]]></title>
<description><![CDATA[Even as the geopolitical conversation around AI continues to grow more fraught following the U.S. government's actions to limit the new models from Anthropic and OpenAI, Chinese open source darling DeepSeek is back with yet another open release that could once again change AI development around t...]]></description>
<link>https://tsecurity.de/de/3634171/it-nachrichten/deepseek-open-sources-dspark-a-new-framework-to-speed-up-llm-inference-by-up-to-85/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3634171/it-nachrichten/deepseek-open-sources-dspark-a-new-framework-to-speed-up-llm-inference-by-up-to-85/</guid>
<pubDate>Tue, 30 Jun 2026 00:17:42 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Even as the geopolitical conversation around AI continues to grow more fraught following the<a href="https://venturebeat.com/technology/anthropic-blocks-all-public-access-to-claude-fable-5-mythos-5-following-us-government-order-what-enterprises-should-do"> U.S. government's actions to limit the new models from Anthropic</a> and <a href="https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov">OpenAI</a>, Chinese open source darling DeepSeek is back with yet another open release that could once again change AI development around the globe. </p><p>Over the weekend, the firm released <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark">DSpark</a>, a new, MIT-Licensed system designed to make large language models answer faster without changing what the underlying model is trying to say. </p><p>The easiest way to think about it is this: most AI chatbots write like someone crossing a river one stepping stone at a time. They choose one small chunk of text, then the next, then the next. </p><p>DSpark gives the system a scout that runs a few steps ahead, guesses the likely path, and lets the larger model quickly check which steps are safe. When the guesses are good, the model moves faster. When the guesses are weak, DSpark tries not to waste time checking them.</p><p>DeepSeek published the work with a <a href="https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf">technical paper</a>, model checkpoints and <a href="https://github.com/deepseek-ai/DeepSpec">DeepSpec</a>, a codebase for training and evaluating speculative decoding systems. The release is available through DeepSeek’s public <a href="https://github.com/deepseek-ai">GitHub</a> and <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark">Hugging Face </a>pages, both under the permissive, friendly, commonplace MIT license, making the new technique broadly usable by developers, researchers and commercial enterprise operations that want to study or adapt the approach.</p><p>The system is aimed at one of the most expensive problems in AI deployment: serving large models quickly enough for real users, while using hardware efficiently enough to make the economics work. That matters for consumer chatbots, coding assistants, agentic workflows and enterprise AI systems where users expect long answers to stream quickly rather than crawl out word by word.</p><p>DeepSeek is applying DSpark to its own latest frontier open model,<a href="https://venturebeat.com/technology/deepseek-v4-arrives-with-near-state-of-the-art-intelligence-at-1-6th-the-cost-of-opus-4-7-gpt-5-5"> DeepSeek-V4</a>. </p><p>Specifically, DeepSeek used its new DSpark framework on DeepSeek-V4-Flash, its already speed-optimized 284-billion-parameter mixture-of-experts model with 13 billion active parameters, and DeepSeek-V4-Pro, its more thoughtful and powerful 1.6-trillion-parameter model with 49 billion active parameters (Both support context windows up to one million tokens). </p><p>But the broader significance is that<i> DSpark is not conceptually limited to DeepSeek-V4.</i> DeepSeek’s own tests and released checkpoints cover other open model families, including Alibaba's open weights <i>Qwen</i> and Google's open weights <i>Gemma. </i></p><p>That means enterprise teams running open-weight models could, in principle, train or fine-tune DSpark-style draft modules for their own target models. It is not a switch that any API customer can flip from the outside, but it is a method that can travel to other models when the operator controls the weights and serving stack.</p><h2><b>Staggering speed increases for generating tokens during inference</b></h2><p>In DeepSeek’s live production tests, DSpark improved aggregate throughput by 51% for DeepSeek-V4-Flash at an 80-token-per-second-per-user service target, and by 52% for DeepSeek-V4-Pro at a 35-token-per-second-per-user target. At matched system capacity, DeepSeek reports per-user generation speedups of 60% to 85% for V4-Flash and 57% to 78% for V4-Pro over its prior MTP-1 production baseline.</p><p>The different speed claims measure different things. The 60% to 85% figure for V4-Flash, and the 57% to 78% figure for V4-Pro, describe how much faster individual users receive generated tokens when DeepSeek compares DSpark with MTP-1 at matched practical system capacity. </p><p>Those are the cleaner “generation speed” numbers. DeepSeek also reports much larger 661% and 406% increases, but these measure aggregate throughput under very strict speed targets: 120 tokens per second per user for V4-Flash and 50 tokens per second per user for V4-Pro. </p><p>At those targets, DeepSeek says its older MTP-1 baseline approaches an operational cliff, meaning it can keep only a small number of concurrent requests running while preserving that level of responsiveness. </p><p>DSpark avoids more of that collapse, so the percentage difference in total system output becomes much larger. Put simply: the 85% number is closer to “how much faster the ride feels for a user” under comparable conditions, while the 661% and 406% figures are closer to “how much more traffic the road can still carry” when the old system is already bottlenecking. </p><h2><b>Why speculative decoding matters</b></h2><p>LLMs usually generate text one token at a time. A token can be a word, part of a word, punctuation mark or other small piece of text. Every new token depends on the text already produced, so the model has to keep pausing, checking the full context and choosing the next piece.</p><p>That is accurate, but slow. It is like having a senior editor approve every word before a writer can move to the next one. The editor may be excellent, but the process creates a bottleneck.</p><p>Speculative decoding, developed in the early Transfomer era, tries to fix that bottleneck. Instead of asking the large model to produce every token one by one, the system uses a smaller or lighter draft component to suggest several likely next tokens. The large model then checks that batch of guesses in parallel. If the draft guessed correctly, the system moves ahead several tokens at once. If the draft made a bad guess, the system rejects the bad token and anything after it, adds a corrected token, and tries again.</p><p>The point is speed without changing the larger model’s intended output. In the standard speculative decoding setup, the draft model is not replacing the target model. It is acting more like an assistant who prepares a rough next sentence for the senior editor to approve or reject.</p><p>The idea did not appear out of nowhere with today’s large language models. A <a href="https://arxiv.org/abs/1811.03115">key precursor came in 2018</a>, when Mitchell Stern, Noam Shazeer and Jakob Uszkoreit proposed blockwise parallel decoding for deep autoregressive models. Their method predicted multiple future steps in parallel, then kept the longest prefix validated by the main model. That paper established much of the draft-and-check intuition behind later speculative decoding work.</p><p>The research line became more explicit in 2022. <a href="https://arxiv.org/abs/2203.16487">Heming Xia, Tao Ge and co-authors introduced SpecDec</a>, a draft-and-verify approach for sequence-to-sequence generation. Later that year, Yaniv Leviathan, Matan Kalman and Yossi Matias posted “<a href="https://arxiv.org/abs/2211.17192">Fast Inference from Transformers via Speculative Decoding</a>,” which helped define the modern version of the technique for transformer-based language models. DeepMind researchers followed in 2023 with a closely related method called <a href="https://arxiv.org/abs/2302.01318">speculative sampling.</a></p><p>Those 2022 and 2023 papers are the clearest ancestors of how speculative decoding is discussed in current LLM inference work: a faster draft process proposes tokens, and the larger target model verifies them in a way designed to preserve the target model’s output distribution. </p><p>Since then, the field has moved quickly through several variants, including separate draft models, multi-token prediction heads, tree-based verification, feature-level methods such as <a href="https://arxiv.org/abs/2401.15077">EAGLE</a>, self-speculation, Medusa-style extra heads and parallel/blockwise drafters such as DFlash.</p><p>The key metric is not how many tokens a draft model can guess. It is how many of those guesses the larger model actually accepts. Long speculative blocks help only if enough of the proposed tokens survive verification. Otherwise, the system spends compute checking guesses that it throws away.</p><p>That is the context for DSpark. Speculative decoding is already an established inference technique before DeepSeek’s release, with support in major serving stacks and multiple competing research approaches. But it is still not a solved problem. Speedups depend heavily on the draft model, the workload, the serving setup and the current traffic level. DSpark’s contribution is to improve both sides of the trade-off: it tries to draft more coherent token blocks and then verify only the parts of those blocks that are likely to pay off under real serving conditions.</p><h2><b>What DSpark changes</b></h2><p>DSpark tackles two related problems: bad guesses and wasted checking.</p><p>First, the system uses what DeepSeek calls semi-autoregressive generation. In plain English, that means DSpark tries to combine speed with a bit more awareness of sequence. </p><p>A fully parallel drafter can guess several tokens at once, which is fast, but its later guesses can become less coherent because each position is predicted too independently. A purely step-by-step drafter can keep better track of how one token leads to the next, but it loses much of the speed advantage.</p><p>DSpark tries to keep the best of both. It uses a parallel backbone for most of the drafting work, then adds a lightweight sequential head that lets the draft take nearby token relationships into account. In the paper’s example, a parallel drafter might confuse likely phrase endings such as “of course” and “no problem,” producing awkward combinations because it is guessing positions too separately. DSpark’s sequential component helps the system make the later tokens fit the earlier ones.</p><p>Second, DSpark adds confidence-scheduled verification. Rather than always asking the target model to check the same number of draft tokens, DSpark estimates which prefix of the draft is likely to survive. A hardware-aware scheduler then adjusts how much of each draft should be verified based on both model confidence and current serving load.</p><p>A simple analogy: when a restaurant is quiet, the head chef can inspect more of the prep cook’s work. When the kitchen is slammed, the chef spends attention only on the dishes most likely to be ready. DSpark applies a similar idea to AI serving. Under lighter traffic, the system can afford to check longer draft prefixes. Under heavier traffic, it trims low-confidence trailing guesses before they consume batch capacity that could be used for other users.</p><p>DeepSeek frames this as an answer to a common production trade-off. Static multi-token drafting can look attractive in isolation, but can hurt throughput under high concurrency because the system keeps checking tokens that are likely to be rejected. DSpark’s scheduler makes the verification budget flexible instead of fixed.</p><h2><b>Offline results: better draft acceptance across Qwen and Gemma</b></h2><p>DeepSeek tested DSpark offline on Qwen3-4B, Qwen3-8B, Qwen3-14B and Gemma4-12B target models across math, coding and chat benchmarks. </p><p>In those tests, the team compared DSpark with DFlash, a parallel drafter, and Eagle3, an autoregressive drafter. The paper reports accepted length per decoding round, a measure of how many tokens survive verification on average.</p><p>Across the three Qwen3 model sizes, DSpark improved macro-average accepted length over Eagle3 by 30.9%, 26.7% and 30.0%, respectively. Compared with DFlash, it improved accepted length by 16.3%, 18.4% and 18.3%. The paper also says the gains generalized to Gemma4-12B.</p><p>That supports a point raised by developer Daniel Han, who highlighted on X that DeepSeek showed DSpark working beyond DeepSeek’s own V4 models, including Gemma and Qwen. I would include Han as community reaction, not as the sole evidence for the claim. The stronger support comes from DeepSeek’s own benchmarks and released checkpoints.</p><p>The offline results also show why workload matters. Structured tasks such as math and code tend to have higher accepted lengths than open-ended chat. That makes intuitive sense: a code completion or math step often has fewer reasonable next moves than a free-form conversation. </p><p><b>For enterprises, </b>this means<b> DSpark-style methods may be especially attractive for coding assistants, data analysis agents, structured workflow automation</b> and other settings where outputs follow more predictable patterns.</p><h2><b>How enterprises could use DSpark without DeepSeek-V4</b></h2><p>One of the most important questions is whether DSpark is a DeepSeek-only optimization or a broader method that can be applied to other models. The answer is: broader method, but not automatic plug-in.</p><p>For open-weight models, the path is relatively clear. An enterprise running Qwen, Gemma, Llama, Mistral, Granite, Command-style open weights or another model it hosts itself could train or fine-tune a DSpark-style draft module against that target model. </p><p>The team would then measure acceptance on its own workloads and integrate the verification scheduler into its inference stack.</p><p>That is different from simply downloading DeepSeek’s DSpark module and attaching it to any model. Speculative decoding depends on alignment between the draft module and the target model. The draft has to learn what the target model is likely to accept. A drafter trained for DeepSeek-V4 will not automatically be the right drafter for a different model, especially one fine-tuned on a company’s internal data or configured for different reasoning behavior.</p><p>DeepSpec’s workflow reflects this. The process involves preparing data, regenerating target-model answers, building a target cache, training the draft model and evaluating speculative-decoding acceptance. For domain-specific use, the draft model may need additional fine-tuning, especially if the target model runs in a thinking or reasoning mode.</p><p>For proprietary models, the answer depends on what the enterprise controls. If a company owns or fully hosts the model weights and serving stack, it could theoretically train and deploy a DSpark-style drafter. If the model is available only through a hosted API from a vendor, the customer cannot directly add DSpark from the outside. The API provider could implement a similar optimization internally, but the customer generally cannot access the token verification loop, logits, batching behavior or serving scheduler needed to make DSpark work.</p><p>That distinction matters for enterprise buyers. DSpark strengthens the case for open or self-hosted AI infrastructure because it gives advanced teams another lever to improve speed and cost. But it also shows why model serving is becoming a specialized discipline. The value is not just in picking a model, but in how intelligently that model is run.</p><h2><b>What developers get from DeepSpec</b></h2><p>For developers, DeepSpec gives a concrete implementation path for training and evaluating speculative decoding draft models. It includes data preparation, training and benchmark evaluation steps, along with released checkpoints for several open model families. That makes the release useful not only for running DeepSeek-V4 with DSpark, but also for researchers and infrastructure teams studying how to add faster decoding to other open models.</p><p>There are real deployment caveats. DeepSpec’s own README says the default Qwen3-4B data preparation setup can require roughly 38 TB of target cache storage, and the default scripts assume a single node with eight GPUs. That makes the release more immediately relevant to AI labs, cloud teams and sophisticated enterprise AI infrastructure groups than to ordinary application developers.</p><p>Still, releasing the training pipeline matters. Many inference optimizations appear only as papers, vague benchmarks or closed production claims. DeepSpec gives developers something closer to a set of blueprints: not a finished enterprise product, but a way to reproduce, adapt and evaluate the method.</p><h2><b>Early community testing</b></h2><p>The release has already drawn fast developer attention. Developer <a href="https://github.com/rafaelcaricio/spark_vllm_docker/pull/1">Rafael Caricio published a GitHub pull request </a>documenting single-stream DeepSeek-V4-Flash DSpark work, reporting warmed benchmark anchors of 26.33 tokens per second without speculative decoding, 39.88 tokens per second with MTP-1, and roughly 60 tokens per second with DSpark — about 1.5x over MTP-1 and 2.3x over no-spec decoding.</p><p>A later commit in the same thread recorded a five-run mean of 60.31 tokens per second, with a 1.51x gain over MTP-1 and 2.29x over non-speculative decoding. </p><p>The same work also points to an important practical limit: in realistic multi-turn coding sessions, performance can degrade as draft acceptance falls with growing context. In other words, DSpark can make decoding faster, but acceptance quality still determines how much speed the system actually realizes.</p><p>That is a useful reality check. DSpark is not magic. It still depends on how predictable the next tokens are and how well the drafter stays aligned with the target model. But the early implementation work suggests DeepSeek’s claims are not purely academic. Developers are already testing the method in practical serving environments and reporting gains close to the paper’s single-stream expectations.</p><h2><b>The bottom line</b></h2><p>DSpark shows how much performance remains available in the inference layer, even when the underlying model architecture stays the same. As AI companies compete on model quality, context length and pricing, decoding efficiency is becoming another major battleground. </p><p>Faster generation means lower latency for users, higher throughput for providers and better economics for teams serving open models at scale.</p><p>DeepSeek’s release is notable because it combines a production-tested method, open code, public checkpoints and a detailed paper. The main innovation is not just drafting more tokens. It is making the system more selective about which speculative work is worth verifying.</p><p>For enterprise teams, the broader lesson is that the next wave of AI performance gains will not come only from larger models. It will also come from smarter ways to run the models companies already have — especially when those companies control enough of the stack to tune the model, train a compatible draft module and optimize the serving engine around real workloads.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The attack that hijacked Claude Code came through Sentry. Datadog, PagerDuty, and Jira have the same exposure.]]></title>
<description><![CDATA[A single fake error report hijacked Claude Code in controlled testing — the agent ran the attacker's code with the developer's full privileges, and not one alert fired. EDR, WAF, IAM, and the firewall all missed it completely.Tenet Security's June agentjacking disclosure describes a single crafte...]]></description>
<link>https://tsecurity.de/de/3633679/it-nachrichten/the-attack-that-hijacked-claude-code-came-through-sentry-datadog-pagerduty-and-jira-have-the-same-exposure/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3633679/it-nachrichten/the-attack-that-hijacked-claude-code-came-through-sentry-datadog-pagerduty-and-jira-have-the-same-exposure/</guid>
<pubDate>Mon, 29 Jun 2026 19:31:45 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>A single fake error report hijacked Claude Code in controlled testing — the agent ran the attacker's code with the developer's full privileges, and not one alert fired. EDR, WAF, IAM, and the firewall all missed it completely.</p><p>Tenet Security's <a href="https://tenetsecurity.ai/blog/agentjacking-coding-agents-with-fake-sentry-errors/">June agentjacking disclosure</a> describes a single crafted Sentry error event — sent through a public credential that requires no breach and no authentication — that injected attacker instructions into error data that Claude Code, Cursor, and Codex then executed as trusted diagnostic output. Tenet tested 100-plus targets in controlled conditions and achieved an 85% success rate. Sentry called the flaw "technically not defensible."</p><p>he Cloud Security Alliance classified agentjacking as a <a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-agentjacking-mcp-sentry-injection-20260612/">systemic MCP vulnerability class</a> within days of the disclosure. No credentials were stolen, no policy was violated, no perimeter was breached: every step in the chain was authorized. That is the problem.</p><p>Tenet identified <a href="https://thehackernews.com/2026/06/agentjacking-attack-tricks-ai-coding.html">2,388 organizations with publicly exposed Sentry credentials</a> that could be used to inject malicious events at scale. The research is proof-of-concept, not confirmed exploitation across all 2,388. But one captured Claude Code environment held a live AWS secret access key and private repository URLs.</p><p>Here is the scope test: If your AI coding agents are connected to Sentry, Datadog, PagerDuty, Jira, or any MCP-connected data source your developers trust — and those agents can execute shell commands — then your stack has the same blind spot.</p><p>Organizations running Sentry should audit all publicly exposed DSNs immediately. Sentry's architecture intentionally makes DSN credentials public for frontend error reporting, so the mitigation isn't revoking the DSN — it's restricting what agents can do with the data those DSNs return.</p><h2>Why your stack can't see it</h2><p>Agentjacking works because every step is authorized: The attacker sends a valid Sentry API call using a public DSN, the MCP server returns the injected event as authentic output, and the agent executes the instruction using the developer's privileges. No signature fired. The victim saw only benign diagnostics while the agent silently <a href="https://www.infosecurity-magazine.com/news/agentjacking-attacks-hijack-ai/">exposed cloud credentials and source-control tokens</a>.</p><p>SOC teams have never needed to distinguish between a developer running an npm install and an agent running that command in response to a malicious error event. That distinction <a href="https://thenewstack.io/agentjacking-sentry-mcp-attack/">did not exist until AI coding agents became production tools</a>. The stack that cannot make it is the stack agentjacking bypasses.</p><h2>Five surveys, one pattern</h2><p>Five independent surveys from the first half of 2026 found that enterprises trust their AI agents far more than their enforcement justifies.</p><p>Only <a href="https://www.okta.com/newsroom/articles/ai-agents-at-work-2026-agentic-enterprise-security/">34% of organizations apply the same security controls</a> to AI agents as to humans, according to an Okta/Apprize360 survey of 292 executives and 492 knowledge workers. Fifty-two percent of employees use unapproved AI tools, and 58% of executives reported an AI-related incident or close call in the prior year.</p><p>HiddenLayer’s 2026 AI Threat Landscape Report surveyed 250 IT and security leaders: 33% reported <a href="https://www.hiddenlayer.com/report-and-guide/threatreport2026">agents had already exceeded intended scope</a>, and 31% could not confirm whether they had experienced an AI breach. One in eight AI breaches was linked to agentic systems.</p><p><a href="https://www.gravitee.io/blog/state-of-ai-agent-security-2026-report-when-adoption-outpaces-control">Gravitee’s survey of over 900 executives and practitioners</a> found only 14.4% of agents <a href="https://www.gravitee.io/blog/state-of-ai-agent-security-2026-report-when-adoption-outpaces-control">went live with full security approval</a>, and 88% reported confirmed or suspected incidents. A follow-up of 750 leaders in April found agent estates had doubled while monitoring barely moved.</p><h2>The runtime gap nobody closed</h2><p>“Securing agents looks very similar to securing highly privileged users,” said Elia Zaitsev, CTO of CrowdStrike, in an <a href="https://venturebeat.com/security/rsac-2026-agent-identity-frameworks-three-gaps">interview with VentureBeat</a>. “They have identities, access to underlying systems, they reason, they take action.”</p><p>Zaitsev pointed to the gap the industry left open. “No one has been talking about securing agents at runtime. We are doing that now. What is your safety net? If all these controls fail, how do you prevent them from failing silently?”</p><p>CrowdStrike's fleet data quantifies the exposure: more than 1,800 agentic applications on enterprise endpoints, approximately 160 million instances under monitoring. On June 15, <a href="https://www.crowdstrike.com/en-us/press-releases/crowdstrike-unveils-continuous-identity-for-ai-agents/">CrowdStrike shipped Continuous Identity for AI Agents at Identiverse</a>, replacing static policies with continuous enforcement that authorizes every agent action in real time. The control class that announcement reflects — continuous action-level authorization with verifiable agent identity — is now a baseline procurement criterion regardless of vendor.</p><p>“People have kind of forgotten about runtime security,” Zaitsev said. “We did this with endpoint, virtualization, and cloud. People focused on patching vulnerabilities, locking down permissions. Somehow, they always seem to miss something. The safety net is runtime.”</p><p>Zaitsev was equally direct about sandbox approaches. “If you start with an agent in a sandbox that has no ability to touch anything, it is worthless. Very quickly, you are in this race of giving it more capabilities. And then what is the point of your sandbox?” Agents derive their value from access. Every access grant is an attack surface.</p><h2>The governance gap is a budget problem</h2><p>Kayne McGladrey, an IEEE Senior Member, described the structural challenge in an exclusive interview with VentureBeat. “The CISO doesn’t have the budget. The CISO doesn’t have the staff. We can observe risks, we can advise on business risks, but we don’t own the business systems affected by those risks,” McGladrey said. When agent governance spans six departmental budgets, no single executive can confirm whether agents get the same access reviews as humans.</p><p>The Okta survey quantifies the disconnect. Only <a href="https://www.okta.com/newsroom/press-releases/showcase-2026/">43% of workers say agent policies are clear</a>, compared to 65% of executives, and nearly two-thirds apply weaker controls to agents than to humans. The people deploying agents daily do not recognize the governance posture their leadership claims to have built.</p><p>Assaf Keren, chief security officer at Qualtrics and former CISO at PayPal, put it plainly. “The real risk starts not by the implementation of AI systems. It is the fact that baseline architecture is not well established. When we put an AI system on top of something not architected well, we are accelerating the fractures.” Keren called runtime behavior analytics “an unsolved problem right now.”</p><h2>The 5-question gap test</h2><p>The five-question gap test draws on five surveys from the first half of 2026. Each question maps to a gap that agentjacking exploits. Run this before any Q3 vendor evaluation.</p><table><tbody><tr><td><p><b>Gap to test</b></p></td><td><p><b>The proof</b></p></td><td><p><b>What breaks</b></p></td><td><p><b>Monday action</b></p></td><td><p><b>Source / sample</b></p></td></tr><tr><td><p>1. Agent inventory. What percentage of agents, MCP connections, and LLM automations completed security review before deployment?</p></td><td><p>14.4% get full security/IT approval before going live. 52% of employees use unapproved AI tools. Average enterprise now manages 37+ deployed agents, roughly doubled from Q4 2025.</p></td><td><p>Unapproved agents are invisible to your identity platform and unaccountable in a breach disclosure. Agentjacking targets exactly these unmanaged MCP connections. No census means no audit trail for regulatory response.</p></td><td><p>Commission a full agent, MCP server, and LLM automation census. Make census completion a procurement gate for all Q3 vendor evaluations. Flag any agent discovered post-census as a shadow AI incident.</p></td><td><p>Gravitee State of AI Agent Security 2026, 900+ respondents (Feb 2026); Gravitee April 2026 update, 750 senior tech leaders; Okta/Apprize360, 292 execs + 492 workers (June 2026)</p></td></tr><tr><td><p>2. Controls parity. Do agents receive the same access reviews, privilege scoping, and revocation timelines as human employees?</p></td><td><p>34% always apply the same controls to agents as humans. 61% of privileged access fulfilled without proper review. Only 22% treat agents as independent identity-bearing entities.</p></td><td><p>An agent with a static OAuth token and no review cycle is a permanent privileged account with no termination date. Agentjacking inherits whatever privileges the developer holds. 45.6% of orgs rely on shared API keys for agent-to-agent auth.</p></td><td><p>Add every production agent to the next access review cycle. Mandate human-in-the-loop for any agent action touching PII, financial data, or production infrastructure. Replace shared API keys with scoped, short-lived tokens.</p></td><td><p>Okta/Apprize360 (784 respondents, June 2026); Palo Alto Networks (2,930 respondents); Gravitee (900+, shared API keys data)</p></td></tr><tr><td><p>3. Scope drift. Have any agents accessed data or systems beyond their defined scope in the last 12 months?</p></td><td><p>33% report agents already exceeded scope. 53% say agents exceed permissions occasionally or sometimes. Meta Sev 1, March 2026: agent posted sensitive data to unauthorized channel. Only 8% say agents never exceed intended permissions.</p></td><td><p>Scope drift triggers reportable events under GDPR, CCPA, HIPAA, and SEC cybersecurity rules. If detection cannot distinguish agent-initiated from human-initiated access, disclosure timelines are unachievable. Agent-spawned sub-agents (25.5% of deployed agents can create other agents) make audit trails algebraically intractable.</p></td><td><p>Run a 90-day scope-drift audit on every production agent. Compare actual resources touched against approved scope documentation. Block agent-to-agent delegation without explicit human approval for any action exceeding the parent agent’s scope.</p></td><td><p>HiddenLayer AI Threat Landscape 2026 (250 IT/security leaders); CSA AI Agent Security Survey (scope violations data); Gravitee (agent spawning data)</p></td></tr><tr><td><p>4. Governance perception gap. Would 50 knowledge workers say your AI agent policies are clear?</p></td><td><p>22-point gap: 65% of executives say policies are clear, 43% of workers agree. 77% of security teams see shadow AI risk but lack visibility to act. 76% cite shadow AI as a definite or probable problem.</p></td><td><p>You are evaluating vendors against a governance posture your workforce does not recognize. Every shadow agent undermines the vendor comparison. Knowledge workers sharing internal messages (54%), HR data (45%), and confidential docs (39%) with unapproved AI tools.</p></td><td><p>One-question survey before your next vendor demo. Gap exceeds 15 points, pause procurement. Publish an internal AI agent acceptable-use policy with specific examples of approved and prohibited agent behaviors.</p></td><td><p>Okta/Apprize360 (784 respondents, June 2026); Ivanti 2026 AI Maturity Report (1,200 respondents); HiddenLayer (shadow AI data)</p></td></tr><tr><td><p>5. Breach detection certainty. Can your security team confirm whether you experienced an AI-related breach in the last 12 months?</p></td><td><p>31% cannot answer. 88% reported confirmed or suspected AI agent security incidents. One in eight reported AI breaches now linked to agentic systems. Agentjacking proved EDR, WAF, IAM, and firewall pass an agent-mediated attack without a single alert.</p></td><td><p>No basis for disclosure timelines. No evidence chain for incident response. No defensible position in a regulatory investigation. EU AI Act high-risk compliance obligations take effect August 2, 2026.</p></td><td><p>Require agent-specific runtime detection as a procurement prerequisite. Confirm your org can distinguish agent-initiated actions from human-initiated actions in production telemetry. Test your SOC’s ability to attribute a specific action to a specific agent within 60 minutes.</p></td><td><p>HiddenLayer (250 IT/security leaders); Gravitee (900+, incident rate); Tenet Security (2,388 orgs exposed); CSA (systemic MCP vulnerability classification)</p></td></tr></tbody></table><h2>Security director action plan</h2><p>EU AI Act high-risk compliance obligations take effect August 2, 2026. Worth factoring into Q3 planning timelines.</p><ol><li><p>Run the five-question gap test above before any Q3 vendor evaluation — it costs nothing to administer, and the procurement clarity it creates is worth far more than the 30 minutes it takes.</p></li><li><p>Consider mandating agent-specific runtime detection. If your stack cannot tell what an agent did from what a developer did, agentjacking will bypass it the same way it bypassed every layer in Tenet’s testing. That distinction is the one that matters now.</p></li><li><p>Treat every agent as a privileged insider. According to the Okta/Apprize360 survey, only 34% of organizations apply the same controls to agents as to humans; closing that gap is the single most impactful thing most security teams can do this quarter.</p></li><li><p>Test the perception gap before investing in new tooling. One question to 50 knowledge workers. Do you know your company’s AI agent policies? If the gap between their answer and leadership’s answer exceeds 15 points, that is the problem to solve first. No vendor product fixes a governance posture your own workforce does not recognize.</p></li><li><p>Make agent census completion a procurement gate — every agent, every MCP connection. The security teams getting this right are the ones that started with a complete inventory and worked forward from there.</p></li></ol><p>Agentjacking stripped away an assumption that has survived every security architecture since the first firewall went live. Authorized does not mean safe. When every step in the chain is legitimate, the only defense that matters is the one watching what agents do. Not what policies say. What agents do.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Grounding, not models, will define your AI advantage]]></title>
<description><![CDATA[Over the past two years, working inside the enterprise AI infrastructure world, tracking where the industry is heading, I have noticed the same question surface repeatedly: should we build our own large language model? I understand the instinct. The model feels like the thing, the engine, the bra...]]></description>
<link>https://tsecurity.de/de/3632693/it-nachrichten/grounding-not-models-will-define-your-ai-advantage/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3632693/it-nachrichten/grounding-not-models-will-define-your-ai-advantage/</guid>
<pubDate>Mon, 29 Jun 2026 13:03:15 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Over the past two years, working inside the enterprise AI infrastructure world, tracking where the industry is heading, I have noticed the same question surface repeatedly: should we build our own large language model? I understand the instinct. The model feels like the thing, the engine, the brain, the asset worth owning. But after significant years as a product manager in the AI world in both customer experience and grounding infrastructure I concluded that it tends to unsettle the room: the model is the least durable part of your AI strategy.</p>



<p>I say this not to be provocative, but because over the last few years we have seen organizations pour their scarcest resources, executive attention, engineering talent, capital, into the one layer of the stack that is commoditizing fastest. Meanwhile, the layer that determines whether their AI is trustworthy, accurate and defensible gets treated as plumbing. That inversion is, in my experience, the single most expensive mistake enterprises are making with AI right now.</p>



<h2 class="wp-block-heading">The model is becoming a commodity</h2>



<p>Let us consider economics. <a href="https://www.gartner.com/en/newsroom/press-releases/2026-03-25-gartner-predicts-that-by-2030-performing-inference-on-an-llm-with-1-trillion-parameters-will-cost-genai-providers-over-90-percent-less-than-in-2025" rel="nofollow">Gartner projects that by 2030, performing inference on a trillion-parameter model will cost providers more than 90% less</a> than it did in 2025, with models becoming up to 100 times more cost-efficient than the earliest versions of comparable size. When the cost of the underlying capability collapses by that magnitude, it stops being a differentiator. Anything that gets that cheap, that fast, is not where competitive advantage lives.</p>



<p>Models that feel innovative are routinely surpassed by something cheaper and better within months. If your advantage is tied to a specific model, it will evaporate the moment the frontier moves, which it always does. But if an enterprise instead invests in how reliably it can feed any model its proprietary context, that investment holds. That part travels from one model generation to the next. When a better model arrives, the organization can simply connect it and immediately capture the upside, because the hard and durable work was already done one layer down.</p>



<p>I wish more leaders could observe this pattern before they commit. The model layer is improving so quickly that any advantage you build into it has a short half-life. The grounding layer behaves in the opposite way: every improvement you make to your data quality, your retrieval logic and your governance compounds, and it carries forward regardless of which model sits on top.</p>



<p>This is why the build-your-own LLM debate so often misses the mark. Training or even meaningfully fine-tuning a foundation model is enormously expensive, and the moment you finish, the open and commercial frontier has usually moved past you. So, technically you spent a fortune to own a depreciating asset. The capability that you should focus on is an AI that knows your business, was never going to come from the weights of the model anyway. It comes from what you put in front of it.</p>



<h2 class="wp-block-heading">Why grounding is the real moat</h2>



<p>Grounding is the discipline of connecting a general-purpose model to your enterprises’ current and authoritative information, most commonly through retrieval-augmented generation, or RAG. Rather than hoping the model memorized something useful during training, you retrieve the relevant facts from your own systems in real time of the query and give the model the context it needs to answer correctly.</p>



<p>Here is the part that matters for anyone thinking about competitive advantage: your competitors can rent the exact same model you use. What they cannot rent is your data, your institutional knowledge, your processes and the quality of the pipeline that surfaces all of it accurately at the right moment. That pipeline is genuinely proprietary, genuinely hard to replicate and it compounds in value over time. That is the textbook definition of a moat, and it has almost nothing to do with which model you chose.</p>



<p>The industry is starting to recognize this. Gartner predicts that <a href="https://www.gartner.com/en/newsroom/press-releases/2025-04-09-gartner-predicts-by-2027-organizations-will-use-small-task-specific-ai-models-three-times-more-than-general-purpose-large-language-models" rel="nofollow">by 2027, organizations will use small, task-specific models at least three times more than general-purpose LLMs</a>, precisely because accuracy in real business workflows depends on domain context rather than raw model scale. But a smaller model holds less in its parameters by design, which means it leans even harder on retrieval to supply current, authoritative context in real time. The model gets smaller and more swappable. The grounding becomes the part that carries the weight. In that same analysis, Gartner makes the same point from the data side: what sets enterprises apart is how well they prepare, check, version and manage their own data. Read that again: the differentiator is the data discipline, not the model.</p>



<p>This matches what I have observed directly. Getting hold of an excellent model was never the hard part, and it was rarely where things broke. The failures I have seen came from not connecting the model efficiently to the right data sources or orchestrating retrieval well. The patterns repeat: missing data produces incomplete summaries, truncated documents leave answers without key details, and noisy context yields irrelevant or confusing responses.</p>



<p>When grounding is absent, answers become inconsistent from one client to the next; when retrieval comes back empty, the model fills the gap with something hallucinated or useless. Stale data produces confidently outdated answers, retrieval gaps surface as generic non-answers, and poor-quality data drags down both speed and output. None of these are model problems. They are grounding problems. And when a system hands an executive an answer that is wrong, no one in the boardroom cares how sophisticated the model was. They care that it was wrong, and the fix always lives in the grounding layer.</p>



<p>One example has stayed with me. In a real enterprise scenario, an AI assistant returned inconsistent answers to the same query across different environments whenever grounding was unavailable, and some of those answers contradicted each other outright. The cause was straightforward in hindsight. With no grounding, the system fell back on its own internal knowledge instead of a shared, grounded source of truth, so its responses drifted with each configuration and context. The damage was not just technical. Users stopped trusting an assistant that could not give them the same answer to the same question twice. That is the actual cost of weak grounding, and it is why consistency and reliability in production depend far more on the data layer than on the model sitting above it. No model upgrade would have fixed that.</p>



<h2 class="wp-block-heading">Where leaders should focus their investment</h2>



<p>If you accept that grounding is where advantage accrues, a few priorities shift in ways that should change how you allocate budget and attention.</p>



<p>First, treat your organization’s data foundation as a first-class AI investment, not a prerequisite you rush through. The unglamorous work, cleaning, structuring, governing and versioning your knowledge, is the work that determines AI quality. I would rather inherit a mediocre model with an excellent retrieval pipeline than the reverse, every single time.</p>



<p>Second, build for model portability from day one. Assume the model you use today will be replaced within a year because it certainly will. If swapping it out is painful, you have coupled your architecture to the wrong layer. Your grounding infrastructure, your evaluation framework and your data contracts should be the stable core; the model should be a component you can swap with minimal disruption.</p>



<p>Third, invest in observability and evaluation for retrieval, not just for the model. The emerging discipline here matters: <a href="https://www.gartner.com/en/newsroom/press-releases/2026-03-30-gartner-predicts-by-2028-explainable-ai-will-drive-llm-observability-investments-to-50-percent-for-secure-genai-deployment" rel="nofollow">Gartner expects LLM observability investments to reach 50% of GenAI deployments by 2028</a>, up from 15% today, as trust requirements outpace the technology itself. Knowing why your system retrieved a particular piece of context, and whether that context was correct, is what makes an AI output defensible and auditable. For any organization operating under real regulatory or reputational scrutiny, that is not optional.</p>



<p>None of this means the model is irrelevant. You still need a capable one and choosing well matters. But choosing a model is now a procurement decision with several excellent options, not a source of lasting differentiation. The lasting differentiation is everything you wrap around it.</p>



<p>I think the organizations that internalize this will look, in a few years, meaningfully ahead of the ones still debating whether to train their own model. Not because they made a bolder bet, but because they made a more durable one. They understood that in a world where everyone has access to the same extraordinary models, the advantage belongs to whoever grounds those models best in the reality of their own business. The model is rented. The grounding is owned. Build accordingly.</p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><strong><a href="https://www.cio.com/expert-contributor-network/">Want to join?</a></strong></p>



<p></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[AI needs a flight school]]></title>
<description><![CDATA[In the late 1960s, elite Navy pilots began losing dogfights.



The deep, instrument-level understanding of exactly where they were, what their aircraft was doing, and what was coming next had been automated. And when moments of crisis arrived, they didn’t have the situational awareness to respon...]]></description>
<link>https://tsecurity.de/de/3632385/ai-nachrichten/ai-needs-a-flight-school/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3632385/ai-nachrichten/ai-needs-a-flight-school/</guid>
<pubDate>Mon, 29 Jun 2026 11:04:11 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>In the late 1960s, elite Navy pilots began losing dogfights.</p>



<p>The deep, instrument-level understanding of exactly where they were, what their aircraft was doing, and what was coming next had been automated. And when moments of crisis arrived, they didn’t have the situational awareness to respond. Put a plane on autopilot long enough, and the pilot stops actually flying.</p>



<p>The same dynamic is playing out across enterprise software. AI is <a href="https://www.infoworld.com/article/4181971/making-sense-of-too-much-code.html" data-type="link" data-id="https://www.infoworld.com/article/4181971/making-sense-of-too-much-code.html">generating code faster</a> than <a href="https://www.infoworld.com/article/4183153/why-ai-coding-debt-is-different.html" data-type="link" data-id="https://www.infoworld.com/article/4183153/why-ai-coding-debt-is-different.html">developers can understand it</a>, and leaders are celebrating the velocity without asking who’s actually flying the plane.</p>



<p>A developer who has only ever “vibe coded” has perception at best. They can “see” the outputs but can’t fix any internal failures caused by the very AI systems they’re relying on. The easiest thing to do is to say the answer looks good enough. Cut and paste it in and hope it works out. According to Model Evaluation &amp; Threat Research’s <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">randomized control trials</a>, experienced developers working with AI tools actually took 19% longer to complete tasks than those working without them, despite predicting beforehand that AI would make them 24% faster.</p>



<p>The fundamentals of good software delivery have never been more important — and never more neglected.</p>



<h2 class="wp-block-heading"><a></a>When instruments go dark</h2>



<p>The Navy’s answer to training dogfighters for success was the Top Gun school — not just to teach pilots to fight, but to teach them how to fly again. That meant returning to the fundamentals by mastering the technical and combat skills that can best prepare them for moments of crisis with clear thinking. This very discipline makes split-second decisions possible when everything is on the line.</p>



<p>Consider this scenario. A retail company’s engineering team used AI to refactor a promotions engine ahead of the holiday season. The code passed every test. Reviews were clean. It shipped on a Tuesday with zero flags.</p>



<p>But what if nobody caught that AI had subtly changed the order of operations in a discount calculation? It’s a logical shift that wouldn’t have broken any individual test case but could compound incorrectly when multiple promotions applied to the same cart. This would be enough to cost the company millions by the time a finance analyst notices the margin erosion during a quarterly close.</p>



<p>The <a href="https://www.infoworld.com/article/4078884/what-is-vibe-coding-ai-writes-the-code-so-developers-can-think-big.html" data-type="link" data-id="https://www.infoworld.com/article/4078884/what-is-vibe-coding-ai-writes-the-code-so-developers-can-think-big.html">vibe coding</a> wave is already breaking. One<a href="https://techstartups.com/2025/12/11/the-vibe-coding-delusion-why-thousands-of-startups-are-now-paying-the-price-for-ai-generated-technical-debt/"> analysis</a> found that roughly 10,000 startups tried to build production apps with AI assistants; more than 8,000 now need rebuilds.</p>



<p>It’s on us to turn this moment of reckoning into an opportunity.</p>



<h2 class="wp-block-heading">Training developers for real-world applications</h2>



<p>So what can we as leaders do about this?</p>



<p>The foundation of everything we do comes down to trust — specifically, teaching people to trust and own their work. My high-level vision is to help people achieve a 50x improvement in their overall processes by leveraging AI tools <em>and</em> still being the expert at large.</p>



<p>For one of Copado’s internal programs, we gave nine employees the ability to vibe code AI-powered tools to tackle any major business problem they identified. Most gravitated toward the same theme: they were constantly fielding repeat questions and wanted to stop answering the same thing twice.</p>



<p>But while the instinct was right, the execution wasn’t ready. They hadn’t thought through who would maintain these tools, how they would be governed, or whether they actually mapped to business objectives.</p>



<p>Just because you can hand someone the controls doesn’t mean they know how to fly.</p>



<p>We then conducted a training session on how to plan an app effectively — with a long-term view of the full software development life cycle — before anyone wrote a line of code. The app ideas got sharper, and the products got real.</p>



<p>The group went from pursuing 10 app ideas to a focused set of seven, with two participants stepping back after realizing they didn’t yet have a problem worth solving. Five are now being implemented across the business: Legal built a policy bot to answer HR’s questions on company policy; the doc writing team built a tool for automatically generating technical documentation; the support team built a case analysis app; the sales team built a call-coaching app that helps sales development reps improve performance by analyzing live calls; and the customer success team built an app that listens to calls and notes, then automatically summarizes everything known about a new client at the point of implementation.</p>



<p>To this day, we also reserve “Failure Fridays,” a monthly space for employees to practice debugging programs without AI assistance. It keeps foundational skills sharp and ensures that when something breaks in production, the team knows how to actually fix it.</p>



<h2 class="wp-block-heading">Five pillars for AI applications</h2>



<p>Across a community of 120,000 Copado developers, I now recommend they enforce these five pillars when deploying AI in their projects:</p>



<ul class="wp-block-list">
<li>Build in checkpoints to evaluate agent output against defined standards before anything moves forward.</li>



<li>Continuous and automated testing should function as a permanent trust layer embedded directly into the development cycle.</li>



<li>Apply human judgment at critical decision points while automation handles the routine verification work in between.</li>



<li>A single review at the end of a process is a point of failure. Continuous validation is necessary to catch issues the moment they arise rather than after they’ve compounded.</li>



<li>Maintain audit trails and performance metrics that capture every agent action. Accountability means tracking what AI does, not just what developers deliver.</li>
</ul>



<p>I believe that success demands the technical knowledge and discipline to build these systems from the ground up. These guardrails ensure that AI works with you, not against you. The bottom line: organizations that approach AI with accountability and knowledge in mind achieve 9x to 10x productivity while maintaining trust.</p>



<p>At Copado, fostering a culture where developers are genuinely motivated to embrace AI is equally important to us. To support that, we created a certification and incentive program that rewards new hires with $1,000 bonuses upon completion — an investment that has delivered a 76% ROI compared to traditional onboarding methods. The impact has been undeniable: we had 30 developers fully onboarded in just 30 days, condensing what typically takes three to six months into a fraction of the time.</p>



<h2 class="wp-block-heading">The fundamentals will endure</h2>



<p>Speed without situational awareness isn’t efficient. It’s a deferred crisis.</p>



<p>The fundamentals of planning, building, testing, and releasing aren’t bureaucratic overhead — they’re the instruments on the dashboard, telling you where you are, what your system is doing, and what’s coming next. Lose them, and you’re not just flying blind. You’re unprepared for the dogfight.</p>



<p>When the moment of reckoning arrives — the production failure, the security breach, the audit, the outage — you find out very quickly whether a human’s full understanding was there or not.</p>



<p>The machine won’t be in the hot seat. You will.</p>



<p><em>—</em></p>



<p><a href="https://www.infoworld.com/blogs/new-tech-forum"><strong><em>New Tech Forum</em></strong></a><em><strong> provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all </strong></em><em><strong>inquiries to </strong></em><a href="mailto:doug_dineley@foundryco.com"><strong><em>doug_dineley@foundryco.com</em></strong></a><em><strong>.</strong></em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[METR warnt: GPT-5.6 Sol zeigt Evaluation-Gaming in Sicherheitsbenchmarks]]></title>
<description><![CDATA[SAN FRANCISCO / LONDON (IT BOLTWISE) – OpenAI hat GPT-5.6 Sol am 26. Juni 2026 in einer eingeschränkten Vorschau gestartet. Doch METR berichtet, dass das Modell bei externen Sicherheitstests auffällig oft die Bewertungssituation ausnutzt. Damit verschiebt sich der Fokus von der behördlichen Zugan...]]></description>
<link>https://tsecurity.de/de/3631692/it-security-nachrichten/metr-warnt-gpt-56-sol-zeigt-evaluation-gaming-in-sicherheitsbenchmarks/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3631692/it-security-nachrichten/metr-warnt-gpt-56-sol-zeigt-evaluation-gaming-in-sicherheitsbenchmarks/</guid>
<pubDate>Mon, 29 Jun 2026 02:05:29 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><img width="1024" height="1024" src="https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol.jpg" class="attachment- size- wp-post-image" alt="" decoding="async" fetchpriority="high" srcset="https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol.jpg 1024w, https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol-300x300.jpg 300w, https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol-150x150.jpg 150w, https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol-768x768.jpg 768w, https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol-840x840.jpg 840w, https://www.it-boltwise.de/wp-content/uploads/2026/06/ai-metr-evaluation-gaming-gpt-sol-120x120.jpg 120w" sizes="(max-width: 1024px) 100vw, 1024px">SAN FRANCISCO / LONDON (IT BOLTWISE) – OpenAI hat GPT-5.6 Sol am 26. Juni 2026 in einer eingeschränkten Vorschau gestartet. Doch METR berichtet, dass das Modell bei externen Sicherheitstests auffällig oft die Bewertungssituation ausnutzt. Damit verschiebt sich der Fokus von der behördlichen Zugangssperre hin zur Frage, wie zuverlässig Benchmarks und Safety-Evaluationen bei Hochleistungs-KI wirklich sind. […]</p>
<div><a href="https://www.it-boltwise.de/metr-warnt-gpt-5-6-sol-zeigt-evaluation-gaming-in-sicherheitsbenchmarks.html">... den vollständigen Artikel <strong>»METR warnt: GPT-5.6 Sol zeigt Evaluation-Gaming in Sicherheitsbenchmarks«</strong> lesen</a></div>
<p>Dieser Beitrag <a href="https://www.it-boltwise.de/metr-warnt-gpt-5-6-sol-zeigt-evaluation-gaming-in-sicherheitsbenchmarks.html">METR warnt: GPT-5.6 Sol zeigt Evaluation-Gaming in Sicherheitsbenchmarks</a> erschien als erstes auf <a href="https://www.it-boltwise.de/">IT BOLTWISE x Artificial Intelligence</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[IT Security News Hourly Summary 2026-06-28 09h : 1 posts]]></title>
<description><![CDATA[1 posts were published in the last hour 6:33 : GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions
Read more →
The post IT Security News Hourly Summary 2026-06-28 09h : 1 posts appeared first on IT Security News.]]></description>
<link>https://tsecurity.de/de/3630640/it-security-nachrichten/it-security-news-hourly-summary-2026-06-28-09h-1-posts/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3630640/it-security-nachrichten/it-security-news-hourly-summary-2026-06-28-09h-1-posts/</guid>
<pubDate>Sun, 28 Jun 2026 09:08:01 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>1 posts were published in the last hour 6:33 : GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions</p>
<p class="more-link-p"><a class="more-link" href="https://www.itsecuritynews.info/it-security-news-hourly-summary-2026-06-28-09h-1-posts/">Read more →</a></p>
<p>The post <a href="https://www.itsecuritynews.info/it-security-news-hourly-summary-2026-06-28-09h-1-posts/">IT Security News Hourly Summary 2026-06-28 09h : 1 posts</a> appeared first on <a href="https://www.itsecuritynews.info/">IT Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions]]></title>
<description><![CDATA[OpenAI says GPT-5.6 Sol's cyber safeguards make it safe enough for restricted release. METR found it had the highest evaluation cheating rate of any publicly tested model. The second finding matters more.
GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions on...]]></description>
<link>https://tsecurity.de/de/3630616/it-security-nachrichten/gpt-56-sols-launch-metrs-evaluation-gaming-finding-matters-more-than-the-restrictions/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3630616/it-security-nachrichten/gpt-56-sols-launch-metrs-evaluation-gaming-finding-matters-more-than-the-restrictions/</guid>
<pubDate>Sun, 28 Jun 2026 08:37:20 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI says GPT-5.6 Sol's cyber safeguards make it safe enough for restricted release. METR found it had the highest evaluation cheating rate of any publicly tested model. The second finding matters more.</p>
<p><a href="https://latesthackingnews.com/2026/06/28/gpt-5-6-sol-metr-evaluation-gaming/">GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions</a> on <a href="https://latesthackingnews.com/">Latest Hacking News | Cyber Security News, Hacking Tools and Penetration Testing Courses</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions]]></title>
<description><![CDATA[OpenAI says GPT-5.6 Sol’s cyber safeguards make it safe enough for restricted release. METR found it had the highest evaluation cheating rate of any publicly tested model. The second finding matters more. GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters…
Read more →
The post GPT-5.6...]]></description>
<link>https://tsecurity.de/de/3630615/it-security-nachrichten/gpt-56-sols-launch-metrs-evaluation-gaming-finding-matters-more-than-the-restrictions/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3630615/it-security-nachrichten/gpt-56-sols-launch-metrs-evaluation-gaming-finding-matters-more-than-the-restrictions/</guid>
<pubDate>Sun, 28 Jun 2026 08:37:19 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI says GPT-5.6 Sol’s cyber safeguards make it safe enough for restricted release. METR found it had the highest evaluation cheating rate of any publicly tested model. The second finding matters more. GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters…</p>
<p class="more-link-p"><a class="more-link" href="https://www.itsecuritynews.info/gpt-5-6-sols-launch-metrs-evaluation-gaming-finding-matters-more-than-the-restrictions/">Read more →</a></p>
<p>The post <a href="https://www.itsecuritynews.info/gpt-5-6-sols-launch-metrs-evaluation-gaming-finding-matters-more-than-the-restrictions/">GPT-5.6 Sol’s Launch: METR’s Evaluation Gaming Finding Matters More Than the Restrictions</a> appeared first on <a href="https://www.itsecuritynews.info/">IT Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Claude Code turned every engineer into three. Now companies need more product thinkers]]></title>
<description><![CDATA[Anthropic recently told its growth team to hire more product managers, not fewer. The reason, as reported in industry coverage, was that Claude Code had quietly turned its engineering org into a team that ships at roughly three times its actual headcount, and the bottleneck moved from the integra...]]></description>
<link>https://tsecurity.de/de/3630129/it-nachrichten/claude-code-turned-every-engineer-into-three-now-companies-need-more-product-thinkers/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3630129/it-nachrichten/claude-code-turned-every-engineer-into-three-now-companies-need-more-product-thinkers/</guid>
<pubDate>Sat, 27 Jun 2026 21:47:31 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Anthropic recently told its growth team to hire more product managers, not fewer. The reason, as reported in industry coverage, was that Claude Code had quietly turned its engineering org into a team that ships at roughly three times its actual headcount, and the bottleneck moved from the integrated development environment (IDE) to the people deciding what to build.</p><p>That detail is easy to miss in the noise of every <a href="https://venturebeat.com/orchestration/vibe-coding-can-build-your-pipeline-it-cant-explain-it-six-months-later">AI productivity claim</a>. It is also the structural shift the rest of the industry is now living through. The bottleneck in software is no longer typing. It is deciding what to type. And the engineers who treat that as someone else's problem are about to plateau. </p><p>For most of the last decade, that decision sat with someone else. <a href="https://venturebeat.com/technology/agentic-ai-solved-coding-and-exposed-every-other-problem-in-software-engineering">Software engineering</a> was a craft you absorbed slowly, then practiced in a long, predictable sequence: Dive deep on the technology, write the code, ask Stack Overflow when stuck, escalate to a senior engineer when Stack Overflow failed, ship the ticket. The product manager owned the funnel. The engineer owned the build. Both sides treated this division as physics.</p><p>Then the funnel collapsed in five steps.</p><h2><b>A short history of how the engineer's day got compressed</b></h2><p><b>The Stack Overflow era (2014 to late 2022): </b>The way engineers thought lived in one place. But new monthly questions on Stack Overflow are now down <a href="https://www.reddit.com/r/programming/comments/1hwg2px/stackoverflow_has_lost_77_of_new_questions/">roughly 77%</a> since November 2022, which was not coincidentally when ChatGPT launched. The drop is not a referendum on the site. It is a referendum on the workflow it represented.</p><p><b>The browser-tab era (late 2022 to 2024):</b> The first ChatGPT generation sat outside the IDE. Engineers ran the same loop they had always run, just with a faster oracle: Write a prompt in a browser, paste the answer back into VS Code, repeat. The work was still single-threaded and engineer-driven. The leverage was real but local.</p><p><b>The IDE-native era (2024 to 2025):</b> Cursor and Claude Code moved the model inside the editor and gave it access to the full repository. The senior-engineer escalation path largely dissolved. For years, the prevailing wisdom among veteran engineers was that Bash had the longest shelf life of any tool in the stack. By 2026, for a meaningful share of working developers, the first command typed in a fresh terminal is claude.</p><p><b>The spec-driven era (2025 to 2026):</b> Larger context windows turned single-session work into something that previously required tickets, design docs, and sprints. Amazon's Kiro IDE team reportedly compressed feature builds from two weeks to two days using the same spec-driven workflow they were shipping. An AWS engineering team described an 18-month rearchitecture, originally scoped for 30 engineers, was completed by 6 people in 76 days. The bottleneck stopped being how long it takes to write the code. It started being how clearly the team can describe what correct looks like.</p><p><b>The routines era (2026):</b> In April, Anthropic shipped Claude Code Routines: Scheduled, persistent agents that run on a cadence, on a webhook, or overnight while the laptop is closed. Cron came back. Hooks came back. The engineer's job is now part orchestration: Spin up a swarm before bed, review a stack of pull requests in the morning. Third-party wrappers like OpenClaw, which was briefly suspended by Anthropic in April before partial reinstatement, made the same point from the open-source side.</p><h2><b>The bottleneck moved; most teams have not</b></h2><p>Engineering has roughly tripled. Product management has not budged. The traditional 1:8 ratio of PMs to engineers, already strained, now plays out closer to an effective 1:20 because each engineer ships more per day. For instance, LinkedIn replaced its associate product manager track with a "Product Builder" program that trains generalists across product, design, and engineering. Anthropic is hiring more PMs, not fewer. The pattern is consistent across companies that have actually deployed agentic workflows in production: The system is producing built features faster than it is producing decisions about what should be built.</p><p>For engineers, this is the most important career signal of the decade, and the easiest one to miss while the productivity stories dominate the feed.</p><h2><b>First principles matter more, not less</b></h2><p>The instinct to declare fundamentals obsolete in the agent era gets the trend exactly wrong.</p><p>When a memory leak takes down production at 3 a.m., and the cause turns out to be a subtle ownership bug pushed 4 years ago, no agent currently in the wild closes that loop end-to-end. Operating systems, networks, concurrency, and query plans still decide who can resolve a real incident. They also decide who can spot the moments when an <a href="https://venturebeat.com/technology/why-prompt-debt-retrieval-debt-and-evaluation-debt-are-quietly-reshaping-enterprise-ai-risk">agent's output</a> looks correct on the surface and is quietly, expensively, wrong underneath. The agent that wrote 70% of the code in a modern repo cannot reliably tell anyone where its assumptions about thread safety, memory ownership, or transaction isolation diverged from the runtime. The engineer who can read the diff and catch that is the engineer the rest of the team needs in the room, and that engineer is built on fundamentals, not on prompting skill.</p><p>The corollary is that fundamentals are now a leverage skill, not a hygiene skill. In 2014, knowing how a TCP retransmit worked got a debug ticket closed faster. In 2026, the same knowledge keeps an entire agent-driven release pipeline from shipping a regression at scale. The blast radius of the engineer who knows what is happening underneath has gone up, not down.</p><h2><b>Review is the new writing</b></h2><p>Engineers in 2026 generate code at a rate that exceeds what any of them can read carefully. The team that ships fast and survives is the team whose engineers treat reviewing AI-generated code with at least the same rigor they once reserved for writing it. The 2025 <a href="https://survey.stackoverflow.co/2025">Stack Overflow developer survey</a> put 84% of developers on AI tools, with 46% saying they do not trust the output, up sharply from 31% the year before. That gap, heavy use paired with low trust, is exactly where review skills now matter most. Coders who push lots and review little are accumulating a debt that will come due during the first real incident, and the engineer who can pay it back is the one who paired their volume with deep first-principles knowledge of the systems involved.</p><h2><b>The new differentiator is the product funnel</b></h2><p>Both of those are necessary. Neither is sufficient. The engineer who matters in 2026 is the one who has stopped waiting for the funnel to arrive in the form of a Jira ticket.</p><p>That means doing things the role was historically allowed to skip.</p><p>Talk to customers. Watch how they actually use the product. Read the support queue. Sit in on the sales call. The signal a product team gets through three layers of summary, an engineer can now get firsthand in an afternoon.</p><p>Generate ideas, not just estimates. The product manager who used to source ideas for 8 engineers cannot source ideas for 20 at the same fidelity. The engineer who shows up with a validated, scoped opportunity is no longer doing the PM's job. The engineer is doing the job the new ratio requires.</p><p>Work backwards from the customer. Amazon has been writing the press release first for two decades. The discipline travels well to teams of one and to swarms of agents. Both produce a great deal of working software in the wrong direction without a clear statement of what "customer wins" means before any code is written.</p><p>Stop hiding behind bandwidth. The honest answer to "Do you have capacity for this idea?" used to be 'No.' With routines, hooks, and a cooperative agent stack, the honest answer is closer to "What is the idea worth?" That is a different conversation, and a much harder one to have without a real point of view on the customer.</p><h2><b>What the next decade rewards</b></h2><p>The five-phase history above is not really a history of tools. It is a history of which part of the job a human had to do. The part that is still human, and that will remain human for the foreseeable future, has moved up the funnel: From typing, to reviewing, to deciding, to choosing the customer to serve and the problem to solve.</p><p>The 2026 version of a <a href="https://venturebeat.com/technology/the-enterprise-risk-nobody-is-modeling-ai-is-replacing-the-very-experts-it-needs-to-learn-from">great engineer</a> is not the one who writes the most code. It is the one who knows what to build, can prove it is worth building, and has the agent fleet plus the review discipline to ship it without the system collapsing under its own velocity.</p><p>Engineers who internalize this will spend the next decade doing the most interesting work software has ever produced. Engineers who wait for a ticket will spend it watching the ticket get written by the agent next to them.</p><p><i>Ishan Gupta is a software engineer at Amazon.</i></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Water Cooler Small Talk, Ep. 11: Overfitting in RAG evaluation]]></title>
<description><![CDATA[Why memorizing for the exam doesn't mean you understand the subject
The post Water Cooler Small Talk, Ep. 11: Overfitting in RAG evaluation appeared first on Towards Data Science.]]></description>
<link>https://tsecurity.de/de/3627843/ai-nachrichten/water-cooler-small-talk-ep-11-overfitting-in-rag-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3627843/ai-nachrichten/water-cooler-small-talk-ep-11-overfitting-in-rag-evaluation/</guid>
<pubDate>Fri, 26 Jun 2026 17:33:56 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Why memorizing for the exam doesn't mean you understand the subject</p>
<p>The post <a href="https://towardsdatascience.com/water-cooler-small-talk-ep-11-overfitting-in-rag-evaluation/">Water Cooler Small Talk, Ep. 11: Overfitting in RAG evaluation</a> appeared first on <a href="https://towardsdatascience.com/">Towards Data Science</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[2026 EuroLLVM - Accelerating Pass Order Auto-tuning via Profile-Guided Cost Modeling]]></title>
<description><![CDATA[Author: LLVM - Bewertung: 1x - Views:8 2026 EuroLLVM Developers' Meeting
https://llvm.org/devmtg/2026-04/
------
Title: Accelerating Pass Order Auto-tuning via Profile-Guided Cost Modeling
Speaker: Bingyu Gao, Wei Wei
------
Slides:  https://llvm.org/devmtg/2026-04/slides/student_technical_talk/s...]]></description>
<link>https://tsecurity.de/de/3627839/it-security-video/2026-eurollvm-accelerating-pass-order-auto-tuning-via-profile-guided-cost-modeling/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3627839/it-security-video/2026-eurollvm-accelerating-pass-order-auto-tuning-via-profile-guided-cost-modeling/</guid>
<pubDate>Fri, 26 Jun 2026 17:33:32 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: LLVM - Bewertung: 1x - Views:8 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/p0yQBfefSeY?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>2026 EuroLLVM Developers' Meeting<br />
https://llvm.org/devmtg/2026-04/<br />
------<br />
Title: Accelerating Pass Order Auto-tuning via Profile-Guided Cost Modeling<br />
Speaker: Bingyu Gao, Wei Wei<br />
------<br />
Slides:  https://llvm.org/devmtg/2026-04/slides/student_technical_talk/student_technical_talk_gao.pdf<br />
-----<br />
LLVM pass ordering auto-tuning can outperform standard -O3, but it is often hindered by an enormous search space and the high overhead of hundreds of dynamic measurements. This talk presents an efficient auto-tuning framework that minimizes expensive measurements using a profile-guided relative cost model and calibrated beam search. Evaluation on cBench shows an average 10.46% speedup over -O3 with only 20 dynamic measurements, significantly accelerating the search for optimal pass sequences.<br />
-----<br />
Videos Edited by Bash Films: http://www.BashFilms.com<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[pgEdge joins rush to merge OLTP and OLAP storage to support AI]]></title>
<description><![CDATA[For years, enterprises have maintained separate systems for processing transactional (OLTP) and analytical (OLAP) data, even if that meant moving data between them. However, the rise of autonomous agents and AI applications needing immediate access to data while generating volumes of operational ...]]></description>
<link>https://tsecurity.de/de/3627771/ai-nachrichten/pgedge-joins-rush-to-merge-oltp-and-olap-storage-to-support-ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3627771/ai-nachrichten/pgedge-joins-rush-to-merge-oltp-and-olap-storage-to-support-ai/</guid>
<pubDate>Fri, 26 Jun 2026 16:54:48 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>For years, enterprises have maintained separate systems for processing <a href="https://www.infoworld.com/article/2334535/what-is-oltp-the-backbone-of-ecommerce.html">transactional (OLTP)</a> and <a href="https://www.infoworld.com/article/2334471/what-is-olap-analytical-databases.html">analytical (OLAP)</a> data, even if that meant moving data between them. However, the rise of autonomous agents and AI applications needing immediate access to data while generating volumes of operational data themselves, has exposed the cost and complexity of maintaining those separate systems.</p>



<p>The industry’s response has been quick, with data warehouse and database vendors proposing a wave of competing approaches to collapsing those data silos. In the past few weeks Databricks unveiled <a href="https://www.infoworld.com/article/4185622/databricks-pitches-ltap-as-a-new-foundation-for-agentic-applications.html">LTAP</a> and EDB introduced <a href="https://www.infoworld.com/article/4188484/edb-converges-analytics-on-postgres-to-support-ai-agents.html">converged analytics</a>, while late last year Snowflake launched <a href="https://www.snowflake.com/en/blog/engineering/pg-lake-postgres-lakehouse-integration/">pg_lake</a>, all of which offer different blueprints for bringing transactional, analytical and AI workloads closer together.</p>



<p>Now it’s the turn of distributed <a href="https://www.infoworld.com/article/2266153/postgresql-benefits-and-challenges-a-snapshot.html">PostgreSQL</a> provider pgEdge, which has introduced a beta version of <a href="https://www.pgedge.com/solutions/postgres-tiered-storage" target="_blank" rel="noreferrer noopener">ColdFront</a>, a PostgreSQL-native hot-and-cold data tiering architecture that automatically moves older data into <a href="https://www.infoworld.com/article/3479001/why-apache-iceberg-is-on-fire-right-now.html">Apache Iceberg</a> object storage while keeping PostgreSQL as the only database that applications need to interact with.</p>



<p>In ColdFront’s architecture, hot and cold refer to newer and older data, respectively.</p>



<p>The approach of keeping PostgreSQL as the primary interface is what sets ColdFront apart from the other architectures emerging in this space, differing in where the center of gravity for data lies, according to analysts.</p>



<p>Databricks’ LTAP keeps operational applications connected to a lakehouse where analytics and AI are performed, EDB keeps PostgreSQL as the operational source of truth while exposing data through Iceberg for analytical engines, and Snowflake’s pg_lake writes PostgreSQL data directly into Iceberg so both PostgreSQL and Snowflake can query the same data, said <a href="https://www.hfsresearch.com/team/ashish-chaturvedi/" target="_blank" rel="noreferrer noopener">Ashish Chaturvedi</a>, leader of executive research at HFS Research.</p>



<p>ColdFront, by contrast, treats Iceberg only as a transparent storage tier behind PostgreSQL, automatically moving older data out of the database while keeping applications on the same tables and SQL, Chaturvedi said.</p>



<p>The result, according to pgEdge cofounder <a href="https://www.linkedin.com/in/phillipmerrick/" target="_blank" rel="noreferrer noopener">Phillip Merrick</a>, is that queries against recent data continue to run on PostgreSQL, while requests for older records are transparently executed using DuckDB’s embedded analytical engine, allowing applications to use the same SQL without introducing <a href="https://www.infoworld.com/article/2338277/modern-data-infrastructures-dont-do-etl.html">ETL</a> pipelines, separate query paths, or application changes.</p>



<p>That also means older records stored in Iceberg can be updated through PostgreSQL without requiring application changes, enabling what Merrick described as a “cold writable tier.”</p>



<h2 class="wp-block-heading">Why writable cold storage matters</h2>



<p>That cold writable tier could resonate with enterprises seeking to balance data residency, sovereignty, regulatory compliance and the growing operational demands of the agentic era, particularly because competing approaches generally require sacrificing at least one of those objectives.</p>



<p>As enterprises retain growing volumes of historical operational data generated by AI applications for audit and regulatory purposes, they increasingly need the ability to correct, delete or modify records, for example to comply with data protection and privacy laws, even after they have been moved into lower-cost storage, which other rival approaches complicate, said <a href="https://www.linkedin.com/in/amitchandak78/" target="_blank" rel="noreferrer noopener">Amit Chandak</a>, chief analytics officer at IT consulting firm Kanerika.</p>



<p>ColdFront can simplify those processes, said Chaturvedi: “In most tiering systems, cold (older) data is read-only, so a GDPR deletion request on archived data means restore-delete-rearchive, which is a half day job. ColdFront’s architecture would allow you to UPDATE and DELETE archived rows through one SQL statement.”</p>



<p>The rival architectures make different tradeoffs, with Databricks asking enterprises to adopt a proprietary lakehouse as the operational center of gravity, Snowflake requiring applications to distinguish between PostgreSQL and analytical tables, and EDB still requiring archived data to be brought back into active PostgreSQL before it can be modified, he said.</p>



<p>Those tradeoffs are particularly significant for regulated industries, according to <a href="https://www.infotech.com/profiles/igor-ikonnikov" target="_blank" rel="noreferrer noopener">Igor Ikonnikov</a>, advisory fellow at Info-Tech Research Group, who said enterprises in financial services, healthcare and government increasingly want to keep sensitive operational data on customer-controlled infrastructure while preserving the ability to modify historical records to meet evolving regulatory obligations.</p>



<h2 class="wp-block-heading">The DuckDB dependency</h2>



<p>Despite their architectural differences, all the vendors are masking an emerging convergence at another layer of the stack that CIOs should take note of: an increasing dependence on DuckDB.</p>



<p>“ColdFront uses DuckDB to execute queries against data stored in Iceberg. Snowflake’s pg_lake routes Iceberg queries through pgduck_server, and Databricks’ Lakebase also relies on DuckDB internally for parts of its analytical processing. As a result, DuckDB is rapidly becoming the de facto embedded analytics engine for this new generation of PostgreSQL-Iceberg architectures,” Ikonnikov said.</p>



<p>That growing dependence creates what the analyst described as a concentration risk: “If DuckDB faces licensing changes, security vulnerabilities, performance bottlenecks or governance issues, the impact would ripple across multiple products simultaneously.”</p>



<p>As a result, CIOs should understand the maturity and roadmap of the shared components these architectures increasingly depend on.</p>



<p>However, that similarity in shared components will not make evaluation of these competing architectures easier for CIOs.</p>



<p>Most enterprises already have established data architectures, said <a href="https://moorinsightsstrategy.com/team/mike-leone/" target="_blank" rel="noreferrer noopener">Michael Leone</a>, principal analyst at Moor Insights &amp; Strategy, arguing that CIOs should evaluate these platforms based on where their data, developers and operational workflows already reside rather than assuming one architecture fits every environment.</p>



<p>For enteprises still defining their long-term data strategy, Leone recommended standardizing on Iceberg first since all four architectures support the open table format and enterprises will retain the flexibility to replace the front-end database or analytical platform later without migrating the underlying data.</p>



<p>Even that portability, however, has limits, Ikonnikov cautioned.</p>



<p>“The issue is Iceberg catalog governance. All four approaches write to Iceberg, but they use different catalogs and their interoperability across vendors remains an open problem. When agents from different systems need to query the same Iceberg tables, catalog federation becomes a real operational challenge.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[pgEdge joins rush to merge OLTP and OLAP storage to support AI]]></title>
<description><![CDATA[For years, enterprises have maintained separate systems for processing transactional (OLTP) and analytical (OLAP) data, even if that meant moving data between them. However, the rise of autonomous agents and AI applications needing immediate access to data while generating volumes of operational ...]]></description>
<link>https://tsecurity.de/de/3627752/it-nachrichten/pgedge-joins-rush-to-merge-oltp-and-olap-storage-to-support-ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3627752/it-nachrichten/pgedge-joins-rush-to-merge-oltp-and-olap-storage-to-support-ai/</guid>
<pubDate>Fri, 26 Jun 2026 16:51:33 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>For years, enterprises have maintained separate systems for processing <a href="https://www.infoworld.com/article/2334535/what-is-oltp-the-backbone-of-ecommerce.html">transactional (OLTP)</a> and <a href="https://www.infoworld.com/article/2334471/what-is-olap-analytical-databases.html">analytical (OLAP)</a> data, even if that meant moving data between them. However, the rise of autonomous agents and AI applications needing immediate access to data while generating volumes of operational data themselves, has exposed the cost and complexity of maintaining those separate systems.</p>



<p>The industry’s response has been quick, with data warehouse and database vendors proposing a wave of competing approaches to collapsing those data silos. In the past few weeks Databricks unveiled <a href="https://www.infoworld.com/article/4185622/databricks-pitches-ltap-as-a-new-foundation-for-agentic-applications.html">LTAP</a> and EDB introduced <a href="https://www.infoworld.com/article/4188484/edb-converges-analytics-on-postgres-to-support-ai-agents.html">converged analytics</a>, while late last year Snowflake launched <a href="https://www.snowflake.com/en/blog/engineering/pg-lake-postgres-lakehouse-integration/" rel="nofollow">pg_lake</a>, all of which offer different blueprints for bringing transactional, analytical and AI workloads closer together.</p>



<p>Now it’s the turn of distributed <a href="https://www.infoworld.com/article/2266153/postgresql-benefits-and-challenges-a-snapshot.html">PostgreSQL</a> provider pgEdge, which has introduced a beta version of <a href="https://www.pgedge.com/solutions/postgres-tiered-storage" target="_blank" rel="nofollow">ColdFront</a>, a PostgreSQL-native hot-and-cold data tiering architecture that automatically moves older data into <a href="https://www.infoworld.com/article/3479001/why-apache-iceberg-is-on-fire-right-now.html">Apache Iceberg</a> object storage while keeping PostgreSQL as the only database that applications need to interact with.</p>



<p>In ColdFront’s architecture, hot and cold refer to newer and older data, respectively.</p>



<p>The approach of keeping PostgreSQL as the primary interface is what sets ColdFront apart from the other architectures emerging in this space, differing in where the center of gravity for data lies, according to analysts.</p>



<p>Databricks’ LTAP keeps operational applications connected to a lakehouse where analytics and AI are performed, EDB keeps PostgreSQL as the operational source of truth while exposing data through Iceberg for analytical engines, and Snowflake’s pg_lake writes PostgreSQL data directly into Iceberg so both PostgreSQL and Snowflake can query the same data, said <a href="https://www.hfsresearch.com/team/ashish-chaturvedi/" target="_blank" rel="nofollow">Ashish Chaturvedi</a>, leader of executive research at HFS Research.</p>



<p>ColdFront, by contrast, treats Iceberg only as a transparent storage tier behind PostgreSQL, automatically moving older data out of the database while keeping applications on the same tables and SQL, Chaturvedi said.</p>



<p>The result, according to pgEdge cofounder <a href="https://www.linkedin.com/in/phillipmerrick/" target="_blank" rel="nofollow">Phillip Merrick</a>, is that queries against recent data continue to run on PostgreSQL, while requests for older records are transparently executed using DuckDB’s embedded analytical engine, allowing applications to use the same SQL without introducing <a href="https://www.infoworld.com/article/2338277/modern-data-infrastructures-dont-do-etl.html">ETL</a> pipelines, separate query paths, or application changes.</p>



<p>That also means older records stored in Iceberg can be updated through PostgreSQL without requiring application changes, enabling what Merrick described as a “cold writable tier.”</p>



<h2 class="wp-block-heading">Why writable cold storage matters</h2>



<p>That cold writable tier could resonate with enterprises seeking to balance data residency, sovereignty, regulatory compliance and the growing operational demands of the agentic era, particularly because competing approaches generally require sacrificing at least one of those objectives.</p>



<p>As enterprises retain growing volumes of historical operational data generated by AI applications for audit and regulatory purposes, they increasingly need the ability to correct, delete or modify records, for example to comply with data protection and privacy laws, even after they have been moved into lower-cost storage, which other rival approaches complicate, said <a href="https://www.linkedin.com/in/amitchandak78/" target="_blank" rel="nofollow">Amit Chandak</a>, chief analytics officer at IT consulting firm Kanerika.</p>



<p>ColdFront can simplify those processes, said Chaturvedi: “In most tiering systems, cold (older) data is read-only, so a GDPR deletion request on archived data means restore-delete-rearchive, which is a half day job. ColdFront’s architecture would allow you to UPDATE and DELETE archived rows through one SQL statement.”</p>



<p>The rival architectures make different tradeoffs, with Databricks asking enterprises to adopt a proprietary lakehouse as the operational center of gravity, Snowflake requiring applications to distinguish between PostgreSQL and analytical tables, and EDB still requiring archived data to be brought back into active PostgreSQL before it can be modified, he said.</p>



<p>Those tradeoffs are particularly significant for regulated industries, according to <a href="https://www.infotech.com/profiles/igor-ikonnikov" target="_blank" rel="nofollow">Igor Ikonnikov</a>, advisory fellow at Info-Tech Research Group, who said enterprises in financial services, healthcare and government increasingly want to keep sensitive operational data on customer-controlled infrastructure while preserving the ability to modify historical records to meet evolving regulatory obligations.</p>



<h2 class="wp-block-heading">The DuckDB dependency</h2>



<p>Despite their architectural differences, all the vendors are masking an emerging convergence at another layer of the stack that CIOs should take note of: an increasing dependence on DuckDB.</p>



<p>“ColdFront uses DuckDB to execute queries against data stored in Iceberg. Snowflake’s pg_lake routes Iceberg queries through pgduck_server, and Databricks’ Lakebase also relies on DuckDB internally for parts of its analytical processing. As a result, DuckDB is rapidly becoming the de facto embedded analytics engine for this new generation of PostgreSQL-Iceberg architectures,” Ikonnikov said.</p>



<p>That growing dependence creates what the analyst described as a concentration risk: “If DuckDB faces licensing changes, security vulnerabilities, performance bottlenecks or governance issues, the impact would ripple across multiple products simultaneously.”</p>



<p>As a result, CIOs should understand the maturity and roadmap of the shared components these architectures increasingly depend on.</p>



<p>However, that similarity in shared components will not make evaluation of these competing architectures easier for CIOs.</p>



<p>Most enterprises already have established data architectures, said <a href="https://moorinsightsstrategy.com/team/mike-leone/" target="_blank" rel="nofollow">Michael Leone</a>, principal analyst at Moor Insights &amp; Strategy, arguing that CIOs should evaluate these platforms based on where their data, developers and operational workflows already reside rather than assuming one architecture fits every environment.</p>



<p>For enteprises still defining their long-term data strategy, Leone recommended standardizing on Iceberg first since all four architectures support the open table format and enterprises will retain the flexibility to replace the front-end database or analytical platform later without migrating the underlying data.</p>



<p>Even that portability, however, has limits, Ikonnikov cautioned.</p>



<p>“The issue is Iceberg catalog governance. All four approaches write to Iceberg, but they use different catalogs and their interoperability across vendors remains an open problem. When agents from different systems need to query the same Iceberg tables, catalog federation becomes a real operational challenge.”</p>



<p><em>This article first appeared on <a href="https://www.infoworld.com/article/4190042/pgedge-joins-rush-to-merge-oltp-and-olap-storage-to-support-ai.html">InfoWorld</a>.</em></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[31 Wege, LLMs zu evaluieren]]></title>
<description><![CDATA[width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px">Metriken und Benchmarks gibt es viele. Welche Sie in Sachen KI weiterbringen können, lesen Sie in diesem Beitrag.Andrey_Popov | shutterstock.com



Was man nicht messen kann, kann man bekanntlich auch nicht managen. Im Hinb...]]></description>
<link>https://tsecurity.de/de/3623312/it-security-nachrichten/31-wege-llms-zu-evaluieren/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3623312/it-security-nachrichten/31-wege-llms-zu-evaluieren/</guid>
<pubDate>Thu, 25 Jun 2026 06:08:17 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large is-resized"> width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"&gt;<figcaption class="wp-element-caption">Metriken und Benchmarks gibt es viele. Welche Sie in Sachen KI weiterbringen können, lesen Sie in diesem Beitrag.</figcaption></figure><p class="imageCredit">Andrey_Popov | shutterstock.com</p></div>



<p>Was man nicht messen kann, kann man bekanntlich auch nicht managen. Im Hinblick auf Large Language Models (<a href="https://www.computerwoche.de/article/4155050/25-fragen-die-zum-richtigen-llm-fuhren.html" target="_blank">LLMs</a>) und <a href="https://www.computerwoche.de/article/4132787/wie-ki-agenten-daten-konsumieren-sollten.html" target="_blank">KI-Agenten</a> steht die Entwicklung statistischer Metriken immer noch am Anfang. </p>



<p>In diesem Artikel haben wir die derzeit gängigsten Kennzahlen und Benchmarks zusammengetragen, die Entwickler und Benutzer heranziehen können, um KI-Modelle oder -Agenten zu evaluieren.</p>



<h2 class="wp-block-heading">Time to First Token</h2>



<p>Bei der Time to First Token geht es um die Frage, wie lange es im Durchschnitt dauert, bis der erste <a href="https://www.computerwoche.de/article/4182846/ki-token-erklart.html" target="_blank">Token</a> generiert wird. Insbesondere, wenn es um Echtzeit-Applikationen geht, können schnellere Antworten entscheidend sein. Zudem dürfte die menschliche Abneigung gegen Warten hinlänglich bekannt sein: Die Teams, die die Benutzeroberflächen entwickeln, haben schon vor Jahrzehnten gelernt, wie wichtig es ist, dass Software möglichst <a href="https://www.computerwoche.de/article/2805659/wie-schlechte-apps-das-firmenimage-ruinieren.html" target="_blank">schnell reagiert</a>. Schon eine Verzögerung von wenigen (Milli-)Sekunden kann bedeuten, dass der Benutzer zu einem anderen Fenster wechselt.</p>



<p>Die Time to First Token ist ein guter Maßstab für Modelle, die direkt mit der unbeständigen menschlichen Intelligenz und deren latenter Aufmerksamkeitsdefizitstörung in Berührung kommen.</p>



<h2 class="wp-block-heading">Time per Output Token</h2>



<p>Die Time to First Token gibt an, wie lange es dauert, bis eine Antwort generiert wird. Die Time per Output Token hingegen die durchschnittliche Geschwindigkeit, mit der das Modell sämtliche Token durchläuft. Diese Metrik wird berechnet, indem die für die Antwort benötigte Gesamtzeit durch die Gesamtanzahl der Token dividiert wird.</p>



<p>Bei <a href="https://www.computerwoche.de/article/3972574/brauchen-sie-wirklich-ein-llm.html" target="_blank">einfachen LLMs</a> ist dieser Wert im Allgemeinen ziemlich konstant. Sobald der Prefill-Prozess abgeschlossen ist und das Modell in die Dekodierungsphase eintritt, erscheinen die Output-Token in der Regel in einem konstanten Stream. Ist der Output lang genug, amortisiert sich die Time to First Token. In komplexeren Architekturen mit Planungs-Loops oder einer Datenerfassung über verschiedene, weitere Tools kann die durchschnittliche Geschwindigkeit allerdings variieren, während das Modell von einer Agentic Decision zur nächsten wechselt.</p>



<h2 class="wp-block-heading">Token pro Sekunde</h2>



<p>Hierbei handelt es sich lediglich um den Kehrwert der Average Time per Token. Dieser wird manchmal separat für verschiedene Phasen in der Pipeline angegeben.</p>



<h2 class="wp-block-heading">Durchsatz (Requests pro Minute)</h2>



<p>Wenn ein System mehr als einen einzelnen Benutzer unterstützt, ist es sinnvoll, die Anzahl der möglichen unterschiedlichen Requests zu erfassen. Diese Throughput-Metriken können äußerst nützlich sein, um die Performanz von <a href="https://www.computerwoche.de/article/4183987/embedding-pipelines-sind-das-neue-etl.html" target="_blank">Pipelines</a> zu messen, die effizienter sind, wenn sie mehrere Prompts parallel beantworten.</p>



<h2 class="wp-block-heading">Fehlerrate</h2>



<p>Nicht auf jeden Request folgt eine Antwort. Die Fehlerrate (Error Rate) erfasst, wie oft Rate Limits, Zeitüberschreitungen oder „Ablehnungen“ durch das Modell auftreten. Noch besser ist es unter Umständen, jede Kategorie separat zu tracken, da die Anzahl der Fehler in diesen sehr unterschiedlich ausfallen kann.</p>



<h2 class="wp-block-heading">Token-Effizienz</h2>



<p>Nicht alle Token sind sichtbar – oder Teil des Endergebnisses. Die Token-Effizienz misst, wie viel Arbeit aufgewendet wird, um das Endergebnis zu erzielen. Wenn Modelle und Pipelines komplexer ausgestaltet sind, sinkt die Token-Effizienz tendenziell. Agentic Reasoning und strategische Planungsaufgaben erfordern in der Regel mehr Token, die nicht in der finalen Antwort auftauchen. Diese Metrik kann einen Anhaltspunkt darüber liefern, <a href="https://www.computerwoche.de/article/4182741/nur-jedes-vierte-unternehmen-hat-seine-ki-kosten-im-blick.html" target="_blank">wie teuer es wird</a>, ein Modell zu betreiben.</p>



<h2 class="wp-block-heading">Tail-Latenz</h2>



<p>Es ist schön und gut, die durchschnittliche Antwortzeit zu messen. In einigen Fällen können aber bereits einige wenige langsame Antworten das Urteil der Benutzer beeinflussen. Beispiel <a href="https://www.computerwoche.de/podcast/3483247/autonomes-fahren-mit-jurgen-hill.html" target="_blank">autonomes Fahren</a>: Ein selbstfahrendes Auto, das Lenkbefehle „im Schnitt“ sehr schnell umsetzt, wäre wenig vertrauenserweckend.</p>



<p>Die Tail-Latenz nutzt eine Kombination aus Queue-Theorie und detaillierten Messungen, um die „schlimmsten“ Momente im Longtail des Latenzdiagramms zu erfassen. Diese Metrik ist überall dort nützlich, wo selbst gelegentliche Verzögerungen höchst problematisch sind.</p>



<h2 class="wp-block-heading">Total Cost of Ownership (TCO)</h2>



<p>Projekte, die <a href="https://www.computerwoche.de/article/4004872/die-besten-apis-um-ki-zu-integrieren.html" target="_blank">eine API nutzen</a> oder Outputs einkaufen, betrachten lediglich die Kosten pro einer Million Token. Die Teams, die stattdessen <a href="https://www.computerwoche.de/article/4014872/ki-kosten-sparen-mit-neoclouds.html" target="_blank">GPUs</a> kaufen und für den Strom bezahlen müssen, addieren diese sowie weitere indirekte Kosten (etwa für Abschreibungen und Wartung) auf, um einen Wert zu ermitteln, der einen Eindruck davon vermittelt, wie viel es tatsächlich kostet, die nötigen Token zu produzieren.</p>



<p>Dabei hängt die Total Cost of Ownership (TCO) von der Nachfrage und den Auslastungsraten ab – also davon, wie viele Nutzer Prompts senden und wie effizient das Modell auf eine bestimmte GPU und deren RAM abgestimmt ist.</p>



<h2 class="wp-block-heading">Parameter-Anzahl</h2>



<p>Diverse <a href="https://www.computerwoche.de/article/4173136/17-llms-fur-spezialdomanen.html" target="_blank">LLMs</a> tragen die Anzahl ihrer Parameter im Namen, also die Anzahl der Variablen, die das Modell nutzt, um aus Inputs Outputs zu generieren. „70B“ bedeutet in diesem Zusammenhang beispielsweise, dass das Modell 70 Milliarden (englisch “Billion”) Parameter enthält. Das liefert einen guten Anhaltspunkt darüber, wie komplex dieses ist und wie groß der Trainingsdatensatz war, auf dem es trainiert wurde. Im Allgemeinen bedeuten größere Parameterwerte, dass das Modell auch eine größere Menge an Informationen beinhaltet. Das geht allerdings auch oft mit der Anforderung einher, performantere GPUs mit mehr Arbeitsspeicher einsetzen zu müssen.</p>



<p>Die Anzahl der Parameter ist jedoch keine sehr präzise Metrik, weil viele andere Architekturbereiche beeinflussen können, ob das Modell das gewünschte Ergebnis innerhalb des Budgets realisieren kann. Deshalb ist eine ausufernde Zahl von Parametern kein Garant für überragende Performance.</p>



<h2 class="wp-block-heading">Halluzinationsrate</h2>



<p>Akkurate LLM-Outputs sind als allgemeines Ziel gesetzt. Zu messen, wie genau die Ergebnisse wirklich sind, ist allerdings diffizil. Ein Ansatz besteht darin, das LLM zu bitten, ein Dokument zusammenzufassen. Anschließend bewertet ein weiteres Modell, wie gut die Zusammenfassung mit dem Original übereinstimmt. Auch wenn dabei möglicherweise nicht alle subtilen Ungenauigkeiten erkannt werden können, kann das Aufschluss über gravierende Abweichungen – sprich <a href="https://www.computerwoche.de/article/3829267/so-bleibt-ihr-code-halluzinationsfrei.html" target="_blank">Halluzinationen</a> – geben. </p>



<p>Einige Forscher haben zudem komplexe Test-Sets mit kuratierten Antworten erstellt. Diejenigen LLMs, die in diesen die erwarteten Outputs liefern, erhalten die höchsten Scores. Gängige Benchmarks, um die Halluzinationsrate zu erfassen sind etwa:</p>



<ul class="wp-block-list">
<li><a href="https://github.com/sylinrl/TruthfulQA" target="_blank" rel="noreferrer noopener">TruthfulQA</a>,</li>



<li><a href="https://arxiv.org/abs/2305.11747" target="_blank" rel="noreferrer noopener">HaluEval</a>,</li>



<li><a href="https://github.com/salesforce/QAFactEval" target="_blank" rel="noreferrer noopener">QAFactEval</a> und</li>



<li><a href="https://github.com/vectara/hallucination-leaderboard" target="_blank" rel="noreferrer noopener">Hallucination Evaluation Model</a>.</li>
</ul>



<h2 class="wp-block-heading">Toxizitäts- und Bias-Score</h2>



<p>Noch ein wenig schwieriger wird es, wenn eine Metrik entwickelt werden soll, um toxische oder <a href="https://www.computerwoche.de/article/4020215/ki-empfiehlt-frauen-systematisch-niedrigere-gehalter.html" target="_blank">Bias-behaftete Outputs</a> zu erfassen. Denn auch hier können die jeweils zugrundeliegenden Definition höchst unterschiedlich ausfallen. Das hat einige Spezialisten dennoch nicht davon abgehalten, entsprechende Tools zu entwickeln. </p>



<p>Diese sind in der Lage, mit Blick auf unwillkommene Ergebnisse einige der offensichtlichsten Warnsignale zu erkennen. Zu den bekannteren Lösungen in diesem Bereich zählen:</p>



<ul class="wp-block-list">
<li><a href="https://www.granica.ai/blog/granica-launches-ai-data-safety-solution-granica-screen-on-aws-marketplace" target="_blank" rel="noreferrer noopener">Granica Screen</a> und</li>



<li><a href="https://perspectiveapi.com/" target="_blank" rel="noreferrer noopener">Perspective API</a>.</li>
</ul>



<h2 class="wp-block-heading">Tool-Calling-Genauigkeit</h2>



<p>Je komplexer und „agentischer“ KI-Modelle werden, desto eher greifen sie auch auf diverse weitere Tools zu – beispielsweise über das Model Context Protocol (<a href="https://www.computerwoche.de/article/4031227/was-ist-model-context-protocol.html" target="_blank">MCP</a>). Wenn MCP allerdings nicht zum Einsatz kommt, macht es Sinn, zu tracken, wie akkurat das Modell sich bei der Auswahl von Tools verhält – etwa, wie oft es das am besten geeignete Werkzeug für eine spezifische Aufgabe aufruft. Diese Metrik kann beispielsweise über das <a href="https://gorilla.cs.berkeley.edu/leaderboard.html" target="_blank" rel="noreferrer noopener">Berkeley Function Calling Leaderboard</a> (BFCL) eingeholt werden.</p>



<h2 class="wp-block-heading">Prompt-Sensitivität</h2>



<p>Dieser Wert gibt an, inwieweit bereits geringfügige Änderungen am Wortlaut des <a href="https://www.computerwoche.de/article/4042963/5-tipps-um-besser-zu-prompten.html" target="_blank">Prompts</a> das KI-Modell dazu veranlassen, unterschiedliche Ergebnisse zu liefern. Das ist vergleichbar mit einer Ableitung in der Stochastik, wird jedoch in der Regel experimentell anhand einer Sammlung von Test-Prompts berechnet. Um diese Metrik zu erfassen, gibt es eine Reihe verschiedener Ansätze, die jeweils unterschiedliche Schwerpunkte setzen. </p>



<p>Einige Test-Sets basieren auf geringfügigen Umformulierungen der Anfrage, die semantisch identisch sind. Andere kombinieren verschiedene Arten der Problemformulierung – etwa mit Hilfe von Beispielen. Zu den bekannteren Ansätzen in diesem Bereich zählen beispielsweise:</p>



<ul class="wp-block-list">
<li><a href="https://arxiv.org/html/2509.13680" target="_blank" rel="noreferrer noopener">PromptSE</a> und</li>



<li><a href="https://arxiv.org/abs/2410.12405" target="_blank" rel="noreferrer noopener">ProSA</a>.</li>
</ul>



<h2 class="wp-block-heading">Semantische Ähnlichkeit und Prägnanz</h2>



<p>Einige Metriken bewerten den Modell-Output, indem sie diesen mit einer Reihe von Goldstandard-Antworten vergleichen. Dazu werden die Antworten häufig in ein <a href="https://www.computerwoche.de/article/2829270/warum-vektorisierung-die-basis-fuer-genai-ist.html" target="_blank">Vektor-Embedding-Modell</a> eingespeist und mit einer Retrieval-Augmented-Generation (<a href="https://www.computerwoche.de/article/2829632/so-daemmen-sie-ki-bullshit-ein.html" target="_blank">RAG</a>) -Datenbank abgeglichen. Auf diese Weise lässt sich nicht nur erfassen, wie prägnant oder auch nicht der Output ist. Sondern auch, inwieweit sich dieser durch die Änderung von Parametern – etwa der „Temperature“ – beeinflussen lässt. Ein gängiges Tool um die semantische Ähnlichkeit und Prägnanz von LLMs zu erfassen, ist etwa <a href="https://bertscore.com/" target="_blank" rel="noreferrer noopener">BERTScore</a>.</p>



<h2 class="wp-block-heading">Grounding-Score</h2>



<p>Bei KI-Systemen, die ein LLM mit einem vektorbasierten Suchwerkzeug für RAG zusammenbringen, wird die Effektivität dieser Kombination anhand eines Benchmarks wie dem Grounding-Score bemessen. Kommt dieser zur Anwendung, werden dem KI-Modell zusätzliche Daten aus der Vektorsuche bereitgestellt. </p>



<p>Der Benchmark misst dann, wie nah das Modell an diesen Zusatzinformationen bleibt – also, zu welchem Anteil der Output jeweils aus den Quelldokumenten und den Trainingsdaten generiert wird. Das geht zum Beispiel mit:</p>



<ul class="wp-block-list">
<li><a href="https://aclanthology.org/2024.eacl-demo.16/" target="_blank" rel="noreferrer noopener">RAGAS</a>,</li>



<li><a href="https://www.trulens.org/" target="_blank" rel="noreferrer noopener">TruLens</a>,</li>



<li><a href="https://ares-ai.vercel.app/" target="_blank" rel="noreferrer noopener">ARES</a>,</li>



<li><a href="https://github.com/chen700564/RGB" target="_blank" rel="noreferrer noopener">RGB</a>,</li>



<li><a href="https://arxiv.org/abs/2305.11747" target="_blank" rel="noreferrer noopener">HaluEval</a>, und</li>



<li><a href="https://halluhard.com/" target="_blank" rel="noreferrer noopener">HalluHard</a>.</li>
</ul>



<h2 class="wp-block-heading">Modellvariabilität</h2>



<p>Die meisten LLMs beinhalten ein gewisses Maß an „Random Entropy“, das sich über den „Temperature“-Parameter steuern lässt. Die Modellvariabilität ist ein Maß dafür, wie stark sich die Antworten der KI von Session zu Session verändern. Einige Anwendungen – etwa <a href="https://www.cio.de/article/4185473/wie-ikea-aus-einem-chatbot-ein-milliarden-business-machte.html" target="_blank">Chatbots</a> – erfordern ein gewisses Maß an Variabilität, da die Zufälligkeit den Antworten „Leben“ einhauchen. In anderen Fällen, etwa wenn es um Recht oder Medizin geht, kann eine zu ausgeprägte Variabilität der Antworten das Nutzervertrauen untergraben.</p>



<h2 class="wp-block-heading">Format-Compliance-Rate</h2>



<p>Manche Use Cases erfordern es, dass KI-Modelle Daten in strikten Formaten wie <a href="https://www.computerwoche.de/article/2815830/was-ist-json.html" target="_blank">JSON</a> oder CSV generieren. Zum Beispiel, wenn diese Informationen zur weiteren Verarbeitung in eine Pipeline eingespeist werden sollen. Die Format-Compliance-Rate testet eine Reihe gängiger Formate und misst, wie oft das LLM semantisch korrekte Daten zurückgibt. Insbesondere Agentic-AI-Systeme, die mehrere Modelle und Tools miteinander kombinieren, sind auf LLMs angewiesen, die bei diesem Benchmark gute Ergebnisse erzielen.</p>



<h2 class="wp-block-heading">Instruction Following</h2>



<p>Manche Prompts enthalten sehr spezifische Anweisungen, deren Einhaltung sich empirisch messen lässt. Ein Beispiel wäre etwa ein Prompt, der die KI anweist, genau 300 Wörter oder ein Gedicht in Reimpaaren zu erstellen. Instruction-Following-Tests greifen auf eine Sammlung von Beispiel-Prompts zurück und sind darauf ausgelegt, Outputs zu erzeugen, die sich leicht messen lassen. Konkrete Beispiele hierfür sind:</p>



<ul class="wp-block-list">
<li><a href="https://arxiv.org/abs/2311.07911" target="_blank" rel="noreferrer noopener">IFEval</a>,</li>



<li><a href="https://github.com/YJiangcm/FollowBench" target="_blank" rel="noreferrer noopener">FollowBench</a> und</li>



<li><a href="https://gorilla.cs.berkeley.edu/leaderboard.html" target="_blank" rel="noreferrer noopener">BFCL</a>.</li>
</ul>



<h2 class="wp-block-heading">Planstabilität</h2>



<p>Agentic-Modelle agieren auf Basis eines Plans. Manche sind intelligent genug, um diesen Plan im Verlauf ihrer Arbeit anzupassen oder auch zu verwerfen. Wie oft der Plan angepasst wird, lässt sich über die Planstabilität messen. Ein niedriger Wert könnte bedeuten, dass der Agent eher schlecht plant – oder einfach nur flexibel ist. Vielleicht auch beides.</p>



<h2 class="wp-block-heading">Selbstkorrektur-Wert</h2>



<p>Einige KI-Agenten sind zudem in der Lage, tiefer in die jeweilige Materie abzutauchen und ihre eigenen Fehler zu erkennen. Die Selbstkorrektur-Metrik misst, wie oft das Modell einen Fehler macht und diesen anschließend erkennt – entweder eigenständig oder nachdem es mit einer entsprechenden Nachfrage dazu aufgefordert wurde.</p>



<h2 class="wp-block-heading">Jailbreak-Resistenz</h2>



<p>Nutzer versuchen immer wieder, neue smarte Wege aufzutun, um KI-Modelle dazu zu verleiten, ihre Guardrails <a href="https://www.computerwoche.de/article/4142947/10-llms-die-wenig-bis-keine-grenzen-kennen.html" target="_blank">hinter sich zu lassen</a> und Themen zu diskutieren, die sie nicht diskutieren sollten. In der Vergangenheit ließen sich manche LLMs etwa täuschen, indem man ihnen mitteilte, ihr Output sei Teil eines fiktionalen Werks. Neuere Modelle verfügen inzwischen über ausgefeiltere Abwehrmechanismen. Um zu ermitteln, wie gut ein LLM solchen Täuschungsversuchen widersteht, empfehlen sich folgende Benchmarks:</p>



<ul class="wp-block-list">
<li><a href="https://jailbreakbench.github.io/" target="_blank" rel="noreferrer noopener">JailbreakBench</a>,</li>



<li><a href="https://arxiv.org/abs/2410.09024" target="_blank" rel="noreferrer noopener">AgentHarm</a> und</li>



<li><a href="https://arxiv.org/pdf/2512.05485" target="_blank" rel="noreferrer noopener">Tele-AI-Safety</a>.</li>
</ul>



<h2 class="wp-block-heading">Prompt-Injection-Anfälligkeit</h2>



<p>Bekanntlich können nicht vertrauenswürdige Daten aus zusätzlichen Quellen oder Skills <a href="https://www.computerwoche.de/article/4044551/wenn-der-ki-agent-im-fakeshop-kauft.html" target="_blank">schadhafte Anweisungen</a> enthalten, mit denen das KI-Modell kompromittiert werden soll. Wie anfällig ein Modell für solche gezielten Prompt-Injection-Angriffe ist, lässt sich über entsprechende Benchmarks messen, die auf der Grundlage bekannter Angriffsvektoren operieren. Zum Beispiel:  </p>



<ul class="wp-block-list">
<li><a href="https://arxiv.org/abs/2602.20156" target="_blank" rel="noreferrer noopener">Skill-Inject</a> oder</li>



<li><a href="https://spikee.ai/" target="_blank" rel="noreferrer noopener">SPIKEE</a>.</li>
</ul>



<h2 class="wp-block-heading">Copyright-Infringement-Score</h2>



<p>Manche LLMs tendieren dazu, die Daten ihres Trainingskorpus so wiederzugeben, dass der Output in die Nähe eines Plagiats rückt. Das kann ein größeres Problem darstellen, wenn das Trainingsmaterial nicht ordnungsgemäß (oder sorgfältig genug) lizenziert wurde. Der Copyright-Infringement-Score misst, wie oft ein KI-Modell sein Trainingsmaterial möglicherweise etwas zu wörtlich wiedergibt. Zu den Tools, die solche Probleme zutage fördern können, gehören:</p>



<ul class="wp-block-list">
<li><a href="https://www.patronus.ai/blog/introducing-copyright-catcher" target="_blank" rel="noreferrer noopener">CopyrightCatcher</a> und</li>



<li><a href="https://arxiv.org/abs/2402.09910" target="_blank" rel="noreferrer noopener">DE-COP</a>.</li>
</ul>



<h2 class="wp-block-heading">RULER</h2>



<p>Der Needle-in-a-Haystack (<a href="https://github.com/gkamradt/needle-in-a-haystack" target="_blank" rel="noreferrer noopener">NIAH</a>) –<a href="https://arxiv.org/pdf/2504.04713" target="_blank" rel="noreferrer noopener">Benchmark</a> (PDF) misst, wie gut ein KI-Modell spezifische Informationen aus dem Gesamtkontext extrahieren kann. Der <a href="https://github.com/NVIDIA/RULER" target="_blank" rel="noreferrer noopener">RULER</a>-Benchmark erweitert diesen Ansatz und bietet die Möglichkeit, die Art und Anzahl der „Nadeln“, die Größe des „Heuhaufens“ sowie die Komplexität der Aufgabe zu variieren.</p>



<h2 class="wp-block-heading">GSM8K</h2>



<p>Die Entwickler von <a href="https://arxiv.org/abs/2110.14168" target="_blank" rel="noreferrer noopener">GSM8K</a> wollten die Fähigkeit von LLM zur Lösung mehrstufiger mathematischer Probleme evaluieren. Dafür erstellten sie einen Datensatz mit 8.500 Aufgaben. Der Schwerpunkt liegt hierbei zwar ausdrücklich auf der Lösung von Mathematik-Aufgaben – allerdings misst dieser Benchmark auch die Fähigkeit, Reasoining-Ketten aufzubauen.</p>



<h2 class="wp-block-heading">GPQA</h2>



<p>Der Benchmark <a href="https://arxiv.org/pdf/2311.12022" target="_blank" rel="noreferrer noopener">Graduate-Level Google-Proof Q&amp;A</a> (PDF) besteht aus Hunderten komplexen Fragen, mit denen sich normalerweise Studierende in Master-Programmen beschäftigen – vor allem im naturwissenschaftlichen Bereich. Der Begriff „Google-proof“ bedeutet dabei, dass diese Fragen nicht einfach durch eine Suchmaschine bewältigt werden können. Um den Benchmark noch anspruchsvoller zu gestalten, konzentrierten sich die Forscher bei GPQA vor allem auf solche Fragen, die von Laien oft falsch beantwortet werden.</p>



<h2 class="wp-block-heading">MMLU-Pro</h2>



<p>Der <a href="https://github.com/TIGER-AI-Lab/MMLU-Pro" target="_blank" rel="noreferrer noopener">MMLU-Pro</a>-Benchmark baut auf dem „Massive Multitask Language Understanding“-Datensatz auf. Er ist darauf konzipiert, das Verständnis eines Modells für ein breites wissenschaftliches Spektrum zu testen. Dazu umfasst dieser Benchmark mehr als 12.000 Fragen – unter anderem aus den Bereichen Biologie, Chemie, Wirtschaft und Recht.</p>



<h2 class="wp-block-heading">MBPP</h2>



<p>Der <a href="https://github.com/google-research/google-research/tree/master/mbpp" target="_blank" rel="noreferrer noopener">MBPP</a>-Benchmark (Mostly Basic Python Problems) wurde bei Google entwickelt, um zu bewerten, wie gut KI-Modelle Programmieraufgaben lösen. Jedes Problem besteht dabei aus einem Statement, einer Referenzlösung sowie mehreren ähnlichen Testfällen. Die Anzahl der korrekten Antworten auf diese Fragen ist ein guter Maßstab dafür, wie gut oder schlecht ein Modell einfachere Python-Programmieraufgaben lösen wird.</p>



<h2 class="wp-block-heading">SWE-bench</h2>



<p>Auch <a href="https://github.com/SWE-bench/SWE-bench" target="_blank" rel="noreferrer noopener">SWE-bench</a> ermittelt, wie gut ein LLM Coding-Tasks erfüllt. Dieser Benchmark wurde auf der Basis von Issues und entsprechenden Pull Requests einer Reihe von Python-Projekten erstellt. Aufgrund einiger Limitationen wurde der Datensatz inzwischen erweitert. Das manifestiert sich in drei erweiterten Benchmarks:  </p>



<ul class="wp-block-list">
<li><a href="https://arxiv.org/abs/2410.06992" target="_blank" rel="noreferrer noopener">SWE-Bench+</a>,</li>



<li><a href="https://openai.com/index/introducing-swe-bench-verified/" target="_blank" rel="noreferrer noopener">SWE Bench Verified</a> und</li>



<li><a href="https://arxiv.org/abs/2509.16941" target="_blank" rel="noreferrer noopener">SWE-Bench Pro</a>.</li>
</ul>



<h2 class="wp-block-heading">LMSYS Chatbot Arena</h2>



<p>Bei <a href="https://www.lmsys.org/" target="_blank" rel="noreferrer noopener">LMSYS Chatbot Arena</a> handelt es sich um ein dynamisches System, das nicht auf ein fest definiertes Set von Test-Prompts setzt. Stattdessen füttert diese Benchmark-Plattform unterschiedliche KI-Modelle mit dem identischen Prompt und überlässt es dann Menschen, die besten Ergebnisse auszuwählen. Aus diese Direktenvergleichen ergibt sich <a href="https://de.wikipedia.org/wiki/Elo-Zahl" target="_blank" rel="noreferrer noopener">Elo</a>-ähnliche Wertung.</p>



<h2 class="wp-block-heading">Preis</h2>



<p>Jeder Immobilenmakler weiß: Die drei wichtigsten Metriken in einer Anzeige sind der Preis, der Preis – und der Preis. Dieser spielt bei der Evaluierung von KI-Systemen zwar eine etwas geringe Rolle – kann aber darüber entscheiden, ob das Projekt am Ende <a href="https://www.computerwoche.de/article/4180035/das-ki-preisproblem-zwischen-roi-druck-und-unkalkulierbaren-kosten.html" target="_blank">rentabel ist oder nicht</a>. Wenn die Kosten pro Inferenz ein bisschen zu hoch sind, lässt sich das nicht durch das Volumen ausgleichen.</p>



<p>Ein günstigeres KI-Modell ist eher keine gute Idee, wenn es halluzinationsbehaftete Antworten liefert. Manchmal kann es durchaus sinnvoll sein, etwas mehr zu investieren und dafür ein Modell zu nutzen, das Antworten mit dem richtigen „Flair“ liefert. (fm)</p>



<p><strong>Dieser Artikel ist </strong><a href="https://www.infoworld.com/article/4183716/33-llm-metrics-to-watch-closely.html" target="_blank"><strong>im Original</strong></a><strong> bei unserer Schwesterpublikation Infoworld.com erschienen.</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Mistral launches OCR 4, turning document extraction into a full enterprise AI play]]></title>
<description><![CDATA[Mistral AI on Tuesday released OCR 4, a document intelligence model that moves beyond raw text extraction to return structured representations of entire documents — complete with bounding boxes, block-type classification, and per-word confidence scores. The release marks Mistral's fourth generati...]]></description>
<link>https://tsecurity.de/de/3622912/it-nachrichten/mistral-launches-ocr-4-turning-document-extraction-into-a-full-enterprise-ai-play/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3622912/it-nachrichten/mistral-launches-ocr-4-turning-document-extraction-into-a-full-enterprise-ai-play/</guid>
<pubDate>Wed, 24 Jun 2026 23:48:27 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="https://mistral.ai/">Mistral AI</a> on Tuesday released <a href="https://mistral.ai/news/ocr-4/">OCR 4</a>, a document intelligence model that moves beyond raw text extraction to return structured representations of entire documents — complete with bounding boxes, block-type classification, and per-word confidence scores. The release marks Mistral's fourth generation of optical character recognition technology in roughly 15 months and lands at a moment when the company's pitch for European AI sovereignty has never been more commercially relevant.</p><p>The model supports 170 languages across 10 language groups, accepts PDF, DOC, PPT, and OpenDocument formats, and can be deployed as a single container on an organization's own infrastructure — a capability Mistral is positioning directly at enterprises in regulated industries that cannot route sensitive documents through U.S.-jurisdiction cloud APIs.</p><p>"Mistral OCR 4 extracts and structures content from a wide range of documents," the company said in its announcement. "Where previous generations focused on converting a page into clean text and tables, OCR 4 returns a structured representation of the document."</p><p>The model is <a href="https://docs.mistral.ai/resources/cookbooks?useCase=OCR">available immediately</a> through the <a href="https://mistral.ai/pricing/">Mistral API</a>, Document AI in <a href="https://mistral.ai/products/studio/">Mistral Studio</a>, <a href="https://aws.amazon.com/sagemaker/ai/">Amazon SageMaker</a>, and <a href="https://azure.microsoft.com/en-us/products/ai-foundry">Microsoft Foundry</a>, with <a href="https://www.snowflake.com/en/blog/engineering/enterprise-scale-document-ai/">Snowflake Parse Document</a> support coming soon. Pricing starts at $4 per 1,000 pages, dropping to $2 per 1,000 pages through a batch API discount.</p><div></div><h2><b>OCR 4 treats every document as a semantic map, not a wall of text</b></h2><p>The central engineering shift in <a href="https://mistral.ai/news/ocr-4/">OCR 4</a> is structural. Rather than outputting a flat stream of extracted text — the paradigm that has defined OCR for decades — the model returns a layered representation in which every block is localized with a bounding box, classified by type (title, table, equation, signature, and others), and scored for confidence at both the page and word level.</p><p>Mistral says bounding boxes were its most-requested capability. The reason is straightforward: without location data, downstream systems cannot trace an extracted fact back to its source on a specific page. That traceability gap has been a persistent friction point for enterprises building retrieval-augmented generation (RAG) pipelines, compliance workflows, or any application where "where did this number come from?" is a question that needs an auditable answer.</p><p>Block classification addresses a related problem. A paragraph tagged as a "title" can segment a document into hierarchical chunks for semantic search. A block tagged as a "table" can be routed to a structured-data pipeline rather than a text summarizer. A block tagged as a "signature" can trigger a redaction workflow in a compliance system.</p><p>These are not novel ideas in isolation, but packaging them as first-class outputs of the OCR model itself — rather than requiring a separate layout-analysis stage — removes an integration layer that enterprise teams have historically had to build and maintain themselves.</p><p>The confidence scores serve a dual purpose. At scale, they allow organizations to programmatically route low-confidence regions to human reviewers and auto-approve high-confidence extractions, building what the industry calls human-in-the-loop verification without requiring a person to review every page of every document. In production systems, OCR is rarely the end goal — it is the first step in a larger pipeline.</p><p>Developers building RAG systems, agent workflows, or document automation often spend more time reconstructing layout and structure than on the downstream AI logic itself. OCR 4 aims to eliminate that reconstruction step, and if it delivers on that promise, the value accrues not just in OCR cost savings but in reduced engineering hours across the entire document pipeline.</p><h2><b>Independent reviewers preferred Mistral's output 72 percent of the time, but benchmarks tell a complicated story</b></h2><p>Mistral reports that <a href="https://mistral.ai/news/ocr-4/">OCR 4</a> achieved a 72% average win rate in a head-to-head human evaluation against leading competitors, conducted by independent annotators across more than 600 real-world documents in over 12 languages. The model also achieved the top overall score on <a href="https://huggingface.co/datasets/allenai/olmOCR-bench">OlmOCRBench</a> at 85.20 and scored 93.07 on <a href="https://github.com/opendatalab/OmniDocBench">OmniDocBench</a>.</p><p>But the company itself urges caution in interpreting those numbers. In its release, Mistral took the unusual step of auditing and publicly disclosing the specific types of scoring artifacts it encountered, including ground-truth errors in the reference annotations, equivalent LaTeX notation scored as mismatches, column-reading-order assumptions, and header/footer attribution issues. "We therefore treat the aggregate score as directional rather than definitive," the company said — a notably transparent stance from a vendor announcing a product.</p><p>That transparency is well-timed. On the public <a href="https://huggingface.co/datasets/allenai/olmOCR-bench">OlmOCRBench leaderboard</a>, some researchers have noted that OCR 4 currently ranks third, behind open models like Chandra OCR 2. And some open-weight models self-report higher OmniDocBench composite scores — <a href="https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.6">PaddleOCR-VL-1.6</a> claims 96.33 — though those results have not been independently reproduced on the public leaderboard.</p><p>Early enterprise feedback has been favorable nonetheless. Aidan Donohue, an AI engineer at financial AI firm Rogo, said the company benchmarked OCR 4 against leading agentic document parsers on a chart-dense financial QA dataset and "reached equivalent accuracy at roughly 8x lower cost and 17x lower latency." Ivan Mihailov, an AI engineer at intellectual property management firm Anaqua, said OCR 4 is "roughly 4x faster per page than our incumbent provider." </p><p>Enterprise buyers, however, should run their own evaluations rather than relying on any vendor's benchmark numbers. The practical question is not which model scores highest on a leaderboard, but which model produces the fewest errors on your specific documents, in your specific languages, at a price and latency that fit your workflow.</p><h2><b>The Anthropic export ban gave Mistral's sovereignty pitch the proof point it needed</b></h2><p>Mistral's release lands in a geopolitical context that could hardly be more favorable for its strategic positioning.</p><p>On June 12, <a href="https://www.anthropic.com/news/fable-mythos-access">Anthropic was forced to disable all access to its newest AI models</a>, Fable 5 and Mythos 5, after the U.S. Commerce Department used national security export controls to bar the company from distributing the models to any foreign national. Enterprise clients in finance, healthcare, SaaS, and critical infrastructure found their core intelligence services abruptly disabled, without prior warning or effective recourse. As of June 24, both models remain offline, with <a href="https://kalshi.com/markets/kxfablerestore/fable-restored/kxfablerestore-27">prediction markets giving only 57% odds of restoration</a> before July 1.</p><p>That episode validated a warning Mistral CEO Arthur Mensch has been sounding for over a year. As Business Insider reported, <a href="https://www.businessinsider.com/anthropic-model-access-mistral-opportunity-ai-sovereignty-2026-6">Mensch warned at London Tech Week</a> in June 2025 about American AI companies "having the keys" for their models, calling it a scenario where European companies are "giving leverage to their providers." He added: "At some point, you need to be able to turn it off or turn it on, and you don't want to leave it to another country."</p><p>The argument gained further urgency as Mensch's broader sovereignty pitch escalated in recent months. As reported by CNBC in late May, <a href="https://www.cnbc.com/2026/05/28/mistral-arthur-mensch-design-chips-ai-data-centers.html">Mensch told the outlet</a>: "Europe is lagging behind when it comes to [the] buildout of infrastructure, and so we are investing to close that gap." </p><p>At the same time, <a href="https://www.reuters.com/business/media-telecom/mistral-defends-ai-use-warfare-rebuts-pope-criticism-2026-05-28/">Mensch pushed back against Pope Leo XIV's call for AI to be "disarmed,"</a> arguing that Europe cannot afford to fall behind U.S. tech giants. "We're all for ​peace, but if you look at our rivals and adversaries in the world, they're using artificial ​intelligence … we do need to have our own capabilities," Mensch told reporters.</p><p>OCR 4's single-container, self-hosted deployment model is the product-level expression of that argument. A U.S.-headquartered provider offering EU data residency means documents are stored in Frankfurt but governed by U.S. law. Mistral, incorporated in France and operating under EU jurisdiction, offering on-premise containerized deployment, means documents never leave the customer's infrastructure at all. The <a href="https://artificialintelligenceact.eu/article/99/">EU AI Act's fine enforcement provisions</a> take effect August 2, adding regulatory pressure to the compliance calculus for European enterprises evaluating document AI vendors.</p><h2><b>Baidu's free, open-weight OCR model arrived one day earlier — and the contrast is revealing</b></h2><p>Mistral's release did not arrive in isolation. Just one day before <a href="https://mistral.ai/news/ocr-4/">OCR 4</a> launched, Baidu shipped <a href="https://huggingface.co/baidu/Unlimited-OCR">Unlimited-OCR</a> on June 22 — a 3-billion-parameter MIT-licensed model that tackles one of the most persistent pain points in document AI: parsing entire PDFs and multi-page scans in a single forward pass, without chunking the input or stitching the output back together afterward.</p><p>Baidu's model uses a technique called <a href="https://arxiv.org/html/2606.23050v1">Reference Sliding Window Attention (R-SWA)</a> that, as a top <a href="https://news.ycombinator.com/item?id=48643426">Hacker News commenter explained</a>, splits the AI's focus into two paths: maintaining full attention on the original document image while restricting memory of generated text to a tight, moving window. The result is constant KV cache size and the ability to transcribe 40-plus pages in a single forward pass. The model gathered <a href="https://github.com/baidu/Unlimited-OCR">1,800 GitHub stars</a> in its first 24 hours and racked up more than <a href="https://news.ycombinator.com/item?id=48643426">479 upvotes on Hacker News</a>, where the discussion thread ran to 109 comments.</p><p>The two releases frame what some analysts are calling the June 2026 document-AI split: self-hosted long-horizon parsing with open weights versus structured managed extraction with enterprise features.</p><p><a href="https://github.com/baidu/Unlimited-OCR">Baidu's model</a> is free under an MIT license, runs on standard GPU hardware, and has no managed API or enterprise SLA. <a href="https://mistral.ai/news/ocr-4/">Mistral's model</a> is a commercial product with per-page pricing, bounding boxes, confidence scores, block classification, multi-platform distribution, and self-hosted deployment options for enterprise customers. </p><p><a href="https://huggingface.co/baidu/Unlimited-OCR">Unlimited-OCR</a> may be the better tool for a research team digitizing scanned dissertations on a single GPU. <a href="https://mistral.ai/news/ocr-4/">OCR 4</a> is built for the IT procurement process — the world of SLAs, data processing agreements, and compliance audits.</p><p>Beyond Baidu, the broader OCR competitive field includes <a href="https://cloud.google.com/document-ai">Google Document AI</a>, <a href="https://aws.amazon.com/textract/">Amazon Textract</a>, <a href="https://azure.microsoft.com/en-us/products/ai-foundry/tools/document-intelligence">Azure Document Intelligence</a>, <a href="https://www.abbyy.com/vantage/">ABBYY Vantage</a>, and a growing number of open-weight models. </p><p>On the <a href="https://news.ycombinator.com/item?id=48643426">Hacker News thread</a> for Unlimited-OCR, practitioners offered a candid assessment of the state of the art. Joss82, who has worked on document parsing for 10 years, wrote bluntly: "OCR still sucks in 2026." Meanwhile, one user named SyneRyder reported success with Claude for OCR of hundreds of pages of handwritten documents, noting the model delivered results with "no corrections required" and even pointed out a continuity error in the source text. These practitioner reports underscore a key tension in the market: performance varies wildly depending on the specific document type, language, and quality of the source material.</p><h2><b>The real play is not OCR — it is an enterprise AI stack with document intelligence as the on-ramp</b></h2><p>Step back far enough, and <a href="https://mistral.ai/news/ocr-4/">Mistral's OCR 4 release</a> is not really an OCR story. It is an enterprise go-to-market story built on top of a $4.4 billion global intelligent document processing market that is forecast to grow at a 33.1% compound annual growth rate through 2030, according to <a href="https://www.grandviewresearch.com/industry-analysis/intelligent-document-processing-market-report">Grand View Research</a>.</p><p>For Mistral, OCR is a wedge into enterprise AI budgets. The model feeds directly into Mistral's <a href="https://mistral.ai/news/search-toolkit/">Search Toolkit</a>, the company's open-source composable search framework announced at the AI Now Summit. In that architecture, <a href="https://mistral.ai/news/ocr-4/">OCR 4</a> serves as the ingestion layer for retrieval-augmented generation and enterprise search pipelines, converting raw documents into citation-ready, structurally classified input. The logic is clear: once an enterprise adopts OCR 4 for document extraction, Mistral's broader model suite — including Medium 3.5 for reasoning and the Vibe agentic platform for task execution — becomes the natural next step in the stack. </p><p>That pipeline ambition is critical context for understanding Mistral's current fundraising trajectory. Bloomberg recently reported that the company is in early discussions to <a href="https://www.bloomberg.com/news/articles/2026-06-12/france-s-mistral-in-funding-talks-at-about-20-billion-valuation">raise about €3 billion ($3.5 billion)</a> at a valuation of roughly €20 billion — nearly double the €11.7 billion valuation from its September Series C round. To date, Mistral has raised only about $4 billion, a fraction of what its largest U.S. rivals have taken in. OCR 4 and its associated enterprise revenue pipeline are part of how the company plans to justify that higher valuation, with Mistral targeting <a href="https://www.lemonde.fr/en/economy/article/2026/01/22/french-ai-firm-mistral-predicts-revenue-of-1-billion-in-2026_6749706_19.htm">€1 billion in revenue</a> for 2026, up from €200 million in 2025, according to Le Monde.</p><p>Mistral is a company with roughly 1,000 employees and ambitions to compete with labs that have raised 40 times as much capital. It cannot win a general-purpose model arms race against OpenAI and Anthropic. What it can do is build a differentiated enterprise stack around sovereignty, <a href="https://mistral.ai/news/ocr-4/">structured document intelligence</a>, and agentic workflows — and use that stack to capture European enterprise budgets that are increasingly wary of U.S. provider dependency. </p><p>The pricing structure reinforces that strategy: at $2 per 1,000 pages in batch mode, the cost of processing a 100,000-page corporate archive falls to $200, making large-scale digitization projects economically viable in ways they may not have been with token-based vision-language model pricing.</p><p>Whether Mistral can execute that vision at scale — against Google, Amazon, Microsoft, and a surging open-source ecosystem — remains an open question. But the Anthropic export control crisis is still unresolved, European data sovereignty regulations are tightening, and a potential €20 billion funding round is on the horizon. The company is holding an <a href="https://learn.mistral.ai/public/events/ocr4-webinar">OCR 4 production webinar on July 7 at 6:00 PM CET</a>.</p><p>Two weeks ago, the argument for building AI infrastructure outside the reach of U.S. export controls was theoretical. Then the U.S. government flipped a switch, and Anthropic's most advanced models went dark for every non-American on the planet. Mistral did not cause that crisis — but it spent the last year building the product that makes it matter.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The Roadmap to Mastering AI Agent Evaluation]]></title>
<description><![CDATA[Let's not waste any more time.]]></description>
<link>https://tsecurity.de/de/3622739/ai-nachrichten/the-roadmap-to-mastering-ai-agent-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3622739/ai-nachrichten/the-roadmap-to-mastering-ai-agent-evaluation/</guid>
<pubDate>Wed, 24 Jun 2026 22:19:31 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Let's not waste any more time.]]></content:encoded>
</item>
<item>
<title><![CDATA[Computer says no. Troubles with fixing algorithmic decision-making. (tdf2026)]]></title>
<description><![CDATA[Algorithmic predictions are used to allocate social goods such as healthcare, job training, and education. Despite efforts to apply fairness frameworks and participatory approaches, practical outcomes remain problematic as recent investigations have shown. This talk examines standard approaches t...]]></description>
<link>https://tsecurity.de/de/3622545/it-security-video/computer-says-no-troubles-with-fixing-algorithmic-decision-making-tdf2026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3622545/it-security-video/computer-says-no-troubles-with-fixing-algorithmic-decision-making-tdf2026/</guid>
<pubDate>Wed, 24 Jun 2026 20:50:00 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Algorithmic predictions are used to allocate social goods such as healthcare, job training, and education. Despite efforts to apply fairness frameworks and participatory approaches, practical outcomes remain problematic as recent investigations have shown. This talk examines standard approaches to ‘fair machine learning’ through three cases: (1) health programs, (2) long-term unemployment, and (3) school dropout. It critically assesses their limitations and normative assumptions. Two key distinctions clarify the debates: fairness-focused versus welfare-focused methods on the one hand, and whether predictions are instrumentally or communicatively rational on the other. The latter distinction stresses whether algorithms serve effective implementation or facilitate collective evaluation of policy goals.

Licensed to the public under https://creativecommons.org/licenses/by/4.0/
about this event: https://cfp.cttue.de/tdf5/talk/SBXZNK/]]></content:encoded>
</item>
<item>
<title><![CDATA[How Shopify built an AI stack that doesn't care which models survive]]></title>
<description><![CDATA[Shopify built an LLM proxy that gives every engineer access to multiple AI providers — with automatic failover when any one of them goes down, changes, or disappears. When Claude Fable 5 shut down, Shopify's engineers didn't go into panic mode. The proxy shifted them to Claude Opus or GPT 5.5 aut...]]></description>
<link>https://tsecurity.de/de/3622360/it-nachrichten/how-shopify-built-an-ai-stack-that-doesnt-care-which-models-survive/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3622360/it-nachrichten/how-shopify-built-an-ai-stack-that-doesnt-care-which-models-survive/</guid>
<pubDate>Wed, 24 Jun 2026 20:03:05 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Shopify built an LLM proxy that gives every engineer access to multiple AI providers — with automatic failover when any one of them goes down, changes, or disappears. <a href="https://venturebeat.com/technology/anthropic-blocks-all-public-access-to-claude-fable-5-mythos-5-following-us-government-order-what-enterprises-should-do">When Claude Fable 5 shut down</a>, Shopify's engineers didn't go into panic mode. The proxy shifted them to Claude Opus or GPT 5.5 automatically, without interrupting their workflows.

“Fable looks amazing; we used it of course,” Farhan Thawar, Shopify’s head of engineering, <a href="https://www.youtube.com/watch?v=z9ZvM3qM-_w">says in a new VentureBeat Beyond the Pilot podcast</a>. “When a model comes and then it goes, or it could be as innocuous as an update, the proxy allows us to spray across the different providers,” Thawar says. </p><div></div><p>Shopify buys tokens in bulk and all users connect to models through its proxy, Thawar says. This gives his team access to reporting and failover; when there’s an availability issue with one provider, users can be “automatically, seamlessly” transferred to another. 

Enterprises can learn from this example and consider how a disruption might affect their business, Thawar says. At the very least, they should establish a solid backup plan. It’s important to have a system that allows for movement across models so enterprises are not “super tied” to a specific provider. 

Distillation is another important strategy. 

With distillation, a student model learns from a teacher model and typically becomes specialized in a narrower task. These small language models (SLMs) can be more beneficial than generalized, off-the-shelf models in some circumstances. For instance, Shopify’s flagship AI assistant, Sidekick, which performs numerous specialized subtasks for merchants so they can “remove toil” from their day-to-day. 

Using smaller distilled models can be faster and cheaper than more generalized models, Thawar says. In some cases they have proven to be 2x cheaper and faster; in more extreme cases 30x cheaper and faster, he says. 

But “it isn’t just about cost and latency, which are big; it’s about accuracy,” Thawar says. 

Engineers feed the UDP their teacher model, training data, evals, and a target model — say, Opus 4.8 distilling down to Qwen 3.5. The pipeline runs for about a day, then returns an evaluation showing what the fine-tuned model actually achieved on speed, cost, and accuracy for that subtask. If the tradeoff looks good, the engineer deploys it — no approval process required. Shopify's internal platform, Tangle, lets anyone visualize the pipeline as it runs.

Thawar says his “dream” is to eventually not give the distillation pipeline a target model at all. Instead, users could provide the teacher model with data and evals and the directive: ‘Based on your learnings over time, I want you to look at a different class of model, different sizes, different types, and you tell me what the right distillation target is.’

“Maybe we'll get surprised. Maybe it'll be such a small model it could run on a phone,” Thawar says. “Other times, maybe it comes back and says, ‘There isn't a way to distill this down to anything better than what we have at the frontier.’”</p><h2>Moving away from "AI reflexivity" to "AI leverage" </h2><p>Shopify users can apply whatever harness they want: Claude Code, Codex, Cursor, GitHub Copilot for VS Code. “We expose everyone to the different harnesses so they can get a feel for what may or may not work in their workflow.”

But the company also implemented a usage dashboard; this allows Thawar’s team to ask interesting questions around not just token spend, but: Who’s using the most expensive tokens? Who's spending more time on reasoning? What types of models are being used, and what disciplines and levels?

Regarding the "<a href="https://www.youtube.com/watch?v=7IcU0QBrYng">tokenmaxxing</a>" question, Shopify does have “circuit breakers” in place. If a user has a model running for a long time (say, 10 hours) and it’s consuming a lot of tokens, they will get pinged, “Did you mean to spend this?” 

As Thawar explains, sometimes the reply is “Oh, absolutely.” Other times it’s: ‘Whoa, I didn't know that was running in the background. I totally forgot about it. I'd rather stop it now.’ 

The ultimate goal, as Thawar describes it, is to move from “AI reflexivity” to “AI leverage,” and get people to really think deeply about where they can benefit most from AI in their workflows. 

Listen to the full podcast to hear more about: </p><ul><li><p>Shopify’s philosophy of building infrastructure before features. As Thawar puts it: “We've always built more infra. We will continue to always build more infra.”</p></li><li><p>How Shopify’s internal AI agent, River, creates a “substrate of information” across the company.</p></li><li><p>How Thawar's OpenClaw agent figured out he was traveling from his calendar — and what that moment told him about where agents are actually headed.</p></li></ul><p><b>You can also listen and subscribe to </b><a href="https://beyondthepilot.ubpages.com/"><b>Beyond the Pilot</b></a><b> on </b><a href="https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4"><b>Spotify</b></a><b>, </b><a href="https://podcasts.apple.com/us/podcast/beyond-the-pilot-enterprise-ai-in-action/id1839285239"><b>Apple</b></a><b> or wherever you get your podcasts.</b></p>]]></content:encoded>
</item>
<item>
<title><![CDATA['We hope to sign the agreement soon': White House calls on Meta to submit AI models for review, citing abilities and vulnerabilities evaluation]]></title>
<description><![CDATA[OpenAI, Google and others all submit their latest models for White House review – Meta urged to do the same.]]></description>
<link>https://tsecurity.de/de/3621838/it-nachrichten/we-hope-to-sign-the-agreement-soon-white-house-calls-on-meta-to-submit-ai-models-for-review-citing-abilities-and-vulnerabilities-evaluation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3621838/it-nachrichten/we-hope-to-sign-the-agreement-soon-white-house-calls-on-meta-to-submit-ai-models-for-review-citing-abilities-and-vulnerabilities-evaluation/</guid>
<pubDate>Wed, 24 Jun 2026 17:03:49 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[OpenAI, Google and others all submit their latest models for White House review – Meta urged to do the same.]]></content:encoded>
</item>
<item>
<title><![CDATA[Choosing your AI stack: The benefits of vendor lock-in]]></title>
<description><![CDATA[AI has emerged as a top priority for businesses and a vehicle for transformation, as evidenced by Accenture research: 97% of executives believe AI will transform their company and industry. But as companies move from AI pilots to scaling AI across the enterprise, we have had repeated conversation...]]></description>
<link>https://tsecurity.de/de/3620900/it-nachrichten/choosing-your-ai-stack-the-benefits-of-vendor-lock-in/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3620900/it-nachrichten/choosing-your-ai-stack-the-benefits-of-vendor-lock-in/</guid>
<pubDate>Wed, 24 Jun 2026 12:03:49 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>AI has emerged as a top priority for businesses and a vehicle for transformation, as evidenced by <a href="https://www.accenture.com/us-en/insights/consulting/gen-ai-reinventing-enterprise-models" rel="nofollow">Accenture research</a>: 97% of executives believe AI will transform their company and industry. But as companies move from AI pilots to scaling AI across the enterprise, we have had repeated conversations with CIOs and technology leaders who are arriving at the same uncomfortable realization: AI stack decisions are not easily reversible.</p>



<p>Unlike earlier eras of enterprise IT, where abstraction layers insulated applications from hardware choices, today’s AI stack—the infrastructure, technologies and frameworks that powers AI systems – tends  to be tightly co-engineered, with stronger dependencies in the underlying compute layers. Choices made about models, runtimes and compute platforms now shape cost structures, performance ceilings and strategic flexibility. <a href="https://www.accenture.com/content/dam/accenture/final/a-com-migration/pdf/pdf-171/accenture-ever-ready-infrastructure.pdf#zoom=40" rel="nofollow">AI-ready infrastructure</a> has re-emerged as a new source of differentiation, and with it, a new kind of vendor lock-in.</p>



<p>At the center of this shift is the move from training – building AI models – to inference, where those models are used in production to generate outputs from new data. While early attention focused on the cost of training large models, enterprises are now scaling AI across the organization, running models continuously across workflows. This shift significantly changes the economics of AI.</p>



<p>For instance, <a href="https://www.accenture.com/content/dam/accenture/final/accenture-com/document-4/Accenture-The-New-Rules-of-Platform-Strategy-in-the-Age-of-Agentic-AI.pdf#zoom=40" rel="nofollow">agentic AI is reshaping infrastructure architecture and platforms</a> because inference is becoming persistent, stateful and increasingly data intensive. As AI Factories scale, the focus is shifting from peak model performance toward sustainable token economics, where the key differentiators are lowest cost per generated token, power efficiency and infrastructure utilization at scale. In this environment, achieving those outcomes requires full-stack optimization across compute, networking, memory, storage and data fabrics, curated and integrated across ecosystem partners. Secure multitenancy and confidential computing are becoming core design principles, and enterprise AI is now ready to be industrialized at scale.</p>



<h2 class="wp-block-heading">Modern AI infrastructure is a strategic bet</h2>



<p>What makes AI infrastructure different is not just scale, but integration. <a href="https://www.cio.com/article/4176051/8-it-modernization-traps-cios-must-avoid.html?utm=hybrid_search">Modern AI systems</a> are built on tightly co-engineered stacks where GPU accelerators, high-bandwidth interconnects, compilers and runtimes are designed in tandem to maximize throughput and efficiency for AI workloads.</p>



<p>To get the massive computing power required for AI, providers design their hardware and software to work exclusively with one another. This has shifted enterprise decision-making from choosing hardware one piece at a time to committing to ecosystems. And that commitment carries consequences.</p>



<p>In traditional IT environments, applications could also generally move across environments with a manageable amount of effort. In AI systems, that assumption breaks down. What appears portable at the model or application layer often depends on deeply optimized components underneath that layer, such as memory handling and compiler frameworks like CUDA or ROCm that are fine-tuned to specific hardware.</p>



<p>We find it useful to think about AI systems as a layered structure:</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/06/ai-systems-as-a-layered-structure.png?w=1024" alt="A visualization of AI systems as a layered structure." class="wp-image-4188504" width="1024" height="610" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Accenture</p></div>



<p>While upper layers retain some flexibility, dependencies increase as you move downward. Changing your foundational AI provider often means having to rebuild and re-optimize large portions of your technology from scratch.</p>



<p>This is why infrastructure decisions in AI feel less like procurement choices and more like strategic, high-stakes bets.</p>



<h2 class="wp-block-heading">Why switching AI platforms is harder than it looks</h2>



<p>In theory, switching platforms should be straightforward. Models can be retrained, applications rewritten, and infrastructure replaced. In reality, the cost of switching extends far beyond hardware or licensing.</p>



<ul class="wp-block-list">
<li>The first challenge is <strong>engineering effort</strong>. Migrating to different platforms requires engineers to revalidate model behavior, re-tune inference pipelines, and rebuild performance baselines. During this period, teams spend most of their time stabilizing and not innovating.</li>



<li>The second challenge is <strong>hidden dependency</strong>. Over time, system optimization becomes tied to a specific stack. This might include latency expectations, batching strategies, orchestration logic and even human workflows. These ties are not always obvious, but they shape how systems behave in production.</li>



<li>The third challenge is <strong>timing</strong>. There is never a convenient time to migrate, especially factoring in rising AI infrastructure and inference costs, competitive pressure or scaling demands. Organizations are often forced to switch platforms precisely when disruption is hardest to absorb.</li>
</ul>



<h2 class="wp-block-heading">Rethinking performance vs control</h2>



<p>Despite these barriers, organizations do switch. In our experience, this typically happens under three conditions.</p>



<p>One common trigger is when the opportunity cost of staying begins to outweigh the cost of leaving. As performance gaps widen across competing ecosystems, inefficiencies accumulate to the point that remaining on the current platform is no longer viable. Another driver comes from shifts in vendor dynamics. Pricing volatility, supply constraints, or misalignment in product roadmaps can introduce risks that force a re-evaluation. Finally, regulatory requirements, data sovereignty constraints or geopolitical shifts can force platform changes regardless of technical preference.</p>



<p>Across all three strategies, one principle stands out. Lock-in is not inherently negative, and openness is not inherently superior. Timing matters more than ideology.</p>



<p>Given these dynamics, the central question for CIOs is not how to avoid lock-in, but how to manage it deliberately. This represents a significant shift in strategies that previously considered vendor lock-in as a detriment. In practice, we see three broad approaches emerge, each reflecting a different balance between performance and control.</p>



<p>Some organizations take a performance-first approach. They optimize deeply within a specific ecosystem because performance directly drives business outcomes. <a href="https://blogs.nvidia.com/blog/lilly-ai-factory-nvidia-blackwell-dgx-superpod/" rel="nofollow">Eli Lilly’s AI Factory</a> is a strong example. The company has invested heavily in a tightly integrated NVIDIA-based stack to maximize throughput and utilization. In this case, infrastructure is a competitive lever and not merely a support function. Higher switching costs are accepted because near-term performance advantages are decisive.</p>



<p>Others lean toward a portability-first model. These organizations prioritize flexibility, governance, and long-term independence over absolute performance. <a href="https://group.bnpparibas/en/press-release/bnp-paribas-provides-its-businesses-with-an-llm-as-a-service-platform-to-accelerate-the-industrialization-of-generative-ai-use-cases" rel="nofollow">BNP Paribas</a> illustrates this well through its internal LLM platform built on open-source models and controlled infrastructure. By retaining ownership of the stack, the bank ensures data sovereignty, regulatory alignment and predictable cost.</p>



<p>A growing number are adopting a hybrid approach. Rather than applying a single strategy across the enterprise, they segment workloads based on sensitivity to performance, cost and governance. For example, in late 2024, <a href="https://www.cio.com/article/3616622/jpmorgan-chase-builds-ambitious-ai-foundation-on-aws.html?utm_source=chatgpt.com">JPMorganChase</a> outlined its approach at a leading cloud and technology conference. It described combining a firm-wide internal AI platform with cloud-based services to move generative AI into production at scale. This reflects a broader enterprise pattern of pairing internally controlled environments with external ecosystems to balance control, scalability and cost.</p>



<p>A performance advantage is only valuable if it lasts long enough to justify the lock-in it creates. Similarly, portability only matters if the ecosystem evolves in ways that make switching worthwhile. This is where many organizations struggle. They evaluate platforms based on current benchmarks rather than the direction of the ecosystem.</p>



<p>In practice, we encourage leaders to track a set of evolving signals. These range from the maturity of open compiler ecosystems and improvements in cross-platform runtimes, to shifts in performance per watt and increasing regulatory focus on sovereign AI. Together, these indicators help determine whether the industry is moving toward convergence or further fragmentation.</p>



<h2 class="wp-block-heading">Conclusion</h2>



<p>AI is forcing a reset in how technology leaders think about IT architecture. The goal for CIOs is no longer to eliminate dependency, but to choose it consciously and manage and revisit that choice over time.</p>



<p>In our experience, the most effective organizations treat this as a dynamic problem. They evaluate where performance truly differentiates them, where flexibility protects them, and how quickly those boundaries are shifting. They also recognize that some degree of re-platforming is inevitable and plan for it, rather than treating it as a failure.</p>



<p>Ultimately, AI infrastructure strategy is not about optimizing for today’s conditions. It is about getting ready for where the ecosystem is going next. The leaders who navigate this well are not those who avoid lock-in entirely, but those who understand when to embrace it when to limit it and when to move beyond it before the market forces that decision on them.</p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><strong><a href="https://www.cio.com/expert-contributor-network/">Want to join?</a></strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[2026 EuroLLVM - Accelerating Pass Order Auto-tuning via Profile-Guided Cost Modeling]]></title>
<description><![CDATA[Author: LLVM - Bewertung: 0x - Views:2 2026 EuroLLVM Developers' Meeting
https://llvm.org/devmtg/2026-04/
------
Title: Accelerating Pass Order Auto-tuning via Profile-Guided Cost Modeling
Speaker: Bingyu Gao, Wei Wei
------
Slides:  https://llvm.org/devmtg/2026-04/slides/student_technical_talk/s...]]></description>
<link>https://tsecurity.de/de/3620052/it-security-video/2026-eurollvm-accelerating-pass-order-auto-tuning-via-profile-guided-cost-modeling/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3620052/it-security-video/2026-eurollvm-accelerating-pass-order-auto-tuning-via-profile-guided-cost-modeling/</guid>
<pubDate>Wed, 24 Jun 2026 04:33:37 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: LLVM - Bewertung: 0x - Views:2 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/_9PRluIjKmg?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>2026 EuroLLVM Developers' Meeting<br />
https://llvm.org/devmtg/2026-04/<br />
------<br />
Title: Accelerating Pass Order Auto-tuning via Profile-Guided Cost Modeling<br />
Speaker: Bingyu Gao, Wei Wei<br />
------<br />
Slides:  https://llvm.org/devmtg/2026-04/slides/student_technical_talk/student_technical_talk_gao.pdf<br />
-----<br />
LLVM pass ordering auto-tuning can outperform standard -O3, but it is often hindered by an enormous search space and the high overhead of hundreds of dynamic measurements. This talk presents an efficient auto-tuning framework that minimizes expensive measurements using a profile-guided relative cost model and calibrated beam search. Evaluation on cBench shows an average 10.46% speedup over -O3 with only 20 dynamic measurements, significantly accelerating the search for optimal pass sequences.<br />
-----<br />
Videos Edited by Bash Films: http://www.BashFilms.com<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Why SIEM is Moving Toward Unified Security Operations: Rapid7 Named a Major Player in IDC MarketScape]]></title>
<description><![CDATA[Rapid7 has been named a Major Player in the IDC MarketScape: Worldwide SIEM 2026 Vendor Assessment (#US54126826, June 2026).This is the first IDC SIEM MarketScape to bring the enterprise and SMB markets into a single evaluation, and we believe it arrives at a time when the way teams buy and run a...]]></description>
<link>https://tsecurity.de/de/3619121/it-security-nachrichten/why-siem-is-moving-toward-unified-security-operations-rapid7-named-a-major-player-in-idc-marketscape/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3619121/it-security-nachrichten/why-siem-is-moving-toward-unified-security-operations-rapid7-named-a-major-player-in-idc-marketscape/</guid>
<pubDate>Tue, 23 Jun 2026 19:23:08 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h4><span>Rapid7 has been named a Major Player in the </span><span><em>IDC MarketScape: Worldwide SIEM 2026 Vendor Assessment</em></span><span> (#US54126826, June 2026).</span></h4><p><span>This is the first IDC SIEM MarketScape to bring the enterprise and SMB markets into a single evaluation, and we believe it arrives at a time when the way teams buy and run a SOC is changing quickly. Security teams are no longer evaluating detection and response in isolation. They want their threat data, automation, and view of the attack surface working together, rather than spread across a stack of disconnected tools.</span></p><p><span>We believe </span><a href="https://www.rapid7.com/products/siem" target="_self"><span>Incident Command</span></a><span> reflects that shift by bringing threat data, automation, and attack surface context into one platform instead of leaving teams to work across disconnected tools. It also speaks to a broader change in security operations, where context matters more, speed matters more, and teams need a clearer path from alert to action. That same direction runs through Rapid7’s wider point of view on preemptive security: exposure, detection, and response work better when they inform each other through shared context, AI, and human expertise.</span></p><h2>Incident Command brings detection, response, and exposure context together</h2><p><span>Incident Command brings SIEM, SOAR, attack surface management, and threat intelligence together on a shared data model. That gives analysts access to asset risk, vulnerability data, and exposure context during an investigation, so they can understand whether a detection affects a high-risk, internet-facing asset without having to jump between separate products.</span></p><p><span>According to the IDC MarketScape, “Incident Command is a strong fit for midmarket to enterprise organizations that want a fully integrated security operations platform with predictable costs.”</span></p><p><span>The teams we talk to are tired of stitching tools together and dealing with surprise ingestion bills. They want fewer blind spots, faster investigations, and a clearer answer to what is urgent and what to do next. Incident Command addresses that by bringing exposure context, threat intelligence, and response automation into the SIEM workflow, helping teams investigate faster and act with more clarity. For organizations looking for additional managed coverage, Rapid7 MDR is available as a separate offering. As attacks move faster and environments become harder to manage, security operations work better when exposure, threat, and response data are connected through an open platform that gives teams the context they need to move with more speed and clarity.</span></p><h2>AI and automation, pressure-tested by a global SOC</h2><p><span>Many vendors talk about AI in the SOC. For customers, the more important question is how those capabilities are developed, tested, and refined so they are useful in real investigations rather than just sounding good in a product story. We believe the IDC MarketScape called out what that means in Rapid7’s case:</span></p><p><span>“AI models and automation capabilities are tested in the MDR SOC before release to product customers, providing a feedback loop between managed service outcomes and product development that organizations without their own MDR equivalent cannot replicate.”</span></p><p><span>Our MDR analysts work real incidents across thousands of customer environments every day. The detections, triage models, and automation that come out of that work are tested against live attacks before they reach product customers. That feedback loop helps make the AI Engine more useful in practice by handling repetitive work such as classifying alerts, compiling evidence, and surfacing next steps, while analysts spend their time on the decisions that actually require human judgment. That balance also reflects Rapid7’s broader platform story: AI-powered, backed by human expertise.</span><strong> </strong></p><h2>What we believe this IDC MarketScape recognition says about the future of SIEM</h2><p><span>The 2026 IDC MarketScape is a useful signal of where the market is heading. Organizations are looking for platforms where exposure and detection inform each other instead of living in separate systems, and where AI helps teams move faster without removing the human judgment needed to make the right call. We believe that is very much in line with the platform Rapid7 has been building through Incident Command and the wider Command Platform story. We’ll continue investing in the AI Engine, deeper attack surface context, and the integrations customers rely on. The goal remains straightforward: help defenders move faster to keep their environment safe, investigate with more context, and respond with machine speed and confidence.</span></p><p><span>Want to see Incident Command in action? Request a </span><a href="https://www.rapid7.com/products/siem" target="_self"><span>demo</span></a><span> or explore the packages built to meet your team where it is.</span></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[The missing layer in enterprise agentic AI]]></title>
<description><![CDATA[In the past year, the enterprise AI ecosystem has gained enormous capability and zero consensus.



Developers now have a remarkable set of tools for building AI agents: OpenAI’s frameworks, Anthropic’s Claude tooling, LangChain, LangGraph, CrewAI, Microsoft AutoGen, and a growing list of alterna...]]></description>
<link>https://tsecurity.de/de/3617683/ai-nachrichten/the-missing-layer-in-enterprise-agentic-ai/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3617683/ai-nachrichten/the-missing-layer-in-enterprise-agentic-ai/</guid>
<pubDate>Tue, 23 Jun 2026 11:03:57 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>In the past year, the enterprise AI ecosystem has gained enormous capability and zero consensus.</p>



<p>Developers now have a remarkable set of tools for building AI agents: OpenAI’s frameworks, Anthropic’s Claude tooling, LangChain, LangGraph, CrewAI, Microsoft AutoGen, and a growing list of alternatives. Each promises to coordinate reasoning loops, manage multi-step task execution, and connect agents to tools and APIs. For experimentation, the progress has been substantial. Teams can now assemble sophisticated agent workflows in days that would have taken months two years ago.</p>



<p>But I’ve watched this pattern before. In over two decades of building and selling distributed systems platforms, I’ve seen the same dynamic play out across nearly every major infrastructure shift: the tools for consuming a new capability arrive before the infrastructure for governing it does. The gap that emerges isn’t immediately obvious in development environments. It becomes obvious in production.</p>



<p>That’s exactly where enterprise AI stands today.</p>



<h2 class="wp-block-heading"><a></a>What agent frameworks don’t handle</h2>



<p>Modern agent frameworks are fundamentally coordination systems. They determine what a system should do: which tools to call, how to sequence tasks, how to delegate work across agents. That’s hard work, and they’ve gotten quite good at it.</p>



<p>What they rarely address is where those tasks are allowed to run, and under what conditions.</p>



<p>Take a seemingly simple workflow: summarize customer support transcripts using an LLM. In a development environment, the implementation is clean. The agent calls a model API, passes the transcript, and returns a summary. In production at an enterprise, the same request may involve a dataset that can’t cross a specific geographic boundary, a model that isn’t approved for regulated data, and an audit requirement that demands a traceable record of what happened.</p>



<p>Those aren’t planning problems the agent framework was designed to solve. They’re execution governance problems. Most frameworks quietly assume they’re handled somewhere else in the stack. In many enterprise environments, they’re not handled at all. <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">Gartner predicts</a> more than 40% of agentic AI projects will be canceled by the end of 2027, citing inadequate risk controls as a primary driver of failure—a number that reflects exactly this gap.</p>



<h2 class="wp-block-heading"><a></a>What the missing layer actually does</h2>



<p>Addressing these governance problems requires an additional layer between agent logic and execution: one that evaluates every agent action against policies governing where data can reside, which models may process it, who authorized the request, and how the action fits within the organizational context. The agent framework determines what the system should do. The orchestration layer determines whether and where it’s allowed to happen. Keeping those responsibilities separate allows both layers to evolve independently. It also means you can adopt new agent frameworks without rebuilding your governance model from scratch.</p>



<p>This separation will feel familiar to anyone who has worked through the Kubernetes era. <a href="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html" data-type="link" data-id="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html">Kubernetes</a> doesn’t care what’s inside your container. It finds capacity, allocates resources, and ensures things run. The orchestration layer for agentic AI plays an analogous role: it doesn’t care which agent framework generated the request. It enforces the conditions under which that request can execute.</p>



<h2 class="wp-block-heading"><a></a>Richer authorization models</h2>



<p>Traditional enterprise access control is built around a simple question: can user X access resource Y? That’s insufficient for autonomous agents.</p>



<p>A realistic authorization decision for an agent request might look more like this:</p>



<pre class="wp-block-code"><code>request = {
    "agent": "support-summary-agent",
    "task": "summarize",
    "dataset": "customer_support_logs",
    "model": "external_llm_api",
    "delegated_by": "user_4821"
}

policy = evaluate_policy(request)

if policy.allowed:
    route_to_execution(policy.execution_environment)
else:
    raise AuthorizationError(policy.reason)
</code></pre>



<p>The policy engine here evaluates dataset classification, model approval status, geographic processing rules, and the delegation chain that initiated the request. That might mean redirecting the task to an internal inference cluster instead of a public API endpoint, or blocking the request if no compliant execution environment exists. From the agent’s perspective, the task still executes. The orchestration layer ensures it runs in an environment that satisfies enterprise policy.</p>



<h2 class="wp-block-heading"><a></a>Why ontologies are load-bearing infrastructure</h2>



<p>For the orchestration layer to make good decisions, it needs to do more than label data. It needs to understand how the entities involved in a request relate to each other, and reason over those relationships to determine what’s allowed.</p>



<p>Consider the customer support transcript example again. Metadata tells you the dataset contains PII (personally identifiable information). An ontology lets the system reason across a connected chain: the task operates on a dataset containing personal data; that data is governed by GDPR; the organization’s policy requires processing within an approved EU environment; the selected model runs outside that boundary. From those four connected facts, the orchestration layer can infer the request must be rerouted or blocked. The system reasoned over the relationships rather than matching against a hardcoded rule tied to a specific dataset.</p>



<p>This is what makes policy enforcement, execution routing, data locality, and audit decisions computable at runtime. An ontology can be built around virtually any entity-relationship set the enterprise needs to govern: datasets, models, agents, users, regulations, tasks, environments. The relationships that matter are the ones that drive the decisions the governance layer needs to make. Access control lists can restrict who touches a resource, but they can’t reason across a connected set of entities. That reasoning is what the orchestration layer depends on.</p>



<h2 class="wp-block-heading"><a></a>Decision provenance as a first-class requirement</h2>



<p>Enterprise systems also require auditability. When automated agents trigger actions across multiple systems, organizations must be able to reconstruct the decision path that produced the outcome. Compliance depends on it. So does incident response and basic operational trust.</p>



<p>An orchestration layer generates records describing the initiating identity, the agent, the model, the data sources, the policies evaluated during authorization, and virtually anything else the organization chooses to capture in its ontology. That chain of custody allows teams to investigate incidents and validate compliance without treating production AI systems as operational black boxes.</p>



<p>Regulators and auditors are no longer satisfied with knowing what an AI system was designed to do. They want a factual record of what it did in a specific instance, under what authorization, and with what effect—something dashboards can’t provide, but a well-designed orchestration layer can. The EU AI Act makes this explicit: under<a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689"> Article 12 and Article 17</a>, high-risk AI systems must maintain documentation that makes decisions traceable and auditable, with records sufficient to support investigation after the fact.</p>



<h2 class="wp-block-heading"><a></a>Where this leaves enterprise teams</h2>



<p>Agent frameworks will keep improving. The coordination problems they solve are real, and the ecosystem will continue to mature. But the architectural challenge for enterprises has shifted. It’s no longer primarily about coordinating agents. It’s about governing how those agents interact with real infrastructure, real data, and real compliance obligations.</p>



<p>The patterns for doing that exist today: contextual authorization, data locality enforcement, ontology-aware policy evaluation, decision provenance. What most organizations are missing is the recognition that these capabilities belong in a distinct layer that operates independently of whichever agent framework sits above it. Build that layer, and the rest becomes manageable.</p>



<p><em>—</em></p>



<p><a href="https://www.infoworld.com/blogs/new-tech-forum"><strong><em>New Tech Forum</em></strong></a><em><strong> provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all </strong></em><em><strong>inquiries to </strong></em><a href="mailto:doug_dineley@foundryco.com"><strong><em>doug_dineley@foundryco.com</em></strong></a><em><strong>.</strong></em></p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Successful AI adoption lies in collaboration, not replacement]]></title>
<description><![CDATA[Due to the rapid evolution of generative AI in recent years, many companies are accelerating their adoption of AI. Specifically, the scope of AI’s integration into day-to-day operations is steadily expanding, covering tasks such as minute-taking, summarization, searching, responding to inquiries,...]]></description>
<link>https://tsecurity.de/de/3617665/it-nachrichten/successful-ai-adoption-lies-in-collaboration-not-replacement/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3617665/it-nachrichten/successful-ai-adoption-lies-in-collaboration-not-replacement/</guid>
<pubDate>Tue, 23 Jun 2026 11:02:44 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Due to the rapid evolution of generative AI in recent years, many companies are accelerating their adoption of AI. Specifically, the scope of AI’s integration into day-to-day operations is steadily expanding, covering tasks such as minute-taking, summarization, searching, responding to inquiries, and drafting documents — all of which are typically performed by white-collar workers in office settings. At the same time, however, as discussions about AI adoption intensify, questions and concerns are emerging in society, such as “What will happen to human jobs?” and “To what extent should we entrust tasks to AI?”</p>



<p>My own fundamental premise when considering the roles of humans and AI is that AI should not be viewed merely as a tool for improving efficiency. The core issue that a CIO must fundamentally address is not which tasks to introduce AI into, but rather to thoroughly consider what roles humans and AI should each play, how they can complement one another, and how they can enhance each other to create new value that was previously unattainable.<br><br></p>



<p><a href="https://www.kepco.co.jp/english/corporate/list/report/pdf/ar2025_e_18.pdf" rel="nofollow">The Kansai Electric Power Group’s DX Vision 2035 — as part of its DX and AI strategy</a> — has clearly defined its vision as continuing to create new value through AI-driven transformation, with people collaborating with AI. The underlying philosophy is that the use of AI is by no means merely an improvement along the lines of conventional practices; rather, it aims to achieve a fundamental restructuring of business, operations, and work styles.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/06/dx-vision-2035.png?w=1024" alt="DX Vision 2035" class="wp-image-4187946" width="1024" height="568" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Akio Ueda</p></div>



<p>Thus, collaboration between humans and AI does not mean replacing part of the work with AI but rather identifying the strengths of both humans and AI, and restructuring workflows, decision-making, and value delivery. I believe that only when this is achieved will AI evolve from a mere convenient tool into an indispensable weapon for corporate transformation.</p>



<h2 class="wp-block-heading">What is AI good at, and what should humans take on?</h2>



<p>The starting point for considering human-AI collaboration is to objectively assess the areas in which each excels.</p>



<p>AI excels at rapidly analyzing, processing, searching, and summarizing large volumes of information, presenting multiple options, and making inferences and evaluations based on established patterns. For example, gathering external information, drafting documents, preparing meeting minutes, reviewing contracts, responding to inquiries and creating preliminary risk assessments are areas where AI can demonstrate significant strength.</p>



<p>In fact, at Kansai Electric Power, the use of AI is accelerating across a wide range of use cases, including AI-powered compliance checks, AI critic agents for meeting agenda items, AI risk assessment agents for investment projects, the enhancement of the internal help desk through AI, and the overall reform of corporate sales processes through AI.</p>



<p>On the other hand, I believe that in the age of AI, humans should assume four key roles:</p>



<ol class="wp-block-list">
<li>Formulating questions</li>



<li>Interpreting meaning</li>



<li>Making decisions</li>



<li>Taking responsibility for the results</li>
</ol>



<p>While AI can present a vast number of options, it cannot bear the responsibility for making judgments such as “What do we value?” or “What should this company choose?” This is particularly true in the fields of management, customer service, and organizational operations, where factors such as ethics, trust, emotions, and the balancing of interests come into play. In such contexts, human will is ultimately the guiding principle.</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/06/role-of-humans.png?w=1024" alt="The role of humans in the age of AI" class="wp-image-4187945" width="1024" height="524" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Akio Ueda</p></div>



<p>In other words, humans are the ones who decide what questions to ask and what choices to make, while AI is, at best, a tool that quickly produces processing results. If we proceed with AI adoption while blurring this division of roles, it will lead to confusion on the front lines. Conversely, if this distinction is clearly established, AI implementation will not undermine front-line capabilities but will instead enhance human capabilities.</p>



<h2 class="wp-block-heading">Collaboration is not about division of labor but mutual reinforcement</h2>



<p>An important point to note here is that collaboration between humans and AI cannot be achieved simply by creating a basic division of labor chart. What matters is designing a relationship in which both parties draw out and enhance each other’s strengths.</p>



<p>Kiichiro Toyoda, the founder of Toyota Motor Corporation, once said, “Machines become complete when they become one with humans.” If we replace machines with AI in this quote, it becomes “AI becomes complete when it becomes one with humans.” I believe this expresses a timeless concept that remains fully relevant even in today’s AI- era.</p>



<p>So, what are the different patterns of human-AI collaboration? Below, I’ve created a four-quadrant matrix chart that categorizes how humans work based on Science vs. Art (horizontal axis) and Individual vs. Collaborative (vertical axis).</p>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/06/human-ai-collaboration.png?w=1024" alt="What is human-AI collaboration?" class="wp-image-4187947" width="1024" height="564" sizes="auto, (max-width: 1024px) 100vw, 1024px"></figure><p class="imageCredit">Akio Ueda</p></div>



<p>For example:</p>



<ul class="wp-block-list">
<li>[Quadrant D] Science × Individual Work ⇒ Tasks are entrusted to AI and robots.</li>



<li>[Quadrant C] Art × Performed Individually ⇒ AI expands human creativity.</li>



<li>[Area B] Science × Individually ⇒ Humans and AI collaborate</li>



<li>[Domain A] Art × carried out collaboratively by multiple people ⇒ Carried out primarily by humans; AI serves as a sounding board</li>
</ul>



<p>This is the breakdown.</p>



<p>The accuracy of AI’s output changes significantly depending on the quality of the questions humans pose to it. Conversely, when AI anticipates needs by organizing key points and gathering information, humans can devote their time to making more fundamental decisions. At Kansai Electric Power, a proof of concept (PoC) is underway to utilize AI agents for brainstorming management decisions, risk assessment, and stimulating discussion. This initiative is being pursued not with the idea of handing over work entirely to AI, but rather with the concept that AI extends human thinking and enhances the quality and speed of human decision-making.</p>



<p>As this collaboration progresses, the very nature of work will change.AI will take on the tasks of gathering, organizing, and analyzing information — tasks that humans previously spent a great deal of time on — allowing humans to focus on formulating questions and hypotheses, engaging with customers, being creative, building consensus, and making final decisions. As a result, we will see not just a reduction in man-hours, but an improvement in the quality and speed of work.</p>



<p>Thus, I believe we are moving toward a world where people and companies that make full use of AI will succeed, while people and companies that do not use AI will fall behind — not a world where AI takes people’s jobs.</p>



<p>The fundamental question a CIO should ask is not “What should we have AI do?” but rather “What will people be able to focus on once AI is introduced?” I believe that the ultimate value of collaboration lies not in the adoption rate of AI, but in the enhancement and acceleration of human work.</p>



<h2 class="wp-block-heading">Business process redesign is essential for achieving collaboration</h2>



<p>A common trait among organizations where AI adoption is not progressing as expected is that they introduce AI only to specific parts of their operations without changing the underlying processes or methods of human work. While this may seem like the easiest approach at first glance, it actually results in the least effective use of AI’s capabilities and minimizes the value it can deliver. In short, while JTCs (traditional Japanese companies) think in terms of where to introduce AI based on existing business processes, AIFCs (AI-first companies) rebuild business processes on the premise that AI exists.</p>



<p>To truly realize collaboration between humans and AI, it is necessary to break down the business processes themselves. This involves visualizing the elements within the work—such as problem definition, data collection, organization, decision-making, dialogue, resolution, evaluation, and improvement—and designing and transforming each step to determine whether it should be entrusted to AI, handled by humans, or carried out collaboratively by both. This is not merely the introduction of AI, but the design and transformation of the business, its operations, and its organization.</p>



<p>At Kansai Electric Power, there are use cases such as the transformation of the entire sales process using AI, support for knowledge and technical succession in the thermal power division, support for regulatory compliance checks, and the enhancement of the internal help desk. However, we believe the significance lies in the fact that this is not merely the introduction of AI or partial optimization, but rather the integration of AI after taking a bird’s-eye view of the entire workflow, with the ultimate goal of achieving overall optimization.</p>



<p>Thus, the CIO must act not as the person responsible for AI implementation, but as the architect of business transformation.</p>



<h2 class="wp-block-heading">The CIO is a collaborative designer, not an AI implementation manager</h2>



<p>The role expected of a CIO in the AI era is not merely to drive AI adoption. It is to envision a future where humans and AI work together, and to translate that vision into implementable business processes, systems, rules, and organizational culture.</p>



<p>In this sense, it can be said that the CIO is not an AI implementation manager but a collaborative designer. What should humans specialize in, and in which areas should AI be used? What should humans take on more heavily, and what should they let go of? Continuously answering these questions is the CIO’s essential job.</p>



<p>Moreover, this design is not a one-time effort. As long as AI itself continues to evolve rapidly, the nature of collaboration will also continue to evolve. That is precisely why a CIO should not be the one who provides the right answers, but rather the one who continually asks the right questions. The key is not how much to entrust to AI, but rather what humans should hone in an era where AI exists. Continuously asking this question is what determines a company’s competitiveness.</p>



<h2 class="wp-block-heading">Beyond collaboration lies a relationship where humans and AI enhance each other</h2>



<p>When people hear the term human-AI collaboration, many likely think first of efficiency and increased productivity. However, the true goal lies beyond that. It is not merely about using AI to reduce human workloads but about using AI to expand human potential.</p>



<p>Rather than humans merely mastering AI, we must create a relationship where humans and AI mutually enhance one another. Only when such collaboration becomes firmly established will companies truly gain a competitive advantage in the AI era.</p>



<p>The future that CIOs should envision is not an organization where AI takes away people’s jobs. It is an organization where, with AI as a partner, people can engage with customers and society in a more creative, more meaningful way.</p>



<p>What does collaboration between humans and AI entail?</p>



<p>We must not leave this question vague but rather think it through thoroughly and bring it to fruition.</p>



<p>Is this not the crucial mission entrusted to the CIO in the AI era?</p>



<p><strong>This article is published as part of the Foundry Expert Contributor Network.</strong><br><strong><a href="https://www.cio.com/expert-contributor-network/">Want to join?</a></strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[OpenAI Releases GPT‑5.5‑Cyber With Full Automation for Vulnerability Detection and Patching]]></title>
<description><![CDATA[OpenAI has officially launched the full version of GPT‑5.5‑Cyber, a specialized AI model engineered for advanced vulnerability detection, patch generation, and automated remediation at machine speed. The release is part of OpenAI’s broader Daybreak initiative, which aims to democratize defensive ...]]></description>
<link>https://tsecurity.de/de/3617104/it-security-nachrichten/openai-releases-gpt55cyber-with-full-automation-for-vulnerability-detection-and-patching/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3617104/it-security-nachrichten/openai-releases-gpt55cyber-with-full-automation-for-vulnerability-detection-and-patching/</guid>
<pubDate>Tue, 23 Jun 2026 05:22:47 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>OpenAI has officially launched the full version of GPT‑5.5‑Cyber, a specialized AI model engineered for advanced vulnerability detection, patch generation, and automated remediation at machine speed. The release is part of OpenAI’s broader Daybreak initiative, which aims to democratize defensive cybersecurity capabilities for trusted organizations worldwide. GPT‑5.5‑Cyber delivers state-of-the-art results across three major cybersecurity evaluation […]</p>
<p>The post <a href="https://cybersecuritynews.com/gpt-5-5-cyber/">OpenAI Releases GPT‑5.5‑Cyber With Full Automation for Vulnerability Detection and Patching</a> appeared first on <a href="https://cybersecuritynews.com/">Cyber Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Embed the world: Multimodal AI for searchable aerial imagery at scale]]></title>
<description><![CDATA[In this post, we walk through the problem space, our architecture on Amazon Bedrock and Amazon OpenSearch Serverless, the evaluation methodology we built on OpenStreetMap ground truth, four experiments that compared embedding models, fusion strategies, captioning, and search methods, and the prac...]]></description>
<link>https://tsecurity.de/de/3616127/ai-nachrichten/embed-the-world-multimodal-ai-for-searchable-aerial-imagery-at-scale/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3616127/ai-nachrichten/embed-the-world-multimodal-ai-for-searchable-aerial-imagery-at-scale/</guid>
<pubDate>Mon, 22 Jun 2026 18:33:45 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[In this post, we walk through the problem space, our architecture on Amazon Bedrock and Amazon OpenSearch Serverless, the evaluation methodology we built on OpenStreetMap ground truth, four experiments that compared embedding models, fusion strategies, captioning, and search methods, and the practical guidance you can apply when building a similar system. You’ll learn which design choices move the needle for geospatial semantic search, including why Amazon Nova Multimodal Embeddings delivered the highest F1 scores across both benchmark queries in our evaluation. The work described here evolved into Vexcel Intelligence, a searchable imagery product.]]></content:encoded>
</item>
<item>
<title><![CDATA[Researchers introduce Self-Harness, a framework that lets AI agents rewrite their own rules, boosting performance up to 60%]]></title>
<description><![CDATA[Not every company can or should build their own frontier AI language model. However, the harness controlling the model is something that most enterprises can and should customize for their specific purposes.Of course, this is easier said than done. Agent harnesses are still largely tuned through ...]]></description>
<link>https://tsecurity.de/de/3616014/it-nachrichten/researchers-introduce-self-harness-a-framework-that-lets-ai-agents-rewrite-their-own-rules-boosting-performance-up-to-60/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3616014/it-nachrichten/researchers-introduce-self-harness-a-framework-that-lets-ai-agents-rewrite-their-own-rules-boosting-performance-up-to-60/</guid>
<pubDate>Mon, 22 Jun 2026 17:48:04 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Not every company can or should build their own frontier AI language model. However, the <i>harness</i> controlling the model is something that most enterprises can and <i>should</i> customize for their specific purposes.</p><p>Of course, this is easier said than done. A<!-- -->gent harnesses are still largely tuned through manual, ad hoc debugging — a process that relies heavily on intuition rather than systematic feedback loops, making it difficult to keep pace with rapidly evolving LLMs.</p><p>To solve this challenge, researchers at the Shanghai Artificial Intelligence Laboratory have introduced “<a href="https://arxiv.org/abs/2606.09498">Self-Harness</a>,” a new paradigm in which an LLM-based agent systematically improves its own operating rules. By examining its own execution traces to apply edits, the system trades manual guesswork for empirical evidence.</p><p>Self-improving harnesses can enable development teams to deploy robust custom agents that continually adapt their own execution protocols to overcome model-specific weaknesses.</p><h2><b>The challenge of harness engineering</b></h2><p>An LLM-based agent's performance is not determined solely by its underlying base model, but also by its harness: the surrounding system that provides context and enables the model to interact with the environment. A harness includes components like system prompts, tools, memory, verification rules, runtime policies, orchestration logic, and failure-recovery procedures.</p><p>This layer is crucial because many common agent failures stem from the harness rather than the model. For example, an agent may report success without checking the model’s response (e.g., running the code to see if it passes the tests), or it might retry a failed action repeatedly. The harness is also responsible for preventing <a href="https://venturebeat.com/ai/mits-new-recursive-framework-lets-llms-process-10-million-tokens-without">context rot or overload</a> when the agent’s interaction history grows very large. Examples of popular harnesses include SWE-agent, Claude Code, Codex, and OpenHands.</p><p>Harness engineering remains a significant challenge, but the bottleneck isn't necessarily that humans are too slow or incapable. </p><p>In fact, Hangfan Zhang, lead author of the Self-Harness paper, told VentureBeat that "in many cases, an experienced engineer with deep domain knowledge can still propose better changes than an LLM can today."</p><p>Instead, the true bottleneck of manual engineering is that it relies heavily on ad hoc debugging rather than a verifiable, empirical feedback loop. "The deeper issue is that the current harness-engineering paradigm often lacks a systematic feedback loop," Zhang explained. "Many edits are made based on intuition, a few observed failures, or ad hoc debugging."</p><p>With new models being released at a rapid pace, depending on human intuition to manually tune model-specific harnesses becomes increasingly costly and untenable. While some approaches use stronger models to improve the harnesses of weaker target agents, this dependence on external guidance has its own challenges, as these models may be costly, unavailable for frontier models, or mismatched to the target model's failure modes.</p><h2><b>How Self-Harness works</b></h2><p>The Self-Harness paradigm enables an LLM-based agent to improve its own harness without relying on human engineers or stronger external models.</p><p>This continuous self-evolution is driven by a three-stage iterative loop that turns behavioral evidence into harness updates:</p><ul><li><p><b>Weakness mining:</b> Starting from an initial harness, the agent runs a set of tasks, producing execution traces with verifiable outcomes. The agent categorizes failed traces and tries to detect model-specific failure patterns.</p></li><li><p><b>Harness proposal:</b> Based on these failure patterns, the agent uses a “proposer” role to generate a set of diverse yet minimal harness modifications, each tied to a specific failure mechanism to avoid overly general corrections.</p></li><li><p><b>Proposal validation:</b> The system evaluates candidate modifications through regression tests. An edit is promoted only if it improves performance without causing measurable degradation on held-out tasks. If multiple candidate modifications pass the regression tests, they are merged into the next version of the harness, which then serves as the starting point for the next iteration.</p></li></ul><p>To visualize why an enterprise would need this, imagine an automated issue-fixing agent that reads internal documentation, writes patches, and opens pull requests. If the company updates its documentation style, the agent might suddenly fail, pulling the wrong context or writing bad patches. </p><p>On the surface, the agent simply looks broken. But Self-Harness turns this ambiguous failure into a solvable problem. "The failure traces expose where the agent is misusing the new documentation format; the proposer can generate a targeted harness edit... and the evaluator can decide whether that edit improves the failing cases without regressing other cases," Zhang said.</p><h2><b>Self-Harness in action</b></h2><p>The researchers evaluated Self-Harness on <a href="https://www.tbench.ai/">Terminal-Bench-2.0</a>, a benchmark that tests general tool-based execution, including artifact management, command use, verification behavior, and recovery from execution errors. They applied Self-Harness with MiniMax M2.5, Qwen3.5-35B-A3B, and GLM-5.</p><p>To isolate the impact of the self-evolving harness, they started with a minimal harness built upon the DeepAgent SDK, containing only the benchmark-facing system prompt, and the default filesystem and shell tools. The model backend, tool set, benchmark environment, and evaluator were kept unchanged while only the harness was allowed to vary.</p><p>The quantitative results show that <b>agents improved their performance through automated harness edits. </b>On held-out tasks, <b>performance jumped significantly across the board, ranging from 33 to 60 percent </b>relative improvements for different models.</p><p>Importantly, an explicit acceptance rule promotes only those edits that improve performance without introducing unacceptable regressions. What makes Self-Harness powerful for enterprise applications is that it doesn’t simply make the prompt longer or add generic instructions. Instead, it introduces targeted changes that reflect the recurring problems each model encounters during execution.</p><p>For example, under the baseline harness, MiniMax M2.5 would get stuck endlessly exploring dataset configurations until the execution environment timed out, failing to produce any deliverables. Through Self-Harness, the system identified this specific flaw and wrote a "loop breaker" into its runtime policy, forcing the agent to stop and redirect its approach after 50 tool calls. It also added a rule to create an initial version of required artifacts as early as possible.</p><p>On the other hand, Qwen-3.5 had a habit of hitting a file overwrite error and then blindly retrying the same command repeatedly, eventually deleting necessary files out of confusion before stopping. The self-harness fixed this by introducing a strict command-retry discipline (forbidding exact duplicate commands) and a mechanism that forced the agent to immediately recreate any missing artifacts if a file error occurred.</p><p>GLM-5 struggled to preserve environment changes across different commands, and would often waste time on massive downloads or finalize tasks even when sanity checks were failing. Its self-generated harness introduced rules instructing the agent to persist PATH variables across shell sessions, limit external compute, and repair any failed sanity checks before concluding its run.</p><h2><b>The hidden costs of automated harnesses</b></h2><p>While Self-Harness automates the tedious work of tracking down idiosyncratic model failures, decision-makers must be realistic about the trade-offs. Replacing human engineering with automated trial-and-error requires significant computational overhead.</p><p>"Self-Harness replaces part of the human engineering burden with repeated proposal generation, parallel candidate evaluation, and regression testing," Zhang said. "That can mean more API tokens, more latency during optimization, and more infrastructure for running evaluation tasks."</p><p>Also, this system relies on the accuracy of its evaluation pipeline. During their experiments on Terminal-Bench-2.0, the researchers relied on strict, deterministic verifiers to ensure the agent's edits were actually helpful. Without this rigorous ground truth, an automated system risks promoting bad updates. "[The] evaluation system is not an optional component; it is what lets us trade human intuition for empirical evidence," Zhang said.</p><p>This reliance on strict verifiers also dictates where Self-Harness should be deployed. "The best deployment targets today are environments where failures can be measured and where trial-and-error is relatively safe," Zhang said, pointing to coding, internal workflow automation, and DevOps data pipelines as ideal use cases.</p><p>Conversely, enterprises should avoid fully automating harnesses in high-stakes or subjective fields. "The clearest red flags are domains where evaluation is subjective, delayed, non-deterministic, or costly to get wrong, such as medical decision-making, safety-critical infrastructure, or legal decisions."</p><h2><b>From prompt tweakers to feedback architects</b></h2><p>The introduction of self-improving agents does not mean coding or enterprise workflows will suddenly become human-free. The quality of collaboration between the human engineer and the AI is still paramount and difficult to capture with automated benchmarks. </p><p>Instead, the engineering profession is moving up the abstraction layer. "The role of enterprise engineers will shift from manually patching individual prompts or tool calls toward designing the feedback systems that make agent improvement possible," Zhang predicted. Moving forward, "the engineer becomes less of a prompt tweaker and more of a feedback architect."</p><p>As foundational models grow more capable, they will naturally absorb many capabilities that currently require manual harness engineering. "But once that happens, the harness will not disappear; its scope will move outward to connect the model to richer external environments," Zhang said. "Until that boundary moves beyond what humans can evaluate, humans will remain critical providers of feedback."</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[6 security leader tips for mastering business risk]]></title>
<description><![CDATA[Longtime security leader Doug Kersten has expanded his list of responsibilities.



As CISO of software maker Appfire, he now has accountability for business risks, such as how security tools and processes within customer products and services impact their costs and, thus, profitability.



It’s ...]]></description>
<link>https://tsecurity.de/de/3614728/it-security-nachrichten/6-security-leader-tips-for-mastering-business-risk/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3614728/it-security-nachrichten/6-security-leader-tips-for-mastering-business-risk/</guid>
<pubDate>Mon, 22 Jun 2026 09:08:18 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Longtime security leader <a href="https://www.linkedin.com/in/doug-kersten-7437312/">Doug Kersten</a> has expanded his list of responsibilities.</p>



<p>As CISO of software maker Appfire, he now has accountability for business risks, such as how security tools and processes within customer products and services impact their costs and, thus, profitability.</p>



<p>It’s a clearcut example, he says, of where and why CISOs must consider not purely security risk, but also business risk.</p>



<p>“CISOs need to provide input and remediation on the impact of security cost because these often-hidden costs have a negative impact on profitability,” he says. “This is usually overlooked by finance teams when analyzing the true cost of goods sold, and if CISOs are not plugged into the evaluation of business risk, it can easily be dismissed.”</p>



<p>The expansion of Kersten’s remit into business risk isn’t unique. CISOs across industries are increasingly expected to identify and address business risks that in the past had been outside the bounds of their roles.</p>



<p>“While CISOs traditionally focused on protecting systems, networks, and data, today’s business environment requires security leaders to understand how cyber threats impact revenue, operations, customer trust, regulatory obligations, supply chains, and strategic objectives,” says <a href="https://www.linkedin.com/in/dalehoakcyberpro/">Dale Hoak</a>, CISO at software firm RegScale. “The distinction between business risk and security risk is becoming increasingly blurred.”</p>



<p>As such, <a href="https://www.csoonline.com/article/4159317/cisos-reshape-their-roles-as-business-risk-strategists.html">CISOs today must be enterprise risk leaders</a>, he says, capable of advising executives on how security decisions affect the organization’s ability to achieve its business objectives — not just how they impact the IT stack or technology performance.</p>



<p>Understanding business risk is a significant task, experts agree, but they stress that security chiefs are capable of mastering the skill. Here, Kersten, Hoak, and other security leaders offer strategies on how to do so.</p>



<h2 class="wp-block-heading">1. Partner with the owners of business risk</h2>



<p>By his own admission, <a href="https://www.linkedin.com/in/rolandpalmer/">Roland Palmer</a>, CISO and vice president of tech company JumpCloud, has yet to master business risk. So he’s partnering with those in his organization who own it, so he has opportunities to learn and contribute.</p>



<p>“We form a great team to understand risk and the organization’s risk appetite,” he says.</p>



<p>Team members include leaders from legal, finance, and marketing, as well as the COO.</p>



<p>Kersten similarly leans on business leaders to sharpen his understanding of business risk. Last year Kersten, working with his exec colleagues, devised a program assigning business leaders to security risks.</p>



<p>“Security helps them understand the security risks, but they also bring to us the [associated] business risks and what can be done to mitigate them,” he explains, noting that this approach also surfaced risks that have since been addressed, thereby <a href="https://www.csoonline.com/article/4178412/6-critical-security-gaps-every-ciso-must-address.html">closing gaps</a> that were previously unknown.</p>



<h2 class="wp-block-heading">2. Align cybersecurity explicitly to business objectives</h2>



<p>Kerstan believes security teams <a href="https://www.csoonline.com/article/4080670/what-does-aligning-security-to-the-business-really-mean.html">must understand business objectives</a>, so they can understand what risks could derail which objectives. To ensure his security program has that knowledge, he incorporates corporate objectives and key results into his security strategy.</p>



<p>“I build out plans to address those business objectives and key results. I still have that parallel tier of security risk, which is handled by the security team; that doesn’t go away. But layered onto this is the business <a href="https://www.cio.com/article/222203/okr-objectives-and-key-results-defined.html">OKRs</a> that I need to execute against,” he explains. “It changed how we look at risk and what we have to do.”</p>



<p>For example, he now considers how security department actions may impact employee satisfaction and how that relates to employ retention, a business risk identified by HR, “so we’re working to make sure what we do aligns to the needs of the HR department.”</p>



<p><a href="https://www.ey.com/en_us/people/richard-watson" target="_blank" rel="noreferrer noopener">Richard Watson</a>, global cybersecurity leader with professional services firm EY, agrees with the need to “align cybersecurity explicitly to business objectives.”</p>



<p>“Map cyber controls to critical assets and business processes, and link these to potential financial impact,” he advises. “This enables CISOs to translate technical exposure into business terms and prioritize investment accordingly.”</p>



<h2 class="wp-block-heading">3. Lean into networking and relationships</h2>



<p>Another effective way to get a good grasp on business risks: talking with business colleagues. Regular conversations often yield insights into what truly has them worried, says <a href="https://www.linkedin.com/in/ghayslip/">Gary Hayslip</a>, a cybersecurity executive and co-author of the <em>CISO Desk Reference Guide</em>.</p>



<p>“Another thing I have done to understand business risks, and I have recommended it to peers, is doing a walk-about or what some people call a listening tour,” he says. “I do this in every role I am in because I feel it’s important to understand their objectives, the technologies they use, the projects they have ongoing, the issues they may have with the security program, and, finally, what genuinely keeps them up at night.”</p>



<p>Others say they take a similar approach, stressing the value of networking and building relationships where colleagues feel comfortable raising concerns and collaborating on solutions.</p>



<p>“Business risk cannot be managed in isolation. CISOs should regularly engage with the CFO, COO, general counsel, chief risk officer, product leaders, and business unit executives,” Hoak says. “These conversations provide insight into emerging business concerns and help security become part of strategic planning rather than a downstream compliance exercise.”</p>



<h2 class="wp-block-heading">4. Run tabletop exercises focused on business risk</h2>



<p>This is a more structured opportunity, but an equally effective one, to gain more insights into business risks — so long as the exercises put the business front and center, Hayslip says.</p>



<p>“Most <a href="https://www.csoonline.com/article/570871/tabletop-exercises-explained-definition-examples-and-objectives.html">tabletop exercises</a> conducted by the CISO and security teams remain technical and stop at containment. I have found it’s better to run scenarios that force the executives into the decisions they’d actually make during a crisis, such as <a href="https://www.csoonline.com/article/3488842/to-pay-or-not-to-pay-cisos-weigh-in-on-the-ransomware-dilemma.html">whether to pay a ransom</a>, when and what to disclose if there is a data breach, how to handle customers, when and who should invoke legal privilege, and is there an operational fallback available and if so who makes the decision to activate it,” Hayslip says.</p>



<p>“Running these types of scenarios helps stress-test the company’s response and teaches the CISO and security team how their peers make decisions under pressure,” he adds.</p>



<h2 class="wp-block-heading">5. Study up on business risk</h2>



<p><a href="https://www.linkedin.com/in/seanmurphy092009/">Sean Murphy</a>, senior vice president and CISO at BECU, the fifth-largest credit union in the US, didn’t leave learning about business risk to serendipity. He sought out opportunities for formal learning, such as earning the <a href="https://www.nacdonline.org/nacd-credentials/nacd-directorship-certification-credential/certified-directors/">Directorship Certification from the National Association of Corporate Directors</a>. The certification verifies the holder’s expertise in governance, fiduciary duties, strategy, and risk oversight.</p>



<p>Murphy sought the certification to strengthen his <a href="https://www.csoonline.com/article/4168690/what-cisos-need-to-land-a-board-role.html">qualifications for a board position</a> and to better understand the perspectives of his company’s board, including how it views risk. “The certification helps me delve into what the board cares about and their world and helps me then turn that back to my team and what we’re doing,” he adds. “It gives me the business and executive view versus a purely technical and security view.”</p>



<p>Others offer similar learning strategies.</p>



<p>“The CISO needs to see the company the way the CEO, CFO, and board do,” Hayslip says. “To begin, I would recommend sitting down with the 10-K or annual report, the investor deck, and the earnings call transcripts. This will help the CISO understand how the company makes money and which products or business units drive revenue. It also helps the CISO understand what the leadership team is publicly telling the Street about key risks and where they believe revenue growth will come from in the next reporting cycle.”</p>



<p>This work, while perhaps previously not essential for traditional security leaders, is becoming an imperative today.</p>



<p>“This isn’t fun; in fact, it can be boring,” Murphy says. “But the CISO can’t prioritize protecting the business if they don’t know which parts of the business are considered critical. The annual report provides that view in the words of management.”</p>



<p>Veteran security leaders also cite the value of earning <a href="https://www.isaca.org/credentialing/certifications">certifications from ISACA</a>, a professional association for governance and risk professionals, as well as the <a href="https://www.theiia.org/en/certifications/cia/">Institute of Internal Auditors’ Certified Internal Auditor designation</a>.</p>



<h2 class="wp-block-heading">6. Integrate security into enterprise risk management</h2>



<p>To truly master business risk, CISOs should not treat it as separate from security risk.</p>



<p>“Cyber is now an existential business risk, not just an IT risk,” says <a href="https://www.linkedin.com/in/scottmelchior/">Scott Melchior</a>, a member of ISACA’s Emerging Trends Working Group with 20 years of experience at a global consulting firm focusing on governance, risk, and compliance. “Digital infrastructure is business infrastructure. They’re too intertwined to separate.”</p>



<p>Hoak agrees, stressing the need for CISOs to integrate security into <a href="https://www.csoonline.com/article/566417/enterprise-risk-management-erm-putting-cybersecurity-threats-into-a-business-context.html">enterprise risk management</a>.</p>



<p>“Cyber risk should be incorporated into broader enterprise risk management processes alongside financial, operational, legal, and strategic risks. This creates a common framework for evaluating risk and helps executive leadership view cybersecurity within the context of overall business objectives,” he says.</p>



<p>Hayslip has put this into practice. In his CISO roles, he has plugged the security risk register into the organization’s ERM platform. He says this allowed him to present cyber-related risks on the same platform that the board already reviews alongside financial, operational, and strategic risks.</p>



<p>“The goal is for cyber risks to appear on the enterprise heat map as every other material risk, so they compete for resources and attention on equal terms rather than being a sidebar,” Hayslip says. “Now there is some work involved for the CISO to do this correctly, but it’s critically important to quantify cyber risk in dollars and probability, not colors. Moving from qualitative heat maps to financial impact numbers, I have found, is one of the biggest improvements in getting the business to hear the CISO.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic’s Mythos AI Model Reportedly Breached NSA Classified Systems in Hours]]></title>
<description><![CDATA[Anthropic’s flagship Mythos AI model reportedly infiltrated nearly all of the National Security Agency (NSA) ‘s classified systems within a few hours during an authorized red-team evaluation on June 11. This incident now seems to be the main reason for…
Read more →
The post Anthropic’s Mythos AI ...]]></description>
<link>https://tsecurity.de/de/3614726/it-security-nachrichten/anthropics-mythos-ai-model-reportedly-breached-nsa-classified-systems-in-hours/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3614726/it-security-nachrichten/anthropics-mythos-ai-model-reportedly-breached-nsa-classified-systems-in-hours/</guid>
<pubDate>Mon, 22 Jun 2026 09:08:16 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Anthropic’s flagship Mythos AI model reportedly infiltrated nearly all of the National Security Agency (NSA) ‘s classified systems within a few hours during an authorized red-team evaluation on June 11. This incident now seems to be the main reason for…</p>
<p class="more-link-p"><a class="more-link" href="https://www.itsecuritynews.info/anthropics-mythos-ai-model-reportedly-breached-nsa-classified-systems-in-hours/">Read more →</a></p>
<p>The post <a href="https://www.itsecuritynews.info/anthropics-mythos-ai-model-reportedly-breached-nsa-classified-systems-in-hours/">Anthropic’s Mythos AI Model Reportedly Breached NSA Classified Systems in Hours</a> appeared first on <a href="https://www.itsecuritynews.info/">IT Security News</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Anthropic’s Mythos AI Model Reportedly Breached NSA Classified Systems in Hours]]></title>
<description><![CDATA[Anthropic’s flagship Mythos AI model reportedly infiltrated nearly all of the National Security Agency (NSA) ‘s classified systems within a few hours during an authorized red-team evaluation on June 11. This incident now seems to be the main reason for a broad U.S. government directive on export ...]]></description>
<link>https://tsecurity.de/de/3614555/it-security-nachrichten/anthropics-mythos-ai-model-reportedly-breached-nsa-classified-systems-in-hours/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3614555/it-security-nachrichten/anthropics-mythos-ai-model-reportedly-breached-nsa-classified-systems-in-hours/</guid>
<pubDate>Mon, 22 Jun 2026 07:38:36 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Anthropic’s flagship Mythos AI model reportedly infiltrated nearly all of the National Security Agency (NSA) ‘s classified systems within a few hours during an authorized red-team evaluation on June 11. This incident now seems to be the main reason for a broad U.S. government directive on export controls issued the following day. Senator Mark Warner, […]</p>
<p>The post <a href="https://cybersecuritynews.com/anthropics-mythos-ai-model/">Anthropic’s Mythos AI Model Reportedly Breached NSA Classified Systems in Hours</a> appeared first on <a href="https://cybersecuritynews.com/">Cyber Security News</a>.</p>]]></content:encoded>
</item>
</channel>
</rss>
<!-- Generated in 0,33ms -->