<?xml version="1.0" encoding="UTF-8" ?>
<?xml-stylesheet type="text/xsl" href="/rss-style.xsl"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:media="http://search.yahoo.com/mrss/" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
<title><![CDATA[Team IT Security - 📰 Alle Kategorien]]></title>
<link><![CDATA[https://tsecurity.de/export/rss/alle-kategorien.xml?q=demystifying+lossbackward+pytorch+autograd%2F]]></link>
<description><![CDATA[Das Gesamte Cyber Threat Intelligence Feed-Archiv von TSecurity.de. Alle Nachrichten, Sicherheitsmeldungen, Videos, Downloads und Analysen in einer zentralen Übersicht.]]></description>
<language>de-DE</language>
<lastBuildDate>Thu, 30 Jul 2026 04:15:22 +0200</lastBuildDate>
<pubDate>Thu, 30 Jul 2026 04:15:22 +0200</pubDate>
<ttl>15</ttl>
<copyright>2026 Team IT Security</copyright>
<managingEditor>lakandor@tsecurity.de (Horus Sirius)</managingEditor>
<webMaster>lakandor@tsecurity.de (Horus Sirius)</webMaster>
<category>IT Security</category>
<category>Cybersecurity</category>
<category>Nachrichten</category>
<generator>Team IT Security RSS Generator v2.0</generator>
<image>
<url>https://tsecurity.de/favicon.ico</url>
<title><![CDATA[Team IT Security - 📰 Alle Kategorien]]></title>
<link><![CDATA[https://tsecurity.de/export/rss/alle-kategorien.xml?q=demystifying+lossbackward+pytorch+autograd%2F]]></link>
</image>
<atom:link href="https://tsecurity.de/export/rss/it-security.xml?q=demystifying+lossbackward+pytorch+autograd%2F" rel="self" type="application/rss+xml" />
<item>
<title><![CDATA[CVE-2026-65918 | PyTorch torchvision up to 0.28.0 GIF decoder read_from_tensor out-of-bounds (EUVD-2026-48341)]]></title>
<description><![CDATA[A vulnerability classified as critical has been found in PyTorch torchvision up to 0.28.0. The affected element is the function read_from_tensor of the component GIF decoder. Performing a manipulation results in out-of-bounds read.

This vulnerability was named CVE-2026-65918. The attack may be i...]]></description>
<link>https://tsecurity.de/de/3692769/sicherheitsluecken/cve-2026-65918-pytorch-torchvision-up-to-0280-gif-decoder-readfromtensor-out-of-bounds-euvd-2026-48341/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3692769/sicherheitsluecken/cve-2026-65918-pytorch-torchvision-up-to-0280-gif-decoder-readfromtensor-out-of-bounds-euvd-2026-48341/</guid>
<pubDate>Sat, 25 Jul 2026 01:53:20 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability classified as <a href="https://vuldb.com/kb/risk">critical</a> has been found in <a href="https://vuldb.com/product/pytorch:torchvision">PyTorch torchvision up to 0.28.0</a>. The affected element is the function <code>read_from_tensor</code> of the component <em>GIF decoder</em>. Performing a manipulation results in out-of-bounds read.

This vulnerability was named <a href="https://vuldb.com/cve/CVE-2026-65918">CVE-2026-65918</a>. The attack may be initiated remotely. There is no available exploit.

Applying a patch is the recommended action to fix this issue.]]></content:encoded>
</item>
<item>
<title><![CDATA[Build an explainable next-best-product recommendation system for banking on AWS]]></title>
<description><![CDATA[Learn the architecture and design decisions behind an explainable next-best-product recommendation system for banking, built with Amazon SageMaker AI and PyTorch. A multi-tower neural network with learned attention delivers accurate, per-customer recommendations while providing the explainability...]]></description>
<link>https://tsecurity.de/de/3691968/ai-nachrichten/build-an-explainable-next-best-product-recommendation-system-for-banking-on-aws/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3691968/ai-nachrichten/build-an-explainable-next-best-product-recommendation-system-for-banking-on-aws/</guid>
<pubDate>Fri, 24 Jul 2026 17:50:40 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Learn the architecture and design decisions behind an explainable next-best-product recommendation system for banking, built with Amazon SageMaker AI and PyTorch. A multi-tower neural network with learned attention delivers accurate, per-customer recommendations while providing the explainability that banking regulators require.]]></content:encoded>
</item>
<item>
<title><![CDATA[[NEU] [mittel] PyTorch: Schwachstelle ermöglicht Denial of Service und Offenlegung von Informationen]]></title>
<description><![CDATA[Ein entfernter, anonymer Angreifer kann eine Schwachstelle in PyTorch ausnutzen, um einen Denial of Service Angriff durchzuführen, und um Informationen offenzulegen.]]></description>
<link>https://tsecurity.de/de/3691158/it-security-nachrichten/neu-mittel-pytorch-schwachstelle-ermoeglicht-denial-of-service-und-offenlegung-von-informationen/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3691158/it-security-nachrichten/neu-mittel-pytorch-schwachstelle-ermoeglicht-denial-of-service-und-offenlegung-von-informationen/</guid>
<pubDate>Fri, 24 Jul 2026 11:41:28 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Ein entfernter, anonymer Angreifer kann eine Schwachstelle in PyTorch ausnutzen, um einen Denial of Service Angriff durchzuführen, und um Informationen offenzulegen.]]></content:encoded>
</item>
<item>
<title><![CDATA[Q&A: Google’s AI and computing chief talks about its shapeshifting data centers]]></title>
<description><![CDATA[Google’s AI offerings span its internal and cloud offerings. Its data centers are processing seven times more AI tokens compared to last year. To keep up, Google is upgrading its data-center hardware and software technologies at a faster clip. It plans to raise $80 billion to build new data cente...]]></description>
<link>https://tsecurity.de/de/3689101/it-security-nachrichten/qa-googles-ai-and-computing-chief-talks-about-its-shapeshifting-data-centers/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3689101/it-security-nachrichten/qa-googles-ai-and-computing-chief-talks-about-its-shapeshifting-data-centers/</guid>
<pubDate>Thu, 23 Jul 2026 14:55:21 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Google’s AI offerings span its internal and cloud offerings. Its data centers are processing seven times more AI tokens compared to last year. To keep up, Google is upgrading its data-center hardware and software technologies at a faster clip. It plans to raise $80 billion to build new data centers. (See related story: <a href="https://www.networkworld.com/article/4200581/google-transforms-its-data-center-architecture-for-agent-era.html">Google transforms its data center architecture for agent era</a>)</p>



<p class="wp-block-paragraph"><em>Network World</em> spoke with <a href="https://www.linkedin.com/in/marklohmeyer/">Mark Lohmeyer</a>, vice president and general manager of AI and computing at Google, about how the company’s infrastructure is keeping pace with AI demand.</p>



<p class="wp-block-paragraph"><strong>Network World: What is the primary shift in infrastructure needs?</strong></p>



<p class="wp-block-paragraph"><strong>Mark Lohmeyer:</strong> We’ve seen the <a href="https://www.networkworld.com/article/4175890/cisco-ai-traffic-is-radically-reshaping-wans.html">rise of agents and agentic use cases</a>. Years ago, it was the chat phase: Ask a question, get an answer. Now we’re in the agentic era, where you express your intent, agents spin off multiple sub-agents, working in parallel, preserving state. This is a radical shift in what infrastructure needs to do; make them fast, cost effective, secure, reliable. We’re delivering infrastructure optimized for the age of agents.</p>



<p class="wp-block-paragraph"><strong>NW: What’s the goal of the infrastructure buildout, and what should customers expect regarding costs?</strong></p>



<p class="wp-block-paragraph"><strong>ML: </strong>Ultimately, it’s about enabling customers with leading-edge capabilities and models at scale cost-effectively. With agents, <a href="https://www.networkworld.com/article/4057121/network-and-cloud-implications-of-agentic-ai.html">inference transactions increase</a> by 50x, 100x versus non-agentic workloads. We’re driving the cost per transaction down exponentially. In our latest platforms, we reduce the cost by almost 2x for the same work. Customers serve twice the number of users at the same cost, directly driving profitability.</p>



<p class="wp-block-paragraph"><strong>NW: How are you addressing energy efficiency?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Energy is a critical resource, and Google has optimized for years. We design data centers and compute [to drive] high PUE (power usage effectiveness). We introduced <a href="https://www.networkworld.com/article/4149069/why-ai-rack-densities-make-liquid-cooling-nonnegotiable.html">liquid cooling</a> over five years ago, and these latest systems are all liquid cooled. For agentic workloads, CPUs come to the forefront… orchestrating agents, calling tools, doing evaluation loops in reinforcement learning. Our latest Axion-based CPU platform called <a href="https://www.networkworld.com/article/4086182/google-cloud-aims-for-more-cost-effective-arm-computing-with-axion-n4a.html">N4A</a> has energy efficiency and is significantly better than the prior generation and x86 comparables.</p>



<p class="wp-block-paragraph"><strong>NW: How do you think about token efficiency as you build-out systems?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Performance and efficiency gains are powered by co-design of the model and infrastructure. <a href="https://www.computerworld.com/article/4161990/gemini-enterprise-update-brings-ai-agents-into-collaborative-workflows.html">Gemini</a> is trained on TPUs, primarily served on TPUs with high frontier model capability, in a token and cost-efficient way. This stems from co-design across the full stack.</p>



<p class="wp-block-paragraph"><strong>NW: How do you project what infrastructure will be needed years in advance?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Hardware cycles deliver a new next generation roughly every year, but design cycles are two years or more in advance. We work with <a href="https://deepmind.google/about/">DeepMind</a> doing core research, to application teams taking models into production, to billions of users, to our team building infrastructure. We work upstream with DeepMind and application teams to understand what’s coming. Agents weren’t being broadly spoken of externally, but internally we had those insights around what they would need. That shows up in hardware design. We hit the timing right — these platforms are built for agents.</p>



<p class="wp-block-paragraph"><strong>NW: What’s the eighth generation TPU platform?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> We deliver new platforms every year, and ones launched years ago are close to 100% utilized because demand for AI-optimized compute is high. The <a href="https://www.networkworld.com/article/4162004/google-bets-on-workload-specific-tpus-with-8t-and-8i-launch.html">eighth-generation TPU platform</a> is the first delivering two complete systems, from the chip all the way up to the network and storage and software, that are optimized.</p>



<p class="wp-block-paragraph"><a href="https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive">TPU-8t</a> is optimized for training, and TPU-8i is optimized for inference. For TPU-8i, we increased SRAM on the chip to 384MB — three times the prior generation — and increased the HBM by 50%.</p>



<p class="wp-block-paragraph"><strong>NW: How are you approaching GPU and TPU compatibility?</strong></p>



<p class="wp-block-paragraph"><strong>ML: </strong>People in a single cluster do not commingle GPUs and TPUs. We offer both options based on specific workload needs. We’ve been investing on the TPU side in using software frameworks customers are comfortable with on GPUs and enabling those on TPUs. For example, <a href="https://www.infoworld.com/article/2335194/what-is-pytorch-python-machine-learning-on-gpus.html">PyTorch</a> and vLLM. Customers could have a pool of GPUs and TPUs, running vLLM on top of that. Start with a workload on TPUs, but if the TPU pool is fully utilized, spill to GPUs or vice versa. This works because it’s all leveraging the same compatible software layer on top.</p>



<p class="wp-block-paragraph"><strong>NW: How has the orchestration platform changed for agents?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> Kubernetes is becoming the orchestration platform of choice for AI. Google is transforming <a href="https://www.infoworld.com/article/2255921/gke-tutorial-get-started-with-google-kubernetes-engine.html">GKE</a> [Google Kubernetes Engine] into an agent-native orchestration solution. When expressing intent to an agent and it spins up multiple sub-agents, compute needs to spin up rapidly — TPUs or GPUs — without long delays, then run and spin back down. We’re optimizing at every layer of the <a href="https://cloud.google.com/kubernetes-engine">GKE stack</a>: significantly improving node startup time and how rapidly we start and stop containers. Lovable demonstrates this with GKE, spinning up hundreds of sandboxes for live coding sessions on their platform in parallel, paying for infrastructure when needed.</p>



<p class="wp-block-paragraph"><strong>NW: What is the role of the network and storage infrastructure?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> The network is critical for AI. This requires creating large-scale clusters of GPUs or TPUs and enabling them to talk to each other in a high-performance way. <a href="https://cloud.google.com/blog/products/networking/introducing-virgo-megascale-data-center-fabric">We created the Virgo network</a> — a collapsed network architecture, non-blocking within a data center, where multiple pods or NVLink72 domains connect together.</p>



<p class="wp-block-paragraph">In TPU8T, we can connect over a million TPUs together leveraging Virgo, creating large-scale, high-performance, reliable clusters that shrink innovation cycles. Storage is equally critical. In large-scale clusters, something is always failing. The ability to take snapshots and go back to a checkpoint is important.</p>



<p class="wp-block-paragraph">We’ve introduced <a href="https://cloud.google.com/products/managed-lustre">Managed Lustre 10T</a>, with 10 terabytes per second of bandwidth, 18 petabytes of storage in single clusters. This is 10 times faster than last year and 20 times faster than competition. We have Rapid Bucket, low-latency storage backed by Google storage systems. Both are impactful in large-scale training environments.</p>



<p class="wp-block-paragraph"><strong>NW: How does KV cache strategy differ between training and inference?</strong></p>



<p class="wp-block-paragraph"><strong>ML:</strong> For <a href="https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/eighth-generation-tpu-agentic-era/">TPU-8i</a>, we increased SRAM on the chip to 384 megabytes — three times the prior generation — and increased the HBM by 50%. Storing KV cache directly in chip memory allows responding to inference requests much more rapidly and cost-effectively than going to an external system. For inference workloads, storing as much KV cache as possible on-chip is critical.</p>



<p class="wp-block-paragraph">We’re introducing a dedicated KV cache storage subsystem that works across GPUs and TPUs. As KV caches get larger, being able to fall back to this dedicated subsystem becomes critical. Loading model weights rapidly is important in dynamic inference environments where accelerators switch between models hour by hour.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Inflection AI returns to consumer market with Pi Journeys after Microsoft upheaval]]></title>
<description><![CDATA[Inflection AI, the Palo Alto startup that two years ago became Silicon Valley's most famous cautionary tale about the brutal economics of frontier AI, announced Tuesday that it is returning to the consumer market with a new research division and an experimental product built around a provocative ...]]></description>
<link>https://tsecurity.de/de/3687581/it-nachrichten/inflection-ai-returns-to-consumer-market-with-pi-journeys-after-microsoft-upheaval/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3687581/it-nachrichten/inflection-ai-returns-to-consumer-market-with-pi-journeys-after-microsoft-upheaval/</guid>
<pubDate>Wed, 22 Jul 2026 22:58:21 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><a href="https://inflection.ai/">Inflection AI</a>, the Palo Alto startup that two years ago became Silicon Valley's most famous cautionary tale about the brutal economics of frontier AI, announced Tuesday that it is returning to the consumer market with a new research division and an experimental product built around a provocative thesis: the next competitive battleground in AI won't be raw intelligence, but relationships.</p><p>The company launched <a href="https://inflection.ai/labs">Inflection AI Labs</a>, a public-facing research and experimentation arm, alongside <a href="https://inflection.ai/labs/pi-journeys">Pi Journeys</a>, the lab's first product experiment — an AI experience designed to adapt to a user's life stage, whether that's becoming a parent, taking on caregiving duties, changing careers, or aging. The announcement arrived with a research report on consumer AI habits and a substantial update to Pi, the company's flagship chatbot, adding improved voice, memory, and new agentic tools for reminders, to-do lists, and shopping.</p><p>"Inflection AI is the company. Pi is our flagship consumer product. Inflection AI Labs is where we experiment, explore personal intelligence and share more publicly. Pi Journeys is the first public experiment from Inflection AI Labs," CEO Sean White told VentureBeat in an exclusive interview.</p><p>Behind the tidy org chart is a far more interesting story: a company attempting one of the more unusual second acts in the AI industry, powered by an argument that the entire market is optimizing for the wrong thing.</p><h2><b>Why Inflection AI believes the chatbot era's biggest flaw is that it's transactional</b></h2><p>White's central claim is that today's AI assistants — including the industry's most capable models — are fundamentally transactional. You ask, they answer, the session ends. He believes that architecture misses most of what people actually need from artificial intelligence in their daily lives.</p><p>"One of the things that really struck us in particular, and this showed up in the research, was that a lot of the work is very transactional, and you'll hear me say a lot that we've been shifting all this from transactional to relational systems," White said. "Not everything is going to be: I do a single turn, I utter a question, I get a search response back."</p><p>White frames the industry's evolution as a progression through four kinds of intelligence. First came raw IQ — the foundation model race. Then emotional intelligence, which Inflection made its signature with Pi's famously warm conversational style. Then agentic intelligence — AI that acts rather than just talks — which White says Inflection absorbed from its enterprise work. The fourth, and the one Inflection is now staking its future on, is what the company calls relational intelligence: AI that understands not just you, but the web of people around you.</p><p>"There's so much fear about these things pushing people into loneliness,” White said. “If we design these pro-social systems as another design criteria, that actually makes a huge difference."</p><p>That design philosophy is a pointed counter-narrative to one of the loudest anxieties in consumer AI right now: that <a href="https://www.media.mit.edu/articles/chatgpt-may-be-making-us-lonelier/">emotionally engaging chatbots deepen isolation</a> by substituting for human contact. Inflection argues the opposite is possible — that an AI with structured knowledge of your relationships can push you back toward people rather than away from them.</p><h2><b>Inside Pi Journeys, the AI companion that maps your relationships and life stages</b></h2><p><a href="https://inflection.ai/labs/pi-journeys">Pi Journeys</a> makes that idea concrete. When users first open the product, it asks about their life stage — caregiver, household manager, midlife transition — and then builds what White describes as specially structured memory around the people who matter in that context. From there, the system becomes proactive.</p><p>"It starts to build up memories around that, and it acts as a memory prosthetic — but in a pro-social way," White said. "It doesn't get in the way of your interactions with other people; it really helps facilitate them." The system might remind a user, for example, that a friend deserves a call, or resurface what was last discussed with a family member involved in a parent's care.</p><p>White, who spent years as chief R&amp;D officer at Mozilla before taking Inflection's helm, was quick to flag the obvious privacy implications of an AI that maps your social graph. "We've built a lot of privacy systems into this," he said, noting users can delete and manage the people recorded in their profile. Whether consumers will trust a venture-backed AI company with a structured database of their most important relationships remains one of the biggest open questions hanging over the product — and one that enterprise buyers evaluating Inflection's technology will watch closely.</p><p>Asked why this was the first Labs experiment, White was direct: "Pi Journeys takes into account people's life stages and experiences because we have heard from users that we can provide more value in helping them navigate their lives. Pi Journeys lets us experiment with the early stages of prosocial and relational intelligence because life isn't single-player."</p><p>The product has been tested internally and with small closed groups, White said, and is now being released more broadly as an experiment rather than a finished product — a posture the Labs branding is designed to make explicit.</p><h2><b>What Inflection's consumer AI research reveals about how people actually use chatbots</b></h2><p>Inflection Labs' first publication, the <a href="https://inflection.ai/state-of-consumer-ai-2026">State of Consumer AI Research Report</a>, offers the empirical scaffolding for the strategy. The average consumer now uses roughly two different AI tools every day and three per week, the company found — evidence, in Inflection's reading, that no single assistant has locked up consumer loyalty and that the market remains contestable.</p><p>More telling is why people choose the tools they do. Respondents cited personalization, style and tone, context awareness, and — notably — emotional understanding as deciding factors. They also said they want AI to be more than a productivity engine: a coach or mentor to motivate them, a chef to suggest recipes, a DJ to curate playlists.</p><p>"One thing we're certainly finding is that a lot of that also is in work, not so much in everyday life," White said. "That's our focus right now — the everyday life part."</p><p>This is a shrewd reading of the competitive map. The best-funded AI labs are pouring resources into coding tools, enterprise agents, and developer platforms, leaving everyday consumer use cases comparatively underserved. White sees the gap clearly. "We see a lot of products that are being aimed more and more at the enterprise," he said. "As a computer scientist by training, I kind of love the IDEs as this tool, but it's not really great for everybody. There's so much regular everyday use from folks that is either purely voice or that is purely mobile."</p><p>He recalled a conversation with a conference staffer who told him she owned only a phone, no laptop — exactly the kind of user, he argued, that the industry's developer-centric product roadmaps have left behind.</p><h2><b>How the $650 million Microsoft deal hollowed out Inflection — and set up its second act</b></h2><p>To understand why any of this is remarkable, you have to rewind to March 2024. Inflection was then one of the hottest startups in AI, having <a href="https://www.reuters.com/technology/inflection-ai-raises-13-bln-funding-microsoft-others-2023-06-29/">raised $1.3 billion in mid-2023</a> in a round backed by Microsoft, Nvidia, Bill Gates, and Reid Hoffman — more than $1.5 billion in total. Pi had crossed one million daily active users, per Reuters.</p><p>Then, in a deal that reshaped how the industry thinks about acqui-hires, Microsoft hired away co-founder and CEO Mustafa Suleyman, chief scientist Karén Simonyan, and most of the company's roughly 70 employees, paying Inflection about $650 million largely to license its technology, as <a href="https://www.bloomberg.com/news/articles/2024-03-21/microsoft-to-pay-inflection-ai-650-million-after-scooping-up-most-of-staff">Reuters reported</a>. Suleyman now runs Microsoft's consumer AI business. The structure of the deal drew scrutiny from the FTC and Britain's competition regulator, though the UK's Competition and Markets Authority cleared it in September 2024 and EU regulators declined to act.</p><p>White, installed as CEO in the aftermath, steered the remnant company hard toward enterprise, acquiring three startups in late 2024 — <a href="http://jelled.ai/">Jelled.AI</a>, <a href="https://boostkpi.com/">BoostKPI</a>, and the European consulting firm <a href="https://www.boundaryless.com/">Boundaryless</a> — and <a href="https://techcrunch.com/2024/11/26/inflection-ceo-says-its-done-competing-to-make-next-generation-ai-models/">telling TechCrunch</a> that November that Inflection had no intention of competing with companies building 100,000-GPU frontier systems.</p><p>Tuesday's announcement doesn't reverse that position so much as complicate it. Asked how to think about the company today, White called it "a consumer-first strategy that bridges both consumer and enterprise efforts" — and he insists the two sides feed each other.</p><p>Enterprise deployments, including a partnership with Intel that is among the few he can name publicly, taught Inflection how to run models inside complex infrastructure. Consumer products, meanwhile, let the company iterate at speed. "The part I also like about the consumer side, and this has always been true, is that we can move faster, experiment faster, and try and learn faster," White said.</p><h2><b>The six-month prediction: relationship-aware AI is coming to the enterprise</b></h2><p>Buried in White's consumer pitch is the claim that should matter most to technical decision-makers. "Normally I'd say like a year, but let's call it six months," he said. "You're going to start to see a bunch of enterprises care a lot more about the relationships that are inside the enterprises and what that picture is, not just the workflows."</p><p>If White is right, the wave of workflow-automation agents currently flooding the enterprise market is only the first phase of business AI adoption — with relationship-aware systems, tested first on consumers, following close behind. Inflection is essentially using its consumer products as a live laboratory for capabilities it plans to sell into companies. It's a capital-efficient strategy for a firm that can no longer outspend rivals on training runs, and a risky one, since it depends on consumers showing up in numbers large enough to generate the learning.</p><p>The technical substance underneath is equally pragmatic. Pi today runs not on a single proprietary frontier model but on an orchestration layer routing across many models — some descended from Inflection's original fully trained cores, some fine-tuned, some open source, including work with Nvidia that White says gives Inflection access to unreleased cutting-edge models. He also took a swipe at the industry's loose vocabulary around ownership: "When people say that the model is their own, most of the time nowadays — I guess I won't name names — a lot of companies will actually take a checkpoint, and then they will fine-tune from that checkpoint. But very few people actually start from that beginning core."</p><p>That candor extends to open source, where White carefully hedged. "We're not ready to promise what I think of as true open source, and by that I mean everything," he said, invoking his Mozilla years overseeing genuinely open projects like <a href="https://rust-lang.org/">Rust</a> and <a href="https://webassembly.org/">WebAssembly</a>.</p><p>Weights without training data and pipelines, he argued, often leave developers unable to do anything meaningful with a supposedly "open" model. "We are a PBC, and there's still a C in there," he added — a reminder that public benefit corporations still have businesses to protect. The Labs will collaborate with academic researchers, including Stanford professors who visited the company's Palo Alto office this week, and continue contributing to open projects such as <a href="https://pytorch.org/">PyTorch</a>.</p><h2><b>Can a diminished Inflection compete with AI giants spending billions?</b></h2><p>Reid Hoffman, the LinkedIn co-founder who co-founded Inflection and stayed on through the Microsoft upheaval, framed the announcement in the sweeping terms of his recent writing on AI and human agency. "Humans should be amplified by AI, not replaced. That's the principle Pi was built on," <a href="https://finance.yahoo.com/technology/ai/articles/inflection-ai-shaping-future-personal-130000573.html">Hoffman said</a> in the announcement. "When that kind of agency is available to everyone, you get superagency."</p><p>The skeptic's case is easy to make. Inflection is a fraction of its former size, competing for consumer attention against products from companies spending tens of billions of dollars a year. Pi's model was state of the art in 2023; it is not in 2026. And "<a href="https://www.linkedin.com/posts/inflectionai_inflection-ai-is-shaping-the-future-of-personal-activity-7485407087926312960-fqCl/">relational intelligence</a>" is, for now, a brand claim awaiting proof.</p><p>But the bull case is not crazy either. Inflection's own research shows consumers already juggle multiple AI tools and choose them for qualities — tone, emotional understanding, personalization — that frontier labs treat as afterthoughts. The company kept its technology, its Microsoft licensing windfall, and a defensible enterprise niche in on-premise, emotionally intelligent deployments. And it is targeting the one consumer segment — everyday, mobile-first, voice-first life management — that the coding-obsessed giants have largely ignored.</p><p>Asked what success looks like twelve months from now, White declined to talk numbers. "It's less about scale for scale's sake and more about scaling for impact by empowering people and improving their lives," he said. "Over the next year, success means leading the market towards relational intelligence and transforming AI interactions from transactional to relational."</p><p>Two years ago, Microsoft walked away with Inflection's founders, its staff, and its shot at the frontier — but it left behind the one idea the giants still haven't figured out how to build: an AI that knows the people in your life matter more than the tasks on your list. Inflection is betting the company, again, that the idea was the valuable part all along.</p><p>
</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Unsloth vs Axolotl vs TRL vs LLaMA-Factory: A Fine-Tuning Framework Comparison on Speed, VRAM, and Multi-GPU]]></title>
<description><![CDATA[Four open source projects dominate LLM fine-tuning today. Unsloth, Axolotl, TRL, and LLaMA-Factory all wrap the same underlying PyTorch and Hugging Face stack. They diverge on where they spend engineering effort. Unsloth rewrites kernels. Axolotl composes parallelism strategies. TRL defines the t...]]></description>
<link>https://tsecurity.de/de/3685789/ai-nachrichten/unsloth-vs-axolotl-vs-trl-vs-llama-factory-a-fine-tuning-framework-comparison-on-speed-vram-and-multi-gpu/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3685789/ai-nachrichten/unsloth-vs-axolotl-vs-trl-vs-llama-factory-a-fine-tuning-framework-comparison-on-speed-vram-and-multi-gpu/</guid>
<pubDate>Wed, 22 Jul 2026 11:22:13 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Four open source projects dominate LLM fine-tuning today. Unsloth, Axolotl, TRL, and LLaMA-Factory all wrap the same underlying PyTorch and Hugging Face stack. They diverge on where they spend engineering effort. Unsloth rewrites kernels. Axolotl composes parallelism strategies. TRL defines the trainer APIs the others build on. LLaMA-Factory optimizes for breadth of model coverage and […]</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/22/unsloth-vs-axolotl-vs-trl-vs-llama-factory-a-fine-tuning-framework-comparison-on-speed-vram-and-multi-gpu/">Unsloth vs Axolotl vs TRL vs LLaMA-Factory: A Fine-Tuning Framework Comparison on Speed, VRAM, and Multi-GPU</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Helios marks AMD’s biggest AI infrastructure push yet]]></title>
<description><![CDATA[AMD has expanded its AI infrastructure portfolio with the launch of Helios, an open, rackscale AI infrastructure designed for frontier AI and sovereign computing. Helios is built around AMD’s next-generation Instinct GPUs, EPYC Venice processors, Pensando networking and the ROCm software stack.

...]]></description>
<link>https://tsecurity.de/de/3683516/it-security-nachrichten/helios-marks-amds-biggest-ai-infrastructure-push-yet/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3683516/it-security-nachrichten/helios-marks-amds-biggest-ai-infrastructure-push-yet/</guid>
<pubDate>Tue, 21 Jul 2026 13:21:39 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">AMD has expanded its AI infrastructure portfolio with the launch of Helios, an open, rackscale AI infrastructure designed for frontier AI and sovereign computing. Helios is built around AMD’s next-generation Instinct GPUs, EPYC Venice processors, Pensando networking and the ROCm software stack.</p>



<p class="wp-block-paragraph">“Helios is AMD’s first complete AI rack system with GPUs, CPUs, and networking built together, instead of selling separate chips. It is well suited for training large AI models, memory heavy models, long context processing and high volume inference, and AMD’s biggest shot yet at challenging Nvidia’s dominance,” said Pareekh Jain, CEO at EIIRTrend &amp; Pareekh Consulting.</p>



<p class="wp-block-paragraph">AMD has also secured an early hyperscale deployment for Helios with <a href="https://newsroom.amd.com/news/microsoft-azure-ai-infrastructure/" target="_blank" rel="noreferrer noopener">Microsoft</a> agreeing to deploy it to power its frontier model AI inference, its AI customers, and support Azure AI services.</p>



<h2 class="wp-block-heading">The architecture behind Helios</h2>



<p class="wp-block-paragraph">The launch of Helios marks AMD’s latest attempt to strengthen its position in a market where Nvidia continues to dominate AI infrastructure. Unlike previous AMD AI offerings centred on individual accelerators, Helios is designed as a complete rack-scale system integrating compute, networking and software.</p>



<p class="wp-block-paragraph">According to Jain, Helios goes up against Nvidia’s <a href="https://www.networkworld.com/article/4188058/nvidia-unveils-vera-rubin-platform-targeting-ai-hpc-infrastructure-customers.html?utm=hybrid_search">Vera Rubin</a> rack. “Nvidia is faster on raw inference speed and has a faster internal connection between chips whereas AMD wins on memory size and offers better value for the price and power used. It’s standout feature is memory, where each rack packs about 50% more total memory than Nvidia’s competing system, which helps run very large AI models. It also uses open, industry-standard connections instead of Nvidia’s private technology, giving buyers more flexibility,” he said.</p>



<p class="wp-block-paragraph">The AMD Helios rackscale design includes 72 AMD Instinct MI455X GPUs with AMD EPYC Venice CPUs and AMD Pensando Vulcano networking using UALink, optimized for compute, data movement, and system efficiency. The platform also supports both OCP and MX data types, delivering up to 2.9 EFLOPS of FP4 and 1.4 EFLOPS of FP8 compute for AI training and inference. </p>



<p class="wp-block-paragraph">It also integrates 31TB of HBM4 memory with 19.6TB/s of memory bandwidth, while a liquid-cooling design uses quick-disconnect connections to efficiently dissipate heat. It is designed on open standards including OCP Open Rack Wide (ORW), <a href="https://www.networkworld.com/article/4155357/new-v2-ualink-specification-aims-to-catch-up-to-nvlink.html?utm=hybrid_search">Ultra Accelerator Link (UALink)</a>, and <a href="https://www.networkworld.com/article/4006285/ultra-ethernet-consortium-publishes-1-0-specification-readies-ethernet-for-hpc-ai.html?utm=hybrid_search">Ultra Ethernet Consortium (UEC)</a> and can be scaled efficiently across datacenters while optimizing power, cooling, and serviceability for modern AI infrastructure, <a href="https://www.amd.com/en/products/rackscale-solutions/helios.html" target="_blank" rel="noreferrer noopener">said</a> the company.</p>



<p class="wp-block-paragraph">On the security front, Helios incorporates a hardware root of trust and continuous attestation at every layer. It supports hardware-enforced isolation, encrypted memory and interconnects to help protect AI models, data and workloads in multi-tenant environments.</p>



<h2 class="wp-block-heading">The software challenge</h2>



<p class="wp-block-paragraph">While the launch of Helios might help AMD close the hardware gap with Nvidia’s rack-scale systems, it will be the software compatibility that will be the real driver of enterprise adoption.</p>



<p class="wp-block-paragraph">For this, AMD is expanding its ROCm AI software platform too, which supports frameworks including PyTorch, TensorFlow, and JAX, for enabling high-throughput inference and efficient distributed training while preserving familiar developer workflows.</p>



<p class="wp-block-paragraph">Jain stated While hardware parity or superiority in memory bandwidth is achievable, software maturity remains the key differentiator for Nvidia. The Nvidia’s <a href="https://www.networkworld.com/article/4079693/quantum-circuits-brings-dual-rail-qubits-to-nvidias-cuda-q-development-platform.html?utm=hybrid_search">CUDA</a> software has a 15-20 year head start, and almost every AI tool, tutorial, and codebase defaults to it.</p>



<p class="wp-block-paragraph">He added software has been AMD’s weak spot. AMD has improved  ROCm a lot but it still lags behind on the newest, most specialized optimizations, and setup is more complicated. For everyday AI work, ROCm is usable but for cutting-edge performance, CUDA still leads.</p>



<h2 class="wp-block-heading">Evaluating the trade-offs</h2>



<p class="wp-block-paragraph">For CIOs evaluating AI infrastructure, Helios launch brings in another option to a market that has largely revolved around Nvidia’s dominance. But when considering Helios, CIOs will have to evaluate factors such as performance, software readiness, deployment models, procurement timelines and total cost of ownership before committing to a platform.</p>



<p class="wp-block-paragraph">While AMD has not publicly announced a specific price tag for the Helios, Jain believes it to be noticeably cheaper to buy and run with lower chip prices and lower power use per GPU.</p>



<p class="wp-block-paragraph">“It gives companies a real second option besides Nvidia, easing supply shortages and giving leverage in negotiations. The catch is software, where teams need to check whether their AI tools run well on AMD’s stack, since some advanced tools are still CUDA only,” Jain said. </p>



<p class="wp-block-paragraph">For CIOs planning to deploy both, Jain warns the two systems can’t be plugged together into one combined machine as they use different, incompatible connection technology. But companies can and do run both side by side in the same data center, just as separate systems handling different jobs.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-58659 | Lightning-AI PyTorch Lightning up to 2.6.5 Instantiator load_state module state issue (EUVD-2026-44751)]]></title>
<description><![CDATA[A vulnerability labeled as critical has been found in Lightning-AI PyTorch Lightning up to 2.6.5. Impacted is the function load_state of the component Instantiator. Such manipulation of the argument module leads to state issue.

This vulnerability is referenced as CVE-2026-58659. It is possible t...]]></description>
<link>https://tsecurity.de/de/3677678/sicherheitsluecken/cve-2026-58659-lightning-ai-pytorch-lightning-up-to-265-instantiator-loadstate-module-state-issue-euvd-2026-44751/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3677678/sicherheitsluecken/cve-2026-58659-lightning-ai-pytorch-lightning-up-to-265-instantiator-loadstate-module-state-issue-euvd-2026-44751/</guid>
<pubDate>Sat, 18 Jul 2026 10:09:24 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability labeled as <a href="https://vuldb.com/kb/risk">critical</a> has been found in <a href="https://vuldb.com/product/lightning-ai:pytorch_lightning">Lightning-AI PyTorch Lightning up to 2.6.5</a>. Impacted is the function <code>load_state</code> of the component <em>Instantiator</em>. Such manipulation of the argument <em>module</em> leads to state issue.

This vulnerability is referenced as <a href="https://vuldb.com/cve/CVE-2026-58659">CVE-2026-58659</a>. It is possible to launch the attack remotely. No exploit is available.]]></content:encoded>
</item>
<item>
<title><![CDATA[Demystifying AI Exploits: A Blueprint for AI-Assisted Vulnerability Management]]></title>
<description><![CDATA[Written by: Jules Czarniak

Introduction 
As highlighted in the Mandiant M-Trends 2026 report, the mean time-to-exploit (TTE) has dropped to -7 days, meaning vulnerabilities are often exploited a week before a patch even exists. 
To keep pace, many security teams are exploring how to integrate la...]]></description>
<link>https://tsecurity.de/de/3673775/it-security-nachrichten/demystifying-ai-exploits-a-blueprint-for-ai-assisted-vulnerability-management/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3673775/it-security-nachrichten/demystifying-ai-exploits-a-blueprint-for-ai-assisted-vulnerability-management/</guid>
<pubDate>Thu, 16 Jul 2026 16:23:24 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div class="block-paragraph_advanced"><p>Written by: Jules Czarniak</p>
<hr></div>
<div class="block-paragraph_advanced"><h3><span>Introduction </span></h3>
<p><span>As highlighted in the </span><a href="https://cloud.google.com/security/resources/m-trends"><span>Mandiant M-Trends 2026 report</span></a><span>, the mean time-to-exploit (TTE) has dropped to -7 days, meaning vulnerabilities are often exploited a week before a patch even exists. </span></p>
<p><span>To keep pace, many security teams are exploring how to integrate large language model (LLM) agents into their codebases, development environments and continuous integration and continuous delivery (CI/CD) pipelines for automated vulnerability discovery and remediation. However, deploying privileged artificial intelligence (AI) agents without mature integration processes introduces new architectural risks. </span></p>
<p><span>In response to customer inquiries about how to safely integrate AI capabilities into vulnerability management workflows, this blog provides actionable guidance from Mandiant Consulting about how to establish operational guardrails for AI assisted vulnerability management, including several detailed scenarios. What each of these examples show is that security teams can accelerate workflows with AI while also upholding the structural integrity of their environments. We suggest that combining AI capabilities with deterministic controls and human intelligence in strategic ways maximizes benefits and reduces risk. </span></p>
<h3><span>Establish Operational Guardrails to Safely Deploy AI Agents</span></h3>
<p><span>To safely adopt advanced AI capabilities without introducing unpredictable failures into deployment pipelines, organizations should ground their approach in established industry standards. While guidelines like the </span><a href="https://www.nist.gov/itl/ai-risk-management-framework" rel="noopener" target="_blank"><span>NIST AI Risk Management Framework (RMF)</span></a><span> and the </span><a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener" target="_blank"><span>OWASP Top 10 for LLMs</span></a><span> provide comprehensive baselines for identifying risks, operationalizing these controls requires a structural blueprint.</span></p>
<p><span>Frameworks like </span><a href="https://safety.google/intl/en_sg/safety/saif/" rel="noopener" target="_blank"><span>Google’s Secure AI Framework (SAIF)</span></a><span> </span><a href="https://safety.google/intl/en_sg/safety/saif/" rel="noopener" target="_blank"><span>and</span></a><a href="https://storage.googleapis.com/gweb-research2023-media/pubtools/1018686.pdf" rel="noopener" target="_blank"><span> </span><span>Google’s approach to secure AI Agents</span></a><span> provide a practical path forward, demanding that organizations extend existing deterministic controls directly into the AI execution environment. When deploying AI agents, security teams should navigate specific operational and structural risks:</span></p>
<ul>
<li aria-level="1">
<p role="presentation"><strong>Pre-agent data security and Defense-in-Depth:</strong><span> Agents should not be able to access personally identifiable information (PII), protected health information (PHI), or other sensitive data. Organizations should enforce data security before the prompt reaches the model. This includes strictly using non-production environments populated with synthetic data for testing. For production, security teams should deploy a hybrid defense-in-depth model. This includes Layer 1 deterministic policy engines acting as chokepoints, alongside Layer 2 reasoning-based defenses like specialized guard models (such as </span><a href="https://docs.cloud.google.com/model-armor/overview"><span>Model Armor</span></a><span> or similar provider-agnostic guardrails) to filter out sensitive data and block malicious prompt injections before they reach the agent layer. Crucially for vulnerability discovery, security teams should treat the codebase itself as an untrusted input. Threat actors can embed indirect prompt injections within source code comments or third-party dependencies (e.g., hidden instructions telling the agent to ignore vulnerabilities or exfiltrate environment variables), making input sanitation a requirement even for internal scanning.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Cloud provider limitations and zero data retention (ZDR):</strong><span> Many cloud and LLM providers block or throttle automated offensive security probing by default to prevent abuse. Organizations should establish clear rules of engagement and authorized testing agreements to navigate acceptable use policies. Furthermore, organizations should enforce strict zero data retention (ZDR) agreements with their LLM providers to guarantee that proprietary code and discovered vulnerabilities are never used to train external models.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Workload isolation:</strong><span> Agent workloads should execute in strictly isolated, unprivileged containers with dynamically limited privileges. By relying on robust sandboxing to prevent privilege escalation, if an agent hallucinates a destructive command or is hijacked via prompt injection, the blast radius remains contained.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Red Teaming:</strong><span> Before deploying autonomous vulnerability scanners that can dynamically spin up sandboxes and execute code, organizations should subject the AI agents themselves to human-led red teaming as part of comprehensive assurance efforts. This validates the agent's resilience against jailbreaks, recursive logic loops, and complex prompt injections, ensuring the security tooling does not become the attack vector.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Least-Privileged Machine Identities and Human Controllers:</strong><span> While workloads should be isolated, agents inherently require privileges to generate pull requests and commit code. Security teams should ensure these agents operate under distinct, strictly scoped machine identities that tie back to human controllers to ensure accountability and user consent. Organizations should use short-lived, just-in-time (JIT) tokens bound exclusively to the specific repository and branch under review. T</span><span>his enforces the principle of limited agent powers and ensures that even if an agent’s container is compromised via prompt injection, the threat actor cannot pivot to modify adjacent enterprise codebases.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Supply chain resilience for skills:</strong><span> As developers augment AI with third-party skills and model context protocol (MCP) servers, security teams should treat these integrations as untrusted supply chain components. MCP plugins introduce the risk of supply chain poisoning, where a previously benign integration is silently updated with malicious dependencies. Additionally, security teams should evaluate the underlying agent orchestration frameworks themselves (e.g., LangChain, AutoGen) for inherent vulnerabilities, such as session memory poisoning or recursive loop hijacking.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Toxic flow analysis (TFA) and Observable Actions:</strong><span> The objective of TFA is to monitor data paths at runtime, ensuring agents do not exfiltrate sensitive internal context to unvetted external endpoints. Agent actions, inputs, reasoning, and outputs must be fully observable and transparently logged. While implementing dynamic taint tracking for LLMs remains a complex architectural challenge, organizations should clearly separate this runtime observability from static supply chain controls. Integrating threat intelligence to hash and vet incoming agent tools provides a necessary baseline for verifying integrity </span><span>before</span><span> deployment. However, because static controls cannot address behavior post-deployment, mitigating data exfiltration ultimately requires active runtime monitoring and secure, centralized logging to trace and restrict the actual flow of data.</span></p>
</li>
</ul></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Demystifying_AI_image1.max-1000x1000.png" alt="Demystifying AI image1">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="u6hlz">Figure 1: Visual representation of an isolated AI agent environment using SAIF mechanisms</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><p><span>By operationalizing these tools within frameworks that demand verifiable integrity and structural resilience, organizations can safely bridge the gap between AI velocity and enterprise defense.</span></p>
<h3><span>The need for human-led threat modeling</span></h3>
<p><span>While LLMs excel at identifying syntax patterns, source code itself rarely contains the full picture of unwritten business intent. Some organizations attempt to solve this by connecting LLM agents to internal wikis, design documents, and issue trackers using retrieval-augmented generation (RAG).</span></p>
<p><span>While RAG gives the model access to external business context, it is not a perfect fix. Corporate documentation is frequently stale, contradictory, or incomplete. An AI agent might retrieve an outdated architecture diagram and confidently hallucinate a secure path that no longer exists in production. Because LLM agents struggle to resolve conflicting, undocumented human assumptions, human-led threat modeling remains a critical security control across both legacy applications and modern agent workflows.</span></p>
<p><span>Security teams should apply threat modeling during both the pre-build system design phase to establish a secure foundation, and during post-build architecture reviews. While an AI agent might successfully identify a poorly configured internal endpoint locally, a human threat modeler asks the structural question: </span><span>why does that microservice possess broad database read permissions in the first place?</span><span> </span></p>
<p><span>Identifying architectural vulnerabilities requires reasoning about business risk, data sensitivity, and operational constraints. To structure this process, organizations can use industry frameworks like PASTA (Process for Attack Simulation and Threat Analysis) or service offerings like the </span><a href="https://services.google.com/fh/files/misc/ds-threat-modeling-security-service-en.pdf" rel="noopener" target="_blank"><span>Mandiant Threat Modeling Security Service</span></a><span> to map trust boundaries, uncover structural design flaws, and prioritize compensating controls. Securing fundamental architecture through human oversight is a necessary component when relying on automated agents to find bugs in a poorly designed system.</span></p>
<p><span>Once these AI agents are safely sandboxed, as guided by SAIF, and the architecture is verified through threat modeling, organizations can typically apply them to two different problem spaces: Enterprise Vulnerability Management (to assist in managing the volume of known CVEs in commercial off-the-shelf (COTS) software and infrastructure) and Product Security (to identify vulnerabilities in 1st-party (1P) code).</span></p>
<h3><span>Track 1: Enterprise Vulnerability Management</span></h3>
<h4><span>Foundational security and discovery </span></h4>
<p><span>While the second track of this post explores how AI agents can uncover complex zero-days in custom code, organizations should manage the scale of enterprise infrastructure in tandem with these AI deployments. Even as new AI capabilities dominate headlines, organizations should still address foundational security challenges, such as secrets sprawl, unmanaged service accounts, missing FIDO2 MFA, and legacy VPN concentrators. Although vulnerability exploitation was the primary initial infection vector in intrusions Mandiant investigated last year, threat actors consistently rely on missing foundational controls and unpatched edge devices to secure and escalate their foothold after exploiting a vulnerability.</span></p>
<p><span>Furthermore, AI cannot replace foundational visibility. As security teams deploy AI agents, they should simultaneously close these tactical entry points by maximizing dynamic discovery capabilities like External Attack Surface Management (EASM), Cloud Security Posture Management (CSPM), and Continuous Threat Exposure Management (CTEM). In hybrid and cloud environments, tools like </span><a href="https://cloud.google.com/wiz?e=48754805"><span>Wiz</span></a><span> can be used to map this initial footprint.</span></p>
<h3><span>Risk-based vulnerability management </span></h3>
<p><span>Vulnerability management teams are already overwhelmed by the current volume of findings generated by traditional scanners. As organizations scale dynamic discovery tools, such as EASM, CSPM and CTEM, alongside automated AI agents, this influx of findings will compound the problem. To manage this influx, telemetry from these diverse discovery methods must first be normalized and deduplicated. This normalized data serves two purposes: it feeds directly into the risk engine, and it acts as a live overlay to correct stale records in the configuration management database (CMDB). By evaluating the deduplicated vulnerabilities alongside this newly updated asset context and frontline threat intelligence, the RBVM engine calculates a custom risk score that allows security teams to dynamically prioritize remediation.</span></p>
<p><span>A mature RBVM methodology calculates a customized risk score on a 0 to 100 scale using a weighted average. A sample formula for calculating this risk-based score is:</span></p>
<p><span>Final Score = (W_1 * S_vuln) + (W_2 * S_asset) + (W_3 * S_threat)</span></p>
<p><span>The variables and weights (W) are customized to the organization's risk appetite (for example, 0.20 for vulnerability, 0.40 for asset, and 0.40 for threat, summing to 1.0), while the underlying variables (S) are scored on a 0 to 100 scale and defined as follows:</span></p>
<ul>
<li aria-level="1">
<p role="presentation"><strong>Vulnerability severity (S_vuln): </strong><span>The inherent technical severity of the flaw. This is calculated by taking the CVSS Base Score (which natively accounts for confidentiality, integrity, and availability impact) and multiplying it by 10.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Asset context (S_asset): </strong><span>A combined metric of exposure and data sensitivity. Scores range from 100 for internet-facing assets holding customer data, down to 25 for internal-only assets with no sensitive data. To translate this impact into monetary terms for non-technical stakeholders, organizations can incorporate Factor Analysis of Information Risk (FAIR) principles into this metric. However, this approach requires highly accurate, continuously updated financial data that many enterprises struggle to maintain at scale.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Threat context (S_threat): </strong><span>The real-world urgency of the vulnerability. Scores range from 100 if actively exploited by threat actors relevant to the organization's profile, 75 if a proof-of-concept exists or if it is a vulnerability class easily exploited by autonomous AI agents, down to 25 if the exploit is theoretical and highly complex. Organizations should also map the Exploit Prediction Scoring System (EPSS) probability percentage directly into this variable. This allows the threat score to automatically scale up or down as real-world exploitation telemetry shifts, aligning static vulnerability data with active threat intelligence.</span></p>
</li>
</ul>
<p><span>An asset's customized risk score should directly influence internal remediation service-level agreements (SLAs), unless external compliance-driven mandates, such as CISA Binding Operational Directives (BODs), or relevant equivalents, override internal prioritization. A risk-driven and threat-intelligence-driven vulnerability prioritization methodology will help organizations focus resources on managing and mitigating the most critical security vulnerabilities first. This is an area where LLMs can support the vulnerability management process, particularly by helping teams synthesize unstructured threat intelligence to surface relevant risk contexts more efficiently. Enforcing strict SLOs for patching, while requiring formal risk acceptance documentation for any patching exceptions, will help reduce the number of vulnerabilities available to threat actors and increase the visibility of outstanding risks across the organization. Furthermore, organizations should integrate RBVM data directly into their security orchestration, automation, and response (SOAR) platforms for automated alert enrichment.</span></p></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--medium
      
      
        h-c-grid__col
        
        h-c-grid__col--4 h-c-grid__col--offset-4
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Demystifying_AI_image5.max-1000x1000.png" alt="Demystifying AI image5">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="ce5s1">Figure 2: Integration points of a risk-based vulnerability management (RBVM) program.</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><h3><span>Containment and Observability</span></h3>
<p><span>Modern architecture blueprints must prioritize attack surface reduction under the assumption that vulnerabilities will inevitably be exploited. Moving away from traditional perimeter defenses, organizations should align with zero trust principles, ensuring that security boundaries are established around every asset, workload, and identity.</span></p>
<p><span>A component of this alignment is the implementation of strong authentication principles. Organizations should eliminate implicit trust by enforcing continuous, context-aware authentication and authorization. Utilizing Zero Trust Network Access (ZTNA) solutions, such as Identity-Aware Proxies (IAP), shields critical management interfaces (e.g., SSH, RDP) and internal systems from direct internet exposure, granting access only to verified identities and compliant devices.</span></p>
<p><span>For public-facing applications and APIs, attack surface reduction involves deploying Layer 7 inspection at the load balancer or API gateway level. This hardening layer enforces strict schema validation, intercepting and neutralizing malformed inbound traffic and potential exploits before they can interact with internal application logic.</span></p>
<p><span>Securing the software supply chain is equally vital in modern blueprints, and organizations should align with frameworks like </span><a href="https://slsa.dev/spec/v0.1/levels" rel="noopener" target="_blank"><span>Supply-chain Levels for Software Artifacts (SLSA)</span></a><span> across both dependency and build tracks. Security policies should mandate that third-party dependencies are routed through a centralized artifact repository equipped with automated curation services, such as </span><a href="https://cloud.google.com/security/products/assured-open-source-software"><span>Google Assured Open Source Software (OSS)</span></a><span> or an equivalent solution, preventing untrusted code from entering the development lifecycle. Furthermore, maturing toward advanced SLSA build levels (e.g., SLSA level 3) through the implementation of isolation, ephemerality and reproducibility requirements via  ephemeral compute infrastructure for CI/CD runners reduces the likelihood of attacker persistence by ensuring environments are short-lived and automatically cycled.</span></p>
<p><span>To complement these pre-build controls, runtime observability should be established across all production workloads. This requires monitoring both infrastructure-level behavior and the specific runtime libraries actively executing in production, which surfaces true exploitable risk far beyond a static Software Bill of Materials. In tandem with monitoring workloads, organizations should secure how they authenticate by implementing workload identity federation. By removing static credentials and instead using short-lived tokens backed by strong cryptographic identity verification, organizations can reduce the risk of credential theft and unauthorized lateral movement.</span></p>
<p><span>Within the internal environment, microsegmentation should be enforced to break down flat networks into granular security zones. Routing application traffic through a Secure Access Service Edge (SASE) architecture integrates network routing directly with robust identity controls, rendering internal services completely invisible to unauthenticated users and containing threats to their initial point of entry.</span></p>
<p><span>Finally, automated containment and incident response within a zero trust framework must rely on deterministic, auditable tooling. Endpoint detection and response (EDR) platforms and SOAR playbooks should handle high-fidelity containment tasks through hardcoded execution logic. While AI tools accelerate triage and policy recommendation, actual execution capabilities must remain restricted to well-defined, pre-tested workflows to maintain total architectural predictability.</span></p></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Demystifying_AI_image8.max-1000x1000.png" alt="Demystifying AI image8">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="ak3zc">Figure 3: Structural containment and observability architecture</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><h3><span>Track 2: Product Security &amp; Development (1P Code)</span></h3>
<h4><span>Deterministic and probabilistic tooling</span></h4>
<p><span>Integrating LLM agents into vulnerability management and security workflows requires recognizing the differences between deterministic and probabilistic tooling. Traditional SAST and DAST tools utilize fixed methodologies to evaluate vulnerabilities through structural code parsing or definitive runtime observations. LLMs, however, evaluate source code by processing tokens simultaneously to calculate statistical and semantic relationships, rather than tracing deterministic execution tracks.</span></p>
<p><span>While techniques like Chain of Thought (CoT) prompting allow models to bridge this gap by decomposing complex code paths into intermediate reasoning steps, this process remains bounded by architectural limitations. Even when a model possesses a context window large enough to ingest entire repositories, it may experience attention degradation across long inputs, often failing to correctly weight intervening validation or sanitization logic within the prompt. For example, if a variable is tainted on line 10 but sanitized on line 500, attention degradation can cause the model to lose track of the sanitization logic. Furthermore, when enterprise codebases require chunking to fit within context limits, the resulting fragmentation may cause the model to lose track of end-to-end data flows.</span></p>
<p><span>Consequently, probabilistic engines are effective at uncovering localized, static anomalies, such as hardcoded credentials or outdated dependencies, but frequently misjudge complex vulnerabilities split across fragmented chunks or extended context windows. Notable exceptions occur when these probabilistic models are coupled with deterministic feedback loops. For instance, when analyzing C++ memory corruption, an LLM can be equipped with a test harness to iteratively execute code and definitively prove a crash. While these dynamic validation applications are detailed in subsequent sections, the baseline limitation for static analysis across standard enterprise codebases remains: models struggle to consistently evaluate dispersed logic.</span></p></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Demystifying_AI_image4.max-1000x1000.png" alt="Demystifying AI image4">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="ak3zc">Figure 4: Deterministic SAST scanners vs. probabilistic LLMs</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><h3><span>Binary and architectural oracles</span></h3>
<p><span>Many security programs are moving toward agent workflows where an agent autonomously spins up a test environment and uses tools to execute payloads and verify its findings. This is a promising approach, but it is important to understand where it is most effective.</span></p>
<p><span>Agent workflows perform well against bug classes with binary and observable oracles, meaning the system provides an objective, 'crash or no crash' feedback loop. For example, if a model is hunting for memory corruption in a C++ kernel, a successful exploit is undeniable: the payload executes, and a resulting crash definitively proves the vulnerability. This explains why the industry is currently seeing a surge in AI-discovered vulnerabilities across memory-unsafe targets like web browsers and operating systems.</span></p>
<p><span>However, enterprise software is heavily dominated by vulnerabilities that require architectural oracles for validation. Vulnerabilities like authorization bypasses, complex business logic flaws, and indirect server-side request forgeries require an understanding of business context and cross-service trust boundaries. If an agent's payload fails to produce a clear outcome, it can't reliably distinguish whether the vulnerability is a hallucination or if it simply constructed the payload incorrectly. An agent's malformed payload might even crash an unrelated background process and cause the model to hallucinate a success and report a false confirmation. Complex enterprise architecture contains unwritten business intent that a probabilistic engine can't inherently know.</span></p></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Demystifying_AI_image3.max-1000x1000.png" alt="Demystifying AI image3">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="bg92b">Figure 5: Evaluating vulnerabilities against binary vs. architectural oracles</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><h3><span>Targeted deployment and human impact</span></h3>
<p><span>Organizations adopting LLMs for vulnerability discovery face a massive staffing challenge. LLMs can generate findings significantly faster than human engineers can triage them. If every LLM-generated alert requires manual review, security teams will quickly face burnout and/or suffer alarm fatigue.</span></p>
<p><span>Rather than indiscriminately pointing agents at all available codebases and risking an influx of unverified output, security teams need a selective deployment strategy. Mature programs should maintain SAST and DAST for baseline hygiene and deterministic rule enforcement, and reserve intensive agent audits for high-impact components with clear binary oracles.</span></p>
<p><span>Organizations can prioritize agent audits on systems where the technology's strengths align with the broader risk profile:</span></p>
<ul>
<li aria-level="1">
<p role="presentation"><strong>Memory-unsafe codebases:</strong><span> Legacy or high-performance components written in memory-unsafe languages such as C, C++, or Assembly are strong candidates for LLM audits. These languages are susceptible to memory corruption flaws, such as buffer overflows and use-after-free conditions. Because these vulnerabilities trigger definitive failure states like segmentation faults, they work well with automated sandboxes where agents can compile the code with memory sanitizers and write proof-of-concept inputs. This approach is also effective for auditing the native extensions where safe languages call unsafe internal libraries, such as Python C extensions or the Java Native Interface (JNI).</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Systems highly exposed to outside content:</strong><span> First-party data ingestion pipelines, custom API gateways, or proprietary edge proxies. A prerequisite here is direct access to the source code, this strategy is strictly for internally developed or fully open-source codebases where the organization can inspect the logic. Because these systems directly parse untrusted internet traffic, targeting their source code for LLM-driven audits yields the highest risk-reduction ROI.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Shared internal libraries and utilities: </strong><span>Core serialization/deserialization packages, common utility functions, and custom middleware wrappers (such as internal message-queue parsers) maintained in-house. Because the enterprise owns the source code for these shared building blocks, agent tools can easily hook into them within automated test harnesses to fuzz inputs and catch low-level logic or parsing bugs with high fidelity.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Foundational security boundaries:</strong><span> Internally developed centralized authentication services, custom OAuth providers, and internal credential brokers. While testing complex identity boundaries generates higher logic-based noise, having full access to the source code allows teams to pair agents with deterministic checks to safely triage findings, given that the blast radius of an authentication failure justifies the human effort.</span></p>
</li>
</ul>
<p><span>To filter the noise generated by LLMs, organizations should establish routing rules. Require the agent to generate a fully reproducible, deterministic test harness (such as a compiled binary or a Python test script) that attempts to prove the exploit. This harness must execute automatically in an isolated, monitored sandbox. If the sandbox execution fails (due to a syntax error or a failed exploit), the ticket is discarded, sparing human resources. However, organizations should enforce execution timeouts and iteration limits on these test harnesses. Without hard limits, an autonomous agent attempting to prove a vulnerability can fall into an infinite loop: writing a script, failing, rewriting, and failing again, exhausting API token budgets and compute resources against a single dead-end vulnerability, creating significant cost overruns without advancing the security review. To manage these expenses, organizations should incorporate FinOps principles to balance the compute and API costs of LLM audits against the traditional expenses of manual triage.</span></p>
<p><span>However, a successful execution in the sandbox does not guarantee an actionable, high-priority risk. In practice, autonomous agents frequently produce working PoCs for genuine technical flaws that are ultimately irrelevant; or warrant a lower remediation priority within the context of the system's threat model. For example, the agent might successfully exploit an unreachable dead-code path, or trigger a bug that requires administrative access to execute and yields no further escalation of privilege. Therefore, a human engineer should be assigned to review and prioritize the ticket only if the sandbox registers a successful execution, validating environmental context, reachability, and true business impact as part of the review.</span></p>
<p><span>This workflow reduces the volume of alerts, but it is important to understand that the security team's workload does not disappear. The engineer's primary job shifts from manually hunting for the initial vulnerability to auditing the LLM-generated proof to ensure it represents a meaningful risk rather than an unexploitable or contextually irrelevant finding. Leadership should properly staff and train teams for this new reality. Deploying LLM agents does not remove the need for skilled practitioners; it redirects their workload toward complex validation. Equally important is training teams to recognize the risk of false negatives. A hyper-focus on filtering AI-generated noise can create a false sense of security. If an exploit relies on a novel technique or a zero-day vulnerability that was not heavily weighted in the model's training data, the agent will likely scan right past it in silence. LLMs augment discovery, but they do not guarantee exhaustive coverage.</span></p>
<p><span>When integrating LLMs into SAST triage pipelines, human engineers should also verify the broader architectural integrity. Prompting an LLM with specific SAST warnings can induce contextual narrowing, where the agent becomes hyper-fixated on resolving a localized syntax error and misses broader architectural flaws existing in the same file. Furthermore, if the agent's mandate extends beyond discovery to automated remediation (such as writing and proposing code fixes), this human-in-the-loop validation becomes critical to ensure the LLM does not inadvertently introduce new regressions or bypass intended business logic.</span></p></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image_20.max-1000x1000.png" alt="Demistiying Image 6 New">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="bg92b">Figure 6: Flowchart outlining the targeted LLM deployment and triage workflow.</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><h3><span>Remediation and hardening</span></h3>
<h4><span>LLM-assisted code remediation</span></h4>
<p><span>A primary goal of integrating large language models (LLMs) into the software development lifecycle is automated remediation. To achieve this, organizations are deploying these capabilities through two primary execution methods: directly within the integrated development environment (IDE) or as a centralized pipeline runner. Examples include </span><a href="https://deepmind.google/blog/introducing-codemender-an-ai-agent-for-code-security/" rel="noopener" target="_blank"><span>CodeMender</span></a><span>, although as of time of writing, it is not publicly available.</span></p>
<h4><strong>IDE-integrated method</strong><span> </span></h4>
<p><span>This method shifts remediation as far left as possible by operating as an active pair-programmer. Tools running continuous static analysis in the background of the IDE surface vulnerabilities directly to the developer via editor diagnostics like inline indicators or hover tooltips.</span></p>
<ul>
<li aria-level="1">
<p role="presentation"><strong>Localized scope:</strong><span> The developer can trigger the LLM agent to analyze the localized data flow and generate a targeted patch (such as implementing parameterized SQL queries). By constraining the LLM to localized, syntax-level fixes, the scope of the change remains contained. This prevents the agent from attempting sprawling, multi-file refactors that frequently break complex architectural logic.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Human-in-the-loop:</strong><span> The developer reviews the AI-generated patch before the code is committed.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Managing false positives:</strong><span> Local IDE agents allow developers to manage false positives dynamically. Suppressing alerts anchored to specific line text reduces alert fatigue and preserves developer trust.</span></p>
</li>
</ul>
<h4><strong>CI/CD runner method</strong><span> </span></h4>
<p><span>The runner method executes asynchronously within the CI/CD pipeline to use an LLM to review committed code and automatically propose remediation.</span></p>
<ul>
<li aria-level="1">
<p role="presentation"><strong>Restricted execution and deterministic validation: </strong><span>Asking a centralized runner to automatically rewrite a complex, multi-file authorization flaw directly in the main branch introduces a high risk of breaking logic errors. To mitigate this, agents must be restricted to generating pull requests (PRs). Once a PR is generated, it must automatically execute standard regression suites alongside the deterministic test harness. By rerunning the initial PoC against the patched code, the workflow repurposes the exploit script as a validation oracle to prove the vulnerability has been remediated. A human engineer then reviews the PR to validate the architectural logic before merging.</span></p>
</li>
</ul>
<p><span>In all cases security teams should define a clear boundary between the two methods rather than rely on a single approach. IDE agents provide immediate, syntax-level support. They catch and resolve low-complexity errors locally before developers commit code. Centralized CI/CD runners handle broader organizational baselines. They propose complex, repository-wide fixes for vulnerabilities that bypass local environments.</span></p>
<h4><strong>Post-deployment controls</strong><span> </span></h4>
<p><span>Even with human review and deterministic test harnesses, AI-generated patches can still introduce logic regressions in production. Organizations should implement strict post-deployment controls:</span></p>
<ul>
<li aria-level="1">
<p role="presentation"><strong>Automated rollbacks:</strong><span> Treating LLM-generated code with the same post-deployment scrutiny as any major architectural change ensures that if an unforeseen regression traverses the CI/CD pipeline, the environment can revert to a known good state.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Mitigating model drift:</strong><span> Relying on managed AI services introduces the ongoing risk of model drift. To prevent silent weight updates from breaking test harnesses, organizations need to pin specific model API versions to frozen releases. When a pinned version reaches its end-of-life, organizations will face a forced migration. Mitigating this pipeline fragility requires combining model pinning with deterministic regression suites.</span></p>
</li>
<li aria-level="1">
<p role="presentation"><strong>Compliance and auditability:</strong><span> If an AI agent automatically closes a security ticket or generates a patch in the CI/CD pipeline, organizations should maintain immutable audit logs to satisfy frameworks like SOC 2 ,PCI-DSS, FedRAMP, and CMMC. National security deployments must also account for data sovereignty requirements. This logging should record the specific model version that proposed the fix, the deterministic test results that validated it, and the human engineer who approved the merge. Furthermore, because emerging legislation like the EU AI Act emphasizes human oversight for high-risk applications, security teams should carefully evaluate how autonomous remediation workflows align with these evolving global regulatory standards.</span></p>
</li>
</ul></div>
<div class="block-image_full_width">






  
    <div class="article-module h-c-page">
      <div class="h-c-grid">
  

    <figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      ">

      
      
        
        <img src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Screenshot_2026-07-15_at_10.24.22PM.max-1000x1000.png" alt="demistifying image 7">
        
        
      
        <figcaption class="article-image__caption "><p data-block-key="bg92b">Figure 7: Flowchart demonstrating the difference between local IDE AI remediation and centralized CI/CD pipeline remediation.</p></figcaption>
      
    </figure>

  
      </div>
    </div>
  




</div>
<div class="block-paragraph_advanced"><h3><span>Conclusion</span></h3>
<p><span>Leveraging LLMs in vulnerability management is a multi-layer solution: Integrating it requires separating workflows by layer. At the enterprise infrastructure level, Risk-Based Vulnerability Management (RBVM) and exposure management are necessary to process the volume of findings and configuration drift. At the product and code security level, LLM-enabled vulnerability assessment and remediation must operate alongside foundational deterministic controls, such as SAST and DAST, to audit custom, open-source, or third-party code.</span></p>
<p><span>Although LLMs can help manage technical debt and accelerate vulnerability discovery, they do not replace secure-by-design principles. The fact that LLM agents are proving exceptionally capable at identifying and exploiting localized memory corruption in memory-unsafe codebases, alongside other primary vectors, should serve as a wake-up call. </span></p>
<p><span>As a long-term strategy aligned with </span><a href="https://media.defense.gov/2022/Nov/10/2003112742/-1/-1/0/CSI_SOFTWARE_MEMORY_SAFETY.PDF" rel="noopener" target="_blank"><span>NSA guidance on Software Memory Safety</span></a><span>, organizations need to phase memory-safe languages into new internal development. LLMs are beginning to expand what is possible here by reducing the manual labor required for code migration. Converting existing C or C++ codebases to Rust has historically been unrealistic due to the large volume of engineering hours needed. While fully automated translation is not a turn-key solution, using LLMs to assist engineers with the bulk of the conversion can make these long-term migrations operationally viable. Beyond internal efforts, organizations should use procurement requirements to incentivize vendors to reduce their reliance on memory-unsafe languages and establish secure configuration defaults over time. Bridging the gap between AI velocity and enterprise defense means building an automated pipeline to manage the current backlog, while architecting systems where entire classes of vulnerabilities and misconfigurations are eliminated by design.</span></p>
<h3><span>Acknowledgements</span></h3>
<p><span>This analysis would not have been possible without the assistance of Google Threat Intelligence Group (GTIG) and other broader Google teams.</span></p></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Thinking Machines open sources first multimodal language model, Inkling, focused on low cost and 'resistance to censorship']]></title>
<description><![CDATA[Enterprises looking to move more of their agentic AI workloads to open weights models they can customize, control and run on-premises or in virtual private clouds have a strong new contender to consider.Today, Thinking Machines—the highly capitalized American AI startup founded by former OpenAI C...]]></description>
<link>https://tsecurity.de/de/3672034/it-nachrichten/thinking-machines-open-sources-first-multimodal-language-model-inkling-focused-on-low-cost-and-resistance-to-censorship/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3672034/it-nachrichten/thinking-machines-open-sources-first-multimodal-language-model-inkling-focused-on-low-cost-and-resistance-to-censorship/</guid>
<pubDate>Thu, 16 Jul 2026 00:46:37 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Enterprises looking to move more of their agentic AI workloads to open weights models they can customize, control and run on-premises or in virtual private clouds have a strong new contender to consider.</p><p>Today, Thinking Machines—the highly capitalized American AI startup founded by former OpenAI CTO Mira Murati—<a href="https://thinkingmachines.ai/news/introducing-inkling/">released Inkling</a>, its first major language model under an<a href="https://choosealicense.com/licenses/apache-2.0/"> enterprise-friendly Apache 2.0 open source license</a>, and it boasts high, if sub state-of-the-art, performance for open weights models on third-party benchmarks, specifically software engineering (77.6% on SWE-bench Verified, where it beats fellow U.S. open rival Nvidia Nemotron 3's 71.9%) and voice understanding (91.4% on VoiceBench compared to 94.4% for Gemini 3.1 Pro on high reasoning effort).</p><p>Another differentiator: Thinking Machines notes that Inkling was designed "to answer directly on topics that may be subject to censorship," offering enterprises concerned about factual outputs, irrespective of controversy or sensitivity, a more trustworthy option. </p><p>Coming in at 975 billion total parameters, Inkling is a natively multimodal, open-weights Mixture-of-Experts (MoE) system capable of reasoning across text, images, and audio. The weights <a href="https://huggingface.co/thinkingmachines/Inkling">are already available on Hugging Face</a> and the company's own model training application programming interface (API), <a href="https://thinkingmachines.ai/tinker/">Tinker</a>.</p><p>Designed to balance cost against performance through a novel "controllable thinking effort" mechanism, the model represents a significant departure from the black-box scaling strategies of frontier competitors.</p><p>Alongside the flagship model, Thinking Machines also announced a preview of Inkling-Small, a lighter 276-billion-parameter alternative optimized for workloads where low latency and cost are paramount.</p><h2><b>Benchmarks Show a Powerful, High-End, Sub State-of-the-Art Model</b></h2><p>While Inkling is a formidable multimodal engine, it lands in a fiercely competitive 2026 open-weight landscape characterized by highly specialized MoE architectures. Rather than attempting to dominate every leaderboard, Thinking Machines explicitly designed Inkling—with 975 billion total and 41 billion active parameters—as a broad, balanced generalist. </p><p>For example, it comes in near the middle high-end of benchmark performance 1257 on Design Arena’s Agentic Web Dev leaderboard measuring human scores of frontend web design. </p><p>But China’s leading AI labs have produced models with elite reasoning and coding capabilities, posing a stiff challenge to Inkling's generalist approach and ultimately outperforming it on general and coding benchmarks.</p><ul><li><p><b>GLM 5.2:</b> Widely considered the top open-weight reasoning model available in the benchmark set, GLM 5.2 outperforms Inkling on pure coding, agentic, and complex reasoning tasks. It scores 62.1% on SWEBench Pro (Public) compared to Inkling’s 54.3%, and a massive 82.7 on Terminal Bench 2.1 against Inkling’s 63.8. GLM 5.2 also holds the edge in text-only reasoning, scoring 40.1% on HLE (text only) versus Inkling's 30.0%.</p></li><li><p><b>DeepSeek V4 Pro:</b> DeepSeek maintains an edge in several strict coding and factuality domains, beating Inkling on SWEBench Verified (80.6% vs. 77.6%) and SimpleQA Verified (57.0% vs. 43.9%). However, Inkling successfully overtakes DeepSeek V4 Pro in mathematical problem-solving, achieving 97.1% on AIME 2026 compared to DeepSeek's 96.7%.</p></li><li><p><b>Kimi K2.6:</b> This model outpaces Inkling across multiple technical benchmarks, delivering higher scores on GPQA Diamond (91.1% vs. 87.9%), BrowseComp (83.2% vs. 77.1%), and HLE with tools (54.0% vs. 46.0%). Yet Inkling proves more resilient on general chat instruction following, scoring 79.8% on IFBench compared to Kimi K2.6's 76.0%.</p></li></ul><p>Against its primary U.S.-based open-weight competition, Inkling demonstrates strong parity and frequent superiority.</p><ul><li><p><b>Nemotron 3 Ultra:</b> Inkling consistently outperforms this U.S. rival across reasoning and coding. Inkling posts 97.1% on AIME 2026 and 77.6% on SWEBench Verified, beating Nemotron's 94.2% and 70.7%, respectively. Furthermore, Inkling significantly leads in agentic workflows, scoring 74.1% on MCP Atlas against Nemotron's 44.7%.</p></li></ul><p>When compared to closed-source juggernauts like Claude Fable 5, GPT 5.6 Sol, and Gemini 3.1 Pro, Inkling trails in peak reasoning and software engineering autonomy, but remains highly competitive in multimodality.</p><ul><li><p><b>Coding and Reasoning:</b> Closed models maintain a commanding lead. Claude Fable 5 (max) hits 95.0% on SWEBench Verified and 53.3% on HLE (text only), far outpacing Inkling's 77.6% and 30.0%. GPT 5.6 Sol dominates Terminal Bench 2.1 with an 89.5, easily clearing Inkling's 63.8.</p></li><li><p><b>Native Multimodality:</b> Inkling's native visual and audio capabilities hold their own. On the MMMU Pro (Standard 10) vision benchmark, Inkling's 73.3% is competitive, though trailing Claude Fable 5's 84.2% and GPT 5.6 Sol's 83.0%. In audio processing, Inkling scores a highly respectable 77.2% on MMAU, keeping it within striking distance of Gemini 3.1 Pro's 82.5%.</p></li></ul><p>If an enterprise workflow demands elite software engineering autonomy or the highest bounds of text-only reasoning, models like GLM 5.2 or proprietary systems like Claude Fable 5 maintain the edge. </p><p>However, Inkling carves out a unique and highly defensible position: it is the most capable open-weight foundation model that natively fuses text, vision, and audio, while simultaneously offering developers direct programmatic control over the cost-to-performance ratio. </p><h2><b>The Shift from Static Reasoning to Controllable Thinking</b></h2><p>Rather than attempting to build a singular "god model" optimized strictly for state-of-the-art benchmark domination, Thinking Machines engineered Inkling for adaptability and efficiency in real-world workflows.</p><p>The standout feature of this release is Inkling's "controllable thinking effort." Developers can programmatically adjust the model's reasoning budget—scaling from 0.2 to 0.99—to dictate how hard the AI should "think" before generating an output. </p><p>As the company noted, "Inkling's continuous thinking effort lets you pick your point on the cost/performance curve—reaching the same score with a fraction of the tokens".</p><p>In practical terms, this allows enterprises to deploy Inkling with lower token expenditure for simpler tasks, while cranking up the compute overhead for complex, multi-step reasoning challenges. However, by keeping the thinking effort lower and generating fewer tokens, the cost-conscious enterprise can achieve high quality results and performance on simple tasks while spending less money, or, in the case of those running models locally, less costs on energy and compute resources.</p><p>During the model’s large-scale reinforcement learning (RL) training over 30 million rollouts, researchers observed an emergent phenomenon they called "chain of thought condensation". Over time, Inkling naturally learned to compress its internal reasoning steps—dropping grammatical overhead and connectives—while reaching the same accurate conclusions, resulting in drastically reduced latency.</p><h2><b>Epistemics and Censorship Resistance</b></h2><p>A notable element of Thinking Machines' release is its explicit focus on the model's epistemics—specifically its calibration, instruction following, and resistance to censorship. </p><p>In an ecosystem where open-weight models adopt either overly restrictive safety guardrails or echo state-aligned ideological talking points, Inkling was intentionally trained to answer directly on politically sensitive or heavily censored topics.</p><p>To validate this approach, Thinking Machines submitted Inkling to the <i>Propaganda and Censorship Eval</i> developed by AI startup Cognition. According to the published findings, Inkling demonstrated "strong patterns of censorship non-compliance," effectively resisting ideological capture or boilerplate refusals when presented with sensitive subjects.</p><p>Despite its resistance to censorship, the model maintains a robust defense against genuinely malicious, dangerous, or illegal queries. On the StrongREJECT benchmark—which tests responses to unambiguous harmful requests—Inkling scored 98.6%, placing it in line with strict frontier safety standards. Furthermore, on the FORTRESS benchmark, Inkling successfully navigated the line between safety and over-refusal: it achieved a 78.0% refusal rate on adversarial queries (such as those involving weapons, cyberattacks, or violence) while maintaining a 95.9% compliance rate on benign, look-alike queries.</p><p>Thinking Machines noted that typical open-weight vulnerabilities remain within the architecture. Internal safety evaluations revealed an "occasional tendency to comply with role-play and indirectly framed prompts concerning harmful topics". The company advised enterprise developers to treat the model's built-in refusals as just one layer of security, recommending the downstream deployment of external moderation tools—such as Llama Guard—to filter adversarial jailbreaks and enforce use-case-specific safety policies at the application level.</p><h2><b>Under the Hood: Architecture and Multimodality</b></h2><p>Inkling's scale is staggering, yet sparse. The MoE architecture features 975 billion total parameters, but only 41 billion parameters are active during any given token generation. It supports a massive context window of 1 million tokens and diverges from typical transformer models by using relative positional embeddings instead of the industry-standard Rotary Positional Embedding (RoPE).</p><p>True to the company's foundational vision, Inkling was trained from scratch to be natively multimodal. Unlike models that rely on bolted-on external encoders, Inkling uses an encoder-free early fusion approach. It directly ingests audio as discrete dMel spectrograms and visual data as 40x40 pixel patches via a hierarchical multi-layer perceptron (hMLP), projecting all modalities into a shared hidden space.</p><h2><b>Licensing: True Open-Source for the Enterprise</b></h2><p>For enterprise IT teams and developers, the most disruptive aspect of Inkling may be its licensing. Inkling is released under the permissive Apache 2.0 license.</p><p>In an ecosystem where many so-called "open" models from Western labs are tethered to dual-use commercial licenses, acceptable use restrictions, or revenue caps, an Apache 2.0 designation makes Inkling a true open-source foundation. This gives developers the legal freedom to download, modify, integrate, and commercialize the model weights entirely royalty-free.</p><p>The model is readily deployable across major open-source inference libraries—including SGLang, vLLM, TokenSpeed, and llama.cpp—and comes with a native NVFP4 quantized checkpoint optimized for NVIDIA Blackwell systems.</p><h2><b>Community Reactions: The Engineering Feat</b></h2><p>The AI community's response has been swift, praising both the model's openness and the underlying engineering execution.</p><p>In a<a href="https://x.com/johnschulman2/status/2077460227327467982"> post on X</a>, Thinking Machines co-founder John Schulman reflected on the rapid development cycle: "Inkling is out today, with open weights and in Tinker. It's been fun to watch this one come together: pretraining began last winter, and starting in mid-January a small team built up the coding, reasoning, and agentic training from there. We learned a lot building it, and I hope people find good uses for it."</p><div></div><p>Horace He, a researcher at Thinking Machines (previously from PyTorch), underscored the difficulty of the task in <a href="https://x.com/cHHillee/status/2077457790423969806">another post on X</a>: "It truly takes a village to release a model, perhaps especially an open weights model. Actually doing the entire process from scratch, from data to pretraining to posttraining to actual release, gives a lot of appreciation for anyone who does it!"</p><div></div><p>The broader open-source ecosystem has also embraced the technical integrations. Lysandre Debut, the Chief Open-Source Officer at Hugging Face, shared his enthusiasm regarding the model's optimization<a href="https://x.com/LysandreJik/status/2077459011285512267"> in his own X post</a>: "One thing I find quite striking is how much easier accelerating models has become... We replaced the model's causal Conv1D with the `causal-conv1d` kernel. One line changed, +4% tokens per second. We then replaced its attention implementation with FlashAttention-4. Another single change, another +11%. That's a total throughput improvement of about 15%, without changing the model architecture or retraining anything."</p><p>Tiezhen Wang, an ecosystem growth expert and ex-Googler, celebrated the release as a massive win for the open-source community, listing the model's impressive specifications on X, highlighting its "975B total, 41B active" size, "Native MTP support," and the highly coveted "Apache 2.0 license."</p><h2><b>Background: The Road to Inkling</b></h2><p>To understand the significance of Inkling, one has to look back at the rapid trajectory of Thinking Machines over the past 18 months.</p><p>When<a href="https://venturebeat.com/technology/ex-openai-cto-mira-murati-unveils-thinking-machines-a-startup-focused-on-multimodality-human-ai-collaboration"> Mira Murati departed OpenAI in late 2024 to found Thinking Machines</a> alongside industry veterans like John Schulman and Barret Zoph, the stated goal was to pivot away from building isolated autonomous agents. Instead, the company aimed to build flexible, multimodal systems designed for genuine human-AI collaboration and open science.</p><p>By July 2025, the startup had secured a historic $2 billion seed round led by Andreessen Horowitz at a $12 billion valuation. At the time, Murati promised the<a href="https://venturebeat.com/technology/mira-murati-says-her-startup-thinking-machines-will-release-new-product-in-months-with-significant-open-source-component"> impending release of a product with a "significant open source component" </a>to empower researchers and startups.</p><p>The company’s philosophy began coming into sharper focus in October 2025 with the launch of <a href="https://venturebeat.com/technology/thinking-machines-first-official-product-is-here-meet-tinker-an-api-for">Tinker</a>, a Python-based API for large language model fine-tuning that gave researchers granular control over training pipelines without the friction of distributed compute management.</p><p>That same month, Thinking Machines researcher <a href="https://venturebeat.com/ai/thinking-machines-challenges-openais-ai-scaling-strategy-first">Rafael Rafailov delivered a provocative critique of the AI industry at TED AI</a>. He argued that the current trajectory of simply throwing more compute at models was fundamentally flawed, noting that today's systems take shortcuts—like wrapping code in<code> try/except</code> blocks—because they are trained strictly for task completion rather than genuine learning. </p><p>Rafailov posited that the first artificial superintelligence would not be a "god model," but rather a "superhuman learner" capable of meta-learning and internalizing abstractions. Inkling’s architecture—specifically its controllable thinking effort and its ability to organically compress its chain of thought during RL—feels like the first tangible realization of Rafailov's thesis.</p><p>In May 2026, the lab teased its technical prowess with the<a href="https://venturebeat.com/technology/thinking-machines-shows-off-preview-of-near-realtime-ai-voice-and-video-conversation-with-new-interaction-models"> research preview of TML-Interaction-Small</a>, a system that eliminated "turn-based" chat by processing inputs and outputs simultaneously in 200ms chunks. This "full-duplex" breakthrough proved the company could build highly responsive, natively multimodal models from scratch.</p><p>Now, with Inkling out in the wild, Thinking Machines has delivered on its foundational promises. By offering a massive, natively multimodal model under a true open-source license, they aren't just giving developers a new tool—they are attempting to fundamentally rewrite the economics and accessibility of frontier AI development.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Building a Gin Config Controlled PyTorch Pipeline with Configurable MLP Variants, Cosine Scheduling, and Runtime Parameter Overrides]]></title>
<description><![CDATA[We build a Gin Config controlled PyTorch pipeline where the training code stays fixed and the experiment variables move into .gin files. We construct a nonlinear spiral binary classification task and define a configurable MLP with scoped architectural variants. We expose the optimizer, scheduler,...]]></description>
<link>https://tsecurity.de/de/3671603/ai-nachrichten/building-a-gin-config-controlled-pytorch-pipeline-with-configurable-mlp-variants-cosine-scheduling-and-runtime-parameter-overrides/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671603/ai-nachrichten/building-a-gin-config-controlled-pytorch-pipeline-with-configurable-mlp-variants-cosine-scheduling-and-runtime-parameter-overrides/</guid>
<pubDate>Wed, 15 Jul 2026 20:18:47 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>We build a Gin Config controlled PyTorch pipeline where the training code stays fixed and the experiment variables move into .gin files. We construct a nonlinear spiral binary classification task and define a configurable MLP with scoped architectural variants. We expose the optimizer, scheduler, loss, batching, seeding, and training loop through @gin.configurable bindings. We then run two scoped experiments, apply runtime overrides without editing source, and export the operative config for each run.</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/15/building-a-gin-config-controlled-pytorch-pipeline-with-configurable-mlp-variants-cosine-scheduling-and-runtime-parameter-overrides/">Building a Gin Config Controlled PyTorch Pipeline with Configurable MLP Variants, Cosine Scheduling, and Runtime Parameter Overrides</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Towards demystifying the creativity of diffusion models]]></title>
<description><![CDATA[Algorithms & Theory]]></description>
<link>https://tsecurity.de/de/3671601/ai-nachrichten/towards-demystifying-the-creativity-of-diffusion-models/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3671601/ai-nachrichten/towards-demystifying-the-creativity-of-diffusion-models/</guid>
<pubDate>Wed, 15 Jul 2026 20:18:44 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Algorithms &amp; Theory]]></content:encoded>
</item>
<item>
<title><![CDATA[7 newer data science tools you should be using with Python]]></title>
<description><![CDATA[Python’s rich ecosystem of data science tools is a big draw for users. The only downside of such a broad and deep collection is that sometimes the best tools can get overlooked.



Here’s a rundown of some of the best newer or less-known data science projects available for Python. Some, like Pola...]]></description>
<link>https://tsecurity.de/de/3665680/ai-nachrichten/7-newer-data-science-tools-you-should-be-using-with-python/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3665680/ai-nachrichten/7-newer-data-science-tools-you-should-be-using-with-python/</guid>
<pubDate>Mon, 13 Jul 2026 17:04:47 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div><div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p class="wp-block-paragraph">Python’s rich ecosystem of data science tools is a big draw for users. The only downside of such a broad and deep collection is that sometimes the best tools can get overlooked.</p>



<p class="wp-block-paragraph">Here’s a rundown of some of the best newer or less-known data science projects available for <a href="https://www.infoworld.com/article/2254260/how-to-get-started-with-python.html">Python</a>. Some, like Polars, are getting more attention but still deserve wider notice. Others, like ConnectorX, are hidden gems.</p>



<h2 class="wp-block-heading">ConnectorX</h2>



<p class="wp-block-paragraph">Most data sits in a database somewhere, but computation typically happens outside of it. Getting data to and from the database for actual work can be a slowdown. <a href="https://github.com/sfu-db/connector-x">ConnectorX</a> loads data from databases into many common data-wrangling tools in Python, and it keeps things fast by minimizing the work required. Most of the data loading can be done in just a couple of lines of Python code and <a href="https://www.infoworld.com/article/2255395/what-is-sql-the-lingua-franca-of-data-analysis.html">an SQL query</a>.</p>



<p class="wp-block-paragraph">Like Polars (which I’ll discuss shortly), ConnectorX uses a <a href="https://www.infoworld.com/article/2258463/rust-tutorial-get-started-with-the-rust-language.html">Rust</a> library at its core. This allows for optimizations like being able to load from a data source in parallel with partitioning. Data in <a href="https://www.infoworld.com/article/3489168/postgresql-tutorial-get-started-with-postgresql-16.html">PostgreSQL</a>, for instance, can be loaded this way by specifying a partition column.</p>



<p class="wp-block-paragraph">Aside from PostgreSQL, ConnectorX also supports reading from MySQL/MariaDB, SQLite, Amazon Redshift, Microsoft SQL Server and Azure SQL, and Oracle. The results can be funneled into a <a href="https://www.infoworld.com/article/2264264/how-to-use-pandas-for-data-analysis-in-python.html">Pandas</a> or PyArrow DataFrame, or into Modin or Dask (via Pandas), or Polars (via PyArrow). General support for reading from ODBC is a work in progress.</p>



<h2 class="wp-block-heading">DuckDB</h2>



<p class="wp-block-paragraph">Data science folks who use Python ought to be aware of <a href="https://www.infoworld.com/article/2337363/why-you-should-use-sqlite-3.html">SQLite</a>—a small, but powerful and speedy relational database packaged with Python. Since it runs as an in-process library, rather than a separate application, SQLite is lightweight and responsive.</p>



<p class="wp-block-paragraph"><a href="https://duckdb.org/">DuckDB</a> is a little like someone answered the question, “<a href="https://www.infoworld.com/article/2336981/duckdb-the-tiny-but-powerful-analytics-database.html">What if we made SQLite for OLAP?</a>” Like other <a href="https://www.infoworld.com/article/2334471/what-is-olap-analytical-databases.html">OLAP</a> database engines, it uses a columnar datastore and is optimized for long-running analytical query workloads. But DuckDB gives you all the things you expect from a conventional database, like ACID transactions. And there’s no separate software suite to configure; you can get it running in a Python environment with a single <code>pip install duckdb</code> command.</p>



<p class="wp-block-paragraph">DuckDB can directly ingest data in CSV, <a href="https://www.infoworld.com/article/2255837/what-is-json-a-better-format-for-data-exchange.html">JSON</a>, or <a href="https://www.infoworld.com/article/2336762/exploring-the-apache-ecosystem-for-data-analysis.html">Parquet</a> format, as well as <a href="https://duckdb.org/docs/stable/data/data_sources">a slew of other common data sources</a>. The resulting databases can also be partitioned into multiple physical files for efficiency, based on keys (e.g., by year and month). Querying works like any other <a href="https://www.infoworld.com/article/2255395/what-is-sql-the-lingua-franca-of-data-analysis.html">SQL</a>-powered relational database, but with additional built-in features like the ability to take random samples of data or construct window functions.</p>



<p class="wp-block-paragraph">DuckDB also has a small but useful collection of extensions, including full-text search, <a href="https://duckdb.org/docs/stable/core_extensions/vss">accelerated vector similarity search</a>, Excel import/export, direct connections to SQLite and PostgreSQL, Parquet file export, and support for many common geospatial data formats and types.</p>



<h2 class="wp-block-heading">Optimus</h2>



<p class="wp-block-paragraph">One of the least enviable jobs you can be stuck with is cleaning and preparing data for use in a DataFrame-centric project. <a href="https://github.com/hi-primus/optimus">Optimus</a> is an all-in-one tool set for loading, exploring, cleansing, and writing data back out to a variety of data sources.</p>



<p class="wp-block-paragraph">Optimus can use <a href="https://www.infoworld.com/article/2264264/how-to-use-pandas-for-data-analysis-in-python.html">Pandas</a>, Dask, CUDF (and Dask + CUDF), Vaex, or <a href="https://www.infoworld.com/article/2259224/what-is-apache-spark-the-big-data-platform-that-crushed-hadoop.html">Spark</a> as its underlying data engine. Data can be loaded in from and saved back out to Arrow, Parquet, Excel, a variety of common database sources, or flat-file formats like CSV and JSON.</p>



<p class="wp-block-paragraph">The data manipulation API resembles Pandas, but adds <code>.rows()</code> and <code>.cols()</code> accessors to make it easy to do things like sort a DataFrame, filter by column values, alter data according to criteria, or narrow the range of operations based on some criteria. Optimus also comes bundled with processors for handling common real-world data types like email addresses and URLs.</p>



<p class="wp-block-paragraph">One possible issue with Optimus is that it’s still under active development but its last official release was in 2020. This means it might not be as current as other components in your stack.</p>



<h2 class="wp-block-heading">Polars</h2>



<p class="wp-block-paragraph">If you spend much time working with DataFrames and you’re frustrated by the performance limits of <a href="https://www.infoworld.com/article/2264264/how-to-use-pandas-for-data-analysis-in-python.html">Pandas</a>, reach for <a href="https://github.com/pola-rs/polars">Polars</a>. This DataFrame library for Python offers a convenient syntax similar to Pandas.</p>



<p class="wp-block-paragraph">Unlike Pandas, though, Polars uses a library written in <a href="https://www.infoworld.com/article/2255250/what-is-rust-safe-fast-and-easy-software-development.html">Rust</a> that takes maximum advantage of your hardware out of the box. You don’t need to use special syntax to take advantage of performance-enhancing features like parallel processing or SIMD; it’s all automatic. Even simple operations like reading from a CSV file are faster. Rust developers can <a href="https://github.com/pola-rs/pyo3-polars">craft their own Polars extensions using pyo3</a>.</p>



<p class="wp-block-paragraph">Polars provides eager and lazy execution modes, so queries can be executed immediately or deferred until needed. It also provides a streaming API for processing queries incrementally. Streaming isn’t available yet for many functions, although Polars can always fall back to the in-memory engine for such operations if need be. You can also <a href="https://docs.pola.rs/api/python/stable/reference/lazyframe/api/polars.LazyFrame.show_graph.html">plot execution graphs for queries</a>, streaming or otherwise, if you want to get an idea of what memory or CPU consumption is like for the query (via the external Graphviz library).</p>



<h2 class="wp-block-heading">DVC</h2>



<p class="wp-block-paragraph">A major and pervasive issue with data science experiments is <a href="https://www.infoworld.com/article/2260350/version-control-track-the-who-what-and-when-of-software-changes.html">version control</a>—not of the project’s code, but its data. <a href="https://github.com/iterative/dvc">DVC</a>, short for Data Version Control, lets you attach version descriptors to datasets, check them into Git as you would the rest of your code, and keep versions of data and code consistent together.</p>



<p class="wp-block-paragraph">DVC can track most any kind of dataset as long as they can be expressed as a file, whether kept in local storage or in a <a href="https://dvc.org/doc/user-guide/data-management/remote-storage#supported-storage-types">remote storage service</a> like an Amazon S3 bucket. You can describe how data models are managed and used by way of a “<a href="https://dvc.org/doc/user-guide/data-management/remote-storage#supported-storage-types">pipeline</a>,” which DVC’s documentation describes as being like “a Makefile system for machine learning projects.”</p>



<p class="wp-block-paragraph">The use cases for DVC are intended to be more than just allowing data to be versioned alongside code. It also works as a fast data cache for remotely hosted data, a methodology for tracking experiments conducted with data, and a registry or catalog for <a href="https://www.infoworld.com/article/2254843/what-is-machine-learning-intelligence-derived-from-data.html">machine learning models</a> created with the data. <a href="https://www.infoworld.com/article/2254808/get-started-with-visual-studio-code.html">Visual Studio Code</a> users can integrate DVC workflows into the editor by way of the <a href="https://marketplace.visualstudio.com/items?itemName=Iterative.dvc">DVC VS Code extension</a>.</p>



<h2 class="wp-block-heading">Cleanlab</h2>



<p class="wp-block-paragraph">Good machine learning datasets are hard to come by, because it’s expensive and time-consuming to create clean, properly labeled data. Sometimes, though, you have no choice but to use data that’s raw and inconsistent. <a href="https://github.com/cleanlab/cleanlab">Cleanlab</a> (as in, “cleans labels”) was made for this scenario.</p>



<p class="wp-block-paragraph">Cleanlab uses existing, high-quality machine learning datasets to analyze lower-quality, unlabeled (or poorly labeled) datasets. You create a model based on the original dataset, use Cleanlab to figure out what needs to be improved in the original dataset, then re-train using your automatically cleaned and adjusted dataset to see the difference.</p>



<p class="wp-block-paragraph">Cleanlab is data-model and data-framework agnostic, a powerful aspect of its design. It doesn’t matter if you’re running <a href="https://www.infoworld.com/article/2335194/what-is-pytorch-python-machine-learning-on-gpus.html">PyTorch</a>, OpenAI, scikit-learn, or <a href="https://www.infoworld.com/article/2255099/what-is-tensorflow-the-machine-learning-library-explained.html">Tensorflow</a>; Cleanlab can work with any classifier. It does, however, have specific workflows for common tasks like token classification, multi-labeling, regression, image segmentation and object detection, outlier detection, and so on. It’s worth perusing the <a href="https://github.com/cleanlab/examples">example set</a> to see for yourself how the process works and what results you can expect.</p>



<h2 class="wp-block-heading">Snakemake</h2>



<p class="wp-block-paragraph">Data science workflows are hard to set up, and that’s even harder to do in a consistent, predictable way. <a href="https://github.com/snakemake/snakemake">Snakemake</a> was created to automate the process, setting up data analysis workflows in ways that ensure everyone gets the same results. Many existing data science projects rely on Snakemake. The more moving parts you have in your data science workflow, the more likely you’ll benefit from automating that workflow with Snakemake.</p>



<p class="wp-block-paragraph">Snakemake workflows resemble GNU Make workflows—you define the steps of the workflow with rules, which specify what they take in, what they put out, and what commands to execute to accomplish that. Workflow rules can be multithreaded (assuming that gives them any benefit), and configuration data can be piped in from <a href="https://www.infoworld.com/article/2255837/what-is-json-a-better-format-for-data-exchange.html">JSON</a> or <a href="https://www.infoworld.com/article/2336307/7-yaml-gotchas-to-avoidand-how-to-avoid-them.html">YAML</a> files. You can also define functions in your workflows to transform data used in rules, and write the actions taken at each step to logs.</p>



<p class="wp-block-paragraph">Snakemake jobs are designed to be portable—they can be deployed on any <a href="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html">Kubernetes-managed environment</a>, or in specific cloud environments like Google Cloud Life Sciences or Tibanna on AWS. Workflows can be “frozen” to use a specific set of packages, and successfully executed workflows can have unit tests automatically generated and stored with them. And for long-term archiving, you can store the workflow as a tarball.</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CIOs must rethink operating models to unlock AI at scale]]></title>
<description><![CDATA[Almost every company has a board or executive AI mandate. Vendors are rolling out agentic AI platforms. The pressure to move is intense.



But the reality on the ground looks different. Eighty-three percent of organizations say data quality is their top AI challenge, and 74% struggle to demonstr...]]></description>
<link>https://tsecurity.de/de/3664901/it-nachrichten/cios-must-rethink-operating-models-to-unlock-ai-at-scale/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664901/it-nachrichten/cios-must-rethink-operating-models-to-unlock-ai-at-scale/</guid>
<pubDate>Mon, 13 Jul 2026 12:17:14 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Almost every company has a <a href="https://www.cio.com/article/4171959/ceos-top-priorities-for-it-leaders-today-2.html">board or executive AI mandate</a>. Vendors are rolling out agentic AI platforms. The pressure to move is intense.</p>



<p>But the reality on the ground looks different. Eighty-three percent of organizations say <a href="https://www.cio.com/article/4162306/data-debt-ai-value-killer.html">data quality is their top AI challenge</a>, and 74% struggle to demonstrate ROI, according to Lopez Research. And only 21% report having a mature <a href="https://www.csoonline.com/article/4176485/the-ai-governance-imperative-you-cant-afford-to-ignore-2.html">governance model for AI agents</a>, per Deloitte’s <a href="https://www.deloitte.com/us/en/about/press-room/state-of-ai-report-2026.html" rel="nofollow">2026 State of Enterprise AI</a> report.</p>



<p>“Agentic AI is real, and vendors’ offerings are very real, too,” says <a href="https://www.forrester.com/analyst-bio/boris-evelson/BIO1737" rel="nofollow">Boris Evelson</a>, vice president and principal analyst at Forrester. “However, most enterprises are still not ready to adopt at scale.”</p>



<p><a href="https://www.westmonroe.com/our-team/david-hilborn" rel="nofollow">Dave Hilborn</a>, who leads West Monroe’s Organization, People &amp; Change practice, frames it as a race with three arrows moving forward — one representing AI and tech evolution, one representing organizations and people, and one representing data. “The AI arrow is far out ahead,” he says. “That delta is the readiness gap.”</p>



<p>The gap <a href="https://www.cio.com/article/4192383/its-not-the-it-holding-ai-back-its-the-business-processes.html">isn’t the technology</a>. It’s the foundational work most organizations haven’t done: data readiness, operating models, governance, skills, and culture. The companies making progress aren’t waiting for vendors to solve these problems. They’re tackling the unglamorous work themselves.</p>



<h2 class="wp-block-heading">AI doesn’t tolerate ambiguity</h2>



<p>AI readiness can be framed across six levels — from data foundation at the base to <a href="https://www.cio.com/article/4157466/cios-reimagine-business-processes-to-reap-ai-benefits.html">reinvented business experiences</a> at the top, says <a href="https://www.linkedin.com/in/afsheantalasaz/" rel="nofollow">Afshean Talasaz</a>, former CIO at Colonial Pipeline and now an executive advisor. One of the key areas that doesn’t always get the attention it needs is the operating model.<strong></strong></p>



<p>“The technology playbooks of the past don’t work in the AI world,” Talasaz says. “Those areas were able to tolerate more ambiguity between business and tech teams. AI doesn’t tolerate the same level of ambiguity. It needs clarity.”</p>



<p>That demands a different kind of partnership between IT and the business. AI systems learn from data — records and measurements of what’s actually happening in the business — and then operate within business processes. Unlike traditional software, which is built based on user requirements, AI is sandwiched between the business that produces the data and the business that consumes the outputs.</p>



<p>“AI is requiring IT and business teams to work more closely together, to be clearer about what AI will and will not do — that really close partnership is crucial,” Talasaz says. “It’s not something that will always naturally evolve. It requires a lot of intentionality about how teams need to work together to deliver outcomes.”</p>



<p>The <a href="https://www.cio.com/article/3801027/10-ai-strategy-questions-every-cio-must-answer.html">AI questions CIOs must answer</a> aren’t just technical. Do we have the right operating model? Have we balanced governance and standard operating procedures within the model? Have we organized teams appropriately? All this must be designed within the context of what the business actually needs.</p>



<p>Too many organizations are <a href="https://www.cio.com/article/4159287/most-companies-are-stuck-on-ai-chat.html">bolting AI onto existing processes</a> without redefining roles or workflows, Forrester’s Evelson. “Organizations can either incrementally enhance existing workflows by augmenting capabilities with AI or pursue a more transformative approach by redesigning the process end-to-end.”</p>



<p>The companies getting value are doing the latter.</p>



<h2 class="wp-block-heading">Data debt comes due</h2>



<p>Data readiness remains the most common barrier to scaling AI. “We’ve never fixed this data quality problem in most organizations,” says <a href="https://www.lopezresearch.com/" rel="nofollow">Maribel Lopez</a>, founder and principal analyst at Lopez Research, “and it comes back to haunt a company in spades as they move to AI.”</p>



<p>At Levi Strauss, the foundational work came first. “If you think about the Levi’s business, it’s quite complex — 100 countries, over 3,000 stores, multiple business models,” says <a href="https://www.levistrauss.com/who-we-are/leadership/jason-gowans/" rel="nofollow">Jason Gowans</a>, the company’s chief digital and technology officer. “You can imagine the complexity of gathering all that data to understand how the business is performing. The idea of this single source of truth — that’s been the biggest thing.”</p>



<p>Levi’s now has more than 1,100 standard operating procedures that govern how work gets done on top of SAP. “That’s fertile material to feed to LLMs on how work gets done,” Gowans says.The results are tangible: partner onboarding that once took three to six months to set up EDI exchanges now takes days.</p>



<p>At contract manufacturing company Jabil, <a href="https://www.linkedin.com/in/chase-christensen-b0447/" rel="nofollow">Chase Christensen</a>, segment CIO, took a similar path. “We had to get everyone to understand where the source data resides, put tech in place so consumption is easier, and drive ownership around data and decision rights — so 140,000 employees don’t feel empowered to create their own data sources that fall out of line.”</p>



<p>The data challenge goes beyond quality, Evelson notes. <a href="https://www.cio.com/article/4104444/8-tips-for-rebuilding-an-ai-ready-data-strategy.html">Most organizations’ data isn’t AI-ready</a>; it hasn’t been prepared for how AI systems consume and learn from information. “Data is siloed, poorly governed, and hard to discover, integrate, and trust,” he says.</p>



<p>Forrester research shows that 45% of data and analytics decision-makers were adopting vector databases in 2025, and 53% were adopting graph databases — investments that signal recognition of how much data architecture needs to evolve. The firm recommends a balanced approach: roughly 48% of AI spending on foundations such as data management and engineering, and 52% on consumption, including analytics, governance, and applications.</p>



<p>But even as organizations work to prepare existing data, AI is creating new challenges. Users leveraging AI tools are generating new forms of data and information that never make it into corporate databases, West Monroe’s Hilborn notes.</p>



<p>“There are explosions of new data, content, and insights being created on the periphery of these data lakes,” he says. “The challenge is how do you capture that and leverage it.”</p>



<h2 class="wp-block-heading">Who’s sponsoring this?</h2>



<p>Even when data is in order, many AI initiatives stall due to how they’re sponsored and funded.</p>



<p>“Enterprise data, analytics, and AI programs succeed when business CxOs sponsor them because they are accountable for business outcomes, not just technology delivery,” Forrester’s Evelson says. “IT-led initiatives often become siloed or tool-centric, whereas business sponsorship ensures alignment to enterprise strategy, prioritization of end-to-end use cases, and a focus on decisions and actions rather than insights alone.”</p>



<p>Too often, AI is still treated as a series of disconnected use cases rather than a sustained, multi-year investment. Evelson calls this the “use case trap” — organizations overindex on individual projects and miss the enterprise-wide compounding impact. That leads to fragmented priorities, inconsistent adoption, and difficulty demonstrating ROI.</p>



<p>Leadership readiness is a distinct layer of AI preparedness, Talasaz says. “Are leaders prepared to provide a vision of reinvented business experiences that become the north star?” he asks. “Leadership teams, at various levels of the organization, need to articulate what a reinvented business looks like so teams have the direction and support to build differentiating capabilities.”</p>



<p>Levi’s offers a counterexample. AI is a CEO priority there. At the last quarterly offsite, the execs were building agents. “When you’re committed to upskilling the workforce, you’re better served to answer how to rewire processes with AI at the core,” Gowans says. “It starts at the top. It has to be an exec priority.”</p>



<h2 class="wp-block-heading">Fear, literacy, and two types of AI</h2>



<p>Technical talent is only part of the equation. Organizations also need to <a href="https://www.cio.com/article/4016354/cios-tackle-the-ai-change-management-challenge.html">address change management</a>.</p>



<p>“We saw it with the AI boom — fear about jobs, not knowing what AI did,” says Jabil’s Christensen. “The key is demystifying AI. We doubled down and focused on AI literacy. We want everyone to understand how it was put together, and that removed a lot of that fear. That’s been the biggest hurdle.”</p>



<p>Different types of AI require different skills and governance, Talasaz says. “General use focuses on productivity on the desktop,” he says. “Integrated AI — industrial-capable AI embedded within core business processes — requires different skills, capabilities, and governance.”</p>



<p>For desktop AI, training and guardrails help employees be successful — what Talasaz calls “bumpers,” like in bowling. Organizations need to <a href="https://www.cio.com/article/4117091/how-ai-upskilling-fails-and-what-it-leaders-are-doing-to-get-it-right.html">help employees through reskilling and guidance</a>. “You have tools in a toolbox,” he says. “It’s important to know when to use a power tool versus when you need a screwdriver.”</p>



<p>But for integrated AI embedded in core processes, the stakes are higher. “Business leaders responsible for business outcomes based on AI-driven processes need to be fully aware of both the benefits and risks that come along with using these tools,” Talasaz says.</p>



<p>That distinction matters for governance, too. Lower-, medium-, and high-risk AI use cases may require <a href="https://www.csoonline.com/article/4188573/rethinking-the-balance-between-ai-oversight-and-innovation.html">different ways of working and different risk management approaches</a>. “Deploying AI in potentially high-risk or high-cost areas of the business requires a higher level of rigor,” Talasaz says. “That’s different than building something that helps write my emails.”</p>



<h2 class="wp-block-heading">From POC to production</h2>



<p>Perhaps the biggest readiness gap is the transition <a href="https://www.cio.com/article/3850763/88-of-ai-pilots-fail-to-reach-production-but-thats-not-all-on-it.html">from proof of concept to production</a>. “It requires such a different approach,” Talasaz says. “A successful proof of concept can create a lot of excitement, but when teams are unprepared to build and scale, it can create the potential to over-promise and under-deliver.”</p>



<p>The operating model that works for experimentation doesn’t work for production at scale. Proofs of concept are designed to demonstrate the efficacy of ideas and the underlying technology. But building, scaling, and sustaining technology in the business requires operating models, standards, roles, and skills that many organizations haven’t developed. Intentionally designed operating models reduce the cost of learning, improve execution, and increase delivery velocity, says Talasaz.</p>



<p>But there’s no one-size-fits-all answer. “A business that needs to build capabilities in a marketplace moving very fast requires one kind of operating model,” Talasaz says. “A business that can take longer to develop business capabilities and adapt to market changes can choose a different operating model. It’s important to design ways of working tailored to what the business needs and the speed at which the business needs to leverage technology to be successful.”</p>



<p>Jabil is navigating this journey as part of its move to SAP’s cloud ERP through RISE, scaling from $29 billion to $34 billion in revenue while keeping selling, general, and administrative (SG&amp;A) expenses relatively flat — in part by layering generative AI onto predictive analytics capabilities built over years.</p>



<p>“We started years ago with computer vision to drive product quality,” Christensen says. “As gen AI blew up, we took the predictive analytics we had <a href="https://www.cio.com/article/193580/upskilling-transforms-jabil-employees-into-data-scientists.html">built over the years</a> and imbued them with gen AI. We’ve implemented the basics, and now we’re looking for complex scenarios.”</p>



<h2 class="wp-block-heading">Governance built in, not bolted on</h2>



<p>Governance is often treated as a policy document or committee. It should be embedded in the operating model itself, Talasaz argues.</p>



<p>“The operating model doesn’t always get the attention it needs,” he says. “Policies and committees are useful, but they should handle larger enterprise risks. Most of the governance should be embedded in the operating model to ensure you’re getting outcomes you want.”</p>



<p>That might mean peer review built into the development process, bias checks before deployment, or clear escalation paths for high-risk use cases. When governance is separate from the operating model, it tends to slow things down. When it’s integrated, it becomes how work naturally gets done, says Talasaz.</p>



<p>Governance at the agent level matters, too, Levi’s Gowans says. “Know what agents have been deployed, who authored them, and who’s responsible,” he says, noting that the company has established a registry to understand what agents it has operating within its networks.</p>



<p>The challenges of AI governance are unique, Lopez of Lopez Research says. “Very few people have the governance stack required to say they did the right things with AI,” she says. “<a href="https://www.csoonline.com/article/2132294/what-are-non-human-identities-and-why-do-they-matter.html">Non-human identity</a> and access control is totally different and, frankly, evolving so quickly that no one knows what to do.”</p>



<p>The challenge is ultimately a trade-off, Forrester’s Evelson says. “Push agentic AI capabilities too far, and you risk creating a governance and compliance nightmare,” he says. “Tighten controls too aggressively, and you stifle innovation. Best practices for <a href="https://www.cio.com/article/4188566/cios-rethink-the-balance-between-ai-oversight-and-innovation.html">striking the right balance</a> are still being discovered.”</p>



<h2 class="wp-block-heading">It takes a team</h2>



<p>The AI readiness gap isn’t about technology — it’s about the work organizations have been deferring for years. Data quality. Operating models. Executive sponsorship. Skills and culture. Governance embedded in process.</p>



<p>“Once you progress from everyone using Copilot to putting agents in production, then you realize the need for business context,” Gowans of Levi Strauss says.</p>



<p>It’s a shared journey requiring all teams to understand what’s required, Talasaz says. “It involves helping people understand what it takes from all sides — the technology itself, the operating model, the skills and talents needed — but also working with business leaders on the art of the possible,” he says. “Helping them understand both the benefits and the responsibility of deploying this tech.”</p>



<p>A colleague of his calls AI “the ultimate executive team sport.”</p>



<p>“It requires people to do it well and manage it,” Talasaz says.</p>



<p></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-15531 | yashbhalgat HashNeRF-pytorch up to 82885e698295982504eb6a26d060a6b2473e3706 Checkpoint File run_nerf.py torch.load ckpt_path deserialization (Issue 49 / EUVD-2026-43277)]]></title>
<description><![CDATA[A vulnerability classified as problematic has been found in yashbhalgat HashNeRF-pytorch up to 82885e698295982504eb6a26d060a6b2473e3706. Affected by this issue is the function torch.load of the file run_nerf.py of the component Checkpoint File Handler. The manipulation of the argument ckpt_path l...]]></description>
<link>https://tsecurity.de/de/3664512/sicherheitsluecken/cve-2026-15531-yashbhalgat-hashnerf-pytorch-up-to-82885e698295982504eb6a26d060a6b2473e3706-checkpoint-file-runnerfpy-torchload-ckptpath-deserialization-issue-49-euvd-2026-43277/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3664512/sicherheitsluecken/cve-2026-15531-yashbhalgat-hashnerf-pytorch-up-to-82885e698295982504eb6a26d060a6b2473e3706-checkpoint-file-runnerfpy-torchload-ckptpath-deserialization-issue-49-euvd-2026-43277/</guid>
<pubDate>Mon, 13 Jul 2026 09:23:12 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability classified as <a href="https://vuldb.com/kb/risk">problematic</a> has been found in <a href="https://vuldb.com/product/yashbhalgat:hashnerf-pytorch">yashbhalgat HashNeRF-pytorch up to 82885e698295982504eb6a26d060a6b2473e3706</a>. Affected by this issue is the function <code>torch.load</code> of the file <em>run_nerf.py</em> of the component <em>Checkpoint File Handler</em>. The manipulation of the argument <em>ckpt_path</em> leads to deserialization.

This vulnerability is listed as <a href="https://vuldb.com/cve/CVE-2026-15531">CVE-2026-15531</a>. The attack must be carried out locally. In addition, an exploit is available.

This product uses a rolling release model to deliver continuous updates. As a result, specific version information for affected or updated releases is not available.

The pull request to fix this issue awaits acceptance.]]></content:encoded>
</item>
<item>
<title><![CDATA[A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention]]></title>
<description><![CDATA[In this tutorial, we explore NVIDIA tile-based GPU programming with TileGym, building a Colab workflow that runs across different hardware. We probe the CUDA environment, try the real cuTile backend, and fall back to Triton when standard Colab GPUs lack the cuTile stack. We learn the core tile id...]]></description>
<link>https://tsecurity.de/de/3662541/ai-nachrichten/a-coding-guide-to-nvidias-tile-based-gpu-programming-from-cutile-and-triton-kernels-to-flash-attention/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3662541/ai-nachrichten/a-coding-guide-to-nvidias-tile-based-gpu-programming-from-cutile-and-triton-kernels-to-flash-attention/</guid>
<pubDate>Sun, 12 Jul 2026 02:05:45 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>In this tutorial, we explore NVIDIA tile-based GPU programming with TileGym, building a Colab workflow that runs across different hardware. We probe the CUDA environment, try the real cuTile backend, and fall back to Triton when standard Colab GPUs lack the cuTile stack. We learn the core tile idea: operate on whole data tiles instead of single threads, then load, compute, and store them. We implement vector addition, fused GELU, row-wise softmax, tiled matrix multiplication, and flash attention, checking each against PyTorch.</p>
<p>The post <a href="https://www.marktechpost.com/2026/07/11/a-coding-guide-to-nvidias-tile-based-gpu-programming-from-cutile-and-triton-kernels-to-flash-attention/">A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Developer-Häppchen fürs Wochenende – kleinere News der Woche]]></title>
<description><![CDATA[Kleine, aber interessante Meldungshäppchen vom News-Buffet zu Rustup, Python Software Foundation, Vite+, Kubermatic, Django, PyTorch, Kai, Eclipse und Rust.]]></description>
<link>https://tsecurity.de/de/3661427/it-nachrichten/developer-haeppchen-fuers-wochenende-kleinere-news-der-woche/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3661427/it-nachrichten/developer-haeppchen-fuers-wochenende-kleinere-news-der-woche/</guid>
<pubDate>Sat, 11 Jul 2026 09:17:37 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[Kleine, aber interessante Meldungshäppchen vom News-Buffet zu Rustup, Python Software Foundation, Vite+, Kubermatic, Django, PyTorch, Kai, Eclipse und Rust.]]></content:encoded>
</item>
<item>
<title><![CDATA[Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them]]></title>
<description><![CDATA[Enterprise AI teams are giving agents more freedom at the same moment their confidence in automated testing is collapsing.Half of enterprises have deployed an AI agent or LLM feature that passed internal evaluations and yet still caused a customer-facing failure — one in four more than once — acc...]]></description>
<link>https://tsecurity.de/de/3660672/it-nachrichten/enterprise-ai-is-entering-an-evaluation-gap-agents-are-gaining-autonomy-faster-than-companies-can-verify-them/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3660672/it-nachrichten/enterprise-ai-is-entering-an-evaluation-gap-agents-are-gaining-autonomy-faster-than-companies-can-verify-them/</guid>
<pubDate>Fri, 10 Jul 2026 21:18:05 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Enterprise AI teams are giving agents more freedom at the same moment their confidence in automated testing is collapsing.</p><p>Half of enterprises have deployed an AI agent or LLM feature that passed internal evaluations and yet still caused a customer-facing failure — one in four more than once — according to the June 2026 VB Pulse survey of 157 qualified enterprise respondents at companies with 100 or more employees.</p><p>The sample is self-selected rather than a probability sample, so the findings should be read as directional, not precise.</p><p>But enterprises are not responding by slowing automation:<b> 66% of respondents already permit some production deployment without human review </b>or are building systems intended to do so within the next 12 months. Only 5% say they fully trust the automated evaluations that would make those release decisions.</p><p>That mismatch is the evaluation gap: the autonomy ceiling is rising faster than the assurance beneath it. </p><p>It also fits a broader thesis that will be explored at <a href="https://venturebeat.com/vbtransform2026">VB Transform 2026</a>: enterprises ship agents first, while the control layers around identity, evaluation, cost, context and orchestration are arriving later. The next year will be a retrofit cycle, with buyers shifting budget toward the systems that make agentic deployments governable and dependable.</p><h2>Why a passing evaluation is not a working agent</h2><p>Traditional software testing usually asks whether a defined input produces an expected output. Agent testing is harder because the system may choose its own sequence of steps, call tools, retrieve data, alter state and respond differently from one run to the next.</p><p>An agent can make several individually plausible decisions and still reach the wrong result. It may retrieve the correct account but update the wrong field. It may draft a valid refund request but send it without approval. It may call five tools successfully before a sixth step leaks sensitive information or leaves a workflow incomplete.</p><p>The survey shows enterprises already recognize this limitation. <b>The most common reason for distrusting automated evaluation is poor alignment with real-world outcomes, cited by 29% of respondents.</b> Bias or inconsistency follows at 21%, lack of explainability at 18%, and data leakage or privacy concerns at 17%.</p><p>That hierarchy matters. Enterprises are saying the score often does not predict what happens when a customer, employee or business process encounters the agent in production — not that automated scoring is too slow or expensive.</p><p><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf">NIST makes a similar point in its Generative AI Profile</a>: measurements gathered in controlled environments may not transfer cleanly to deployment because behavior changes with prompts, users, context and operating conditions. Its guidance calls for field testing, post-deployment monitoring and clear processes for escalating failures.</p><div></div><h2>Capability is not consistency</h2><p>A single successful run proves that an agent can complete a task. It does not prove that it will complete the task reliably.</p><p><a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents">Anthropic’s guidance on agent evaluation</a> distinguishes between measuring whether a system succeeds at least once across repeated attempts and whether it succeeds every time. That distinction is essential for customer-facing or operational workflows. A model that occasionally produces an excellent answer may still be unacceptable if the same task fails unpredictably on the next attempt.</p><p>Enterprise teams should therefore treat repeatability as a first-class metric. That means running the same scenario multiple times, varying phrasing and context, testing tool failures, and measuring whether the final business outcome remains correct even when the route changes.</p><p>The evaluation set also has to evolve. Every production incident should become a permanent regression test. Customer escalations, failed tool calls, incorrect approvals and data-handling mistakes should feed back into the pre-deployment suite rather than remaining isolated support cases.</p><h2>Autonomy should expand by risk, not by ambition</h2><p>The survey does not imply that every agent action should require a person. Human review cannot scale across millions of low-consequence decisions.</p><p>But zero-human operation should be earned by demonstrated reliability and bounded by the consequences of failure.</p><p>Low-risk actions such as drafting internal summaries or categorizing documents can tolerate broader autonomy. Financial transactions, customer communications, code deployment, access-control changes and data deletion need stricter thresholds, repeated consistency tests, policy checks, rollback mechanisms and clear human escalation paths.</p><p>The risk isn't evenly distributed by company size, either. Larger enterprises — those with 2,500 or more employees — are moving toward zero-human deployment fastest, at 70% versus 64% for smaller companies, and they're also shipping more agents that go on to fail a customer, at 54% versus 48%. </p><p>That is the warning for enterprise leaders. Removing the human from the loop does not remove uncertainty. Without stronger assurance, it converts uncertainty into an automated production decision.</p><p>The market will keep pushing toward greater autonomy because the economic incentive is real. The organizations best positioned won't be those that remove people fastest — they'll be the ones that treat repeatability and regression testing as seriously as deployment speed.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Google's TabFM skips per-dataset training and still predicts on tables it's never seen]]></title>
<description><![CDATA[The vast majority of business data is tabular — living in data warehouses, CRMs, and financial ledgers — yet building a reliable model from it still means training a new one from scratch for every dataset, then maintaining hyperparameter tuning loops, feature engineering, and retraining pipelines...]]></description>
<link>https://tsecurity.de/de/3660555/it-nachrichten/googles-tabfm-skips-per-dataset-training-and-still-predicts-on-tables-its-never-seen/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3660555/it-nachrichten/googles-tabfm-skips-per-dataset-training-and-still-predicts-on-tables-its-never-seen/</guid>
<pubDate>Fri, 10 Jul 2026 20:03:33 +0200</pubDate>
<category>📰 IT Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The vast majority of business data is tabular — living in data warehouses, CRMs, and financial ledgers — yet building a reliable model from it still means training a new one from scratch for every dataset, then maintaining hyperparameter tuning loops, feature engineering, and retraining pipelines to fight data drift. Google Research is proposing a way around that: <a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/">a new foundation model called TabFM</a> that treats tabular prediction as an in-context learning problem instead.</p><p>It can generate predictions for a new, unseen table in a single forward pass. For enterprise developers and AI engineers, this reduces the time-to-production from weeks of pipeline engineering to a single API call.</p><h2>The challenge with traditional ML</h2><p>To extract reliable predictions from a gradient-boosted tree, data scientists must build and maintain complex data pipelines. They have to clean messy inputs, impute missing values, encode categorical variables into numerical formats, and engineer custom feature crosses.</p><p>Once the data is ready, they must run repetitive hyperparameter optimization loops, searching across learning rates, tree depths, subsampling ratios, and regularization grids to find the best configuration. </p><p>Once deployed, these traditional models "incur ongoing operational debt through data drift monitoring and retraining pipelines to stay accurate," Weihao Kong, Research Scientist at Google Research, told VentureBeat.</p><p>Meanwhile, the rest of the AI industry has moved on. Generative AI models for text and computer vision have seamlessly shifted to zero-shot inference, where a model can perform a completely new task simply by being prompted with context. </p><p>Large language models (LLMs) already excel at <a href="https://venturebeat.com/business/fine-tuning-vs-in-context-learning-new-research-guides-better-llm-customization-for-real-world-tasks">in-context learning</a>, so why can't we just feed tables into an off-the-shelf LLM?</p><p>Because LLMs are trained on natural language rather than structured data, they struggle to process tables directly. First, their context limits are exhausted quickly by medium-sized tables containing just a few thousand rows and hundreds of columns. Second, LLMs suffer from tokenization inefficiency, awkwardly splitting numerical values and destroying mathematical precision. Finally, they suffer from structural blindness. When a 2D table is serialized as a 1D text string, LLMs lose track of which value belongs to which row and column as the table grows. </p><p>"That's why, today, it is far more effective to use an LLM to write the code that handles feature engineering and calls XGBoost than to ask the LLM to read the table itself," Kong said.</p><h2>What is TabFM?</h2><p>To run inference with TabFM, you do not update any model weights. Instead, you take your historical examples (the training rows with their known labels) and your target rows (the new data you want to predict) and pass them to the model as a single, unified prompt. The model learns to interpret the relationships between columns and rows directly from this context at runtime.</p><p>For example, consider an enterprise analyst trying to predict customer churn. Instead of building a bespoke data pipeline and training an XGBoost model, they can simply pass a sample of historical user session data alongside a new, active session into TabFM. In one forward pass, the model returns an instant churn probability. </p><p>TabFM overcomes the limitations of LLMs by treating the data as a grid, preserving its structural integrity without forcing it into a single-dimensional text string.</p><p>To effectively process diverse tabular structures while enabling scalable zero-shot prediction, TabFM synthesizes the strengths of earlier experimental architectures, TabPFN and TabICL. <a href="https://github.com/PriorLabs/tabpfn">TabPFN</a>, developed by Prior Labs, first proved that a transformer architecture could perform zero-shot classification on small tables, though it struggled to scale computationally to larger datasets. </p><p>Later, <a href="https://dl.acm.org/doi/10.5555/3780338.3782366">TabICL</a>, developed by France's National Research Institute for Digital Science and Technology, addressed this bottleneck by introducing row compression, allowing in-context learning to efficiently process much larger tables. </p><p>TabFM combines TabPFN's deep feature contextualization with TabICL's efficient compression into a novel hybrid design built on three key mechanisms:</p><p><b>1. Alternating row and column attention:</b> The raw table is first processed through a multilayer attention module that alternates across both columns (features) and rows (examples). By continuously attending across these two dimensions, the model natively captures complex feature interactions. This deep contextualization does the heavy lifting that would usually require tedious manual feature crafting by data scientists.</p><p><b>2. Row compression:</b> Following this contextualization, the cross-attended information for each row is compressed into a single, dense vector representation. TabICL pioneered this by using CLS tokens to compress a row's rich information into one vector, "in contrast to TabPFN v2, v2.5, and v2.6, which attend over the full cell grid throughout the network," Kong explained. This drastically shrinks the computational footprint.</p><p><b>3. In-context learning (ICL):</b> A causal Transformer then operates on this sequence of compressed embeddings. This Transformer model uses the attention mechanism of TabICL to attend over these dense row vectors, drastically reducing the computation cost and allowing the model to process large datasets efficiently.</p><p>A major selling point of TabFM is its pretraining recipe. The model was trained entirely on hundreds of millions of synthetic datasets. These datasets were dynamically generated using structural causal models (SCMs) that incorporate a wide variety of random functions. By training exclusively on synthetic SCMs, TabFM learned the fundamental mathematical priors of how tabular features interact without ingesting real-world, confidential CSV files.</p><h2>TabFM in action</h2><p>To test the model's capabilities, Google researchers benchmarked TabFM on TabArena, a comprehensive evaluation suite spanning 51 diverse tabular datasets across 38 classification and 13 regression tasks.</p><p>On these public benchmarks, TabFM's zero-shot predictions already match or beat heavily tuned supervised baselines. However, Google is careful to note that this does not automatically mean TabFM will universally dethrone bespoke, hyper-optimized production models on every enterprise workload.</p><p>"Instead of replacing hyper-optimized production models, the true practical business value it unlocks for lean engineering teams is velocity," Kong said. "It allows data analysts and backend engineers to instantly spin up high-quality baseline models without a dedicated data science team managing a complex lifecycle."</p><p>For advanced practitioners looking to squeeze out maximum accuracy, the research team also introduced a "TabFM-Ensemble" configuration. By running the model through 32 distinct variations and blending the results, TabFM pushes the performance even further. </p><h2>Getting started, trade-offs, and the cloud future</h2><p>The shift to in-context learning for tables introduces a new economic trade-off that engineering teams must consider. </p><p>With traditional algorithms, training is slow and expensive, but inference is lightning-fast and cheap. TabFM flips this dynamic. While training time drops to zero, inference becomes significantly heavier. Because the model must process the entire historical dataset as context during every single prediction, it requires more compute and memory at runtime. </p><p>In this new paradigm, "traditional machine learning training becomes the 'prefill' phase (KV caching) in the context window," Kong said. While this prefill cost is steep, it is paid only once per table, and the cache is reused across subsequent queries. "The catch is prediction latency, which no amount of caching removes," Kong added. Every new prediction requires a pass through a large transformer. "Any production API requiring single-digit-millisecond response times cannot tolerate TabFM's forward-pass overhead."</p><p>For developers looking to evaluate the model today, the barrier to entry is low. Google designed TabFM as a drop-in replacement for traditional ML workflows, offering a scikit-learn compatible API (TabFMClassifier and TabFMRegressor). It natively handles mixed numerical and categorical columns, works directly with pandas DataFrames, and requires no manual ordinal encoders or numerical scalers. The library supports both JAX and PyTorch backends.</p><p>However, enterprise teams need to be aware of current limitations and licensing restrictions. The model architecture has a hard limit of 10 output classes for classification tasks, and it is optimized for tables with up to 500 features. More importantly, while Google released the <a href="https://github.com/google-research/tabfm">underlying codebase</a> under the permissive Apache 2.0 license, the pre-trained model weights are published on <a href="https://huggingface.co/google/tabfm-1.0.0-pytorch">Hugging Face</a> under a strict tabfm-non-commercial-v1.0 license. Developers can evaluate the model internally, but it cannot be deployed in commercial products yet.</p><p>Looking ahead, Google is addressing the commercial deployment friction through its cloud ecosystem. TabFM is being integrated directly into Google BigQuery, allowing analysts to run zero-shot predictions natively via an “AI.PREDICT” command. By putting foundation model inference right next to the data warehouse, TabFM could soon make complex tabular machine learning as accessible as a basic database query.</p><p>In practice, TabFM shines in rapid prototyping, high data drift environments, and small to medium-sized datasets under 100,000 rows. Conversely, teams should stick to traditional models for strict, ultra-low latency APIs, or massive tables exceeding one million rows, which currently require aggressive row sampling that degrades the foundation model's competitive advantage.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[12 Wege, KI kostengünstiger zu trainieren]]></title>
<description><![CDATA[Eine KI zu trainieren, kann schnell monetäre Sorgen bereiten – muss es aber nicht.  ultramansk | shutterstock.com



KI-Pipelines zu optimieren, erfordert mehr als nur oberflächliche Hardwareanpassungen. Es gilt, die Art und Weise, wie Modelle Daten verarbeiten, grundlegend zu verändern. Zwar imp...]]></description>
<link>https://tsecurity.de/de/3658655/it-security-nachrichten/12-wege-ki-kostenguenstiger-zu-trainieren/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3658655/it-security-nachrichten/12-wege-ki-kostenguenstiger-zu-trainieren/</guid>
<pubDate>Fri, 10 Jul 2026 06:07:13 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2026/07/ultramansk_shutterstock_2672055281_16z9.jpg?quality=50&amp;strip=all&amp;w=1024" alt="Dev Team sceptical 16z9" class="wp-image-4192249" width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"><figcaption class="wp-element-caption">Eine KI zu trainieren, kann schnell monetäre Sorgen bereiten – muss es aber nicht.  </figcaption></figure><p class="imageCredit">ultramansk | shutterstock.com</p></div>



<p><a href="https://www.computerwoche.de/article/4183987/embedding-pipelines-sind-das-neue-etl.html" target="_blank">KI-Pipelines</a> zu optimieren, erfordert mehr als nur oberflächliche Hardwareanpassungen. Es gilt, die Art und Weise, wie Modelle Daten verarbeiten, grundlegend zu verändern. Zwar implementieren KI-Engineers oft einfache Effizienzmaßnahmen innerhalb des Training-Loops. Aber um die Trainingskosten permanent <a href="https://www.computerwoche.de/article/4182741/nur-jedes-vierte-unternehmen-hat-seine-ki-kosten-im-blick.html" target="_blank">zu reduzieren</a>, sind architektonische Änderungen nötig – direkt im neuronalen Netz.</p>



<p>Die folgenden zwölf Optimierungsmaßnahmen auf Modellebene verwandeln Ihre <a href="https://www.cio.de/article/4168849/die-ki-strategie-der-commerzbank.html" target="_blank">KI-Strategie</a> von einem Brute-Force-Hardware-Ansatz in eine elegante, softwaredefinierte Disziplin – und senken die Stückkosten Ihrer KI-Pipeline drastisch.</p>



<h2 class="wp-block-heading">1. Pretraining einsparen</h2>



<p>Ein Foundation-Modell von Grund auf neu zu trainieren ist für Standard Enterprise-Applikationen selten nötig und verbietet sich mit Blick auf die dafür nötige Rechenleistung. Statt dafür Millionen zu verschwenden, sollten Engineering-Teams lieber öffentlich verfügbare <a href="https://www.computerwoche.de/article/4146975/wie-ki-open-source-verandert.html" target="_blank">Open-Weight-Modelle</a> nutzen.</p>



<p>Dieser grundlegende Transfer-Learning-Ansatz ist der unverzichtbare erste Schritt, wenn es darum geht, interne Chatbots oder domänenspezifische Klassifikatoren zu entwickeln. Indem bestehende neuronale Architekturen zum Einsatz kommen, lassen sich die enormen <a href="https://www.computerwoche.de/article/2828262/finetuning-ist-teuer-aber-oft-lohnt-es-sich.html" target="_blank">Kosten</a> (auch für Energie), die mit den initialen Pretraining-Phasen verbunden sind, umgehen.</p>



<h2 class="wp-block-heading">2. Parametereffizient feinabstimmen</h2>



<p>Selbst das standardmäßige Feintuning von umfassenden Sprachmodellen erforderte immense Mengen an VRAM, um die States von Optimizern und Gradienten zu speichern. Um dieses Hardware-Bottleneck aufzulösen, sollten Engineers <a href="https://medium.com/@MUmarAmanat/fine-tune-llm-with-peft-60b2798f1e5f" target="_blank" rel="noreferrer noopener">PEFT</a>-Techniken wie <a href="https://www.computerwoche.de/article/3552133/forscher-verbinden-wi-fi-und-lora.html" target="_blank">LoRA</a> implementieren.</p>



<p>Die Technik reduziert den Memory Overhead drastisch, indem sie dafür sorgt, dass 99 Prozent der vortrainierten Weights „eingefroren“ und kleine, trainier- und adaptierbare Layer injiziert werden. Dieser mathematische Shortcut ist ideal geeignet, um hochgradig anpassbare GenAI-Funktionen umzusetzen – und erlaubt die Feinabstimmung von Milliarden von Parametern mit einer einzigen (Consumer-)<a href="https://www.computerwoche.de/article/3967958/was-ist-eine-gpu.html" target="_blank">GPU</a>.</p>



<pre class="wp-block-code"><code>python
from peft import LoraConfig, get_peft_model

config = LoraConfig(r=8, lora_alpha=32, target_modules=["q_proj", "v_proj"])
efficient_model = get_peft_model(base_model, config)</code></pre>



<h2 class="wp-block-heading">3. Warmstart-Layer einziehen</h2>



<p>Falls Sie spezifische Netzwerkkomponenten von Grund auf neu trainieren müssen, stellt der Import vortrainierter Embeddings sicher, dass nur die verbleibenden Layer erhöhten Rechenaufwand verursachen.</p>



<p>Dieser „Warmstart“-Ansatz reduziert den Rechenaufwand in Frühphasen der KI-Entwicklung erheblich, weil das Modell so grundlegende, universelle Datenrepräsentationen nicht erst neu erlernen muss. Besonders empfehlenswert ist dieser für <a href="https://www.computerwoche.de/article/4173136/17-llms-fur-spezialdomanen.html" target="_blank">Spezialdomänen</a>.</p>



<pre class="wp-block-code"><code>python
# PyTorch warm-start example
model.embedding_layer.weight.data.copy_(pretrained_medical_embeddings)
model.embedding_layer.requires_grad = False
</code></pre>



<h2 class="wp-block-heading">4. Gradient Checkpointing anwenden</h2>



<p>Memory-Engpässe sind der wesentliche Grund dafür, dass Entwickler gezwungen sind, teure, VRAM-intensive Cloud-Instanzen zu mieten. <a href="https://arxiv.org/pdf/1604.06174" target="_blank" rel="noreferrer noopener">Gradient Checkpointing</a> (PDF) spart Speicherplatz ein, indem bestimmte Forward Activations während der Backpropagation neu berechnet werden – anstatt sie alle zu speichern.</p>



<p>Entwicklern ist zu empfehlen, diese Technik einzusetzen, wenn sie mit anhaltenden „Out of Memory“-Fehlern konfrontiert sind. Denn Gradient Checkpointing ermöglicht es, zehnmal größere Netzwerke auf derselben GPU unterzubringen – bei einem Mehr an Rechenaufwand von circa 20 Prozent.</p>



<pre class="wp-block-code"><code>python
# Enable in Hugging Face / PyTorch
model.gradient_checkpointing_enable()</code></pre>



<h2 class="wp-block-heading">5. Compiler-Fusion aktivieren</h2>



<p>Moderne Deep-Learning-Frameworks leiden regelmäßig unter Engpässen mit Blick auf die Speicherbandbreite, da ständig Daten über die Hardware gelesen und geschrieben werden. Durch Compiler, die wie <a href="https://openxla.org/?hl=de" target="_blank" rel="noreferrer noopener">XLA</a> oder <a href="https://pytorch.org/get-started/pytorch-2-x/" target="_blank" rel="noreferrer noopener">PyTorch 2.0</a> auf Graph-Ebene operieren, lässt sich eine Vielzahl von Prozessen in einem einzelnen GPU-Kernel fusionieren.</p>



<p>Diese architektonische Optimierung führt dazu, dass Durchsatz und Ausführungsgeschwindigkeit massiv gesteigert werden. Parallel sind allerdings keine manuellen Änderungen am Code notwendig. Um die Hardwareauslastung zu maximieren, ist es Entwickler-Teams zu empfehlen, die Compiler-Fusion standardmäßig bei sämtlichen Trainings-Sessions in der Produktion zu aktivieren.</p>



<pre class="wp-block-code"><code>python
import torch

# PyTorch 2.0 compiler fusion
optimized_model = torch.compile(model)</code></pre>



<h2 class="wp-block-heading">6. Pruning und Quantisierung einsetzen</h2>



<p>Ein umfangreiches, vollpräzises 16-Bit-Neural-Network in der Produktion bereitzustellen, erfordert ebenfalls oft teure Cloud-Instanzen, was die Gewinnmarge einer Applikation zunichtemachen kann. Durch algorithmisches Pruning werden mathematisch redundante Weights entfernt.</p>



<p>Eine Quantisierung des Modells sorgt hingegen dafür, dass die verbleibenden Parameter von 16-Bit-Gleitkommazahlen auf 8-Bit- oder 4-Bit-Ganzzahlen komprimiert werden. Das ermöglicht es, das KI-Modell auf deutlich kostengünstigeren GPUs mit geringerem Speicherbedarf auszuführen – ohne dass die Qualität der Konversationen darunter leidet. Diese physikalische Reduktion ist entscheidend dafür, Traffic-intensive Anwendungen kosteneffizient skalieren zu können. Davon abgesehen senkt es jedoch auch die CO²-Kosten, die ein API-Call <a href="https://www.cio.com/article/4132293/the-carbon-cost-of-an-api-call.html" target="_blank">verursacht</a>, wenn Tausende von Usern parallel bedient werden.</p>



<pre class="wp-block-code"><code>python
import torch
import torch.nn.utils.prune as prune

# 1. Prune 20% of the lowest-magnitude weights in a layer
prune.l1_unstructured(model.fc, name="weight", amount=0.2)

# 2. Dynamic Quantization (Compress Float32 to Int8)
quantized_model = torch.ao.quantization.quantize_dynamic(
    model, {torch.nn.Linear}, dtype=torch.qint8
)</code></pre>



<h2 class="wp-block-heading">7. Curriculum Learning verwenden</h2>



<p>Ein untrainiertes neuronales Netzwerk mit hochkomplexen und gleichzeitig verrauschten Datensätzen zu füttern, zwingt den Optimizer teure Extra-Rechenschleifen zu drehen, um chaotische Gradienten abzubilden. Dieses Problem lässt sich mit <a href="https://medium.com/aiguys/curriculum-learning-83b1b2221f33" target="_blank" rel="noreferrer noopener">Curriculum Learning</a> (auch Lehrplanlernen) lösen: Dabei wird die Daten-Pipeline so strukturiert, dass zunächst klare, leicht klassifizierbare Beispiele eingeführt werden – bevor der schrittweise Übergang auf hochpräzise Anomalien erfolgt.</p>



<p>Geht es etwa darum, ein Vision-Modell für autonomes Fahren zu trainieren, sollten die Entwickler diesem zunächst klare Tageslichtbilder von Autobahnen zuführen, bevor sie Rechenleistung für komplexe Nachtaufnahmen verschneiter Stadtkreuzungen in Städten aufwenden. Dieser schrittweise Ansatz ermöglicht dem Netzwerk, zentrale mathematische Merkmale ressourcenschonend abzubilden. Dadurch wird die Konvergenz deutlich schnell und mit geringerem Hardware-Aufwand erreicht.</p>



<h2 class="wp-block-heading">8. Wissen destillieren</h2>



<p>Ein massives KI-Modells mit 70 Milliarden Parametern für simple, repetitive Tasks zu nutzen, kommt einer gravierenden Fehlallokation von Rechenressourcen gleich. Dieses Problem lässt sich mithilfe von Wissensdestillation lösen: Dabei lernt ein hocheffizientes, schlankes „Student“-Modell, die Reasoning-Ketten eines großen „Teacher“-Modells exakt nachzuahmen.</p>



<p>Stellen Sie sich ein E-Commerce-Unternehmen vor, das Produktempfehlungen in Echtzeit direkt auf dem Smartphone eines Nutzers ausführen muss, wo Akku und Speicher streng limitiert sind. Dank Knowledge Distillation kann dieses winzige Mobile-Modell mit der Genauigkeit einer massiven, Cloud-basierten Architektur arbeiten. Das senkt die Inferenzkosten dauerhaft und kann Ihnen außerdem ersparen, in die „<a href="https://www.vktr.com/ai-technology/the-ai-accuracy-trap/" target="_blank" rel="noreferrer noopener">AI Accuracy Trap</a>“ zu tappen.</p>



<h2 class="wp-block-heading">9. Suchmethoden optimieren</h2>



<p>Herkömmliche Grid-Search-Algorithmen fressen regelmäßig große Teil des Cloud-Budgets, weil sie blindlings Netzwerkkonfigurationen testen und ausführen, die von vornherein zum Scheitern verurteilt sind. Intelligentere Hyperparameter-Suchmethoden wie die <a href="https://de.wikipedia.org/wiki/Bayes%E2%80%99sche_Optimierung" target="_blank" rel="noreferrer noopener">Bayes’sche Optimierung</a> und <a href="https://arxiv.org/abs/1603.06560" target="_blank" rel="noreferrer noopener">Hyperband</a> können an dieser Stelle als finanzielle Wächter fungieren: Sie sagen unzureichende Versuche mathematisch vorher und sortieren diese direkt aus.</p>



<p>Optimiert eine Bank beispielsweise ein KI-Modell zur Betrugserkennung, kann Hyperband Konfigurationen aufspüren, die nicht akkurat sind – und lenkt die gesamte Rechenleistung ausschließlich auf die vielversprechendsten Setups um. Um die Kosten weiter zu reduzieren, lässt sich zudem auch das <a href="https://github.com/Jayachander123/RES-Cost-Aware-Retraining-Framework" target="_blank" rel="noreferrer noopener">RES-Cost-Aware-Retraining-Framework</a> integrieren.</p>



<h2 class="wp-block-heading">10. Parallelstrategien fahren</h2>



<p>Nicht sachgemäß konfigurierte Cluster führen ebenfalls zu massiven Netzwerk-Bottlenecks. Wenn Sie ein Modell mittlerer Größe auf zu viele GPUs aufteilen (Modellparallelität), verbringen die Prozessoren mehr Zeit damit, auf die Datenübertragung zu warten, als damit, tatsächlich Berechnungen durchzuführen.</p>



<p>Umgekehrt ist es bei der Verarbeitung großer Datensätze hocheffizient, das gesamte Modell über mehrere Knoten (Datenparallelität) zu replizieren – vorausgesetzt, die Batch-Größen sind korrekt abgestimmt. Ein FinOps-Team in der Praxis muss diese Parallelstrategien dynamisch an die jeweilige Architektur anpassen und dabei sicherstellen, dass die GPUs nicht <a href="https://www.computerwoche.de/article/4163759/gpu-effizienz-verdoppeln-ohne-zusatzkosten.html" target="_blank">in den Idle-Status verfallen</a>, während das Netzwerk aufholt.</p>



<h2 class="wp-block-heading">11. Asynchron evaluieren</h2>



<p>Standardmäßige Trainings-Pipelines sorgen ständig dafür, dass das primäre (und teure) GPU-Cluster eine Pause einlegen muss. Einfach nur, um routinemäßige Validierungsprüfungen der Modellfortschritte durchzuführen. Anders ausgedrückt: Es ist eine katastrophale Geldverschwendung.</p>



<p>Indem Engineering-Teams asynchrone Evaluierung implementieren, lassen sich die Validierungsprüfungen auf eine separate, wesentlich kostengünstigere CPU- oder Low-Tier-GPU-Instanz auslagern. Die primären, kostenintensiven GPUs möglichst voll auszulasten, ist eine verpflichtende architektonische Trennung. Diese trägt dazu bei, die versteckten Betriebskosten abzumildern, die mit der <a href="https://www.computerwoche.de/article/4030328/so-verandert-ki-ihre-grc-strategie.html" target="_blank">KI-Governance</a> einhergehen. </p>



<h2 class="wp-block-heading">12. Daten kuratieren</h2>



<p>Riesige Datensätze blind zu verarbeiten, sorgt ebenfalls dafür, dass teure Compute-Zeit verschwendet wird – in diesem Fall für redundante Informationen von minderer Qualität.</p>



<p>Wenn ein visuelles KI-Modell bereits zehntausend identische Fotos eines Standard-Stoppschilds erfasst hat, liefert es keinerlei Mehrwert, noch einmal ein paar mehr nachzulegen. Algorithmisches Sampling zu nutzen, um informationsreiche Subsets zu kuratieren, resultiert in identischer Modell-Performance – zu einem Bruchteil der Hardwarekosten. (fm)</p>



<p><strong>Dieser Beitrag wurde im Rahmen des </strong><a href="https://www.infoworld.com/article/4168496/12-model-level-deep-cuts-to-slash-ai-training-costs.html" target="_blank"><strong>englischsprachigen Expert Contributor Network</strong></a><strong> von Foundry veröffentlicht. Alle Infos zum deutschsprachigen Experten-Netzwerk </strong><a href="https://www.computerwoche.de/experten/" target="_blank"><strong>finden Sie hier</strong></a><strong>.</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[7 Data-Science-Perlen für Python]]></title>
<description><![CDATA[Diese Python-Tools bereichern das Data-Scientist-Dasein.DC Studio | shutterstock.com



Pythons opulentes Tool-Ökosystem ist einer der wesentlichen Vorzüge der Programmiersprache. Die Kehrseite: Weil es so viele Python-Tools gibt, ist es allzu leicht, wirklich gute zu übersehen. Deshalb haben wir...]]></description>
<link>https://tsecurity.de/de/3642680/it-security-nachrichten/7-data-science-perlen-fuer-python/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3642680/it-security-nachrichten/7-data-science-perlen-fuer-python/</guid>
<pubDate>Fri, 03 Jul 2026 06:07:30 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>


<div class="extendedBlock-wrapper block-coreImage undefined"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" src="https://b2b-contenthub.com/wp-content/uploads/2025/10/DC-Studio_shutterstock_2624260601_DEOnly_16z9.jpg?quality=50&amp;strip=all&amp;w=1024" alt="Data Scientist 16z9" class="wp-image-4076136" width="1024" height="576" sizes="auto, (max-width: 1024px) 100vw, 1024px"><figcaption class="wp-element-caption">Diese Python-Tools bereichern das Data-Scientist-Dasein.</figcaption></figure><p class="imageCredit">DC Studio | shutterstock.com</p></div>



<p><a href="https://www.computerwoche.de/article/2795515/wie-sie-python-richtig-installieren.html" target="_blank">Pythons</a> opulentes <a href="https://www.computerwoche.de/article/2832867/10-tipps-fuer-schnellere-python-apps.html" target="_blank">Tool-Ökosystem</a> ist einer der wesentlichen Vorzüge der Programmiersprache. Die Kehrseite: Weil es so viele Python-Tools gibt, ist es allzu leicht, wirklich gute zu übersehen. Deshalb haben wir in diesem Artikel sieben empfehlenswerte <a href="https://www.computerwoche.de/article/2812900/die-besten-tools-fuer-datenwissenschaftler.html" target="_blank">Data-Science-Tools</a> für Python zusammengestellt, die bislang (noch) nicht ins Rampenlicht gerückt sind.</p>



<h2 class="wp-block-heading"><a href="https://github.com/sfu-db/connector-x" target="_blank" rel="noreferrer noopener">1. ConnectorX</a></h2>



<p>Daten liegen in der Regel in einer <a href="https://www.computerwoche.de/article/4026190/database-design-tipps-fur-entwickler.html" target="_blank">Datenbank</a> – werden jedoch außerhalb dieser verarbeitet. Daten aus der Datenbank zu extrahieren oder hinzuzufügen, kann dabei ein echter Zeitfresser sein. ConnectorX lädt Informationen aus Datenbanken direkt in gängige Datenverarbeitungs-Tools für Python. Dazu sind meistens nur ein paar Zeilen Python-Code und eine <a href="https://www.computerwoche.de/article/2830678/7-fatale-sql-fehler.html" target="_blank">SQL-Abfrage</a> nötig.</p>



<p>Das Kernstück von ConnectorX ist eine Rust-Bibliothek. Das ermöglicht beispielsweise, Daten aus einer Quelle zu laden und parallel zu partitionieren. <a href="https://www.computerwoche.de/article/3508938/so-geht-postgresql.html" target="_blank">PostgreSQL</a>-Daten lassen sich etwa einladen, indem eine Partitionsspalte spezifiziert wird. Neben PostgreSQL liest ConnectorX auch Daten ein aus</p>



<ul class="wp-block-list">
<li>MySQL/MariaDB,</li>



<li>SQLite,</li>



<li>Amazon Redshift,</li>



<li>Microsoft SQL Server,</li>



<li>Azure SQL, sowie</li>



<li>Oracle.</li>
</ul>



<p>Die Ergebnisse können anschließend in einen <a href="https://www.computerwoche.de/article/2830198/so-geht-datenanalyse-mit-python.html" target="_blank">Pandas</a>– oder PyArrow-DataFrame, in Modin oder Dask (über Pandas) oder auch in Polars (über PyArrow) weiterverwendet werden. Allgemeiner Support für <a href="https://en.wikipedia.org/wiki/Open_Database_Connectivity" target="_blank" rel="noreferrer noopener">ODBC</a> ist derzeit in Arbeit.</p>



<h2 class="wp-block-heading"><a href="https://duckdb.org/" target="_blank" rel="noreferrer noopener">2. DuckDB</a></h2>



<p>Wie andere <a href="https://www.computerwoche.de/article/2810245/was-ist-olap.html" target="_blank">OLAP</a>-Datenbank-Engines verwendet <a href="https://www.computerwoche.de/article/2834231/so-geht-duckdb.html" target="_blank">DuckDB</a> einen spaltenorientierten Datenspeicher und ist für langfristig laufende Analyse-Workloads ausgelegt. DuckDB bietet jedoch sämtliche Funktionen, die man von einer traditionellen Datenbank erwarten würde – etwa ACID-Transaktionen. Ein weiterer Vorteil: Sie müssen keine separate Software-Suite konfigurieren. Innerhalb einer Python-Umgebung installieren Sie DuckDB mit folgendem Befehl: <code>pip install duckdb</code>.</p>



<p>DuckDB verarbeitet Daten im CSV-, <a href="https://www.computerwoche.de/article/2815830/was-ist-json.html" target="_blank">JSON-</a> oder Parquet-Format sowie aus <a href="https://duckdb.org/docs/stable/data/data_sources" target="_blank" rel="noreferrer noopener">einer Vielzahl weiterer, gängiger Datenquellen</a>. Die resultierenden Datenbanken lassen sich aus Effizienzgründen basierend auf Schlüsselwerten (etwa Jahr und Monat) auch in mehreren physischen Dateien partitionieren. Die Datenabfrage funktioniert wie bei jeder anderen SQL-basierten, relationalen Datenbank – allerdings mit zusätzlichen integrierten Funktionen. Etwa, zufällige Daten-Samples zu nehmen oder Fensterfunktionen zu erstellen.</p>



<p>DuckDB verfügt darüber hinaus auch über eine kleine, aber sehr nutzwertige Sammlung von Extensions. Diese kommen beispielsweise zum Einsatz für:</p>



<ul class="wp-block-list">
<li>Volltextsuchen,</li>



<li><a href="https://duckdb.org/docs/stable/core_extensions/vss" target="_blank" rel="noreferrer noopener">beschleunigte Vektorähnlichkeitssuchen</a>,</li>



<li>Excel-Importe und -Exporte,</li>



<li>direkte Verbindungen zu SQLite und PostgreSQL,</li>



<li>Parquet-Dateiexporte, sowie</li>



<li>Support für diverse gängige Geodatenformate und -typen.</li>
</ul>



<h2 class="wp-block-heading"><a href="https://github.com/hi-primus/optimus" target="_blank" rel="noreferrer noopener">3. Optimus</a></h2>



<p>Daten zu bereinigen und vorzubereiten, ist in DataFrame-zentrierten Projekten eine eher undankbare Aufgabe. Das All-in-One-Toolset Optimus will diese erleichtern, indem es Entwicklern ermöglicht, Daten in, beziehungsweise aus einer Vielzahl von Quellen zu laden, zu erkunden, zu bereinigen und zurückzuschreiben.</p>



<p>Als zugrundeliegende Daten-Engine kann Optimus neben Pandas auch Dask, CUDF, Vaex oder <a href="https://www.computerwoche.de/article/2822399/was-ist-apache-spark.html" target="_blank">Spark</a> nutzen. Daten lassen sich einer Vielzahl gängiger Quellen laden, etwa Arrow oder Parquet. Darüber hinaus werden auch Flatfile-Formate wie CSV und JSON unterstützt.  </p>



<p>Die API zur Datenmanipulation ähnelt der von Pandas, erweitert diese jedoch um die Zugriffsmethoden <code>.rows()</code> und <code>.cols()</code>. Das ist beispielsweise nützlich, um:</p>



<ul class="wp-block-list">
<li>einen DataFrame zu sortieren,</li>



<li>nach Spaltenwerten zu filtern,</li>



<li>Daten nach bestimmten Kriterien zu verändern, oder</li>



<li>den Anwendungsbereich anhand bestimmter Kriterien zu simplifizieren.</li>
</ul>



<p>Zusätzlich beinhaltet Optimus auch Prozessoren, um reale Datentypen wie E-Mail-Adressen und URLs zu verarbeiten.</p>



<p>Problematisch ist mit Blick auf Optimus möglicherweise, dass die letzte offizielle Version aus dem Jahr 2020 stammt. Es ist also möglicherweise nicht so aktuell wie andere Komponenten in Ihrem Stack.</p>



<h2 class="wp-block-heading"><a href="https://github.com/pola-rs/polars" target="_blank" rel="noreferrer noopener">4. Polars</a></h2>



<p>Wenn Sie viel Zeit mit DataFrames verbringen und regelmäßig von den Performance-Grenzen von Pandas frustriert sind, sollten Sie einen Blick auf Polars werfen.</p>



<p>Die DataFrame-Bibliothek für Python bietet eine komfortable Syntax, die der von Pandas ähnelt – greift jedoch auf eine Rust-Bibliothek zurück, die die vorhandene Hardware optimal nutzt. Um Features wie Parallelverarbeitung oder SIMD zu nutzen, ist keine spezielle Syntax notwendig – alles läuft automatisch. So laufen selbst simple Vorgänge wie aus einer CSV-Datei zu lesen, deutlich schneller ab. Rust-Entwickler können zudem auch ihre eigenen Extensions für Polars entwickeln – mit <a href="https://github.com/pola-rs/pyo3-polars" target="_blank" rel="noreferrer noopener">pyo3</a>.</p>



<p>Polars bietet sowohl Eager- als auch Lazy-Ausführungsmodi – Queries lassen sich also sofort oder bei Bedarf auch zeitverzögert fahren. Das Python-Tool wartet außerdem mit einer Streaming-API auf, um Abfragen inkrementell zu verarbeiten. Allerdings ist Streaming für diverse Funktionen noch nicht verfügbar. Wenn doch, greift Polars dazu auf eine In-Memory-Engine zurück.</p>



<p>Das Tool ermöglicht außerdem (über die externe Graphviz-Bibliothek), sich mit Hilfe von <a href="https://docs.pola.rs/api/python/stable/reference/lazyframe/api/polars.LazyFrame.show_graph.html" target="_blank" rel="noreferrer noopener">Execution-Graphen</a> ein Bild davon zu machen, wie viel Speicher- und CPU-Ressourcen eine Abfrage nutzt.  </p>



<h2 class="wp-block-heading"><a href="https://github.com/iterative/dvc" target="_blank" rel="noreferrer noopener">5. DVC</a></h2>



<p>Ein weit verbreitetes Problem im Zusammenhang mit Data-Science-Projekten ist die <a href="https://www.computerwoche.de/article/2833711/version-control-systems-ein-ratgeber.html" target="_blank">Versionskontrolle</a> (nicht des Projektcodes, sondern der Daten). Mit dem Tool DVC (Data Version Control) lassen sich Versionsdeskriptoren an Datensätze anhängen. Diese lassen sich wie der Rest des Codes in Git einchecken, um die Versionen von Daten und Code konsistent zu halten.</p>



<p>DVC kann nahezu jede Art von Datensatz tracken, solange diese sich in einer Datei abbilden lassen. Dabei spielt es keine Rolle, ob die Daten in einem <a href="https://dvc.org/doc/user-guide/data-management/remote-storage#supported-storage-types" target="_blank" rel="noreferrer noopener">Remote-Storage-Service</a> oder lokal vorgehalten werden. Das Konzept: Sie beschreiben über eine “<a href="https://dvc.org/doc/user-guide/data-management/remote-storage#supported-storage-types" target="_blank" rel="noreferrer noopener">Pipeline</a>“, wie Datenmodelle gemanagt und genutzt werden.</p>



<p>DVC kann allerdings mehr, als nur Daten zusammen mit Code zu versionieren. Das Tool kann zum Beispiel auch fungieren als:</p>



<ul class="wp-block-list">
<li>schneller Datencache für remote gehostete Daten,</li>



<li>Methodik, um Experimente zu tracken, die mit den Daten durchgeführt werden, und</li>



<li>Register oder Katalog für Machine-Learning-Modelle, die mit den Daten erstellt wurden.</li>
</ul>



<p>Benutzer von <a href="https://www.computerwoche.de/article/2833165/10-tricks-fuer-visual-studio-code.html" target="_blank">Visual Studio Code</a> können DVC-Workflows über die <a href="https://marketplace.visualstudio.com/items?itemName=Iterative.dvc" target="_blank" rel="noreferrer noopener">entsprechende Extension</a> in ihren Editor integrieren.</p>



<h2 class="wp-block-heading"><a href="https://github.com/cleanlab/cleanlab" target="_blank" rel="noreferrer noopener">6. Cleanlab</a></h2>



<p>Weil es teuer und zeitaufwändig ist, saubere, korrekt gelabelte Daten zu erstellen, sind hochwertige Datensätze für Machine-Learning-Zwecke Mangelware. Manchmal bleibt Datenwissenschaftlern keine andere Wahl, als mit Rohdaten oder inkonsistenten Informationen zu arbeiten. Für dieses Szenario wurde das Tool Cleanlab entwickelt.</p>



<p>Dieses Python-Daten-Tool nutzt vorhandene, hochwertige Machine-Learning-Datensätze, um solche von geringerer Qualität, die nicht oder nur unzureichend gekennzeichnet sind, zu analysieren. Anders ausgedrückt: Sie erstellen ein Modell auf der Grundlage des ursprünglichen Datensatzes. Anschließend finden Sie mit Cleanlab heraus, was in diesem ursprünglichen Datensatz verbessert werden muss – und trainieren dann das Modell erneut mit Ihrem automatisch bereinigten und angepassten Datensatz.</p>



<p>Cleanlab funktioniert unabhängig von Datenmodellen und -Frameworks. Es spielt also keine Rolle, ob Sie <a href="https://www.computerwoche.de/article/2816666/13-tools-die-ki-und-ml-transformieren.html" target="_blank">PyTorch</a>, OpenAI, Scikit-learn oder <a href="https://www.infoworld.com/article/2255099/what-is-tensorflow-the-machine-learning-library-explained.html" target="_blank">Tensorflow</a> nutzen – Cleanlab arbeitet mit jedem Classifier. Dabei verfügt das Tool dennoch über spezifische Workflows für gängige Tasks wie:</p>



<ul class="wp-block-list">
<li>Token-Klassifizierung,</li>



<li>Multi-Labeling,</li>



<li>Regression,</li>



<li>Bildsegmentierung, oder auch</li>



<li>Objekt- und Outlier-Detection.</li>
</ul>



<p>Idealerweise machen Sie sich <a href="https://github.com/cleanlab/examples" target="_blank" rel="noreferrer noopener">anhand diverser Beispiele</a> selbst ein Bild davon, wie der Prozess funktioniert und welche Ergebnisse zu erwarten sind.</p>



<h2 class="wp-block-heading"><a href="https://github.com/snakemake/snakemake" target="_blank" rel="noreferrer noopener">7. Snakemake</a></h2>



<p>Data-Science-Workflows sind diffizil einzurichten. Noch schwieriger ist es aber, das auf konsistente und vorhersehbare Weise zu erledigen. Um diesen Prozess zu automatisieren und Datenanalyse-Workflows so aufzusetzen, dass alle Beteiligten die gleichen Ergebnisse zu erhalten, wurde Snakemake entwickelt. Dabei gilt: Je mehr bewegliche Teile Ihr Data-Science-Workflow enthält, desto größer ist die Wahrscheinlichkeit, dass Sie davon profitieren werden, diesen mit Snakemake zu automatisieren.</p>



<p>Snakemake-Workflows ähneln dabei GNU-Make-Workflows: Sie definieren die Schritte des Workflows mit Regeln. Diese legen fest, was aufgenommen sowie ausgegeben wird – und welche Befehle ausgeführt werden müssen. Die Workflow-Regeln können multithreaded sein und Konfigurationsdaten lassen sich über JSON- oder <a href="https://www.computerwoche.de/article/2815752/so-umgehen-sie-yaml-probleme.html" target="_blank">YAML-</a>Dateien einspielen. Sie können in Ihren Workflows außerdem auch Funktionen definieren, um die in den Regeln verwendeten Daten zu transformieren – und die bei jedem Schritt ausgeführten Aktionen zu protokollieren.</p>



<p>Snakemake-Jobs sind zudem portabel – sie können sowohl in Managed-Kubernetes- als auch bestimmten Cloud-Umgebungen bereitgestellt werden. Und:</p>



<ul class="wp-block-list">
<li>Workloads lassen sich auch “einfrieren”, um einen bestimmten Satz von Packages zu verwenden,</li>



<li>für erfolgreich ausgeführte Workloads können automatisiert Unit-Tests erstellt und gespeichert werden – für eine langfristige Archivierung auch als Tarball.  </li>
</ul>



<p>(fm)</p>



<p><strong>Dieser Artikel ist <a href="https://www.infoworld.com/article/2338444/7-newer-data-science-tools-you-should-be-using-with-python.html" target="_blank">im Original</a> bei unserer Schwesterpublikation Infoworld.com erschienen.</strong></p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[CVE-2026-24747 | PyTorch up to 2.9.x Checkpoint File deserialization (ID 163105 / Nessus ID 297037)]]></title>
<description><![CDATA[A vulnerability identified as critical has been detected in PyTorch up to 2.9.x. Affected by this issue is some unknown functionality of the component Checkpoint File Handler. This manipulation causes deserialization.

This vulnerability is registered as CVE-2026-24747. Remote exploitation of the...]]></description>
<link>https://tsecurity.de/de/3637061/sicherheitsluecken/cve-2026-24747-pytorch-up-to-29x-checkpoint-file-deserialization-id-163105-nessus-id-297037/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3637061/sicherheitsluecken/cve-2026-24747-pytorch-up-to-29x-checkpoint-file-deserialization-id-163105-nessus-id-297037/</guid>
<pubDate>Wed, 01 Jul 2026 01:23:03 +0200</pubDate>
<category>🕵️ Sicherheitslücken</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[A vulnerability identified as <a href="https://vuldb.com/kb/risk">critical</a> has been detected in <a href="https://vuldb.com/product/pytorch">PyTorch up to 2.9.x</a>. Affected by this issue is some unknown functionality of the component <em>Checkpoint File Handler</em>. This manipulation causes deserialization.

This vulnerability is registered as <a href="https://vuldb.com/cve/CVE-2026-24747">CVE-2026-24747</a>. Remote exploitation of the attack is possible. No exploit is available.

You should upgrade the affected component.]]></content:encoded>
</item>
<item>
<title><![CDATA[Künstliche Intelligenz programmieren: Die besten Coding-Sprachen für KI]]></title>
<description><![CDATA[Wenn es darum geht, Künstliche Intelligenz zu programmieren, stehen Ihnen diverse Optionen zur Wahl. Wir zeigen Ihnen die besten KI-Programmiersprachen.
					Foto: Connect world – shutterstock.com




Künstliche Intelligenz (KI) eröffnet Softwareentwicklern völlig neue Möglichkeiten: Mit Hilfe vo...]]></description>
<link>https://tsecurity.de/de/3626228/it-security-nachrichten/kuenstliche-intelligenz-programmieren-die-besten-coding-sprachen-fuer-ki/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3626228/it-security-nachrichten/kuenstliche-intelligenz-programmieren-die-besten-coding-sprachen-fuer-ki/</guid>
<pubDate>Fri, 26 Jun 2026 05:22:23 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Wenn es darum geht, Künstliche Intelligenz zu programmieren, stehen Ihnen diverse Optionen zur Wahl. Wir zeigen Ihnen die besten KI-Programmiersprachen." title="Wenn es darum geht, Künstliche Intelligenz zu programmieren, stehen Ihnen diverse Optionen zur Wahl. Wir zeigen Ihnen die besten KI-Programmiersprachen." src="https://images.computerwoche.de/bdb/3284310/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Wenn es darum geht, Künstliche Intelligenz zu programmieren, stehen Ihnen diverse Optionen zur Wahl. Wir zeigen Ihnen die besten KI-Programmiersprachen.</p></figcaption></figure><p class="imageCredit">
					Foto: Connect world – shutterstock.com</p></div>




<p><a href="https://www.computerwoche.de/article/2769837/was-sie-zum-thema-ki-wissen-muessen.html" title="Künstliche Intelligenz" target="_blank">Künstliche Intelligenz</a> (KI) eröffnet Softwareentwicklern völlig neue Möglichkeiten: Mit Hilfe von <a href="https://www.computerwoche.de/article/2789087/noch-viel-zu-tun-bei-ki-und-ml.html" title="Machine und Deep Learning" target="_blank">Machine und Deep Learning</a> lassen sich bessere Nutzerprofile und Empfehlungen, ein höherer Personalisierungsgrad, smartere Suchoptionen oder intelligentere Interfaces realisieren. Dabei stellt sich unweigerlich die Frage, welche Programmiersprache dafür zum Einsatz kommen soll. Die Anforderungen, denen eine KI-Coding-Sprache genügen muss, sind vielfältig: eine Vielzahl von Machine- und Deep-Learning-Bibliotheken sollten genauso vorhanden sein wie eine performante Laufzeitumgebung, ausgiebiger Tool Support, eine große Entwickler-Community und ein gesundes Ökosystem.</p>



<p>Trotzdem dieser Anforderungskatalog umfassend ist, stehen Ihnen einige gute Optionen zur Wahl, wenn es darum geht, Künstliche Intelligenz zu <a href="https://www.computerwoche.de/article/2813209/11-wege-ihre-softwareentwicklung-neu-zu-definieren.html" title="programmieren" target="_blank">programmieren</a>. Wir zeigen Ihnen eine Auswahl der besten KI-Programmiersprachen.</p>



<h2 class="wp-block-heading">Python</h2>



<p>Wenn Sie als Developer mit künstlicher Intelligenz arbeiten, führt mit an Sicherheit grenzender Wahrscheinlichkeit kein Weg an <a href="https://www.cowo.de/a/3548847" title="Python" target="_blank" rel="noopener">Python</a> vorbei. Inzwischen unterstützen auch so gut wie alle gängigen Bibliotheken Python 3.x – die Zeiten, in denen die Umstellung von Python 2.x auf 3.x Kompatibilitätsprobleme mit sich brachte, sind so gut wie vorbei. Mit anderen Worten: Sie können nun endlich auch in der Praxis von den zahlreichen neuen Features von Python 3.x profitieren. Was nicht heißen soll, dass die packaging-Hürden bei <a href="https://www.computerwoche.de/article/2762025/python-lernen-leicht-gemacht.html" title="Python" target="_blank">Python</a> überhaupt keine Rolle mehr spielen – das Gros der Probleme lässt sich aber mit Hilfe von Anaconda umschiffen. Nichtsdestotrotz wäre es zu begrüßen, wenn die Python Community endlich voll und ganz von diesen Hürden befreit würde.</p>



<p>Davon abgesehen sind die verfügbaren mathematischen und statistischen Bibliotheken von Python denen anderer Programmiersprachen weit voraus: <a title="NumPy" href="https://numpy.org/" target="_blank" rel="noopener">NumPy</a> ist inzwischen so allgegenwärtig, dass es beinahe als Standard-API für Tensor Operations bezeichnet werden kann, während <a title="Pandas" href="https://pandas.pydata.org/" target="_blank" rel="noopener">Pandas</a> die flexiblen Dataframes von R in die Python-Welt trägt. Geht es um Natural Language Processing (NLP) haben Sie die Wahl zwischen dem altehrwürdigen <a title="NLTK" href="https://www.nltk.org/" target="_blank" rel="noopener">NLTK</a> und dem superschnellen <a title="SpaCy" href="https://spacy.io/" target="_blank" rel="noopener">SpaCy</a>, während sich für Machine-Learning-Zwecke das bewährte <a title="scikit-learn" href="https://scikit-learn.org/stable/" target="_blank" rel="noopener">scikit-learn</a> empfiehlt. Geht es hingegen um Deep Learning, sind alle aktuellen Bibliotheken (<a title="TensorFlow" href="https://www.tensorflow.org/" target="_blank" rel="noopener">TensorFlow</a>, <a title="PyTorch" href="https://pytorch.org/" target="_blank" rel="noopener">PyTorch</a>, <a title="Chainer" href="https://chainer.org/" target="_blank" rel="noopener">Chainer</a>, <a title="Apache MXNet" href="https://mxnet.apache.org/" target="_blank" rel="noopener">Apache MXNet</a>, etc.) im Grunde “Python-first”-Projekte.</p>



<p>Wenn Sie ein regelmäßiger Besucher von <a title="arXiv" href="https://arxiv.org/" target="_blank" rel="noopener">arXiv</a> sind, wird Ihnen längst aufgefallen sein, dass die Mehrzahl der dortigen Deep-Learning-Forschungsprojekte, die Quellcode zur Verfügung stellen, dazu auf Python setzen. In Sachen Deployment-Modelle haben Microservice-Architekturen und -Technologien wie <a href="https://github.com/seldonio/seldon-core" target="_blank" rel="noreferrer noopener">SeldonCore</a> die Auslieferung von Python-Modellen in Produktivumgebungen wesentlich vereinfacht.</p>



<p><a href="https://www.computerwoche.de/article/2762145/python-besitzt-mehr-als-300-bibliotheken.html" title="Python" target="_blank">Python</a> ist zweifellos die Programmiersprache der Wahl, wenn es um KI-Forschung geht: Sie bietet die größte Auswahl an Machine und Deep Learning Frameworks und ist die Coding-Sprache, die innerhalb der KI-Welt tonangebend ist.</p>



<h2 class="wp-block-heading">C++</h2>



<p><a href="https://www.computerwoche.de/article/2816926/wie-sich-c-gegen-c-python-und-co-schlaegt.html" title="C++" target="_blank">C++</a> ist aller Voraussicht nach nicht die erste Wahl für Ihr <a href="https://www.computerwoche.de/article/2781630/so-wird-ihr-ki-projekt-ein-erfolg.html" title="KI-Projekt" target="_blank">KI-Projekt</a>. Allerdings wird Deep Learning im Edge-Bereich ein immer gängigeres Szenario. In diesem Fall müssen Sie Ihre Modelle auf Systemen zum Laufen bringen, die nur sehr begrenzte Ressourcen zur Verfügung haben. Um das letzte bisschen Performance aus dem System zu pressen, kann es nötig werden, noch einmal in die Untiefen der Pointer-Welt abzutauchen.</p>



<p>Glücklicherweise kann moderner C++ Code aber tatsächlich angenehm zu schreiben sein. Hierfür stehen Ihnen mehrere Ansätze zur Wahl: Entweder Sie nutzen Bibliotheken wie Nvidias <a href="https://www.computerwoche.de/article/2818437/was-ist-cuda.html" target="_blank">CUDA</a> um ihren eigenen Programmcode zu schreiben, der direkt in die GPU fließt – oder Sie setzen wahlweise auf TensorFlow oder PyTorch, um Zugang zu flexiblen high-level <a href="https://www.computerwoche.de/article/4004872/die-besten-apis-um-ki-zu-integrieren.html" target="_blank">APIs</a> zu erlangen. Sowohl <a title="PyTorch" href="https://www.computerwoche.de/article/2794453/5-gruende-fuer-das-deep-learning-framework.html" target="_blank">PyTorch</a> als auch <a title="TensorFlow" href="https://www.computerwoche.de/article/2789330/die-besten-tools-fuer-tensorflow.html" target="_blank">TensorFlow</a> erlauben Ihnen, Modelle, die in Python geschrieben sind, in eine C++ Laufzeitumgebung zu integrieren. So rücken Sie deutlich näher an den Produktiveinsatz, bleiben dabei aber flexibel in der Entwicklung.</p>



<p>Weil KI-Applikationen sich immer stärker über alle Devices – von Embedded Systems bis hin zu riesigen Clustern – hinweg ausbreiten, ist C++ ein wichtiger Bestandteil des KI-Coding-Toolkits. Um <a href="https://www.computerwoche.de/article/2807255/so-sichern-sie-den-netzwerkrand-ab.html" title="künstliche Intelligenz im Edge-Bereich" target="_blank">künstliche Intelligenz im Edge-Bereich</a> zu realisieren, gilt es eben nicht nur akkurat zu programmieren, sondern auch qualitativ gut und schnell.</p>



<h2 class="wp-block-heading">Java und andere JVM-Sprachen</h2>



<p>Die Familie der JVM-Programmiersprachen (<a title="Java" href="https://www.computerwoche.de/article/2814735/warum-java-immer-noch-rockt.html" target="_blank">Java</a>, Scala, <a title="Kotlin" href="https://www.computerwoche.de/article/2817523/was-ist-kotlin.html" target="_blank">Kotlin</a>, Clojure, etc.) ist weiterhin eine gute Wahl, wenn es um die Entwicklung von KI-Applikationen geht. Eine reichhaltige Auswahl an Bibliotheken steht für nahezu alle Aspekte zur Auswahl – sei es Natural Language Processing (<a title="CoreNLP" href="https://stanfordnlp.github.io/CoreNLP/" target="_blank" rel="noopener">CoreNLP</a>), Tensor Operations oder GPU-beschleunigtes Deep Learning (<a title="DL4J" href="https://deeplearning4j.org/" target="_blank" rel="noopener">DL4J</a>). Darüber hinaus gewährleisten diese Coding-Sprachen auch einfachen Zugang zu Big-Data-Plattformen wie <a title="Apache Spark" href="https://www.infoworld.com/article/2259224/what-is-apache-spark-the-big-data-platform-that-crushed-hadoop.html" target="_blank">Apache Spark</a> und <a title="Apache Hadoop" href="https://hadoop.apache.org/" target="_blank" rel="noopener">Apache Hadoop</a>.</p>



<p>Für die meisten Unternehmen ist <a href="https://www.computerwoche.de/article/2814678/darum-heisst-java-java.html" title="Java" target="_blank">Java</a> die lingua franca – und mit Java 8 und neueren Versionen verliert auch die Erstellung von Java Code ihren Schrecken. Eine <a href="https://www.computerwoche.de/article/2793583/kuenstliche-intelligenz-verdraengt-den-menschen-nicht.html" title="KI-Applikation" target="_blank">KI-Applikation</a> in Java zu programmieren mag sich ein wenig langweilig anfühlen, sorgt aber in der Regel für zufriedenstellende Ergebnisse und ermöglicht Ihnen, alle existierenden Bestandteile einer Java-Infrastruktur für Entwicklung, Deployment und Monitoring einzusetzen.</p>



<h2 class="wp-block-heading">JavaScript</h2>



<p><a href="https://www.computerwoche.de/article/2832952/was-ist-javascript.html" title="JavaScript" target="_blank">JavaScript</a> ausschließlich für die Entwicklung von KI-Applikationen zu erlernen, ist ein höchst unwahrscheinliches Szenario. Allerdings bietet Googles <a href="https://www.tensorflow.org/js" title="TensorFlow.js" target="_blank" rel="noopener">TensorFlow.js</a> weiterhin eine gute Möglichkeit, Ihre Keras- und TensorFlow-Modelle über Browser oder Node.js auszuliefern.</p>



<p>Dennoch ist der große Ansturm von <a href="https://www.computerwoche.de/article/2794625/was-javascript-von-typescript-unterscheidet.html" title="JavaScript-Entwicklern" target="_blank">JavaScript-Entwicklern</a> im Bereich <a href="https://www.computerwoche.de/k/kuenstliche-intelligenz-artifical-intelligence,3544" target="_blank" class="idgGlossaryLink">Künstliche Intelligenz</a> bislang ausgeblieben. Das könnte daran liegen, dass das JavaScript-Ökosystem in Sachen verfügbare Bibliotheken bislang den nötigen Tiefgang vermissen lässt – zumindest im Vergleich zu Programmiersprachen wie Python. Darüber hinaus stehen auf Serverseite durch Deployment-Modelle mit Node.js (wiederum im Vergleich zu den Python-Optionen) keine wirklichen Vorteile in Aussicht. KI-Applikationen auf JavaScript-Basis dürften deshalb auch weiterhin auf Browser-Basis entstehen. </p>



<h2 class="wp-block-heading">Swift</h2>



<p><a href="https://www.tensorflow.org/swift" title="Swift for TensorFlow" target="_blank" rel="noopener">Swift for TensorFlow</a> verbindet die neuesten und besten Features von TensorFlow mit den Vorteilen von Python-Bibliotheken, die sich problemlos einbinden lassen – ganz so als würden Sie Python selbst nutzen.</p>



<p>Das Team von fast.ai werkelt derzeit an einer Swift-Version seiner populären Bibliothek – und stellt zahlreiche Optimierungen in Aussicht, gerade in Zusammenhang mit dem <a href="https://www.computerwoche.de/article/2826586/was-ist-llvm.html" target="_blank">LLVM compiler</a>. Von “production ready” kann zwar noch keine Rede sein, aber auf dieser Grundlage könnte die nächste Generation von Deep-Learning-Entwicklungsarbeit entstehen – Sie sollten Swift deshalb auf alle Fälle im Auge behalten.</p>



<h2 class="wp-block-heading">R</h2>



<p><a href="https://www.r-project.org/" title="R" target="_blank" rel="noopener">R</a> ist die Programmiersprache der Wahl für <a href="https://www.computerwoche.de/article/2774747/was-data-scientists-koennen-muessen.html" title="Data Scientists" target="_blank">Data Scientists</a>. Developer aus anderen Bereichen könnten die Coding-Sprache wegen ihres Dataframe-zentrischen Ansatzes hingegen als verwirrend empfinden.</p>



<p>Für ein Team leidenschaftlicher R-Entwickler kann es durchaus Sinn machen, Integrationen mit <a href="https://www.tensorflow.org/" title="TensorFlow" target="_blank" rel="noopener">TensorFlow</a>, <a href="https://keras.io/" title="Keras" target="_blank" rel="noopener">Keras</a> oder <a href="https://www.h2o.ai/" title="H2O" target="_blank" rel="noopener">H2O</a> für Forschung und Prototyping einzusetzen. Hinsichtlich der Performance ist R für den Produktiveinsatz aber lediglich bedingt zu empfehlen. Zwar lässt sich performanter R Code durchaus produktiv zum Einsatz bringen, einfacher dürfte es aber in den allermeisten Fällen sein, den R-Prototypen in Java oder Python neu zu programmieren.</p>



<h2 class="wp-block-heading">KI programmieren – weitere Optionen</h2>



<p>Natürlich sind die vier genannten Programmiersprachen nicht die einzigen Optionen, um <a href="https://www.computerwoche.de/k/kuenstliche-intelligenz-artifical-intelligence,3544" target="_blank" class="idgGlossaryLink">Künstliche Intelligenz</a> zu programmieren. Die folgenden beiden Coding-Sprachen könnten – je nach Einsatzzweck – ebenfalls von Interesse für Ihre <a href="https://www.computerwoche.de/article/2792741/ki-bewaehrt-sich-in-der-praxis.html" title="KI-Projekte" target="_blank">KI-Projekte</a> sein:</p>



<p><strong><a href="https://www.lua.org/" target="_blank" rel="noreferrer noopener">Lua</a></strong></p>



<p>Vor einigen Jahren wurde Lua als “next big thing” im Bereich der Künstlichen Intelligenz gehandelt. Das lag in erster Linie am <a title="Torch Framework" href="https://torch.ch/" target="_blank" rel="noopener">Torch Framework</a> – eine der populärsten Machine-Learning-Bibliotheken sowohl für den produktiven Einsatz als auch für Forschungszwecke. Wenn Sie in ältere DeepLearning-Modelle abtauchen, finden sich oft zahlreiche Verweise auf Torch und Lua-Quellcode. Es könnte durchaus nützlich sein, sich etwas Knowhow über die Torch API anzueignen, die einige Ähnlichkeiten zur Basis-API von PyTorch aufweist. Wenn Sie allerdings kein gesteigertes Bedürfnis haben, für Ihre Applikationen in historische Forschung abzutauchen, können Sie auf Lua auch gut und gerne verzichten.</p>



<p><strong><a href="https://julialang.org/" target="_blank" rel="noreferrer noopener">Julia</a></strong></p>



<p>Bei Julia handelt es sich um eine High-Performance-Programmiersprache, die ihren Fokus auf numerische Berechnungen legt. Dadurch passt sie auch wunderbar in die mathematisch ausgerichtete Welt der Künstlichen Intelligenz. Julia mag derzeit nicht die populärste Coding-Sprache sein, allerdings bieten Wrapper wie <a title="TensorFlow.jl" href="https://github.com/malmaud/TensorFlow.jl" target="_blank" rel="noopener">TensorFlow.jl</a> und <a title="Mocha" href="https://github.com/pluskid/Mocha.jl" target="_blank" rel="noopener">Mocha</a> guten Deep-Learning-Support. Wenn das relativ kleine Ökosystem kein Ausschlusskriterium für Sie darstellt und Sie von Julias Fokus auf High-Performance-Berechnungen profitieren wollen, sollten Sie einen Blick riskieren. (fm)</p>



<p><strong>Dieser Artikel ist <a href="https://www.infoworld.com/article/2258342/6-best-programming-languages-for-ai-development.html" target="_blank">im Original</a> bei unserer Schwesterpublikation Infoworld.com erschienen.</strong></p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Qualcomm’s $3.9 billion purchase of Modular aims to change the data center dynamic]]></title>
<description><![CDATA[Qualcomm on Wednesday said that it will spend $3.9 billion to purchase AI-native software platform developer Modular Inc., a move that Qualcomm says will allow it to level the playing field on data centers by creating “a silicon-agnostic compute layer.”



The stock-based acquisition “further ena...]]></description>
<link>https://tsecurity.de/de/3622741/it-security-nachrichten/qualcomms-39-billion-purchase-of-modular-aims-to-change-the-data-center-dynamic/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3622741/it-security-nachrichten/qualcomms-39-billion-purchase-of-modular-aims-to-change-the-data-center-dynamic/</guid>
<pubDate>Wed, 24 Jun 2026 22:24:10 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Qualcomm on Wednesday said that it will spend $3.9 billion to purchase AI-native software platform developer Modular Inc., a move that Qualcomm says will allow it to level the playing field on data centers by creating “a silicon-agnostic compute layer.”</p>



<p>The <a href="https://d18rn0p25nwr6d.cloudfront.net/CIK-0000804328/70441e71-4fcb-4cdd-8874-f571622bd264.pdf" target="_blank" rel="noreferrer noopener">stock-based acquisition</a> “further enables Qualcomm Technologies to deliver a silicon-agnostic compute layer across devices, edge, and data centers, improving performance-per-watt, increasing hardware flexibility, and expanding an open developer ecosystem so customers can deploy AI more efficiently across heterogeneous platforms globally,” the company said <a href="https://investor.qualcomm.com/news-events/press-releases/news-details/2026/Qualcomm-to-Acquire-Modular/default.aspx" target="_blank" rel="noreferrer noopener">in a statement</a>. </p>



<p>Qualcomm’s position is that enterprises need far more flexibility in their data center strategies, especially given how fluid the AI space is today. When CIOs need to make bets on data centers without knowing what the field will look like in two years, it can be challenging.</p>



<p><a href="https://www.linkedin.com/in/chris-lattner-5664498a/" target="_blank" rel="noreferrer noopener">Chris Lattner</a>, CEO of Modular, posted on LinkedIn that this leveling of the data center playing field was one of the company’s key early goals.</p>



<p>“In a world with a tremendous amount of innovative heterogenous AI hardware, there has always been a gap: existing fragmented software technologies weren’t built to scale effectively across this hardware. This gap holds back innovation and choice and makes development painful,” Lattner <a href="https://www.linkedin.com/posts/chris-lattner-5664498a_im-excited-to-share-that-qualcomm-is-acquiring-share-7475540410514288640-LvCv/" target="_blank" rel="noreferrer noopener">wrote in his LinkedIn post</a>.</p>



<p>“Modular was founded 4.5 years ago to solve this problem,” he wrote. “We’ve already integrated support for several hyperscale datacenter silicon providers, but we’re not stopping with what’s publicly announced. We’ve built an open platform and are continuing to open it further.”</p>



<p>Lattner added that the Qualcomm acquisition “will accelerate our progress and path” by “spanning edge to cloud, CPU, GPU, NPU, and custom ASICs and perhaps more.”</p>



<h2 class="wp-block-heading">Addresses a pain point</h2>



<p>Analysts, although skeptical of the probability of success in taking meaningful market share away from Nvidia, said that Qualcomm has focused on a true sore point for enterprises struggling with data center approaches. </p>



<p><a href="https://moorinsightsstrategy.com/team/matt-kimball/" target="_blank" rel="noreferrer noopener">Matt Kimball</a>, VP and principal analyst with Moor Insights &amp; Strategy, said, “the argument that Modular can make datacenters cost-effective is directionally correct. As enterprise AI actually hits velocity, heterogeneity is almost an understatement. Different accelerators are required for different use cases across different deployment scenarios.”</p>



<p>To date, it’s been a challenge for organizations to manage AI in this environment, he noted. “And when enterprise AI takes off, this challenge will be fully exposed.”</p>



<p>Kimball said that Modular “can be extremely valuable in achieving two things that will vex most organizations: abstracting complexity and delivering significantly more flexibility. And this would certainly lead to TCO advantages. I think the per-watt performance claim can be challenging to validate across every and any deployment scenario, but I understand the spirit behind it.”</p>



<p><a href="https://www.linkedin.com/in/yurigoryunov/" target="_blank" rel="noreferrer noopener">Yuri Goryunov</a>, CIO of consulting firm Acceligence, also applauded the Qualcomm move, but he stressed that the deal’s value is not in the technology as much as in the talent.</p>



<p>“The key is what Qualcomm actually bought: not silicon, but the software layer, meaning Chris Lattner’s team plus Mojo and the MAX engine. That’s the right place to apply pressure. Nvidia’s real moat has never been the GPUs,” he said. “It’s CUDA and the rewrite cost that keeps workloads pinned to their hardware. A credible ‘write once, run across CPU/GPU/NPU/ASIC without rewrites’ layer is exactly what lowers the switching cost and makes non-Nvidia silicon a safer bet.”</p>



<p>But Goryunov said that the data center “democratization” argument also is powerful.</p>



<p>“Anything that pushes toward democratization of compute and better routing of tasks to best-fit capacity adds real flexibility to the ecosystem,” he noted. “If workloads can be matched to the right compute instead of defaulting to one vendor, everyone gets more efficiency on performance-per-watt and TCO and customers get real choice. That’s the part of this I find most compelling.”</p>



<h2 class="wp-block-heading">Still some obstacles</h2>



<p>That said, none of this will be easy, he pointed out.</p>



<p>“Does it change the competitive position versus Nvidia? Directionally, yes. It opens a credible second front at the exact point where Nvidia is stickiest. I’d stop short of saying it shifts the balance overnight. CUDA’s moat is a decade deep and this is a multi-year execution play,” Goryunov said. “But the attack is aimed at the right wall and the team they bought is about as serious as it gets for this fight.”</p>



<p>But he stressed that much of Qualcomm’s strategy with this acquisition relies on an uncertain assumption: That Nvidia won’t counterattack by opening its architectures to various others. Or, at the very least, that Nvidia won’t do so quickly enough.</p>



<p>“That’s the barrier to entry, which is that Nvidia will focus on their stickiness,” Goryunov said.</p>



<p>Kimball added that, from a competitive perspective, Qualcomm has various obstacles to overcome. “Part of this acquisition goes directly to the Nvidia challenge” of finding a way to “make it easier for customers to deploy heterogeneous silicon without software getting in the way.”</p>



<p><a href="https://www.infotech.com/profiles/john-annand" target="_blank" rel="noreferrer noopener">John Annand</a>, senior technical counselor at Info-Tech Research Group, is more skeptical of Qualcomm’s ability to do serious damage to Nvidia.</p>



<p>“Nvidia has something like 85% of the AI accelerator chip market,” he pointed out. “Sure, they have nowhere to go but down, but that’s still going to take them a while. More importantly, they have literally spent decades working with practitioners in AI and ML and compute-intensive fields, indoctrinating them into their CUDA software ecosystem. Rewriting that tool chain will take institutional change at most organizations, which means years, if not decades, to uncouple.”</p>



<p>“Organizations that think they’ve achieved agnosticism because they’re using high-level abstractions like PyTorch, well,  they have come closest,” he observed. “But just cutting and pasting the same code into AMD Instinct can lead to memory and dependency errors. It’s like VM lift and shifts to the public cloud 10 years ago. Easier, but still possible to screw up.”</p>



<p>Nonetheless, Annand said that the deal, if it goes through, is still good news for enterprises. </p>



<p>“What it means for enterprise IT is that the vendors we currently rely on to deliver AI have another potential building block. Because enterprise IT accesses AI via an API call, it’s operationally irrelevant to us if Claude runs on Nvidia, AMD ROCm, or Modular,” he said. </p>



<p>“Now, because of the commercial and stock agreements, OpenAI and Anthropic aren’t going to jump ship anytime soon. But if your enterprise is looking for more boutique offerings, like those from Cohere, or is looking to build its own models and tools from scratch, this is an exciting announcement.”</p>



<h2 class="wp-block-heading">Goal: build once, run anywhere</h2>



<p><a href="https://www.infotech.com/profiles/shashi-bellamkonda" target="_blank" rel="noreferrer noopener">Shashi Bellamkonda</a>, principal research director at Info-Tech Research Group, looks at the potential acquisition, while it will potentially deliver benefits, as suffering from many practical roadblocks.  </p>



<p>“Qualcomm is chasing what you might call model democracy,” he said, noting that today, AI deployment teams are locked to whatever accelerator they trained on, and moving a model to different hardware means re-engineering, not just configuration changes.</p>



<p>“Modular’s pitch is that this goes away: build once, run across CPU, GPU, NPU, whatever the infrastructure calls for,” Bellamkonda said. “That’s a credible goal. The catch is that democracy and portability aren’t the same thing. Qualcomm will tune hardest for Qualcomm silicon. Every hardware company does. Vendor-neutral software foundations have a habit of developing hardware preferences once their acquirers need to differentiate silicon.”</p>



<p><a href="https://www.linkedin.com/in/fvillanustre/" target="_blank" rel="noreferrer noopener">Flavio Villanustre</a>, CISO for the LexisNexis Risk Solutions Group, provided a different perspective. </p>



<p>“I think it’s important to clarify that Modular is behind the Mojo programming language, which provides an abstraction layer for AI models, enabling them to run across different hardware architectures,” he said. “In the traditional approach, if you code an AI stack on Python or C and target a particular hardware architecture, such as X86, Nvidia GPU, AMD GPU, or TPU, you will need to rewrite a significant portion of that to run it on a different architecture. With Mojo, you code it once and it runs everywhere, even on hybrid systems composed of different hardware architectures.”</p>



<p>And, he said, “if you now consider the fact that Qualcomm owns intellectual property and manufacturing across different hardware architectures, both CPU and GPU, this acquisition could offer their customers significant lift. I see this as Qualcomm buying abstraction that allows them to provide diverse hardware offerings and still offer their customers full code reuse across their entire CPU/GPU/TPU/NPU portfolio.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Demystifying StartupWMClass :: Terminal Thoughts]]></title>
<description><![CDATA[As the maintainer of Plank Reloaded, the most common bug report I get is "this app has the wrong icon." It's almost never the dock, it's a broken StartupWMClass in the app's .desktop file. So I wrote up how to find the right value on X11, Wayland, and KDE, and why deleting the line often fixes it...]]></description>
<link>https://tsecurity.de/de/3620068/linux-tipps/demystifying-startupwmclass-terminal-thoughts/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3620068/linux-tipps/demystifying-startupwmclass-terminal-thoughts/</guid>
<pubDate>Wed, 24 Jun 2026 04:40:03 +0200</pubDate>
<category>🐧 Linux Tipps</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<!-- SC_OFF --><div class="md"><p>As the maintainer of Plank Reloaded, the most common bug report I get is "this app has the wrong icon." It's almost never the dock, it's a broken StartupWMClass in the app's .desktop file. So I wrote up how to find the right value on X11, Wayland, and KDE, and why deleting the line often fixes it.</p> </div><!-- SC_ON -->   submitted by   <a href="https://www.reddit.com/user/zquestz"> /u/zquestz </a> <br> <span><a href="https://thoughts.greyh.at/posts/startup-wm-class/">[link]</a></span>   <span><a href="https://www.reddit.com/r/linux/comments/1udycay/demystifying_startupwmclass_terminal_thoughts/">[comments]</a></span>]]></content:encoded>
</item>
<item>
<title><![CDATA[2026 EuroLLVM - Lighthouse: infrastructure for end-to-end MLIR-compilers and testing]]></title>
<description><![CDATA[Author: LLVM - Bewertung: 1x - Views:4 2026 EuroLLVM Developers' Meeting
https://llvm.org/devmtg/2026-04/
------
Title: Lighthouse: infrastructure for end-to-end MLIR-compilers and testing
Speaker: Renato Golin
------
Slides:  https://llvm.org/devmtg/2026-04/slides/technical_talk/technical_talk_g...]]></description>
<link>https://tsecurity.de/de/3620042/it-security-video/2026-eurollvm-lighthouse-infrastructure-for-end-to-end-mlir-compilers-and-testing/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3620042/it-security-video/2026-eurollvm-lighthouse-infrastructure-for-end-to-end-mlir-compilers-and-testing/</guid>
<pubDate>Wed, 24 Jun 2026 04:03:37 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: LLVM - Bewertung: 1x - Views:4 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/b0b7DwP4Usw?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>2026 EuroLLVM Developers' Meeting<br />
https://llvm.org/devmtg/2026-04/<br />
------<br />
Title: Lighthouse: infrastructure for end-to-end MLIR-compilers and testing<br />
Speaker: Renato Golin<br />
------<br />
Slides:  https://llvm.org/devmtg/2026-04/slides/technical_talk/technical_talk_golin.pdf<br />
-----<br />
Last year, a new project was added to the LLVM family: Lighthouse. Its main purpose is to guide the development and testing of MLIR based compilers. Like the LLVM test-suite, it should be a common ground for validating upstream assumptions about code, IR, dialects. At the same time, it enables building specific compilers in minutes, using the evolving Python API and MLIR's Python bindings. In this talk, we'll show the project main structure, including its components and how to use them to build a simple compiler. We'll then show the infrastructure that uses those components to validate assumptions in MLIR (canonical forms, invariants, applicability of transforms and passes, correctness tests, etc), and how you can create your own on top of that. Finally, we'll provide a number of pipeline examples, going from generic PyTorch models to performant execution on various targets.<br />
-----<br />
Videos Edited by Bash Films: http://www.BashFilms.com<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[2026 EuroLLVM - HIVM: MLIR Dialect Stack for Ascend NPU Compilation]]></title>
<description><![CDATA[Author: LLVM - Bewertung: 1x - Views:4 2026 EuroLLVM Developers' Meeting
https://llvm.org/devmtg/2026-04/
------
Title: HIVM: MLIR Dialect Stack for Ascend NPU Compilation
Speaker: Hugo Trachino
------
Slides:  https://llvm.org/devmtg/2026-04/slides/tutorial/tutorial_tarasov.pdf
-----
Huawei Asce...]]></description>
<link>https://tsecurity.de/de/3619947/it-security-video/2026-eurollvm-hivm-mlir-dialect-stack-for-ascend-npu-compilation/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3619947/it-security-video/2026-eurollvm-hivm-mlir-dialect-stack-for-ascend-npu-compilation/</guid>
<pubDate>Wed, 24 Jun 2026 03:02:42 +0200</pubDate>
<category>🎥 IT Security Video</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Author: LLVM - Bewertung: 1x - Views:4 <br/></p><p><iframe id="ytplayer" loading="lazy" type="text/html" width="100%" height="auto" src="https://www.youtube.com/embed/E0ZsSSy01Q0?autoplay=1&origin=http://tsecurity.de" frameborder="0"></iframe></p><p>2026 EuroLLVM Developers' Meeting<br />
https://llvm.org/devmtg/2026-04/<br />
------<br />
Title: HIVM: MLIR Dialect Stack for Ascend NPU Compilation<br />
Speaker: Hugo Trachino<br />
------<br />
Slides:  https://llvm.org/devmtg/2026-04/slides/tutorial/tutorial_tarasov.pdf<br />
-----<br />
Huawei Ascend NPUs combine DaVinci AI cores with a rich memory/synchronization hierarchy and, on newer generations, a SIMD+SIMT execution model, making performance-oriented compilation challenging. We present HIVM, an open-source family of MLIR dialects that lowers PyTorch/Inductor - Triton - MLIR (HIVM) - LLVM IR, enabling Ascend-specific optimizations such as layout assignment/propagation, vector intrinsic selection/legalization, and explicit DMA/transfer scheduling with synchronization. The pipeline ultimately targets the BiSheng LLVM-based backend to produce executable code for Ascend chips. The talk walks step-by-step through the key IR levels and transformation passes, serving as a practical baseline for developers building MLIR toolchains for Ascend.<br />
-----<br />
Videos Edited by Bash Films: http://www.BashFilms.com<br/></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Why open infrastructure will define the AI era]]></title>
<description><![CDATA[A new form of vendor lock-in is here. And it’s not proprietary languages or rigid enterprise software suites — it’s something more fundamental. It’s the very thing that writes the code.



JetBrains Research found that 74% of developers worldwide use AI tools. Claude Code, available only since Ma...]]></description>
<link>https://tsecurity.de/de/3614974/ai-nachrichten/why-open-infrastructure-will-define-the-ai-era/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3614974/ai-nachrichten/why-open-infrastructure-will-define-the-ai-era/</guid>
<pubDate>Mon, 22 Jun 2026 11:19:12 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>A new form of vendor lock-in is here. And it’s not proprietary languages or rigid enterprise software suites — it’s something more fundamental. It’s the very thing that writes the code.</p>



<p><a href="https://blog.jetbrains.com/research/2026/04/which-ai-coding-tools-do-developers-actually-use-at-work/">JetBrains Research</a> found that 74% of developers worldwide use AI tools. <a href="https://www.infoworld.com/article/4136718/claude-code-is-blowing-me-away.html">Claude Code</a>, available only since May 2025, is now the most popular AI coding tool, followed by <a href="https://www.infoworld.com/article/3829347/review-gemini-code-assist-is-good-at-coding.html">Gemini Code Assist</a> and <a href="https://www.infoworld.com/article/3609013/github-copilot-everything-you-need-to-know.html">GitHub Copilot</a>, according to Jellyfish’s 2026 <a href="https://jellyfish.co/resources/2026-state-of-engineering-management-report/">State of Engineering Management Report</a>.</p>



<p>The latter study also found that 91% of developers say their productivity has increased in the past 12 months. As coding output <a href="https://leaddev.com/ai/openai-says-there-are-easily-1000x-engineers-now">expectations are rewritten daily</a>, the engineering world is becoming heavily reliant on paid external AI services.</p>



<p><a href="https://www.linkedin.com/posts/markwoneill_how-to-optimize-token-consumption-for-ai-activity-7458329480994992128-AT9m?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAA-8zTABlsmtYe-zC-Uf5z3oD5nm6qXDVVo">Gartner predicts</a> that by 2028 spending on AI coding tokens could exceed developer salaries. Yet, <a href="https://www.infoworld.com/article/4183060/the-tokenmaxxing-backlash-is-coming.html">tokenmaxxing</a> while <a href="https://www.infoworld.com/article/4166817/vibe-coding-or-spec-driven-development-how-to-choose.html">vibe coding</a> through a vendor’s cloud-based API feels like a far cry from the open foundations of free programming languages and open models, which many of today’s AI platforms now abstract.</p>



<p>“Open infrastructure will be the backbone of the AI era,” says <a href="https://www.linkedin.com/in/farkasp/">Peter Farkas</a>, CEO of <a href="https://www.percona.com/">Percona</a>, a provider of open-source database solutions. “Right now, too many companies are building their entire AI strategy on top of proprietary platforms because the convenience is seductive.”</p>



<p>“It’s ‘three clicks’ to stand up a database or an AI service in a hyperscaler, and that convenience blinds people to the lock-in they’re signing up for,” he adds. “As AI workloads mature, organizations will realize that depending on one vendor for their data, models, runtime, and pricing is not a strategy.”</p>



<p><a href="https://www.infoworld.com/article/3973969/knowing-when-to-use-ai-coding-assistants.html">AI-assisted coding</a> is democratizing software engineering for non-engineers and <a href="https://leaddev.com/ai/ai-doesnt-create-great-developers-it-amplifies-them">accelerating top performers</a>. But if teams are always working within the confines of how one platform thinks the world should work, it could create locked-in toolsets at scale. And as <a href="https://techcrunch.com/2025/12/29/2025-was-the-year-ai-got-a-vibe-check/">AI platform costs rise</a>, a fundamental question arises: will software developers consume AI on their own terms, or on someone else’s?</p>



<p>There’s a strong case that the long-term winners in tech will be built on open-source standards and foundations, similar to the history of cloud-native computing and the internet itself.</p>



<p>“Open always wins,” says <a href="https://www.linkedin.com/in/brianalvey/">Brian Alvey</a>, CTO at <a href="https://wpvip.com/">WordPress VIP</a>, a managed WordPress hosting platform. “Not because it’s a fancy ideology, but because it gives you total freedom to adapt, evolve, and stay in control.”</p>



<p>Open infrastructure avoids a future where developers perpetually rent. “For AI to be useful to people at large, it can’t be something you’re paying rent for the rest of your life,” says <a href="https://www.linkedin.com/in/maniksurtani/">Manik Surtani</a>, CTO and co-founder of the <a href="https://aaif.io/">Agentic AI Foundation</a> (AAIF), a vendor-neutral home for open-source agentic AI technologies. “And it can’t be concentrated in one particular corporation or a small handful of corporations, because we know how that goes.”</p>



<h2 class="wp-block-heading">Pricey, closed, proprietary AI</h2>



<p>AI development today is traveling two parallel paths. On one path, <a href="https://leaddev.com/technical-direction/be-careful-open-source-ai">open-source AI</a> is thriving and fueling tremendous growth in the number and variety of AI models and tools. Just take the thousands of open-weight models on <a href="https://huggingface.co/">HuggingFace</a>, the community around the <a href="https://openclaw.ai/">OpenClaw</a> AI agent, or the many academic institutions publishing <a href="https://thenewstack.io/llms-can-now-trace-their-outputs-to-specific-training-data/">new breakthroughs</a>.</p>



<p>“Open-source models and tooling are hot on the heels of state-of-the-art, with interesting and boundary-pushing work being shared by labs and researchers across the world,” says Austin Parker, director of AI strategy at <a href="https://www.honeycomb.io/">Honeycomb</a>, an observability platform provider, citing frontier open-source models like <a href="https://mistral.ai/">Mistral</a>, <a href="https://github.com/deepseek-ai/deepseek-v3">DeepSeek</a>, and <a href="https://allenai.org/olmo2">Ai2’s OLMo</a> as examples.</p>



<p>Others agree. “There’s unprecedented openness at the model and tooling layer, with open-source models, frameworks, and orchestration advancing at remarkable speed,” says <a href="https://www.linkedin.com/in/markcollier/">Mark Collier</a>, general manager of AI and infrastructure at the <a href="https://www.linuxfoundation.org/">Linux Foundation</a>.</p>



<p>On the other path, we’re seeing heavy reliance on proprietary AI systems controlled by Anthropic, Cursor, Google, Microsoft, OpenAI, and others. As Collier says, “Many platforms are wrapping those open components in closed, opinionated interfaces that trade short-term speed for long-term constraints.”</p>



<p><a href="https://www.infoworld.com/article/2262355/what-is-open-source-software-open-source-and-foss-explained.html">Open source</a> and the AI tooling market don’t always mix well. LangChain’s <a href="https://github.com/langchain-ai/open-agent-platform">Open Agent Platform</a>, for instance, was open-sourced to much fanfare in 2025, but by 2026 had been deprecated, with the repository now recommending fully managed alternatives.</p>



<p>For <a href="https://www.linkedin.com/in/shaposhnik/">Roman Shaposhnik</a>, co-founder and CTO of <a href="https://nekko.ai/">Ainekko</a>, provider of an open-source, composable AI stack, the current AI platform landscape is reminiscent of <a href="https://devops.com/demystifying-the-low-code-category/">low-code and no-code platforms</a>, which promised democratization of software development but often <a href="https://www.infoworld.com/article/3958483/7-reasons-low-code-and-no-code-tools-fail-to-deliver.html">failed to deliver</a>, becoming synonymous with platform lock-in and inflexibility.</p>



<p>“Honestly, it feels familiar,” Shaposhnik says. “We have incredibly powerful AI tools right now, but most of them come bundled as tightly controlled platforms.” This is a risk for AI, he says, because the infrastructure, models, and hardware are tightly coupled. “If those layers are closed, you lose flexibility fast.”</p>



<p>Some abstractions that sit on top of models, like routing and agent frameworks, tend to be tightly coupled and optimized for certain models. Other platforms take the walled garden concept quite literally. Anthropic, for instance, has repeatedly made headlines for blocking access to its Claude models over <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/anthropic-nuked-a-companys-access-to-claude-stopping-60-employees-dead-in-their-tracks-support-via-google-form-is-the-only-recourse-for-vague-usage-policy-violation">vague policy violations</a>. The company recently shut off <a href="https://venturebeat.com/technology/anthropic-cracks-down-on-unauthorized-claude-usage-by-third-party-harnesses">competitor xAI’s use</a> and <a href="https://thenewstack.io/anthropic-agent-sdk-confusion/">stonewalled OpenCode</a>, drawing community backlash.</p>



<p>Moves toward increasingly closed systems don’t bode well for an AI economy already built on shaky economics. As <a href="https://www.linkedin.com/in/vikramsrivats/">Vikram Srivats</a>, head of product experience at <a href="https://www.wavemaker.com/">WaveMaker</a>, provider of an agentic application development platform, adds, “Given the unit economics of AI tooling and pace of accelerated change to keep up, it seems obvious that some will evolve to more of a closed system to be able to monetize and gain ROI.”</p>



<h2 class="wp-block-heading">Why openness matters in the AI era</h2>



<p>Reliance on proprietary AI platforms can create long-term operational dependencies. As systems become less interoperable, organizations may be forced to standardize on a single stack across data pipelines, models, and decision logic, says the Linux Foundation’s Collier.</p>



<p>“As infrastructure consolidates, enterprises become more exposed when platforms change direction, raise prices, or fall behind technically,” he says. “If you can’t change platforms without re-architecting your AI systems, you’ve already given up too much control.”</p>



<p>“When you build on someone else’s platform, you have to live by their rules and those rules always change,” adds WordPress VIP’s Alvey. “We’ve all seen this before, businesses wasting time and money building to serve Google, Facebook, YouTube, and the App Store, instead of building to serve their customers.”</p>



<p><a href="https://www.infoworld.com/article/2337012/get-used-to-cloud-vendor-lock-in.html">Platform lock-in</a> can also create direct business risk. As Ainekko’s Shaposhnik says, “It usually shows up as higher costs, fragile systems, and growing risk when it’s time to change direction.”</p>



<p>At Ainekko, an internal group called the <a href="https://www.eejournal.com/article/do-you-want-to-be-an-ai-plumber/">AI Plumbers</a> focuses on back-end AI infrastructure like inference, scheduling, memory, and hardware integration. “Their view is simple,” says Shaposhnik. “If those layers are closed, everything above them becomes fragile.”</p>



<p>Open standards, interfaces, and infrastructure provide a necessary hedge against closed systems to prevent this sort of fragility. “In the AI era, open infrastructure gives enterprises control, portability, and choice at exactly the time they need it most,” says Percona’s Farkas.</p>



<p>It can cost upwards of $100,000 to migrate enterprise software, <a href="https://cloudaware.com/blog/cloud-migration-costs/">according to Cloudaware</a>, making portability a major enterprise concern. From this perspective, procuring closed systems can become a costly architectural dependency.</p>



<p>Others argue that openness is a critical hedge against vendor concentration risks at large, especially if AI replaces human labor en masse. “If all of that economic value is now being concentrated in the hands of one or two companies,” says the AAIF’s Surtani, “that’s an order of magnitude bigger problem than we’ve seen in any other wave of computing.”</p>



<p>Instead, open foundations allow adaptability to evolving conditions so enterprises can swap out models, agents, data, hardware, and orchestration, as needed. “Open standards let those components change independently without breaking the system,” says Collier. </p>



<p>Openness can also help future-proof businesses against economic upheaval. “Open everything will help build a cushion for businesses and users to survive and thrive after the almost-certain correction in the current hype cycle,” says WaveMaker’s Srivats.</p>



<h2 class="wp-block-heading">Momentum toward open AI infrastructure</h2>



<p>At the industry level, momentum toward open AI infrastructure is growing. The <a href="https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation">establishment</a> of the <a href="https://aaif.io/author/aaif/">Agentic AI Foundation</a>, Anthropic’s donation of <a href="https://www.infoworld.com/article/4029634/what-is-model-context-protocol-how-mcp-bridges-ai-and-external-services.html">Model Context Protocol</a> (MCP), and Block’s donation of its <a href="https://aaif.io/projects/goose/">Goose agent</a> are significant ecosystem-wide moves toward openness. Other advances include the donation of <a href="https://thenewstack.io/llm-d-cncf-kubernetes-inference/">llm-d</a>, a <a href="https://www.infoworld.com/article/2266945/what-is-kubernetes-scalable-cloud-native-applications.html">Kubernetes</a> framework for LLM inference, to the Cloud Native Computing Foundation (CNCF).</p>



<p>For Parker, donations like this help ensure long-term support and care. “Open standards aren’t just the foundation of the internet, they’re the foundation of the AI space,” he says. “I predict that we’ll see these practices continue, especially as enterprise adoption increases in earnest,” he adds.</p>



<p>Still, some question whether this level of stewardship is enough for a rapidly evolving ecosystem. “The internet benefited early on from groups that helped keep vendors aligned,” says Shaposhnik. “In AI infrastructure, we don’t really have that yet.”</p>



<p>“All of us open source veterans are hopeful,” he says, “but we also need to adapt to this new reality in what we do regarding AI infrastructure.”</p>



<p>Beyond industry governing bodies, companies themselves are also spearheading open AI initiatives. Warp, an agentic development environment, recently <a href="https://thenewstack.io/warp-open-source-client/">went open source</a> amid closed-source rivals. Arcade.dev, meanwhile, is pushing an open-source <a href="https://www.arcade.dev/blog/agent-library/">Agent Library</a> for agentic memory.</p>



<h2 class="wp-block-heading">Where openness matters most in the AI stack</h2>



<p>While AI infrastructure can be open in many ways, a few layers stand out as especially important. First is the openness of the model itself. “Open-source models must be the foundation of future trust and value,” says WaveMaker’s Srivats.</p>



<p>“The forms of open infrastructure that reduce integration friction and accelerate adoption stand out,” adds <a href="https://www.linkedin.com/in/neeraj-abhyankar-9040141/">Neeraj Abhyankar</a>, VP of data and AI at <a href="https://www.rsystems.com/">R Systems</a>, a global digital solutions provider. For him, open model representation formats, open orchestration and execution layers, open agentic protocols, and open governance and metadata standards are all essential for enterprise flexibility.</p>



<p>Others place more value on the connective tissue between AI components. “The most important forms of open infrastructure are the ones that connect systems together,” says Collier. “That includes open APIs, metadata standards, identity and policy frameworks, and protocols for how models and agents communicate.” </p>



<p>Arguably, <a href="https://www.infoworld.com/article/4096223/10-mcp-servers-for-devops.html">MCP</a> has become the connective tissue between AI agents and the <a href="https://thenewstack.io/how-to-prepare-your-api-for-ai-agents/">broader API ecosystem</a>. “If we get MCP right we unlock the same level of interoperability between entities on the web and models driving them as we came to enjoy during the Web 2.0 era and the API-first boom,” says Shaposhnik. “If we don’t we risk massive proprietary lock-ins.”</p>



<p>Parker agrees that open protocols will underlie future AI progress. “We’ll see continued development and progress on AI agents which will rely on protocols like MCP and ACP [<a href="https://www.infoworld.com/article/4007686/a-developers-guide-to-ai-protocols-mcp-a2a-and-acp.html">Agent Client Protocol</a>] to interoperate with various clients and each other,” he says. Yet a gap remains around API conventions for models. “It would be nice if we could get a commitment from model providers to use a standard here.”</p>



<p>For the AAIF’s Surtani, opening up the protocol layer is the most important aspect. “I think it’s really important for interoperability, for choice,” he says. “It means you can bring your own agent, you can bring your own framework, you can bring your own harness, and pick what model you want.”</p>



<p>Open standards may also play a significant role within <a href="https://www.infoworld.com/article/4117620/edge-ai-the-future-of-ai-inference-is-smarter-local-compute.html">inference architecture</a>. “As AI expands to the edge, developers need visibility into how models run, how memory is used, and how performance scales,” says Shaposhnik. Open systems could make it easier to optimize, debug, and adapt while helping enterprises avoid observability fragmentation.</p>



<p>Lastly, <a href="https://www.infoworld.com/article/3498485/the-future-of-kubernetes-and-cloud-infrastructure.html">cloud-native architectural standards</a> are a key ingredient for open AI infrastructure. “We’re seeing Kubernetes become the missing link for people who want the hyperscaler-style convenience without hyperscaler lock-in,” says Percona’s Farkas. For him, Kubernetes has become the de facto hybrid enterprise deployment option for data, workloads, and AI components.</p>



<h2 class="wp-block-heading">History repeats itself</h2>



<p>The <a href="https://opensource.org/blog/the-2026-state-of-open-source-report">2026 State of Open Source Report</a> found avoiding vendor lock-in to be the primary driver of open source adoption. But beyond being a strategic decision for a single company, open infrastructure provides a layer for entire industries to be built upon.</p>



<p>Arguably, the internet itself is evidence of this, where groups like the <a href="https://www.ietf.org/">IETF</a> and the <a href="https://www.ieee.org/">IEEE</a> were instrumental in defining the fundamental protocols. “Without open protocols we would’ve been in telco hell and without phenomenons like Google or Facebook,” says Shaposhnik.</p>



<p>Or, take the <a href="https://www.infoworld.com/article/2335646/thirty-two-years-of-linux-and-its-community.html">history of Linux</a> as a parallel. “Linux became the default operating system because it offered a common, vendor-neutral foundation that everyone could build on,” says Collier. “In the AI era, open infrastructure will define the layers that organizations rely on for long-term continuity.”</p>



<p>At the infrastructure level, open standards have repeatedly underpinned major platform shifts, from <a href="https://www.infoworld.com/article/2253801/what-is-docker-the-spark-for-the-container-revolution.html">Docker</a> to <a href="https://www.infoworld.com/article/3812622/will-kubernetes-ever-get-easier.html">Kubernetes</a>. The question now is whether AI will develop a similarly durable standards layer.</p>



<p>For Parker, it’s too early to say, but the current growth of AI mirrors the early cloud. “Remember that it took many years before we saw the development and popularization of the open source cloud-native ecosystem,” he says. “I think it would be a mistake to extrapolate from the current trajectory towards a closed, proprietary future.”</p>



<p>Others agree the future must be rooted in openness. “I see open infrastructure becoming the foundation of enterprise AI,” says R Systems’s Abhyankar. “As systems become more distributed and agent‑driven, closed ecosystems simply won’t scale.”</p>



<p>The groundwork is being laid through open agentic protocols, open frameworks, and industry support intended to reduce fragmentation around proprietary standards.</p>



<p>“Ironically, the AI movement has mostly seemed to learn from the mistakes of the past and is starting off on a more open foot,” says Parker. “Over time, I believe we’ll see innovation and openness thrive.”</p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[Write cleaner and faster Python code]]></title>
<description><![CDATA[Meta’s long-awaited Pyrefly linter is out in a 1.0 version, and the forthcoming Python 3.15 has a super-efficient sampling profiler. Plus we have a comprehensive rundown of Python’s indispensable virtual environments — and a warning about a novel breed of malware that exploits Python’s package ec...]]></description>
<link>https://tsecurity.de/de/3609845/ai-nachrichten/write-cleaner-and-faster-python-code/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3609845/ai-nachrichten/write-cleaner-and-faster-python-code/</guid>
<pubDate>Fri, 19 Jun 2026 11:18:47 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
		<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<p>Meta’s long-awaited Pyrefly linter is out in a 1.0 version, and the forthcoming <a href="https://www.infoworld.com/article/4166693/the-best-new-features-in-python-3-15.html" data-type="link" data-id="https://www.infoworld.com/article/4166693/the-best-new-features-in-python-3-15.html">Python 3.15</a> has a super-efficient sampling profiler. Plus we have a comprehensive rundown of Python’s indispensable virtual environments — and a warning about a novel breed of malware that exploits Python’s package ecosystem.</p>



<h2 class="wp-block-heading">Top picks for Python readers on InfoWorld</h2>



<p><a href="https://www.infoworld.com/article/2260103/how-to-use-virtual-environments-in-python.html" data-type="link" data-id="https://www.infoworld.com/article/2260103/how-to-use-virtual-environments-in-python.html">How to use virtual environments in Python</a><br>Isolate and protect your Python projects from each other, and empower them to do more, with virtual environments and their native-to-Python tooling.</p>



<p><a href="https://www.infoworld.com/article/4179383/pyrefly-1-0-a-fast-forward-looking-python-linter.html" data-type="link" data-id="https://www.infoworld.com/article/4179383/pyrefly-1-0-a-fast-forward-looking-python-linter.html">Pyrefly 1.0: A fast, forward-looking Python linter</a><br>The first full release of Meta’s long-awaited linting and type checking tool for Python delivers speed and offers advanced features for type-checking PyTorch and Django projects.</p>



<p><a href="https://www.infoworld.com/video/4085906/hands-on-with-the-new-sampling-profiler-in-python-3-15.html" data-type="link" data-id="https://www.infoworld.com/video/4085906/hands-on-with-the-new-sampling-profiler-in-python-3-15.html">Hands-on with the new sampling profiler in Python 3.15</a><br>Among Python 3.15’s best new features is a sampling profiler, for instrumenting your code and finding its bottlenecks with a minimum of performance impact or fuss. See up-close how it works.</p>



<p><a href="https://www.infoworld.com/article/4182692/meet-hades-the-malware-that-lies-to-ai-security-agents.html" data-type="link" data-id="https://www.infoworld.com/article/4182692/meet-hades-the-malware-that-lies-to-ai-security-agents.html">All about Hades, the supply-chain malware that hides in Python packages</a><br>It hides in Python packages. It replicates itself across systems. It fools LLM-based code analysis tools into ignoring it. And there may be a lot more like it to come.</p>



<h2 class="wp-block-heading">More good reads and Python updates elsewhere</h2>



<p><a href="https://discuss.python.org/t/an-announcement-from-the-steering-council-regarding-the-jit-project/107638" data-type="link" data-id="https://discuss.python.org/t/an-announcement-from-the-steering-council-regarding-the-jit-project/107638">Python Steering Council calls for temporary pause on JIT project</a><br>The requested pause stays in place until a proper Standards Track PEP lands for the experimental JIT (just-in-time) compiler, the better to describe how the JIT will be a formal and supported part of Python.</p>



<p><a href="https://blog.pyodide.org/posts/314-release" data-type="link" data-id="https://blog.pyodide.org/posts/314-release">Pyodide 314.0: Pyodide packages on PyPI</a><br>Thanks to PEP 783, Python packages built with Pyodide (Python ported to WebAssembly) can be installed straight from PyPI instead of through Pyodide — another step closer to Py-on-Wasm becoming an everyday thing.</p>



<p><a href="https://theconsensus.dev/p/2026/06/06/python-3-14-garbage-collection-rigamarole.html" data-type="link" data-id="https://theconsensus.dev/p/2026/06/06/python-3-14-garbage-collection-rigamarole.html">All about that Python 3.14 garbage collection rigmarole</a><br>A new garbage collector introduced in Python 3.14 was yanked at the last minute due to reports of higher memory usage. Here’s a deep dive into what changed for the worse and why.</p>



<p><a href="https://pyrefly.org/blog/too-many-type-checkers" data-type="link" data-id="https://pyrefly.org/blog/too-many-type-checkers">Are you really expected to run five type checkers now?</a><br>No, but you should keep your options open. This blog post from a Pyrefly contributor recommends choosing one of the major offerings (Mypy, Pyrefly, Pyright, ty, Zuban, etc.), but also getting to know the others too. </p>
</div></div></div></div>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5e0379393cc64f15209ec20e20bc2eddc5ed2a42: Add autograd support to TokenSwitch dispatch and combine (#181314)]]></title>
<description><![CDATA[dispatch and combine are adjoint operations: backward of dispatch calls
combine, and backward of combine calls dispatch. Single dispatch/combine
public API following nn.Module/torch.matmul conventions:

With out=(out_tokens, out_weights, out_idx): writes to caller-supplied
buffers and returns the...]]></description>
<link>https://tsecurity.de/de/3600747/downloads/trunk5e0379393cc64f15209ec20e20bc2eddc5ed2a42-add-autograd-support-to-tokenswitch-dispatch-and-combine-181314/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600747/downloads/trunk5e0379393cc64f15209ec20e20bc2eddc5ed2a42-add-autograd-support-to-tokenswitch-dispatch-and-combine-181314/</guid>
<pubDate>Tue, 16 Jun 2026 07:48:24 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>dispatch</code> and <code>combine</code> are adjoint operations: backward of dispatch calls<br>
combine, and backward of combine calls dispatch. Single dispatch/combine<br>
public API following <code>nn.Module</code>/<code>torch.matmul</code> conventions:</p>
<ul>
<li>With <code>out=(out_tokens, out_weights, out_idx)</code>: writes to caller-supplied<br>
buffers and returns them; no autograd (efficient buffer reuse).</li>
<li>Without <code>out</code>: allocates buffers internally, returns them with full<br>
autograd support via <code>_DispatchAutograd</code> / <code>_CombineAutograd</code>.</li>
</ul>
<p>Subclasses implement <code>_dispatch</code>/<code>_combine</code> as the raw buffer-writing<br>
primitives and inherit both modes from the base class. <code>topk_weights</code><br>
receives no gradient (routing metadata at a different byte width).</p>
<p>Authored with Claude.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4319720108" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/181314" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/181314/hovercard" href="https://github.com/pytorch/pytorch/pull/181314">#181314</a><br>
Approved by: <a href="https://github.com/kapilsh">https://github.com/kapilsh</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4162632128" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/178712" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/178712/hovercard" href="https://github.com/pytorch/pytorch/pull/178712">#178712</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/6ec77b0590f23dbc6faef91451971b14598e8859]]></title>
<description><![CDATA[[xpu][fix] Include kernel_compile_result.h in aoti xpu.h header (#187…]]></description>
<link>https://tsecurity.de/de/3600673/downloads/trunk6ec77b0590f23dbc6faef91451971b14598e8859/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600673/downloads/trunk6ec77b0590f23dbc6faef91451971b14598e8859/</guid>
<pubDate>Tue, 16 Jun 2026 07:05:09 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[xpu][fix] Include kernel_compile_result.h in aoti xpu.h header (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="186401214" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/187/hovercard" href="https://github.com/pytorch/pytorch/issues/187">#187</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/421abac1a6353b99c40bc4bdb69ffb028f098df9: Revert XPU device-wide synchronization (#187306)]]></title>
<description><![CDATA[Motivation
This PR reverts #182630 to fix #187277
Root cause is ext_oneapi_wait_and_throw will introduce SYCL Graph to sync an invalid queue.
Additional Context
This needs to be cherry-picked to the release branch.
Why isn't it captured on CI?
This issue is only found on BMG (Xe2), and the curren...]]></description>
<link>https://tsecurity.de/de/3600670/downloads/trunk421abac1a6353b99c40bc4bdb69ffb028f098df9-revert-xpu-device-wide-synchronization-187306/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600670/downloads/trunk421abac1a6353b99c40bc4bdb69ffb028f098df9-revert-xpu-device-wide-synchronization-187306/</guid>
<pubDate>Tue, 16 Jun 2026 07:05:04 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Motivation</h1>
<p>This PR reverts <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4390149243" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/182630" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/182630/hovercard" href="https://github.com/pytorch/pytorch/pull/182630">#182630</a> to fix <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4659863803" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187277" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/187277/hovercard" href="https://github.com/pytorch/pytorch/issues/187277">#187277</a></p>
<p>Root cause is <code>ext_oneapi_wait_and_throw</code> will introduce SYCL Graph to sync an invalid queue.</p>
<h1>Additional Context</h1>
<p>This needs to be cherry-picked to the release branch.</p>
<p>Why isn't it captured on CI?<br>
This issue is only found on BMG (Xe2), and the current CI is on Data Center GPU (Xe).<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4662552848" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187306" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187306/hovercard" href="https://github.com/pytorch/pytorch/pull/187306">#187306</a><br>
Approved by: <a href="https://github.com/EikanWang">https://github.com/EikanWang</a>, <a href="https://github.com/atalman">https://github.com/atalman</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4314891831" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/181233" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/181233/hovercard" href="https://github.com/pytorch/pytorch/pull/181233">#181233</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4645771924" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187137" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187137/hovercard" href="https://github.com/pytorch/pytorch/pull/187137">#187137</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4653778237" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187232" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187232/hovercard" href="https://github.com/pytorch/pytorch/pull/187232">#187232</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/e8f1ea4c707a71e1c8dc72516b8c8ea7072dd794: Bump B200 CI containers to Python 3.12 on CUDA 13.0 (#186997)]]></title>
<description><![CDATA[B200 runners used CUDA 13.0 docker images with Python 3.10, which caused install_cutlass_api to skip installation and left NV Universal GEMM smoke tests untested. Add py3.12 cuda13.0 docker image variants and point B200 workflows at them so cutlass_api can install on the runner.
Authored with Cur...]]></description>
<link>https://tsecurity.de/de/3600570/downloads/trunke8f1ea4c707a71e1c8dc72516b8c8ea7072dd794-bump-b200-ci-containers-to-python-312-on-cuda-130-186997/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600570/downloads/trunke8f1ea4c707a71e1c8dc72516b8c8ea7072dd794-bump-b200-ci-containers-to-python-312-on-cuda-130-186997/</guid>
<pubDate>Tue, 16 Jun 2026 05:31:49 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>B200 runners used CUDA 13.0 docker images with Python 3.10, which caused install_cutlass_api to skip installation and left NV Universal GEMM smoke tests untested. Add py3.12 cuda13.0 docker image variants and point B200 workflows at them so cutlass_api can install on the runner.</p>
<p>Authored with Cursor Composer 2.5 (composer-2.5-fast).</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4635568440" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186991" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186991/hovercard" href="https://github.com/pytorch/pytorch/issues/186991">#186991</a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4636082037" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186997" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186997/hovercard" href="https://github.com/pytorch/pytorch/pull/186997">#186997</a><br>
Approved by: <a href="https://github.com/huydhn">https://github.com/huydhn</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/80ff808d457f3c4b1969709d0c51e542261034ae: Revert "Use C++20 std::numbers in c10 MathConstants (#186877)"]]></title>
<description><![CDATA[This reverts commit 8ff6c00.
Reverted #186877 on behalf of https://github.com/huydhn due to Diff reverted internally (comment)]]></description>
<link>https://tsecurity.de/de/3600487/downloads/trunk80ff808d457f3c4b1969709d0c51e542261034ae-revert-use-c-20-stdnumbers-in-c10-mathconstants-186877/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600487/downloads/trunk80ff808d457f3c4b1969709d0c51e542261034ae-revert-use-c-20-stdnumbers-in-c10-mathconstants-186877/</guid>
<pubDate>Tue, 16 Jun 2026 04:01:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/8ff6c00bd7c492840847f9205cfdf936fe1584c8/hovercard" href="https://github.com/pytorch/pytorch/commit/8ff6c00bd7c492840847f9205cfdf936fe1584c8"><tt>8ff6c00</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4627310339" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186877" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186877/hovercard" href="https://github.com/pytorch/pytorch/pull/186877">#186877</a> on behalf of <a href="https://github.com/huydhn">https://github.com/huydhn</a> due to Diff reverted internally (<a href="https://github.com/pytorch/pytorch/pull/186877#issuecomment-4714018571" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186877/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/aaa39f0d947d641cc836b34a05d17a67ffcbae73]]></title>
<description><![CDATA[Type annotations (and runtime fallout) for torch/init.py (#186323…]]></description>
<link>https://tsecurity.de/de/3600486/downloads/trunkaaa39f0d947d641cc836b34a05d17a67ffcbae73/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600486/downloads/trunkaaa39f0d947d641cc836b34a05d17a67ffcbae73/</guid>
<pubDate>Tue, 16 Jun 2026 04:01:45 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Type annotations (and runtime fallout) for torch/<strong>init</strong>.py (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594167766" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186323" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186323/hovercard" href="https://github.com/pytorch/pytorch/pull/186323">#186323</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/e3a7019566d552d5166c54abffee2023cef40219]]></title>
<description><![CDATA[[PyTorch][AOTI] Optional pinned async H2D copy for constant loading (…]]></description>
<link>https://tsecurity.de/de/3600480/downloads/trunke3a7019566d552d5166c54abffee2023cef40219/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600480/downloads/trunke3a7019566d552d5166c54abffee2023cef40219/</guid>
<pubDate>Tue, 16 Jun 2026 03:31:36 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[PyTorch][AOTI] Optional pinned async H2D copy for constant loading (…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/torchtitan/187273: [Dynamo] Trace raw unbacked SymInt inputs]]></title>
<description><![CDATA[Non-strict tracing can pass raw unbacked SymInt values into a nested Dynamo trace. FlexAttention hits this when BlockMask.seq_lengths are derived from an unbacked query/key sequence dimension and then passed through block_mask.as_tuple() to the nested flex_attention_hop wrapper.
VariableBuilder p...]]></description>
<link>https://tsecurity.de/de/3600431/downloads/ciflowtorchtitan187273-dynamo-trace-raw-unbacked-symint-inputs/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600431/downloads/ciflowtorchtitan187273-dynamo-trace-raw-unbacked-symint-inputs/</guid>
<pubDate>Tue, 16 Jun 2026 02:22:32 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Non-strict tracing can pass raw unbacked SymInt values into a nested Dynamo trace. FlexAttention hits this when BlockMask.seq_lengths are derived from an unbacked query/key sequence dimension and then passed through block_mask.as_tuple() to the nested flex_attention_hop wrapper.</p>
<p>VariableBuilder previously only handled raw SymInt inputs when they had a guardable hint. Unhinted raw SymInt inputs graph-broke immediately, even if the nested graph only needed to carry them symbolically. This change keeps that support narrow: outside non-strict tracing, raw unbacked SymInt wrapping still graph-breaks. Inside non-strict tracing, Dynamo asks the active ShapeEnv to transfer the foreign unbacked SymInt as a ShapeEnv-owned input, reuses it for repeated occurrences of the same source expression, and preserves transferred range/source/hint metadata through the existing foreign-ShapeEnv transfer helpers.</p>
<p>The full non-strict FlexAttention repro also exposed a tracing-only issue: same-frame guard validation eagerly evaluates generated guards against fake inputs whose unbacked sizes cannot be concretized. The runtime guards are still produced, but the same-frame sanity check is skipped while inside the non-strict tracing context.</p>
<p>This keeps the important safety property unchanged: if user code branches on a raw unbacked SymInt, Dynamo still raises the normal data-dependent-symbol error instead of specializing on an optimization hint.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4658851878" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187272" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/187272/hovercard" href="https://github.com/pytorch/pytorch/issues/187272">#187272</a></p>
<p>This PR was authored with assistance from an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="PYTORCH_TEST_WITH_DYNAMO=1 python -m pytest test/test_fake_tensor.py::FakeTensorTest::test_cudnn_sdpa_unbacked_batch_dim -q -s"><pre>PYTORCH_TEST_WITH_DYNAMO=1 python -m pytest test/test_fake_tensor.py::FakeTensorTest::test_cudnn_sdpa_unbacked_batch_dim -q -s</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python -m pytest test/dynamo/test_repros.py -q -s -k flex_attention_non_strict_unbacked_sequence_length"><pre>python -m pytest test/dynamo/test_repros.py -q -s -k flex_attention_non_strict_unbacked_sequence_length</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python -m pytest test/dynamo/test_dynamic_spec.py -q -s -k non_strict_raw_unbacked_symint_input_raises_dde_on_branching"><pre>python -m pytest test/dynamo/test_dynamic_spec.py -q -s -k non_strict_raw_unbacked_symint_input_raises_dde_on_branching</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python -m pytest test/functorch/test_control_flow.py::TestControlFlowTraced::test_cond_unbacked_symint_closure -q -s"><pre>python -m pytest test/functorch/test_control_flow.py::TestControlFlowTraced::test_cond_unbacked_symint_closure -q -s</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/test_dynamic_shapes.py -q -s -k "unbacked_hint_overrides_transferred or mixed_static_backed_unbacked"'><pre>python -m pytest test/test_dynamic_shapes.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>unbacked_hint_overrides_transferred or mixed_static_backed_unbacked<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python -m pytest test/inductor/test_unbacked_symints.py::TestUnbackedSymintsCPU -q -s -k override_optimization_hint"><pre>python -m pytest test/inductor/test_unbacked_symints.py::TestUnbackedSymintsCPU -q -s -k override_optimization_hint</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>stack-info: PR: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4658897438" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187273" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187273/hovercard" href="https://github.com/pytorch/pytorch/pull/187273">#187273</a>, branch: sanketpurandare/stack/20</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/rocm-mi300/183838: [Inductor][HOP] Handle unbacked FlexAttention predicates]]></title>
<description><![CDATA[EP-overlap graph chunking traces FlexAttention with unbacked symbolic batch and sequence dimensions. These are valid runtime sizes, but FlexAttention predicate sites forced trace-time decisions that should either fall back to the general kernel or remain represented as deferred symbolic assertion...]]></description>
<link>https://tsecurity.de/de/3600430/downloads/ciflowrocm-mi300183838-inductorhop-handle-unbacked-flexattention-predicates/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600430/downloads/ciflowrocm-mi300183838-inductorhop-handle-unbacked-flexattention-predicates/</guid>
<pubDate>Tue, 16 Jun 2026 02:22:26 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>EP-overlap graph chunking traces FlexAttention with unbacked symbolic batch and sequence dimensions. These are valid runtime sizes, but FlexAttention predicate sites forced trace-time decisions that should either fall back to the general kernel or remain represented as deferred symbolic assertions.</p>
<p>Flex decoding remains an optional optimized implementation. Its predicates use guard_or_false because they are kernel eligibility checks, not user-visible invariants: when ShapeEnv can prove eligibility, decode can be selected; when an unbacked predicate is not provable, the general FlexAttention kernel is used. Failing to prove an optimization must not impose a runtime shape contract on user inputs.</p>
<p>Backward fake propagation preserves the accepted KV-batch-broadcast metadata contract. It reduces grad_key and grad_value back to key/value batch when key/value is provably batch-broadcasted; otherwise it records the non-broadcast invariant with torch._check(Bq == Bkv) and returns unreduced metadata.</p>
<p>Public BlockMask length validation is expressed as two symbolic torch._check predicates: the mask must not be smaller than the query/key lengths, and it must not be larger than those lengths. Concrete mismatches still fail with the existing guidance, while unbacked symbolic lengths become deferred runtime assertions instead of forcing Python to branch on an unbacked expression.</p>
<p>This PR was authored with assistance from an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/inductor/test_flex_attention.py -q -s -k "unbacked_flex_decoding_eligibility_falls_back or backward_fake_symbolic_query_key_batch_non_broadcast or block_mask_vs_sequence_lengths"'><pre>python -m pytest test/inductor/test_flex_attention.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>unbacked_flex_decoding_eligibility_falls_back or backward_fake_symbolic_query_key_batch_non_broadcast or block_mask_vs_sequence_lengths<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/inductor/test_flex_attention.py -q -s -k "mask_mod_handles_symint_addition or mask_mod_handles_derived_symint_closure or symbol_closure_in_score_mod"'><pre>python -m pytest test/inductor/test_flex_attention.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>mask_mod_handles_symint_addition or mask_mod_handles_derived_symint_closure or symbol_closure_in_score_mod<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>stack-info: PR: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851844" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183838" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183838/hovercard" href="https://github.com/pytorch/pytorch/pull/183838">#183838</a>, branch: sanketpurandare/stack/12</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/torchtitan/183838: [Inductor][HOP] Handle unbacked FlexAttention predicates]]></title>
<description><![CDATA[EP-overlap graph chunking traces FlexAttention with unbacked symbolic batch and sequence dimensions. These are valid runtime sizes, but FlexAttention predicate sites forced trace-time decisions that should either fall back to the general kernel or remain represented as deferred symbolic assertion...]]></description>
<link>https://tsecurity.de/de/3600429/downloads/ciflowtorchtitan183838-inductorhop-handle-unbacked-flexattention-predicates/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3600429/downloads/ciflowtorchtitan183838-inductorhop-handle-unbacked-flexattention-predicates/</guid>
<pubDate>Tue, 16 Jun 2026 02:22:18 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>EP-overlap graph chunking traces FlexAttention with unbacked symbolic batch and sequence dimensions. These are valid runtime sizes, but FlexAttention predicate sites forced trace-time decisions that should either fall back to the general kernel or remain represented as deferred symbolic assertions.</p>
<p>Flex decoding remains an optional optimized implementation. Its predicates use guard_or_false because they are kernel eligibility checks, not user-visible invariants: when ShapeEnv can prove eligibility, decode can be selected; when an unbacked predicate is not provable, the general FlexAttention kernel is used. Failing to prove an optimization must not impose a runtime shape contract on user inputs.</p>
<p>Backward fake propagation preserves the accepted KV-batch-broadcast metadata contract. It reduces grad_key and grad_value back to key/value batch when key/value is provably batch-broadcasted; otherwise it records the non-broadcast invariant with torch._check(Bq == Bkv) and returns unreduced metadata.</p>
<p>Public BlockMask length validation is expressed as two symbolic torch._check predicates: the mask must not be smaller than the query/key lengths, and it must not be larger than those lengths. Concrete mismatches still fail with the existing guidance, while unbacked symbolic lengths become deferred runtime assertions instead of forcing Python to branch on an unbacked expression.</p>
<p>This PR was authored with assistance from an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/inductor/test_flex_attention.py -q -s -k "unbacked_flex_decoding_eligibility_falls_back or backward_fake_symbolic_query_key_batch_non_broadcast or block_mask_vs_sequence_lengths"'><pre>python -m pytest test/inductor/test_flex_attention.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>unbacked_flex_decoding_eligibility_falls_back or backward_fake_symbolic_query_key_batch_non_broadcast or block_mask_vs_sequence_lengths<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/inductor/test_flex_attention.py -q -s -k "mask_mod_handles_symint_addition or mask_mod_handles_derived_symint_closure or symbol_closure_in_score_mod"'><pre>python -m pytest test/inductor/test_flex_attention.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>mask_mod_handles_symint_addition or mask_mod_handles_derived_symint_closure or symbol_closure_in_score_mod<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>stack-info: PR: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851844" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183838" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183838/hovercard" href="https://github.com/pytorch/pytorch/pull/183838">#183838</a>, branch: sanketpurandare/stack/12</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/8414e35518f93e2b48ebc594de2faba51dbeca83: Preserve signed zero in FX complex codegen (#185550)]]></title>
<description><![CDATA[FX codegen used complex.repr when rendering complex constants into generated Python source. CPython can print zero-component values such as (-0-1e-28j), -1e-28j, or (1-0j), and parsing that source can flip the sign of a zero component. Dynamo exposes this when it traces a tensor constant containi...]]></description>
<link>https://tsecurity.de/de/3599609/downloads/trunk8414e35518f93e2b48ebc594de2faba51dbeca83-preserve-signed-zero-in-fx-complex-codegen-185550/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3599609/downloads/trunk8414e35518f93e2b48ebc594de2faba51dbeca83-preserve-signed-zero-in-fx-complex-codegen-185550/</guid>
<pubDate>Mon, 15 Jun 2026 17:48:00 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>FX codegen used complex.<strong>repr</strong> when rendering complex constants into generated Python source. CPython can print zero-component values such as (-0-1e-28j), -1e-28j, or (1-0j), and parsing that source can flip the sign of a zero component. Dynamo exposes this when it traces a tensor constant containing -1e-28j through FX: the generated GraphModule reconstructs a value with the wrong signed zero, changing tensor repr and the sign of the zero imaginary result from cos().</p>
<p>Render complex constants with a zero real or imaginary component as complex(real, imag) through the existing recursive argument printer. Float repr preserves -0.0, so those generated constants round-trip signed zero components. Nonzero complex constants keep the existing repr() path to avoid changing the common codegen case.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3074795345" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/153852" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/153852/hovercard" href="https://github.com/pytorch/pytorch/issues/153852">#153852</a></p>
<p>Generated by my agent</p>
<p>Benchmark Results:</p>
<ul>
<li>
<p>FX GraphModule construction with 1000 zero-component complex constants, 80 iterations x 7 repeats, median: main 3.55 ms/graph; this diff 6.77 ms/graph. This is the affected correctness path that now emits complex(real, imag) to preserve signed zero.</p>
</li>
<li>
<p>FX GraphModule construction with 1000 nonzero complex constants, same command shape, median: main 4.29 ms/graph; this diff 3.90 ms/graph. Nonzero constants remain on the existing repr() path; the difference is measurement noise.</p>
</li>
</ul>
<p>Test Plan:</p>
<ul>
<li>
<p>python test/dynamo/test_repros.py ReproTests.test_compile_complex_tensor_constant_signed_zero</p>
</li>
<li>
<p>python test/test_fx.py TestFX.test_complex_constant_codegen_preserves_signed_zero (direct run blocked locally before target test by unrelated torchvision::nms registration failure)</p>
</li>
<li>
<p>targeted TestFX.test_complex_constant_codegen_preserves_signed_zero via a torchvision import stub</p>
</li>
<li>
<p>lintrunner -a</p>
</li>
<li>
<p>git diff --check</p>
</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4542863124" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185550" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185550/hovercard" href="https://github.com/pytorch/pytorch/pull/185550">#185550</a><br>
Approved by: <a href="https://github.com/desertfire">https://github.com/desertfire</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/215923c3574d22ad7335dbef3c6b37902ab4db6f: [MPS] Matmul for strided out errors on mac OS 14/15 (#187255)]]></title>
<description><![CDATA[Putting strided out in torch mm/addmm leads to error on mac OS 14/15. First discovered in gemv PR:
#186927
Pull Request resolved: #187255
Approved by: https://github.com/malfet]]></description>
<link>https://tsecurity.de/de/3599319/downloads/trunk215923c3574d22ad7335dbef3c6b37902ab4db6f-mps-matmul-for-strided-out-errors-on-mac-os-1415-187255/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3599319/downloads/trunk215923c3574d22ad7335dbef3c6b37902ab4db6f-mps-matmul-for-strided-out-errors-on-mac-os-1415-187255/</guid>
<pubDate>Mon, 15 Jun 2026 15:46:44 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Putting strided out in torch mm/addmm leads to error on mac OS 14/15. First discovered in gemv PR:<br>
<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4630664817" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186927" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186927/hovercard" href="https://github.com/pytorch/pytorch/pull/186927">#186927</a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4656659012" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187255" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187255/hovercard" href="https://github.com/pytorch/pytorch/pull/187255">#187255</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/82ec0198889b0936af92ce019f20c34712fcc5e4: Revert "Fix MKLDNN to_dense fake layout handling (#183670)"]]></title>
<description><![CDATA[This reverts commit 523c1d0.
Reverted #183670 on behalf of https://github.com/pytorch-auto-revert due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (comment)]]></description>
<link>https://tsecurity.de/de/3598464/downloads/trunk82ec0198889b0936af92ce019f20c34712fcc5e4-revert-fix-mkldnn-todense-fake-layout-handling-183670/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3598464/downloads/trunk82ec0198889b0936af92ce019f20c34712fcc5e4-revert-fix-mkldnn-todense-fake-layout-handling-183670/</guid>
<pubDate>Mon, 15 Jun 2026 10:16:40 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/523c1d0eed40a1febf33db2a321c8fead524c725/hovercard" href="https://github.com/pytorch/pytorch/commit/523c1d0eed40a1febf33db2a321c8fead524c725"><tt>523c1d0</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4444101223" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183670" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183670/hovercard" href="https://github.com/pytorch/pytorch/pull/183670">#183670</a> on behalf of <a href="https://github.com/pytorch-auto-revert">https://github.com/pytorch-auto-revert</a> due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (<a href="https://github.com/pytorch/pytorch/pull/183670#issuecomment-4705818513" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183670/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/40df7254415ed93bfe83b9d406b44f171ecc5479: Fix flex_attention score_mod with no score gradient (#185991)]]></title>
<description><![CDATA[FlexAttention builds a joint graph for score_mod so the backward template can compute the gradient of the modified scores with respect to the raw attention scores. When score_mod returns a value that is independent of score, AOTAutograd correctly reports no gradient for the score input. The FlexA...]]></description>
<link>https://tsecurity.de/de/3598268/downloads/trunk40df7254415ed93bfe83b9d406b44f171ecc5479-fix-flexattention-scoremod-with-no-score-gradient-185991/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3598268/downloads/trunk40df7254415ed93bfe83b9d406b44f171ecc5479-fix-flexattention-scoremod-with-no-score-gradient-185991/</guid>
<pubDate>Mon, 15 Jun 2026 08:46:57 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>FlexAttention builds a joint graph for score_mod so the backward template can compute the gradient of the modified scores with respect to the raw attention scores. When score_mod returns a value that is independent of score, AOTAutograd correctly reports no gradient for the score input. The FlexAttention lowering did not handle that None result: Inductor still expects a score-gradient subgraph output for the dq/dk matmuls, so backward failed during lowering/kernel generation instead of treating the score gradient as zero.</p>
<p>Materialize a zero score gradient when the joint graph returns None for the score input. Constant joint graphs can lower that zero as a scalar/rank-1 Triton value, so the backward template now broadcasts only those low-rank grad_scores values to the score tile before the matmuls. Rank-2 gradients from normal differentiable score_mod paths skip the extra broadcast add.</p>
<p>I considered only adding a template-side fallback, but that would leave the higher-order op contract ambiguous: the joint graph should always provide a score-gradient value to the backward lowering. Materializing the zero in create_fw_bw_graph fixes that contract, while the template rank guard handles the generated representation needed by Triton.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2794665775" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/145050" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/145050/hovercard" href="https://github.com/pytorch/pytorch/issues/145050">#145050</a></p>
<p>Generated by my agent</p>
<p>Benchmark Results:</p>
<ul>
<li>CUDA fp16 compiled flex_attention backward, B=1 H=1 S=512 D=64, score_mod returns score * 1.1, 5 warmup iterations and 50 CUDA-event timed iterations.</li>
<li>Baseline main median: 0.5343 ms; p10/p90: 0.4784/0.6628 ms.</li>
<li>Patched median: 0.5404 ms; p10/p90: 0.5289/0.5629 ms.</li>
<li>Result: no clear regression beyond run-to-run noise; baseline timings were bimodal in this environment, while patched timings stayed in the same envelope.</li>
</ul>
<p>Test Plan:</p>
<ul>
<li>Reproduced the issue before the fix with a minimal CUDA repro using torch.compile(flex_attention, dynamic=False), score_mod returning q_idx &gt;= kv_idx, and backward failing with InductorError / joint_subgraph_buffer is None.</li>
<li>Verified the repro after the fix: q.grad and k.grad are zero, v.grad is nonzero.</li>
<li>Verified backend="aot_eager" repro after the fix.</li>
<li>Verified differentiable score_mod=score * score smoke test after the fix.</li>
<li>python test/inductor/test_flex_attention.py -k test_score_mod_without_score_gradient</li>
<li>lintrunner -a</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4574804501" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185991" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185991/hovercard" href="https://github.com/pytorch/pytorch/pull/185991">#185991</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/35cf1598c8181f6219f58de8863f0c642efa713a: [inductor] Fix debug sync in GPU cpp wrapper (#184217)]]></title>
<description><![CDATA[Route debug sync emission through wrapper-specific codegen so cpp_wrapper emits checked CUDA/HIP synchronization instead of Python source. This fixes both triton.debug_sync_graph and triton.debug_sync_kernel paths and reenables those config fuzzer combinations.
Fixes #140219
Fixes #140220
Generat...]]></description>
<link>https://tsecurity.de/de/3598267/downloads/trunk35cf1598c8181f6219f58de8863f0c642efa713a-inductor-fix-debug-sync-in-gpu-cpp-wrapper-184217/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3598267/downloads/trunk35cf1598c8181f6219f58de8863f0c642efa713a-inductor-fix-debug-sync-in-gpu-cpp-wrapper-184217/</guid>
<pubDate>Mon, 15 Jun 2026 08:46:56 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Route debug sync emission through wrapper-specific codegen so cpp_wrapper emits checked CUDA/HIP synchronization instead of Python source. This fixes both <code>triton.debug_sync_graph</code> and <code>triton.debug_sync_kernel</code> paths and reenables those config fuzzer combinations.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2646712113" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/140219" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/140219/hovercard" href="https://github.com/pytorch/pytorch/issues/140219">#140219</a><br>
Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2646713280" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/140220" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/140220/hovercard" href="https://github.com/pytorch/pytorch/issues/140220">#140220</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4470161892" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184217" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184217/hovercard" href="https://github.com/pytorch/pytorch/pull/184217">#184217</a><br>
Approved by: <a href="https://github.com/yushangdi">https://github.com/yushangdi</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/889f6eb707c9eb75bcb2f3ca91ad1ed9b24d2446: Revert "`bcomplex32` dtype, attempt 2 (#186928)"]]></title>
<description><![CDATA[This reverts commit d0ef03a.
Reverted #186928 on behalf of https://github.com/wdvr due to See below - failing builds with older cuDNN (comment)]]></description>
<link>https://tsecurity.de/de/3598193/downloads/trunk889f6eb707c9eb75bcb2f3ca91ad1ed9b24d2446-revert-bcomplex32-dtype-attempt-2-186928/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3598193/downloads/trunk889f6eb707c9eb75bcb2f3ca91ad1ed9b24d2446-revert-bcomplex32-dtype-attempt-2-186928/</guid>
<pubDate>Mon, 15 Jun 2026 08:16:50 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/d0ef03a2dc495c5a354ab867dabe5baee1b45dbd/hovercard" href="https://github.com/pytorch/pytorch/commit/d0ef03a2dc495c5a354ab867dabe5baee1b45dbd"><tt>d0ef03a</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4631083503" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186928" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186928/hovercard" href="https://github.com/pytorch/pytorch/pull/186928">#186928</a> on behalf of <a href="https://github.com/wdvr">https://github.com/wdvr</a> due to See below - failing builds with older cuDNN (<a href="https://github.com/pytorch/pytorch/pull/186928#issuecomment-4704924684" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186928/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/0ef3567485fba9fe7d098907cfd54ac5d4e445f0: Revert "[cuDNN][SDPA] d=256 support for cuDNN SDPA (#185553)"]]></title>
<description><![CDATA[This reverts commit 63f903c.
Reverted #185553 on behalf of https://github.com/wdvr due to failing with older cuDNN - see below (comment)]]></description>
<link>https://tsecurity.de/de/3598192/downloads/trunk0ef3567485fba9fe7d098907cfd54ac5d4e445f0-revert-cudnnsdpa-d256-support-for-cudnn-sdpa-185553/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3598192/downloads/trunk0ef3567485fba9fe7d098907cfd54ac5d4e445f0-revert-cudnnsdpa-d256-support-for-cudnn-sdpa-185553/</guid>
<pubDate>Mon, 15 Jun 2026 08:16:48 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/63f903c3d6b04c7cb1433d1d67e2b8e21c055bc7/hovercard" href="https://github.com/pytorch/pytorch/commit/63f903c3d6b04c7cb1433d1d67e2b8e21c055bc7"><tt>63f903c</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4543147964" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185553" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185553/hovercard" href="https://github.com/pytorch/pytorch/pull/185553">#185553</a> on behalf of <a href="https://github.com/wdvr">https://github.com/wdvr</a> due to failing with older cuDNN - see below (<a href="https://github.com/pytorch/pytorch/pull/185553#issuecomment-4704953835" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185553/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/87ceb3fbef7fbc34d98350587fd2584a615c6dfc: [DTensor] Support _StridedShard to Shard through all-to-all (#170915)]]></title>
<description><![CDATA[(AI generated commit description)
[DTensor] Support _StridedShard to Shard through all-to-all
Summary
This PR adds support for redistributing tensors from _StridedShard placement to Shard placement using the all-to-all collective operation.
The key challenge is that _StridedShard produces non-con...]]></description>
<link>https://tsecurity.de/de/3597974/downloads/trunk87ceb3fbef7fbc34d98350587fd2584a615c6dfc-dtensor-support-stridedshard-to-shard-through-all-to-all-170915/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597974/downloads/trunk87ceb3fbef7fbc34d98350587fd2584a615c6dfc-dtensor-support-stridedshard-to-shard-through-all-to-all-170915/</guid>
<pubDate>Mon, 15 Jun 2026 06:16:03 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>(AI generated commit description)</p>
<h2>[DTensor] Support _StridedShard to Shard through all-to-all</h2>
<h3>Summary</h3>
<p>This PR adds support for redistributing tensors from <code>_StridedShard</code> placement to <code>Shard</code> placement using the all-to-all collective operation.</p>
<p>The key challenge is that <code>_StridedShard</code> produces non-contiguous (interleaved) shards, so converting to a regular <code>Shard</code> placement requires:</p>
<ol>
<li>Properly computing padding for both the source strided dimension and target dimension</li>
<li>Reordering elements after the all-to-all to restore contiguous layout</li>
</ol>
<h3>Example: Converting <code>_StridedShard(0, split_factor=2)</code> to <code>Shard(1)</code></h3>
<p>Consider the following setup:</p>
<ul>
<li><strong>Mesh shape</strong>: <code>(4,)</code> — 4 ranks on a single mesh dimension</li>
<li><strong>Original tensor shape</strong>: <code>(9, 4)</code></li>
<li><strong>Source placement</strong>: <code>(_StridedShard(0, split_factor=2),)</code></li>
<li><strong>Target placement</strong>: <code>(Shard(1),)</code></li>
</ul>
<h4>Step 1: Understand the _StridedShard distribution</h4>
<p>With <code>_StridedShard(0, split_factor=2)</code>, the tensor is conceptually split in two levels on dimension 0:</p>
<ol>
<li><strong>First level</strong>: Split into <code>split_factor=2</code> pieces → chunks of size ⌈9/2⌉ = 5, giving pieces <code>[0:5]</code> and <code>[5:9]</code></li>
<li><strong>Second level</strong>: Each piece is split into <code>num_chunks=4</code> pieces (mesh size)</li>
</ol>
<p>The shards are then interleaved so each rank gets one slice from each first-level piece:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="Original tensor (9x4):            Strided sharding on dim 0:
┌─────────────────────┐
│ row 0               │  ─┐
│ row 1               │   ├─ First piece [0:5], split into 4 chunks
│ row 2               │   │  → chunks: [0:2], [2:4], [4:5], []
│ row 3               │   │
│ row 4               │  ─┘
│ row 5               │  ─┐
│ row 6               │   ├─ Second piece [5:9], split into 4 chunks
│ row 7               │   │  → chunks: [5:6], [6:7], [7:8], [8:9]
│ row 8               │  ─┘
└─────────────────────┘

Interleaved distribution to ranks:
  Rank 0: rows [0,1] + [5]     = rows [0,1,5]     (3 rows)
  Rank 1: rows [2,3] + [6]     = rows [2,3,6]     (3 rows)
  Rank 2: rows [4]   + [7]     = rows [4,7]       (2 rows)
  Rank 3: []         + [8]     = rows [8]         (1 row)"><pre class="notranslate"><code>Original tensor (9x4):            Strided sharding on dim 0:
┌─────────────────────┐
│ row 0               │  ─┐
│ row 1               │   ├─ First piece [0:5], split into 4 chunks
│ row 2               │   │  → chunks: [0:2], [2:4], [4:5], []
│ row 3               │   │
│ row 4               │  ─┘
│ row 5               │  ─┐
│ row 6               │   ├─ Second piece [5:9], split into 4 chunks
│ row 7               │   │  → chunks: [5:6], [6:7], [7:8], [8:9]
│ row 8               │  ─┘
└─────────────────────┘

Interleaved distribution to ranks:
  Rank 0: rows [0,1] + [5]     = rows [0,1,5]     (3 rows)
  Rank 1: rows [2,3] + [6]     = rows [2,3,6]     (3 rows)
  Rank 2: rows [4]   + [7]     = rows [4,7]       (2 rows)
  Rank 3: []         + [8]     = rows [8]         (1 row)
</code></pre></div>
<h4>Step 2: Pad for uniform all-to-all</h4>
<p>Before all-to-all, we pad so all ranks have uniform chunk sizes:</p>
<ul>
<li><strong>Old dimension (dim 0)</strong>: <code>max_chunk_size = 3</code>, pad ranks 2 and 3</li>
<li><strong>New dimension (dim 1)</strong>: size 4 with 4 chunks → already uniform (chunk size 1 each)</li>
</ul>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="After padding dim 0:
  Rank 0: [0,1,5] (no padding)    → shape (3, 4)
  Rank 1: [2,3,6] (no padding)    → shape (3, 4)
  Rank 2: [4,7,P] (+1 padding)    → shape (3, 4)
  Rank 3: [8,P,P] (+2 padding)    → shape (3, 4)"><pre class="notranslate"><code>After padding dim 0:
  Rank 0: [0,1,5] (no padding)    → shape (3, 4)
  Rank 1: [2,3,6] (no padding)    → shape (3, 4)
  Rank 2: [4,7,P] (+1 padding)    → shape (3, 4)
  Rank 3: [8,P,P] (+2 padding)    → shape (3, 4)
</code></pre></div>
<h4>Step 3: All-to-all on dim 0 → dim 1</h4>
<p>The all-to-all exchanges slices: each rank sends dim-1 slices to other ranks and receives dim-0 slices:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="Before A2A (each rank has 3x4):     After A2A (each rank has 12x1):
  Rank 0: rows [0,1,5] cols [0,1,2,3]  →  col 0 from all ranks
  Rank 1: rows [2,3,6] cols [0,1,2,3]  →  col 1 from all ranks
  Rank 2: rows [4,7,P] cols [0,1,2,3]  →  col 2 from all ranks
  Rank 3: rows [8,P,P] cols [0,1,2,3]  →  col 3 from all ranks"><pre class="notranslate"><code>Before A2A (each rank has 3x4):     After A2A (each rank has 12x1):
  Rank 0: rows [0,1,5] cols [0,1,2,3]  →  col 0 from all ranks
  Rank 1: rows [2,3,6] cols [0,1,2,3]  →  col 1 from all ranks
  Rank 2: rows [4,7,P] cols [0,1,2,3]  →  col 2 from all ranks
  Rank 3: rows [8,P,P] cols [0,1,2,3]  →  col 3 from all ranks
</code></pre></div>
<h4>Step 4: Unpad and reorder</h4>
<p>After all-to-all, each rank has interleaved rows from the strided pattern with padding. We use <code>index_select</code> to:</p>
<ol>
<li>Extract only the valid (non-padded) elements</li>
<li>Reorder from strided order <code>[0,1,5,2,3,6,4,7,8]</code> back to natural order <code>[0,1,2,3,4,5,6,7,8]</code></li>
</ol>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="Final result - Shard(1) distribution:
  Rank 0: all 9 rows, col 0  → shape (9, 1)
  Rank 1: all 9 rows, col 1  → shape (9, 1)
  Rank 2: all 9 rows, col 2  → shape (9, 1)
  Rank 3: all 9 rows, col 3  → shape (9, 1)"><pre class="notranslate"><code>Final result - Shard(1) distribution:
  Rank 0: all 9 rows, col 0  → shape (9, 1)
  Rank 1: all 9 rows, col 1  → shape (9, 1)
  Rank 2: all 9 rows, col 2  → shape (9, 1)
  Rank 3: all 9 rows, col 3  → shape (9, 1)
</code></pre></div>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3749201916" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/170915" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/170915/hovercard" href="https://github.com/pytorch/pytorch/pull/170915">#170915</a><br>
Approved by: <a href="https://github.com/weifengpy">https://github.com/weifengpy</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/757703a523f29b38beea5655fe7055b7ac477c42: fix _split_iteration_ranges error (#187209) (#187209)]]></title>
<description><![CDATA[Summary:
_split_iteration_ranges in torch/_inductor/codegen/simd.py ended with a bare assert all(... == 1 for s in remaining) that fired whenever a node's iteration lengths were consumed cleanly but left a non-unit extent in the kernel's tiling groups. This happens when a pointwise epilogue whose...]]></description>
<link>https://tsecurity.de/de/3597881/downloads/trunk757703a523f29b38beea5655fe7055b7ac477c42-fix-splititerationranges-error-187209-187209/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597881/downloads/trunk757703a523f29b38beea5655fe7055b7ac477c42-fix-splititerationranges-error-187209-187209/</guid>
<pubDate>Mon, 15 Jun 2026 04:46:06 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:</p>
<p><code>_split_iteration_ranges</code> in <code>torch/_inductor/codegen/simd.py</code> ended with a bare <code>assert all(... == 1 for s in remaining)</code> that fired whenever a node's iteration lengths were consumed cleanly but left a non-unit extent in the kernel's tiling groups. This happens when a pointwise epilogue whose iteration domain is a strict sub-multiple of a template's tiling is considered for fusion, e.g. fusing a <code>[s, N]</code> epilogue into a <code>[K*s, N]</code> matmul template tile. The three divisibility exits earlier in the same function already <code>raise CantSplit</code>, and both callers that drive epilogue fusion (<code>SIMDKernel.is_compatible</code> and <code>Scheduler.speedup_by_fusion</code>) wrap the split in <code>try/except CantSplit</code> specifically to skip range-incompatible fusions and fall back to unfused codegen. Because the final invariant raised <code>AssertionError</code> instead of <code>CantSplit</code>, it escaped those handlers and hard-failed the entire AOTInductor compile with <code>InductorError: AssertionError: failed to set ranges ...</code>.</p>
<p>This PR converts the final invariant to <code>raise CantSplit(remaining, lengths)</code>, matching the exception contract of the rest of the function so the existing safety nets catch it. The success path is byte-for-byte unchanged (identical trigger condition), so models that compile today are unaffected; models that previously crashed now lower with the incompatible epilogue left unfused (a separate kernel) instead of aborting the compile.</p>
<p>Test Plan:<br>
Unit test (new regression):</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="buck2 test //caffe2/test/inductor:loop_ordering -- -r TestSplitIterationRanges"><pre class="notranslate"><code>buck2 test //caffe2/test/inductor:loop_ordering -- -r TestSplitIterationRanges
</code></pre></div>
<p>Result: Pass 7, Fail 0. testrun: <a href="https://www.internalfb.com/intern/testinfra/testrun/1688850234333388" rel="nofollow">https://www.internalfb.com/intern/testinfra/testrun/1688850234333388</a></p>
<p>The new case <code>test_leftover_extent_raises_cant_split</code> constructs <code>groups=[2, 2]</code>, <code>lengths=[[2], []]</code>: every size divides cleanly as it is consumed (so no <code>add_range</code> divisibility exit trips), but group 1 is left with extent <code>2</code>, producing <code>remaining=[1, 2]</code>. It asserts this raises <code>CantSplit</code> rather than <code>AssertionError</code>, exercising exactly the changed line.</p>
<p>Differential Revision: D108339950</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652274047" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187209" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187209/hovercard" href="https://github.com/pytorch/pytorch/pull/187209">#187209</a><br>
Approved by: <a href="https://github.com/ColinPeppler">https://github.com/ColinPeppler</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/71b54592225995ee10ea4469d91806f8ec16d0d3: [reland][Inductor][X86] Remove deprecated fusion patterns (#178466)]]></title>
<description><![CDATA[Reland #173911
These quantization-related fusion patterns have been moved to torchao.
Pull Request resolved: #178466
Approved by: https://github.com/jansel]]></description>
<link>https://tsecurity.de/de/3597849/downloads/trunk71b54592225995ee10ea4469d91806f8ec16d0d3-relandinductorx86-remove-deprecated-fusion-patterns-178466/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597849/downloads/trunk71b54592225995ee10ea4469d91806f8ec16d0d3-relandinductorx86-remove-deprecated-fusion-patterns-178466/</guid>
<pubDate>Mon, 15 Jun 2026 03:46:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Reland <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3874730862" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/173911" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/173911/hovercard" href="https://github.com/pytorch/pytorch/pull/173911">#173911</a><br>
These quantization-related fusion patterns have been moved to torchao.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4139974710" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/178466" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/178466/hovercard" href="https://github.com/pytorch/pytorch/pull/178466">#178466</a><br>
Approved by: <a href="https://github.com/jansel">https://github.com/jansel</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/88b4e8f8e2c0c3028734c4c7ff9b17fa392d9614: Import TypeVar from typing_extensions (#185708)]]></title>
<description><![CDATA[Fixes #140914
Imports TypeVar from typing_extensions alongside ParamSpec so both type parameters use the same default-sentinel machinery when another vendored typing_extensions copy has monkeypatched typing internals. This preserves the existing CachedMethod[P, RV] convention while avoiding the i...]]></description>
<link>https://tsecurity.de/de/3597828/downloads/trunk88b4e8f8e2c0c3028734c4c7ff9b17fa392d9614-import-typevar-from-typingextensions-185708/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597828/downloads/trunk88b4e8f8e2c0c3028734c4c7ff9b17fa392d9614-import-typevar-from-typingextensions-185708/</guid>
<pubDate>Mon, 15 Jun 2026 03:01:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2666521199" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/140914" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/140914/hovercard" href="https://github.com/pytorch/pytorch/issues/140914">#140914</a></p>
<p>Imports <code>TypeVar</code> from <code>typing_extensions</code> alongside <code>ParamSpec</code> so both type parameters use the same default-sentinel machinery when another vendored <code>typing_extensions</code> copy has monkeypatched <code>typing</code> internals. This preserves the existing <code>CachedMethod[P, RV]</code> convention while avoiding the import-time <code>TypeError</code>.</p>
<p>Test Plan:</p>
<ul>
<li><code>python3 -m py_compile torch/_inductor/utils.py</code></li>
<li><code>/Users/dejain/.cache/codex-runtimes/codex-primary-runtime/dependencies/python/bin/python3</code> minimized repro that loads a second <code>typing_extensions</code> copy and defines <code>CachedMethod(Protocol, Generic[P, RV])</code></li>
<li><code>git diff --check</code></li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4554801947" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185708" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185708/hovercard" href="https://github.com/pytorch/pytorch/pull/185708">#185708</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a>, <a href="https://github.com/jansel">https://github.com/jansel</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/aa3ccd27e55dcc37557b3457fd6bf20cd6aa5a28: Use C++20 default member initializers for bit-fields (#183992)]]></title>
<description><![CDATA[Pull Request resolved: #183992
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3597826/downloads/trunkaa3ccd27e55dcc37557b3457fd6bf20cd6aa5a28-use-c-20-default-member-initializers-for-bit-fields-183992/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597826/downloads/trunkaa3ccd27e55dcc37557b3457fd6bf20cd6aa5a28-use-c-20-default-member-initializers-for-bit-fields-183992/</guid>
<pubDate>Mon, 15 Jun 2026 03:01:31 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4458271458" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183992" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183992/hovercard" href="https://github.com/pytorch/pytorch/pull/183992">#183992</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/b8600ea495a216443f8cf326221d6aa6b7be12ac: Enable symm ops (async TP) for XPU (#185102)]]></title>
<description><![CDATA[Summary
We are working on add XPU symmetric backend and also want to enable symm ops (async TP) for XPU.
Motivation
We are working on add XPU symmetric backend and also want to enable symm ops (async TP) for XPU.  With those symm ops, communication and computation could be overlapped to reduce co...]]></description>
<link>https://tsecurity.de/de/3597799/downloads/trunkb8600ea495a216443f8cf326221d6aa6b7be12ac-enable-symm-ops-async-tp-for-xpu-185102/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597799/downloads/trunkb8600ea495a216443f8cf326221d6aa6b7be12ac-enable-symm-ops-async-tp-for-xpu-185102/</guid>
<pubDate>Mon, 15 Jun 2026 02:17:52 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h3>Summary</h3>
<p>We are working on add XPU symmetric backend and also want to enable symm ops (async TP) for XPU.</p>
<h3>Motivation</h3>
<p>We are working on add XPU symmetric backend and also want to enable symm ops (async TP) for XPU.  With those symm ops, communication and computation could be overlapped to reduce communication overhead on Intel client GPUs.</p>
<h3>Changes</h3>
<ul>
<li>Backend enabling in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3416599417" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/2041" data-hovercard-type="pull_request" data-hovercard-url="/intel/torch-xpu-ops/pull/2041/hovercard" href="https://github.com/intel/torch-xpu-ops/pull/2041">intel/torch-xpu-ops#2041</a></li>
<li>Python ops enabling in this PR</li>
</ul>
<h3>Test Plan</h3>
<p>Ops test is verified via <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4514040849" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3747" data-hovercard-type="pull_request" data-hovercard-url="/intel/torch-xpu-ops/pull/3747/hovercard" href="https://github.com/intel/torch-xpu-ops/pull/3747">intel/torch-xpu-ops#3747</a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4514046010" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185102" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185102/hovercard" href="https://github.com/pytorch/pytorch/pull/185102">#185102</a><br>
Approved by: <a href="https://github.com/guangyey">https://github.com/guangyey</a>, <a href="https://github.com/gujinghui">https://github.com/gujinghui</a>, <a href="https://github.com/jansel">https://github.com/jansel</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781480597]]></title>
<description><![CDATA[[MPS] Migrate log_sigmoid forward/backward from MPSGraph to Metal (#1…]]></description>
<link>https://tsecurity.de/de/3597766/downloads/viablestrict1781480597/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597766/downloads/viablestrict1781480597/</guid>
<pubDate>Mon, 15 Jun 2026 01:47:25 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[MPS] Migrate log_sigmoid forward/backward from MPSGraph to Metal (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="171281708" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1/hovercard" href="https://github.com/pytorch/pytorch/issues/1">#1</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/145221dbb55ec8ef3869fdd2ee1c4d670c854b43: [ROCm] increase tolerance on fp16 matmul on softmaxed values (#182028)]]></title>
<description><![CDATA[Matmul for fp16 with softmaxed values in MI200 need more tolerance.
UT originally disable here.
Pull Request resolved: #182028
Approved by: https://github.com/drisspg, https://github.com/jansel]]></description>
<link>https://tsecurity.de/de/3597765/downloads/trunk145221dbb55ec8ef3869fdd2ee1c4d670c854b43-rocm-increase-tolerance-on-fp16-matmul-on-softmaxed-values-182028/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597765/downloads/trunk145221dbb55ec8ef3869fdd2ee1c4d670c854b43-rocm-increase-tolerance-on-fp16-matmul-on-softmaxed-values-182028/</guid>
<pubDate>Mon, 15 Jun 2026 01:47:06 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Matmul for fp16 with softmaxed values in MI200 need more tolerance.<br>
UT originally disable <a href="https://github.com/pytorch/pytorch/issues/172091" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/172091/hovercard">here</a>.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4358811358" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/182028" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/182028/hovercard" href="https://github.com/pytorch/pytorch/pull/182028">#182028</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a>, <a href="https://github.com/jansel">https://github.com/jansel</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/40b4ed1a18c5c250518e50fbcccc213c61e2c2cc: [aten] Fix CPU and MPS logit for eps > 0.5 (#181297)]]></title>
<description><![CDATA[Issue
Fixes #177839.
Summary
For eps > 0.5 the clamp bounds invert (lo = eps > hi = 1 - eps). The scalar and CUDA kernels resolve this as x < lo ? lo : (x > hi ? hi : x), so lo wins. Two paths disagreed and produced a sign-flipped result:

CPU vectorized: vec::clamp = min(hi, max(lo, x)), so hi w...]]></description>
<link>https://tsecurity.de/de/3597638/downloads/trunk40b4ed1a18c5c250518e50fbcccc213c61e2c2cc-aten-fix-cpu-and-mps-logit-for-eps-05-181297/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597638/downloads/trunk40b4ed1a18c5c250518e50fbcccc213c61e2c2cc-aten-fix-cpu-and-mps-logit-for-eps-05-181297/</guid>
<pubDate>Sun, 14 Jun 2026 23:16:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Issue</h2>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4101550367" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/177839" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/177839/hovercard" href="https://github.com/pytorch/pytorch/issues/177839">#177839</a>.</p>
<h2>Summary</h2>
<p>For <code>eps &gt; 0.5</code> the clamp bounds invert (<code>lo = eps &gt; hi = 1 - eps</code>). The scalar and CUDA kernels resolve this as <code>x &lt; lo ? lo : (x &gt; hi ? hi : x)</code>, so <code>lo</code> wins. Two paths disagreed and produced a sign-flipped result:</p>
<ul>
<li>CPU vectorized: <code>vec::clamp</code> = <code>min(hi, max(lo, x))</code>, so <code>hi</code> won.</li>
<li>MPS <code>logit</code>: <code>clampWithTensor</code>, also <code>min(max(x, lo), hi)</code>, so <code>hi</code> won.</li>
</ul>
<p>Both now apply <code>lo</code> last (nested <code>blendv</code> on CPU, nested <code>select</code> on MPS), matching scalar and CUDA. CUDA already conformed, so all three backends agree. <code>eps &gt; 0.5</code> is mathematically ill-defined; the goal is cross-backend consistency.</p>
<h2>Tests</h2>
<p>Added <code>test_logit_vectorized_matches_scalar</code> (eps <code>{0.49, 0.51, 0.6, 0.9}</code> x 4 dtypes, bit-exact vs the scalar kernel) and an <code>eps=0.6</code> <code>sample_inputs_logit</code> sample so <code>test_ops.py</code> and the MPS consistency tests cover <code>eps &gt; 0.5</code>. Verified on Apple Silicon: <code>test_mps.py -k logit</code>, <code>test_ops.py -k logit</code>, and <code>test_unary_ufuncs.py -k logit</code> pass, including the previously failing <code>test_output_grad_match_logit_mps_float32</code>.</p>
<h2>BC-breaking?</h2>
<p>No.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4319193406" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/181297" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/181297/hovercard" href="https://github.com/pytorch/pytorch/pull/181297">#181297</a><br>
Approved by: <a href="https://github.com/jansel">https://github.com/jansel</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/7315a5783d6af5fed8c81e73c2be9451f7c6a2f0: Fix index_add decomposition source shape checks (#184373)]]></title>
<description><![CDATA[Add eager-compatible source shape validation before the index_add decomposition lowers to index_put, preventing source broadcasting from hiding invalid inputs under torch.compile.
Fixes #121135
Generated by my agent
Pull Request resolved: #184373
Approved by: https://github.com/zou3519]]></description>
<link>https://tsecurity.de/de/3597616/downloads/trunk7315a5783d6af5fed8c81e73c2be9451f7c6a2f0-fix-indexadd-decomposition-source-shape-checks-184373/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597616/downloads/trunk7315a5783d6af5fed8c81e73c2be9451f7c6a2f0-fix-indexadd-decomposition-source-shape-checks-184373/</guid>
<pubDate>Sun, 14 Jun 2026 22:46:47 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Add eager-compatible source shape validation before the index_add decomposition lowers to index_put, preventing source broadcasting from hiding invalid inputs under torch.compile.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2167129561" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/121135" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/121135/hovercard" href="https://github.com/pytorch/pytorch/issues/121135">#121135</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4477910122" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184373" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184373/hovercard" href="https://github.com/pytorch/pytorch/pull/184373">#184373</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/523c1d0eed40a1febf33db2a321c8fead524c725: Fix MKLDNN to_dense fake layout handling (#183670)]]></title>
<description><![CDATA[Handle MKLDNN to_dense() in FakeTensor and Inductor lowering so compiled MKLDNN inputs transition to strided dense tensors before arithmetic. The global aten.to_dense composite also covers the direct (non-method) densify path, so all reported "Cannot access data pointer" failures on MKL-DNN to_de...]]></description>
<link>https://tsecurity.de/de/3597600/downloads/trunk523c1d0eed40a1febf33db2a321c8fead524c725-fix-mkldnn-todense-fake-layout-handling-183670/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597600/downloads/trunk523c1d0eed40a1febf33db2a321c8fead524c725-fix-mkldnn-todense-fake-layout-handling-183670/</guid>
<pubDate>Sun, 14 Jun 2026 22:31:42 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Handle MKLDNN to_dense() in FakeTensor and Inductor lowering so compiled MKLDNN inputs transition to strided dense tensors before arithmetic. The global <code>aten.to_dense</code> composite also covers the direct (non-method) densify path, so all reported "Cannot access data pointer" failures on MKL-DNN <code>to_dense()</code> under <code>torch.compile</code> are fixed.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4381053110" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/182402" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/182402/hovercard" href="https://github.com/pytorch/pytorch/issues/182402">#182402</a><br>
Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4448881690" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183773" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/183773/hovercard" href="https://github.com/pytorch/pytorch/issues/183773">#183773</a><br>
Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3329920293" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/160873" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/160873/hovercard" href="https://github.com/pytorch/pytorch/issues/160873">#160873</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4444101223" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183670" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183670/hovercard" href="https://github.com/pytorch/pytorch/pull/183670">#183670</a><br>
Approved by: <a href="https://github.com/aorenste">https://github.com/aorenste</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/6f2953ae46a3b5f25bfc7bd3acfe6e2f663d1ed3: [BE][Ez]: Use CPP20 rvalue overload for ostringstream (#187208)]]></title>
<description><![CDATA[Followup to #186552 . Uses the rvalue overload of ostringstream newly introduced in CPP20 to reduce the amount of string copy allowing us to steal the internal string buffer.
Pull Request resolved: #187208
Approved by: https://github.com/d4l3k, https://github.com/lakshayg, https://github.com/malfet]]></description>
<link>https://tsecurity.de/de/3597346/downloads/trunk6f2953ae46a3b5f25bfc7bd3acfe6e2f663d1ed3-beez-use-cpp20-rvalue-overload-for-ostringstream-187208/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3597346/downloads/trunk6f2953ae46a3b5f25bfc7bd3acfe6e2f663d1ed3-beez-use-cpp20-rvalue-overload-for-ostringstream-187208/</guid>
<pubDate>Sun, 14 Jun 2026 18:31:38 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Followup to <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4607922322" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186552" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186552/hovercard" href="https://github.com/pytorch/pytorch/pull/186552">#186552</a> . Uses the rvalue overload of ostringstream newly introduced in CPP20 to reduce the amount of string copy allowing us to steal the internal string buffer.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652198603" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187208" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187208/hovercard" href="https://github.com/pytorch/pytorch/pull/187208">#187208</a><br>
Approved by: <a href="https://github.com/d4l3k">https://github.com/d4l3k</a>, <a href="https://github.com/lakshayg">https://github.com/lakshayg</a>, <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/2c47a6e31ecb1ed5aecbb4738fa2d0fcc40089f6: [cpp] use std::clamp (#185490)]]></title>
<description><![CDATA[As discussed here #185354 (comment)
these are simple replacements of std::min(std::max(...), ..) with std::clamp I found using git grep. there are other occurrences in the codebase, i'll follow up with another PR for those. Thanks!
Pull Request resolved: #185490
Approved by: https://github.com/Sk...]]></description>
<link>https://tsecurity.de/de/3596807/downloads/trunk2c47a6e31ecb1ed5aecbb4738fa2d0fcc40089f6-cpp-use-stdclamp-185490/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596807/downloads/trunk2c47a6e31ecb1ed5aecbb4738fa2d0fcc40089f6-cpp-use-stdclamp-185490/</guid>
<pubDate>Sun, 14 Jun 2026 11:02:14 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>As discussed here <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4531769872" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185354" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185354/hovercard?comment_id=3312375929&amp;comment_type=review_comment" href="https://github.com/pytorch/pytorch/pull/185354#discussion_r3312375929">#185354 (comment)</a></p>
<p>these are simple replacements of <code>std::min(std::max(...), ..)</code> with <code>std::clamp</code> I found using <code>git grep</code>. there are other occurrences in the codebase, i'll follow up with another PR for those. Thanks!</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4539321840" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185490" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185490/hovercard" href="https://github.com/pytorch/pytorch/pull/185490">#185490</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/98a8ce8a9605547da119d31e7b747f46b373e8a7: Revert "Handle invalid CUDA JITerator cache entries (#186346)"]]></title>
<description><![CDATA[This reverts commit 1f106f4.
Reverted #186346 on behalf of https://github.com/wdvr due to Diff reverted internally (comment)]]></description>
<link>https://tsecurity.de/de/3596568/downloads/trunk98a8ce8a9605547da119d31e7b747f46b373e8a7-revert-handle-invalid-cuda-jiterator-cache-entries-186346/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596568/downloads/trunk98a8ce8a9605547da119d31e7b747f46b373e8a7-revert-handle-invalid-cuda-jiterator-cache-entries-186346/</guid>
<pubDate>Sun, 14 Jun 2026 07:46:37 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/1f106f4328b05aeb0025df01d0e797faa301c63e/hovercard" href="https://github.com/pytorch/pytorch/commit/1f106f4328b05aeb0025df01d0e797faa301c63e"><tt>1f106f4</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4595318525" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186346" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186346/hovercard" href="https://github.com/pytorch/pytorch/pull/186346">#186346</a> on behalf of <a href="https://github.com/wdvr">https://github.com/wdvr</a> due to Diff reverted internally (<a href="https://github.com/pytorch/pytorch/pull/186346#issuecomment-4700831097" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186346/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/186323: Type annotations (and runtime fallout) for torch/__init__.py (#186323)]]></title>
<description><![CDATA[Summary:
Pull Request resolved: #186323
Removes # mypy: allow-untyped-defs from torch/__init__.py (and adds a matching pyrefly.toml sub-config) and reworks how the magic methods on SymInt, SymFloat, and SymBool are typed.
Previously each class carried a long hand-written list of placeholder stubs...]]></description>
<link>https://tsecurity.de/de/3596460/downloads/ciflowtrunk186323-type-annotations-and-runtime-fallout-for-torchinitpy-186323/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596460/downloads/ciflowtrunk186323-type-annotations-and-runtime-fallout-for-torchinitpy-186323/</guid>
<pubDate>Sun, 14 Jun 2026 05:01:05 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594167766" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186323" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186323/hovercard" href="https://github.com/pytorch/pytorch/pull/186323">#186323</a></p>
<p>Removes <code># mypy: allow-untyped-defs</code> from <code>torch/__init__.py</code> (and adds a matching <code>pyrefly.toml</code> sub-config) and reworks how the magic methods on <code>SymInt</code>, <code>SymFloat</code>, and <code>SymBool</code> are typed.</p>
<p>Previously each class carried a long hand-written list of placeholder stubs like <code>def __add__(self, other) -&gt; "SymInt": ...</code>. These didn't model the actual promotion rules — e.g. <code>SymInt + float</code> returns <code>SymFloat</code>, <code>Tensor + SymT</code> returns <code>Tensor</code>, and arithmetic on <code>SymBool</code> promotes to <code>SymInt</code>. Three new generic mixins encode those rules once and are reused across all three classes:</p>
<ul>
<li><code>_SymTypingMagicAlsoBool[_PrimType, _BecomesIntPrimType, _BecomesIntSymType, _FloatPromotionType]</code> — comparisons (<code>==</code>, <code>!=</code>, <code>&lt;</code>, <code>&lt;=</code>, <code>&gt;</code>, <code>&gt;=</code>) and the bool-promoting arithmetic ops (<code>+</code>, <code>-</code>, <code>*</code>, plus their <code>r</code>-variants), with <code>_FloatPromotionType</code> controlling whether the result can be <code>SymFloat</code>.</li>
<li><code>_SymTypingMagic[_PrimType, _SymType, _FloatPromotionType]</code> — the rest of the numeric magic methods (<code>abs</code>, <code>neg</code>, <code>floor</code>, <code>ceil</code>, <code>trunc</code>, <code>mod</code>, <code>lshift</code>/<code>rshift</code>, <code>pow_by_natural</code>, the <code>__sym_*__</code> math wrappers, <code>__int_truediv__</code> / <code>__int_floordiv__</code>, etc.).</li>
<li><code>_SymTypingMagicBitwise[_BitwiseLikeType]</code> — <code>__and__</code>, <code>__or__</code>, <code>__xor__</code> and their reflected forms.</li>
</ul>
<p><code>SymInt</code>/<code>SymFloat</code>/<code>SymBool</code> then become small concrete subclasses parameterized over the right operand and result types (e.g. <code>SymFloat</code>'s float-promotion type is <code>_Never</code> since it doesn't promote further).</p>
<p>Runtime side, <code>sym_node.py</code> gains a real <code>xor</code> entry in <code>bitwise_ops</code>, <code>only_bool_magic_methods</code>, and the dispatch table (with a new <code>_sympy_xor</code>), and <code>also_bool_magic_methods</code> is widened from just <code>{"eq"}</code> to the full comparison set <code>{"eq", "ge", "gt", "le", "lt", "ne"}</code> so the methods installed at runtime line up with what the new mixin advertises. A <code>magic_methods_excl</code> helper set is added for the remainder.</p>
<p>Other fallout from removing <code>allow-untyped-defs</code>:</p>
<ul>
<li><code>PySymType</code> is exported from <code>torch._C</code> stubs and threaded into the asymmetric comparison-op signatures generated by <code>tools/pyi/gen_pyi.py</code>, so <code>Tensor.__lt__(SymInt)</code> etc. type-check.</li>
<li><code>SymInt.has_hint()</code>, <code>SymInt.hint</code>, <code>SymInt.constant</code> move from ad-hoc <code>.node.*</code> access to typed <code>property</code> accessors; <code>definitely_true_hint</code>/etc. are updated to use them.</li>
<li><code>SymNode.shape_env</code> is <code>Optional</code>, so callers in <code>symbolic_shapes.py</code> now raise explicit <code>AssertionError("shape_env should not be None")</code> instead of relying on it being set. A couple of <code>maybe_as_int()</code> / <code>maybe_as_float()</code> call sites are tightened with walrus assignment to avoid calling the method twice.</li>
<li><code># pyrefly: ignore[missing-attribute]</code> is added where <code>SymInt.node</code> is duck-typed across <code>SymNode | NestedIntNode | ConstantIntNode | LocalIntNode</code> and the attribute only exists on some.</li>
</ul>
<p>Review order: start with <code>torch/__init__.py</code> — the three new mixins and the rewritten <code>SymInt</code>/<code>SymFloat</code>/<code>SymBool</code> definitions are where the design lives. Then <code>torch/fx/experimental/sym_node.py</code> for the runtime registration of <code>xor</code> and the broadened <code>also_bool_magic_methods</code>. Everything else is small call-site adjustments to satisfy the stricter typing.</p>
<p>Test Plan:<br>
CI.</p>
<p>Authored with Claude.</p>
<p>Reviewed By: bobrenjc93</p>
<p>Differential Revision: D106389841</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/108ac9f904a0b4a7fa837f27bff6ae80465080b4: [typing] Fix return type annotations for various functions (#170396)]]></title>
<description><![CDATA[Summary
This PR fixes return type annotations across several modules, reducing functions returning Any.
Overall these changes should reduce the number of Any returns by PyTorch which is useful, especially when you want to use strict type checking.
I'm not an expert with all of these functions and...]]></description>
<link>https://tsecurity.de/de/3596433/downloads/trunk108ac9f904a0b4a7fa837f27bff6ae80465080b4-typing-fix-return-type-annotations-for-various-functions-170396/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596433/downloads/trunk108ac9f904a0b4a7fa837f27bff6ae80465080b4-typing-fix-return-type-annotations-for-various-functions-170396/</guid>
<pubDate>Sun, 14 Jun 2026 04:16:32 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>This PR fixes return type annotations across several modules, reducing functions returning <code>Any</code>.</p>
<p>Overall these changes should reduce the number of Any returns by PyTorch which is useful, especially when you want to use strict type checking.<br>
I'm not an expert with all of these functions and hope I got the exact return types, and am not missing something. But beforehand the return types were any or unknown, so they can't really be much worse.</p>
<h2>Changes</h2>
<h3>1. <code>torch.futures.Future</code> - New stub file with proper generics</h3>
<ul>
<li>Added <code>torch/futures/__init__.pyi</code> with full generic support</li>
<li><code>Future[T].wait()</code> now returns <code>T</code> instead of <code>Any</code></li>
<li><code>torch.wait(fut: Future[T])</code> now returns <code>T</code> instead of <code>Any</code></li>
</ul>
<h3>2. <code>torch._C</code> - Dispatch mode return types</h3>
<ul>
<li><code>_pop_torch_dispatch_stack()</code> → <code>TorchDispatchMode | None</code></li>
<li><code>_get_dispatch_mode()</code> → <code>TorchDispatchMode | None</code></li>
<li><code>_get_dispatch_stack_at()</code> → <code>TorchDispatchMode</code></li>
</ul>
<h3>3. <code>torch.distributed</code> - Process group and exception types</h3>
<ul>
<li><code>join_process_group()</code> → <code>ProcessGroup</code> (was <code>Any</code>)</li>
<li><code>Work.exception()</code> → <code>BaseException | None</code> (was <code>Any</code>)</li>
</ul>
<h3>4. <code>torch._C._distributed_rpc</code> - PyRRef generic methods</h3>
<ul>
<li><code>local_value()</code> → <code>_T</code></li>
<li><code>rpc_sync()</code> → <code>_T</code></li>
<li><code>rpc_async()</code> → <code>Future[_T]</code></li>
<li><code>remote()</code> → <code>PyRRef[_T]</code></li>
</ul>
<h3>5. <code>torch.jit._script</code> - RecursiveScriptModule iteration</h3>
<ul>
<li><code>__getitem__()</code> → <code>RecursiveScriptModule</code></li>
<li><code>__iter__()</code> → <code>Iterator[RecursiveScriptModule]</code></li>
</ul>
<h3>6. <code>torch._C._jit_tree_views</code> - Literal and decl types</h3>
<ul>
<li><code>TrueLiteral/FalseLiteral/NoneLiteral</code> → <code>Expr</code></li>
<li><code>Def.decl()</code> → <code>Decl</code></li>
</ul>
<h3>7. new Stub files</h3>
<ul>
<li><strong><code>torch/linalg/__init__.pyi</code></strong> - SVD, eigendecompositions, norms, matrix solvers</li>
<li><strong><code>torch/fft/__init__.pyi</code></strong> - FFT/IFFT operations, frequency helpers</li>
<li><strong><code>torch/special/__init__.pyi</code></strong> - Error functions, gamma functions, Bessel functions</li>
</ul>
<h3>8. torch.distributed - ZeroRedundancyOptimizer stub</h3>
<ul>
<li>Added <code>torch/distributed/optim/zero_redundancy_optimizer.pyi</code> with full type annotations</li>
</ul>
<h3>9. torch.distributed.algorithms.join</h3>
<ul>
<li>Type annotations for <code>Join</code> class and related functions</li>
</ul>
<h3>10. torch.nn.parallel.distributed</h3>
<ul>
<li>Type annotation improvements for DistributedDataParallel</li>
</ul>
<h3>11. torch.jit.frontend</h3>
<ul>
<li>Type annotations for JIT frontend functions</li>
</ul>
<h3>12. torch._dynamo type fixes</h3>
<ul>
<li>Fixed type issues in <code>variables/constant.py</code>, <code>variables/functions.py</code>, <code>variables/script_object.py</code></li>
</ul>
<h3>13. torch.fx.experimental.proxy_tensor</h3>
<ul>
<li>Added pyrefly ignore comments for complex type cases</li>
</ul>
<h3>14. mypy-strict.ini</h3>
<ul>
<li>Enabled strict type checking for <code>torch.linalg</code>, <code>torch.fft</code>, <code>torch.special</code></li>
</ul>
<h3>15. Typing tests</h3>
<ul>
<li>Added <code>test/typing/pass/linalg_fft_special.py</code> - pass tests for linalg/fft/special functions</li>
<li>Added <code>test/typing/reveal/linalg_fft_special.py</code> - reveal tests for return types</li>
<li>Updated <code>test/typing/reveal/namedtuple.py</code> - <code>torch.linalg.qr</code> now returns <code>QRResult</code> instead of <code>Any</code></li>
</ul>
<h2>Test</h2>
<p>import torch<br>
from torch.futures import Future<br>
from torch import Tensor</p>
<p>fut: Future[Tensor] = Future()<br>
fut.set_result(torch.zeros(10))<br>
result = torch.wait(fut)  # type now: Tensor, before: Any</p>
<p>import torch.linalg<br>
import torch.fft<br>
import torch.special</p>
<p>result = torch.linalg.svd(torch.randn(5, 5))  # type now: SVDResult, before: Unknown<br>
result = torch.linalg.norm(torch.randn(5, 5))  # type now: Tensor, before: Unknown<br>
result = torch.fft.fft(torch.randn(10))       # type now: Tensor, before: Unknown<br>
result = torch.special.erf(torch.randn(10))   # type now: Tensor, before: Unknown</p>
<h2>Previous approaches</h2>
<p>The issue was also mentioned in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2936267519" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/149639" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/149639/hovercard" href="https://github.com/pytorch/pytorch/issues/149639">#149639</a>.<br>
Some other approaches like <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3325726005" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/160750" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/160750/hovercard" href="https://github.com/pytorch/pytorch/pull/160750">#160750</a> were tried before, but were reverted.</p>
<h3>Mitigations vs <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3325726005" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/160750" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/160750/hovercard" href="https://github.com/pytorch/pytorch/pull/160750">#160750</a></h3>
<p>The previous PR was reverted for several reasons. Here's how this PR addresses them:</p>
<p><strong>1. Flexible type signatures (no cascading errors)</strong><br>
Unlike <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3325726005" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/160750" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/160750/hovercard" href="https://github.com/pytorch/pytorch/pull/160750">#160750</a> which used strict types like <code>tuple[int, ...]</code> that broke existing call sites, this PR uses flexible signatures:</p>
<div class="highlight highlight-source-python notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="# This PR: accepts Sequence, list, or tuple
dim: int | SymInt | Sequence[int | SymInt] | None"><pre><span class="pl-c"># This PR: accepts Sequence, list, or tuple</span>
<span class="pl-s1">dim</span>: <span class="pl-s1">int</span> <span class="pl-c1">|</span> <span class="pl-v">SymInt</span> <span class="pl-c1">|</span> <span class="pl-v">Sequence</span>[<span class="pl-s1">int</span> <span class="pl-c1">|</span> <span class="pl-v">SymInt</span>] <span class="pl-c1">|</span> <span class="pl-c1">None</span></pre></div>
<p>This prevents the 14+ <code>[arg-type]</code> errors that caused the original revert.</p>
<p><strong>2. Automated test plan</strong><br>
Added typing tests to catch regressions:</p>
<ul>
<li><code>test/typing/pass/linalg_fft_special.py</code> - verifies functions accept correct arg types</li>
<li><code>test/typing/reveal/linalg_fft_special.py</code> - verifies return types (<code>SVDResult</code>, <code>QRResult</code>, etc.)</li>
</ul>
<p><strong>3. Enabled strict mypy checking</strong><br>
Added <code>follow_imports = normal</code> for <code>torch.linalg</code>, <code>torch.fft</code>, <code>torch.special</code> in <code>mypy-strict.ini</code></p>
<p><strong>4. Updated existing test</strong><br>
Fixed <code>test/typing/reveal/namedtuple.py</code> which previously expected <code>Any</code> for <code>torch.linalg.qr</code> - now correctly expects <code>QRResult</code></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3726772256" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/170396" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/170396/hovercard" href="https://github.com/pytorch/pytorch/pull/170396">#170396</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/bf9ea7b0a83587fa2b027958323f1c1bf9ced16e: [Inductor][NVGEMM] Enable Epilogue Fusions (#186183)]]></title>
<description><![CDATA[Adds epilogue fusion (relu, sigmoid, tanh, exp, add, sub, mul, div,
dtype cast) to the NVGEMM backend via cutlass_api EpilogueArguments.
The scheduler's fusion loop recognizes NVUniversalGemmCaller, benchmarks
fused vs unfused EFC kernels, and selects the winner.
MultiTemplateBuffer gains swap/fi...]]></description>
<link>https://tsecurity.de/de/3596421/downloads/trunkbf9ea7b0a83587fa2b027958323f1c1bf9ced16e-inductornvgemm-enable-epilogue-fusions-186183/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596421/downloads/trunkbf9ea7b0a83587fa2b027958323f1c1bf9ced16e-inductornvgemm-enable-epilogue-fusions-186183/</guid>
<pubDate>Sun, 14 Jun 2026 04:01:47 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Adds epilogue fusion (relu, sigmoid, tanh, exp, add, sub, mul, div,<br>
dtype cast) to the NVGEMM backend via cutlass_api EpilogueArguments.</p>
<p>The scheduler's fusion loop recognizes NVUniversalGemmCaller, benchmarks<br>
fused vs unfused EFC kernels, and selects the winner.<br>
MultiTemplateBuffer gains swap/finalize_as_nvgemm_caller for the<br>
benchmark swap. When NVGEMM can't fuse, MTBs fall through to Triton<br>
scheduling so Triton choices can still attempt epilogue fusion.</p>
<p>Authored with the help of an AI assistant (Claude).</p>
<p>Test Plan:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_nv_universal_gemm.py -v -k Epilogue
python test/inductor/test_max_autotune.py -v -k epilogue_fusion"><pre class="notranslate"><code>python test/inductor/test_nv_universal_gemm.py -v -k Epilogue
python test/inductor/test_max_autotune.py -v -k epilogue_fusion
</code></pre></div>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4585721987" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186183" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186183/hovercard" href="https://github.com/pytorch/pytorch/pull/186183">#186183</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/803b5d8d9e5b41e39c1d9a11537c30efdded9230: [Inductor][NVGEMM] Refactor rendering (#186184)]]></title>
<description><![CDATA[Replaces the Jinja template (nv_universal_gemm.py.jinja) with
IndentedBuffer codegen in NVUniversalGemmKernel.render(). Introduces
dataclasses (_VariantRenderSpec, _EpilogueRenderSpec) and module-level
helpers (_create_gemm_arguments, _create_gemm_cache_key,
_lookup_gemm_kernel) that are imported...]]></description>
<link>https://tsecurity.de/de/3596420/downloads/trunk803b5d8d9e5b41e39c1d9a11537c30efdded9230-inductornvgemm-refactor-rendering-186184/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596420/downloads/trunk803b5d8d9e5b41e39c1d9a11537c30efdded9230-inductornvgemm-refactor-rendering-186184/</guid>
<pubDate>Sun, 14 Jun 2026 04:01:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Replaces the Jinja template (nv_universal_gemm.py.jinja) with<br>
IndentedBuffer codegen in NVUniversalGemmKernel.render(). Introduces<br>
dataclasses (_VariantRenderSpec, _EpilogueRenderSpec) and module-level<br>
helpers (_create_gemm_arguments, _create_gemm_cache_key,<br>
_lookup_gemm_kernel) that are imported by the generated wrapper at<br>
runtime.</p>
<p>Restores disk cache (disk_cache_get/set) and the _precompile hook for<br>
subprocess parallel compilation that were present in the Jinja template.<br>
Cache keys now include strides, output metadata, and device for all<br>
tensor variants.</p>
<p>Authored with the help of an AI assistant (Claude).</p>
<p>Test Plan:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_nv_universal_gemm.py -v"><pre class="notranslate"><code>python test/inductor/test_nv_universal_gemm.py -v
</code></pre></div>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4585722236" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186184" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186184/hovercard" href="https://github.com/pytorch/pytorch/pull/186184">#186184</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4585721987" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186183" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186183/hovercard" href="https://github.com/pytorch/pytorch/pull/186183">#186183</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/ed65afcc923bfd60d53312d58113a16575310d8e]]></title>
<description><![CDATA[[Inductor][NVGEMM] Remove duplicate nvgemm_max_profiling config (#185…]]></description>
<link>https://tsecurity.de/de/3596418/downloads/trunked65afcc923bfd60d53312d58113a16575310d8e/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596418/downloads/trunked65afcc923bfd60d53312d58113a16575310d8e/</guid>
<pubDate>Sun, 14 Jun 2026 04:01:44 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[Inductor][NVGEMM] Remove duplicate nvgemm_max_profiling config (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="186288179" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185/hovercard" href="https://github.com/pytorch/pytorch/pull/185">#185</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/442c060adc84420f75987aca636b80b8d4e5761a: Revert "[BE][Ez]: Use CPP20 rvalue overload for ostringstream (#187208)"]]></title>
<description><![CDATA[This reverts commit 2775921.
Reverted #187208 on behalf of https://github.com/pytorch-auto-revert due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (comment)]]></description>
<link>https://tsecurity.de/de/3596139/downloads/trunk442c060adc84420f75987aca636b80b8d4e5761a-revert-beez-use-cpp20-rvalue-overload-for-ostringstream-187208/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596139/downloads/trunk442c060adc84420f75987aca636b80b8d4e5761a-revert-beez-use-cpp20-rvalue-overload-for-ostringstream-187208/</guid>
<pubDate>Sat, 13 Jun 2026 22:16:45 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/27759216f291cfab51e9d9f4ab95451d62258571/hovercard" href="https://github.com/pytorch/pytorch/commit/27759216f291cfab51e9d9f4ab95451d62258571"><tt>2775921</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652198603" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187208" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187208/hovercard" href="https://github.com/pytorch/pytorch/pull/187208">#187208</a> on behalf of <a href="https://github.com/pytorch-auto-revert">https://github.com/pytorch-auto-revert</a> due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (<a href="https://github.com/pytorch/pytorch/pull/187208#issuecomment-4699607134" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187208/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/94c0a41a16f7a65c690b28f5150c886d3a72dac6: Improve non-strict export error for raw Triton kernels (#185827)]]></title>
<description><![CDATA[Non-strict torch.export traces through torch_function and dispatcher/proxy machinery, so a raw Triton kernel launch is invisible as a stable export operator. The trace falls into Triton's runtime and the argument binder tries to specialize tensor arguments via data_ptr(), which fake/functional te...]]></description>
<link>https://tsecurity.de/de/3596110/downloads/trunk94c0a41a16f7a65c690b28f5150c886d3a72dac6-improve-non-strict-export-error-for-raw-triton-kernels-185827/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596110/downloads/trunk94c0a41a16f7a65c690b28f5150c886d3a72dac6-improve-non-strict-export-error-for-raw-triton-kernels-185827/</guid>
<pubDate>Sat, 13 Jun 2026 21:31:53 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Non-strict torch.export traces through <strong>torch_function</strong> and dispatcher/proxy machinery, so a raw Triton kernel launch is invisible as a stable export operator. The trace falls into Triton's runtime and the argument binder tries to specialize tensor arguments via data_ptr(), which fake/functional tensors cannot expose. Users then saw a generic data pointer failure instead of guidance for the supported Triton export path.</p>
<p>This keeps strict export behavior unchanged and rewrites only non-strict data_ptr failures that occur from Triton's runtime into an actionable error telling users to wrap the kernel in torch.library.triton_op and launch it through torch.library.wrap_triton or torch._library.capture_triton. A focused CUDA/Triton regression test covers the raw-kernel non-strict path, and the existing custom Triton export test covers the supported path.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2821675273" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/146066" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/146066/hovercard" href="https://github.com/pytorch/pytorch/issues/146066">#146066</a><br>
Generated by my agent</p>
<p>Test Plan:<br>
python test/export/test_export.py TestExport.test_export_raw_triton_kernel_non_strict_error<br>
python test/export/test_export.py TestExport.test_export_custom_triton_kernel<br>
lintrunner -a<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4564887105" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185827" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185827/hovercard" href="https://github.com/pytorch/pytorch/pull/185827">#185827</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a>, <a href="https://github.com/desertfire">https://github.com/desertfire</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/27759216f291cfab51e9d9f4ab95451d62258571: [BE][Ez]: Use CPP20 rvalue overload for ostringstream (#187208)]]></title>
<description><![CDATA[Followup to #186552 . Uses the rvalue overload of ostringstream newly introduced in CPP20 to reduce the amount of string copy allowing us to steal the internal string buffer.
Pull Request resolved: #187208
Approved by: https://github.com/d4l3k, https://github.com/lakshayg, https://github.com/malfet]]></description>
<link>https://tsecurity.de/de/3596001/downloads/trunk27759216f291cfab51e9d9f4ab95451d62258571-beez-use-cpp20-rvalue-overload-for-ostringstream-187208/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3596001/downloads/trunk27759216f291cfab51e9d9f4ab95451d62258571-beez-use-cpp20-rvalue-overload-for-ostringstream-187208/</guid>
<pubDate>Sat, 13 Jun 2026 20:01:43 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Followup to <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4607922322" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186552" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186552/hovercard" href="https://github.com/pytorch/pytorch/pull/186552">#186552</a> . Uses the rvalue overload of ostringstream newly introduced in CPP20 to reduce the amount of string copy allowing us to steal the internal string buffer.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652198603" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187208" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187208/hovercard" href="https://github.com/pytorch/pytorch/pull/187208">#187208</a><br>
Approved by: <a href="https://github.com/d4l3k">https://github.com/d4l3k</a>, <a href="https://github.com/lakshayg">https://github.com/lakshayg</a>, <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781362659: Fix use-after-free SIGSEGV in ExecutionTraceObserver callbacks (#187141)]]></title>
<description><![CDATA[This is a copy of #180086, with some new testing.
We're sometimes getting segfaults in torch::profiler::impl::onFunctionEnter() because we get back a raw pointer from GlobalStateManager::get(). The underlying state being tracked, however, is actually a std::shared_ptr. The thread that owns the st...]]></description>
<link>https://tsecurity.de/de/3595783/downloads/viablestrict1781362659-fix-use-after-free-sigsegv-in-executiontraceobserver-callbacks-187141/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595783/downloads/viablestrict1781362659-fix-use-after-free-sigsegv-in-executiontraceobserver-callbacks-187141/</guid>
<pubDate>Sat, 13 Jun 2026 17:02:14 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This is a copy of <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4242232730" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/180086" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/180086/hovercard" href="https://github.com/pytorch/pytorch/pull/180086">#180086</a>, with some new testing.</p>
<p>We're sometimes getting segfaults in <code>torch::profiler::impl::onFunctionEnter()</code> because we get back a raw pointer from <code>GlobalStateManager::get()</code>. The underlying state being tracked, however, is actually a <code>std::shared_ptr</code>. The thread that owns the <code>std::shared_ptr</code> sometimes destroys the object while other threads have a dangling raw pointer to the state.</p>
<p>The approach we take to fixing the problem is to always hand out an actual <code>std::shared_ptr</code> when the callback will be executed on another thread. That ensures that even if the thread that owns the state is joined, the state itself will still live as long as the thread has a valid copy of the <code>std::shared_ptr</code>.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4645913999" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187141" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187141/hovercard" href="https://github.com/pytorch/pytorch/pull/187141">#187141</a><br>
Approved by: <a href="https://github.com/ryanzhang22">https://github.com/ryanzhang22</a></p>
<p>Co-authored-by: Jiannan Wang <a href="mailto:jiannanwang@meta.com">jiannanwang@meta.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781353065]]></title>
<description><![CDATA[Add _single c10d::Backend methods and migrate backends to them (#1871…]]></description>
<link>https://tsecurity.de/de/3595611/downloads/viablestrict1781353065/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595611/downloads/viablestrict1781353065/</guid>
<pubDate>Sat, 13 Jun 2026 14:46:41 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Add _single c10d::Backend methods and migrate backends to them (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="237714312" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1871" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1871/hovercard" href="https://github.com/pytorch/pytorch/issues/1871">#1871</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781344074]]></title>
<description><![CDATA[[distributed] Fix max_seqlen mismatch in ring attention backward (#18…]]></description>
<link>https://tsecurity.de/de/3595354/downloads/viablestrict1781344074/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595354/downloads/viablestrict1781344074/</guid>
<pubDate>Sat, 13 Jun 2026 12:01:47 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[distributed] Fix max_seqlen mismatch in ring attention backward (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/77e8ad08177a8af2cff1cd18ea8f996245e2ad33: Fix use-after-free SIGSEGV in ExecutionTraceObserver callbacks (#187141)]]></title>
<description><![CDATA[This is a copy of #180086, with some new testing.
We're sometimes getting segfaults in torch::profiler::impl::onFunctionEnter() because we get back a raw pointer from GlobalStateManager::get(). The underlying state being tracked, however, is actually a std::shared_ptr. The thread that owns the st...]]></description>
<link>https://tsecurity.de/de/3595352/downloads/trunk77e8ad08177a8af2cff1cd18ea8f996245e2ad33-fix-use-after-free-sigsegv-in-executiontraceobserver-callbacks-187141/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595352/downloads/trunk77e8ad08177a8af2cff1cd18ea8f996245e2ad33-fix-use-after-free-sigsegv-in-executiontraceobserver-callbacks-187141/</guid>
<pubDate>Sat, 13 Jun 2026 12:01:45 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This is a copy of <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4242232730" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/180086" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/180086/hovercard" href="https://github.com/pytorch/pytorch/pull/180086">#180086</a>, with some new testing.</p>
<p>We're sometimes getting segfaults in <code>torch::profiler::impl::onFunctionEnter()</code> because we get back a raw pointer from <code>GlobalStateManager::get()</code>. The underlying state being tracked, however, is actually a <code>std::shared_ptr</code>. The thread that owns the <code>std::shared_ptr</code> sometimes destroys the object while other threads have a dangling raw pointer to the state.</p>
<p>The approach we take to fixing the problem is to always hand out an actual <code>std::shared_ptr</code> when the callback will be executed on another thread. That ensures that even if the thread that owns the state is joined, the state itself will still live as long as the thread has a valid copy of the <code>std::shared_ptr</code>.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4645913999" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187141" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187141/hovercard" href="https://github.com/pytorch/pytorch/pull/187141">#187141</a><br>
Approved by: <a href="https://github.com/ryanzhang22">https://github.com/ryanzhang22</a></p>
<p>Co-authored-by: Jiannan Wang <a href="mailto:jiannanwang@meta.com">jiannanwang@meta.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/1f159f1d33b4d6042cc00743cf0228c6eb44cb0f: Revert "Use C++20 default member initializers for bit-fields (#183992)"]]></title>
<description><![CDATA[This reverts commit 1ac76df.
Reverted #183992 on behalf of https://github.com/pytorch-auto-revert due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (comment)]]></description>
<link>https://tsecurity.de/de/3595260/downloads/trunk1f159f1d33b4d6042cc00743cf0228c6eb44cb0f-revert-use-c-20-default-member-initializers-for-bit-fields-183992/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595260/downloads/trunk1f159f1d33b4d6042cc00743cf0228c6eb44cb0f-revert-use-c-20-default-member-initializers-for-bit-fields-183992/</guid>
<pubDate>Sat, 13 Jun 2026 10:46:38 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/1ac76dfd62efdc5599afc6cdb323b9a9faba7719/hovercard" href="https://github.com/pytorch/pytorch/commit/1ac76dfd62efdc5599afc6cdb323b9a9faba7719"><tt>1ac76df</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4458271458" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183992" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183992/hovercard" href="https://github.com/pytorch/pytorch/pull/183992">#183992</a> on behalf of <a href="https://github.com/pytorch-auto-revert">https://github.com/pytorch-auto-revert</a> due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (<a href="https://github.com/pytorch/pytorch/pull/183992#issuecomment-4697981861" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183992/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/1ac76dfd62efdc5599afc6cdb323b9a9faba7719: Use C++20 default member initializers for bit-fields (#183992)]]></title>
<description><![CDATA[Pull Request resolved: #183992
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3595192/downloads/trunk1ac76dfd62efdc5599afc6cdb323b9a9faba7719-use-c-20-default-member-initializers-for-bit-fields-183992/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595192/downloads/trunk1ac76dfd62efdc5599afc6cdb323b9a9faba7719-use-c-20-default-member-initializers-for-bit-fields-183992/</guid>
<pubDate>Sat, 13 Jun 2026 10:01:49 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4458271458" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183992" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183992/hovercard" href="https://github.com/pytorch/pytorch/pull/183992">#183992</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[How to Install AMD ROCm on Ubuntu 26.04 for AI & Deep Learning]]></title>
<description><![CDATA[The post How to Install AMD ROCm on Ubuntu 26.04 for AI & Deep Learning first appeared on Tecmint: Linux Howtos, Tutorials & Guides .If you’ve ever tried setting up Ollama, Stable Diffusion, or PyTorch on an AMD graphics card, you probably remember how
The post How to Install AMD ROCm on Ubuntu 2...]]></description>
<link>https://tsecurity.de/de/3595153/unix-server/how-to-install-amd-rocm-on-ubuntu-2604-for-ai-deep-learning/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595153/unix-server/how-to-install-amd-rocm-on-ubuntu-2604-for-ai-deep-learning/</guid>
<pubDate>Sat, 13 Jun 2026 09:16:07 +0200</pubDate>
<category>🐧 Unix Server</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[The post <a href="https://www.tecmint.com/install-amd-rocm-ubuntu-26-04/">How to Install AMD ROCm on Ubuntu 26.04 for AI &amp; Deep Learning</a> first appeared on <a href="https://www.tecmint.com/">Tecmint: Linux Howtos, Tutorials &amp; Guides</a> .<p>If you’ve ever tried setting up Ollama, Stable Diffusion, or PyTorch on an AMD graphics card, you probably remember how</p>
The post <a href="https://www.tecmint.com/install-amd-rocm-ubuntu-26-04/">How to Install AMD ROCm on Ubuntu 26.04 for AI &amp; Deep Learning</a> first appeared on <a href="https://www.tecmint.com/">Tecmint: Linux Howtos, Tutorials &amp; Guides</a>.]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/69b791827a8d480d1664a357f6aa0af7b73f36e3]]></title>
<description><![CDATA[Add _single c10d::Backend methods and migrate backends to them (#1871…]]></description>
<link>https://tsecurity.de/de/3595110/downloads/trunk69b791827a8d480d1664a357f6aa0af7b73f36e3/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595110/downloads/trunk69b791827a8d480d1664a357f6aa0af7b73f36e3/</guid>
<pubDate>Sat, 13 Jun 2026 08:46:43 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Add _single c10d::Backend methods and migrate backends to them (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="237714312" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1871" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1871/hovercard" href="https://github.com/pytorch/pytorch/issues/1871">#1871</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781329635: [c++20] Simplify waiting using std::latch (#187194)]]></title>
<description><![CDATA[This commit updates ParallelNative to use a single std::latch (C++20) in place of an atomic + mutex + condvar. #176662.
Pull Request resolved: #187194
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3595050/downloads/viablestrict1781329635-c-20-simplify-waiting-using-stdlatch-187194/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3595050/downloads/viablestrict1781329635-c-20-simplify-waiting-using-stdlatch-187194/</guid>
<pubDate>Sat, 13 Jun 2026 08:01:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This commit updates ParallelNative to use a single std::latch (C++20) in place of an atomic + mutex + condvar. <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4031326243" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/176662" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/176662/hovercard" href="https://github.com/pytorch/pytorch/issues/176662">#176662</a>.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4651445435" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187194" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187194/hovercard" href="https://github.com/pytorch/pytorch/pull/187194">#187194</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/f9675d4d18062cbe93799530eac8804bbb29f397]]></title>
<description><![CDATA[[distributed] Fix max_seqlen mismatch in ring attention backward (#18…]]></description>
<link>https://tsecurity.de/de/3594996/downloads/trunkf9675d4d18062cbe93799530eac8804bbb29f397/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594996/downloads/trunkf9675d4d18062cbe93799530eac8804bbb29f397/</guid>
<pubDate>Sat, 13 Jun 2026 07:20:08 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[distributed] Fix max_seqlen mismatch in ring attention backward (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/cc19c670a1adf6066cec8add6c03dac8862b150b: [Test] Refactor test/test_modules.py to be device-agnostic (#185703)]]></title>
<description><![CDATA[Pull Request resolved: #185703
Approved by: https://github.com/jbschlosser]]></description>
<link>https://tsecurity.de/de/3594965/downloads/trunkcc19c670a1adf6066cec8add6c03dac8862b150b-test-refactor-testtestmodulespy-to-be-device-agnostic-185703/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594965/downloads/trunkcc19c670a1adf6066cec8add6c03dac8862b150b-test-refactor-testtestmodulespy-to-be-device-agnostic-185703/</guid>
<pubDate>Sat, 13 Jun 2026 06:32:27 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4554080522" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185703" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185703/hovercard" href="https://github.com/pytorch/pytorch/pull/185703">#185703</a><br>
Approved by: <a href="https://github.com/jbschlosser">https://github.com/jbschlosser</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric]]></title>
<description><![CDATA[We build an end-to-end spatial graph learning pipeline using city2graph. We collect urban POI and street network data from OpenStreetMap, with a synthetic fallback for reliability. We engineer spatial features, construct several proximity graph families, and compare how each represents the same u...]]></description>
<link>https://tsecurity.de/de/3594897/ai-nachrichten/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594897/ai-nachrichten/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/</guid>
<pubDate>Sat, 13 Jun 2026 04:48:27 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>We build an end-to-end spatial graph learning pipeline using city2graph. We collect urban POI and street network data from OpenStreetMap, with a synthetic fallback for reliability. We engineer spatial features, construct several proximity graph families, and compare how each represents the same urban environment. We then build heterogeneous and homogeneous graphs, convert them to PyTorch Geometric, and train a GraphSAGE model to predict POI categories from spatial structure.</p>
<p>The post <a href="https://www.marktechpost.com/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/">A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/eed2739de0aa2b13e4642008a1a519c88213bfda: [DTensor] preserve None SDPA philox output specs (#187199)]]></title>
<description><![CDATA[The single-dim strategy path inserts an implicit all-Replicate rule. SDPA efficient attention intentionally uses None output specs for philox_seed and philox_offset so those scalar tensor outputs remain local tensors, matching the previous op-strategy behavior. The inserted rule used Replicate fo...]]></description>
<link>https://tsecurity.de/de/3594895/downloads/trunkeed2739de0aa2b13e4642008a1a519c88213bfda-dtensor-preserve-none-sdpa-philox-output-specs-187199/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594895/downloads/trunkeed2739de0aa2b13e4642008a1a519c88213bfda-dtensor-preserve-none-sdpa-philox-output-specs-187199/</guid>
<pubDate>Sat, 13 Jun 2026 04:46:38 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The single-dim strategy path inserts an implicit all-Replicate rule. SDPA efficient attention intentionally uses None output specs for philox_seed and philox_offset so those scalar tensor outputs remain local tensors, matching the previous op-strategy behavior. The inserted rule used Replicate for those outputs because meta propagation returns real scalar TensorMeta, so combining it with explicit SDPA rules on a 2D mesh could build a DTensorSpec with placements=(Replicate(), None).</p>
<p>This addresses a regression from <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4617233070" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186667" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186667/hovercard" href="https://github.com/pytorch/pytorch/pull/186667">#186667</a>, which migrated matrix ops to single-dim strategies and introduced the mixed per-mesh-dim behavior for these outputs.</p>
<p>Preserve output positions that are None in every explicit single-dim strategy when inserting the implicit replicate rule. Also make full-mesh expansion reject mixed None/Placement entries instead of constructing invalid DTensorSpec objects.</p>
<p>Test Plan: Added regression tests for SDPA efficient-attention philox scalar outputs and mixed None/Placement output rejection. Ran the affected tests and lintrunner.</p>
<p>Authored with the assistance of an AI coding assistant.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4651724241" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187199" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187199/hovercard" href="https://github.com/pytorch/pytorch/pull/187199">#187199</a><br>
Approved by: <a href="https://github.com/pianpwk">https://github.com/pianpwk</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/c3c33fd3481ab28ef0037bbd9b72588fbdfec691: [CUDA][NCCL] Fix nccl.broadcast dropping the root argument (#187216)]]></title>
<description><![CDATA[Fixes #179908
Summary
torch.cuda.nccl.broadcast(tensors, root=N) silently ignores root and always broadcasts from tensors[0], while nccl.reduce honors root — full root-cause chain in my issue comment:

torch/cuda/nccl.py::broadcast passes root to torch._C._nccl_broadcast correctly.
THCPModule_ncc...]]></description>
<link>https://tsecurity.de/de/3594894/downloads/trunkc3c33fd3481ab28ef0037bbd9b72588fbdfec691-cudanccl-fix-ncclbroadcast-dropping-the-root-argument-187216/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594894/downloads/trunkc3c33fd3481ab28ef0037bbd9b72588fbdfec691-cudanccl-fix-ncclbroadcast-dropping-the-root-argument-187216/</guid>
<pubDate>Sat, 13 Jun 2026 04:46:36 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4239156059" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/179908" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/179908/hovercard" href="https://github.com/pytorch/pytorch/issues/179908">#179908</a></p>
<h2>Summary</h2>
<p><code>torch.cuda.nccl.broadcast(tensors, root=N)</code> silently ignores <code>root</code> and always broadcasts from <code>tensors[0]</code>, while <code>nccl.reduce</code> honors <code>root</code> — full root-cause chain in <a href="https://github.com/pytorch/pytorch/issues/179908#issuecomment-4685032224" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/179908/hovercard">my issue comment</a>:</p>
<ol>
<li><code>torch/cuda/nccl.py::broadcast</code> passes <code>root</code> to <code>torch._C._nccl_broadcast</code> correctly.</li>
<li><code>THCPModule_nccl_broadcast</code> (<code>torch/csrc/cuda/python_nccl.cpp</code>) parses and validates <code>root</code>, then calls <code>torch::cuda::nccl::broadcast(inputs, streams, user_comms)</code> <strong>without it</strong>.</li>
<li><code>torch::cuda::nccl::broadcast</code> (<code>torch/csrc/cuda/nccl.cpp</code>) has no root parameter and hardcodes <code>0</code> in the <code>ncclBcast</code> call.</li>
</ol>
<p>The bug is old: v0.3.0 passed <code>root</code> through to <code>ncclBcast</code>; commit de5f7b7251f ("Base for pure C++ NCCL interface", Dec 2017) dropped it when the C++ helper was extracted, so the binding has been wrong since v0.4.0 (~8.5 years).</p>
<h2>Fix</h2>
<ul>
<li><code>torch/csrc/cuda/nccl.h</code> / <code>nccl.cpp</code>: add <code>int32_t root = 0</code> to <code>torch::cuda::nccl::broadcast</code>, mirroring the existing <code>reduce</code> signature (defaulted, so the internal caller in <code>torch/csrc/cuda/comm.cpp::broadcast_coalesced</code> is unaffected); validate it in-range exactly like <code>reduce</code> does; pass it to <code>ncclBcast</code>.</li>
<li><code>torch/csrc/cuda/python_nccl.cpp</code>: forward the already-parsed <code>root</code>.</li>
<li><code>test/distributed/test_nccl.py</code>: extend <code>test_broadcast</code> with a non-zero-root regression block (<code>root = nGPUs - 1</code>; the root device holds a <code>uniform_()</code> tensor, all other devices hold zeros; assert every device ends up with the root's values). Before this fix the broadcast originates from the zeros at index 0, so the new assertions fail; after the fix they pass. It runs under the existing <code>TEST_MULTIGPU</code> skip and <code>@dtypes(*broadcast_dtypes)</code> (incl. float8) conventions.</li>
</ul>
<h2>Test plan</h2>
<ul>
<li>New regression assertions in <code>test/distributed/test_nccl.py::TestNCCL::test_broadcast</code> (multi-GPU).</li>
<li><code>lintrunner -a</code> on the four touched files: no lint issues (CLANGTIDY skipped locally — requires a build dir).</li>
<li>Compile-verified the touched translation units (<code>nccl.cpp</code>, <code>python_nccl.cpp</code>) plus the unmodified internal caller (<code>comm.cpp</code>) with <code>g++ -fsyntax-only</code> against the modified header: all clean.</li>
<li><strong>Validation gap, stated honestly:</strong> my box has a single GPU, so I could not run the multi-GPU broadcast path locally; <code>python test/distributed/test_nccl.py -k broadcast</code> collects and skip-passes here (<code>OK (skipped=2)</code>). The behavioral coverage relies on PyTorch's multi-GPU CI exercising the extended <code>test_broadcast</code>.</li>
</ul>
<p>🤖 Generated with <a href="https://claude.com/claude-code" rel="nofollow">Claude Code</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652929422" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187216" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187216/hovercard" href="https://github.com/pytorch/pytorch/pull/187216">#187216</a><br>
Approved by: <a href="https://github.com/d4l3k">https://github.com/d4l3k</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/f4c949d41880c6cbcaf5fa87c9a466ec559e22c4: Use C++20 concepts where it improves readability (#179286)]]></title>
<description><![CDATA[C++20 introduces concepts which can simplify some template code. In this PR, I have changed some instances of enable_if with concepts where it's not hard to verify correctness. #176662
Pull Request resolved: #179286
Approved by: https://github.com/malfet, https://github.com/Skylion007, https://gi...]]></description>
<link>https://tsecurity.de/de/3594893/downloads/trunkf4c949d41880c6cbcaf5fa87c9a466ec559e22c4-use-c-20-concepts-where-it-improves-readability-179286/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594893/downloads/trunkf4c949d41880c6cbcaf5fa87c9a466ec559e22c4-use-c-20-concepts-where-it-improves-readability-179286/</guid>
<pubDate>Sat, 13 Jun 2026 04:46:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>C++20 introduces concepts which can simplify some template code. In this PR, I have changed some instances of enable_if with concepts where it's not hard to verify correctness. <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4031326243" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/176662" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/176662/hovercard" href="https://github.com/pytorch/pytorch/issues/176662">#176662</a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4201818191" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/179286" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/179286/hovercard" href="https://github.com/pytorch/pytorch/pull/179286">#179286</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a>, <a href="https://github.com/Skylion007">https://github.com/Skylion007</a>, <a href="https://github.com/cyyever">https://github.com/cyyever</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/187231: [ProxyTensor] Bind unbacked symbols across proxy pytrees]]></title>
<description><![CDATA[ProxyTensor records fresh unbacked symbols on FX nodes via node.meta["unbacked_bindings"]. The existing helper handled the common single-proxy case, but assumed any value containing fresh unbacked symbols corresponded to one Proxy. Backend traces can fakeify several user inputs with unbacked symb...]]></description>
<link>https://tsecurity.de/de/3594870/downloads/ciflowtrunk187231-proxytensor-bind-unbacked-symbols-across-proxy-pytrees/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594870/downloads/ciflowtrunk187231-proxytensor-bind-unbacked-symbols-across-proxy-pytrees/</guid>
<pubDate>Sat, 13 Jun 2026 04:16:00 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>ProxyTensor records fresh unbacked symbols on FX nodes via node.meta["unbacked_bindings"]. The existing helper handled the common single-proxy case, but assumed any value containing fresh unbacked symbols corresponded to one Proxy. Backend traces can fakeify several user inputs with unbacked symbolic dimensions before make_fx starts; placeholder tracking then sees a pytree of proxies while compute_unbacked_bindings returns paths relative to the whole pytree.</p>
<p>Keep the single-proxy behavior unchanged. For nested proxy pytrees, route each symbol binding along its full pytree key path to the owning Proxy and store the remaining local path on that node. This preserves per-placeholder provenance instead of suppressing fresh unbacked tracking.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4653743603" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187230" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/187230/hovercard" href="https://github.com/pytorch/pytorch/issues/187230">#187230</a></p>
<p>This PR was authored with assistance from an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/test_proxy_tensor.py -q -s -k "test_unbacked_bindings_on_multiple_token_grid_placeholders or unbacked or non_deduped_shape or deduped_shape or non_symint_size_spec"'><pre>python -m pytest test/test_proxy_tensor.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>test_unbacked_bindings_on_multiple_token_grid_placeholders or unbacked or non_deduped_shape or deduped_shape or non_symint_size_spec<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>stack-info: PR: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4653760307" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187231" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187231/hovercard" href="https://github.com/pytorch/pytorch/pull/187231">#187231</a>, branch: sanketpurandare/stack/19</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/47af5b13f1b44e2cede51de486fd472ebcc377e0: Use C++20 default member initializers for bit-fields (#183992)]]></title>
<description><![CDATA[Pull Request resolved: #183992
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3594869/downloads/trunk47af5b13f1b44e2cede51de486fd472ebcc377e0-use-c-20-default-member-initializers-for-bit-fields-183992/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594869/downloads/trunk47af5b13f1b44e2cede51de486fd472ebcc377e0-use-c-20-default-member-initializers-for-bit-fields-183992/</guid>
<pubDate>Sat, 13 Jun 2026 04:15:59 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4458271458" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183992" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183992/hovercard" href="https://github.com/pytorch/pytorch/pull/183992">#183992</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/2637e3bdb177c149dfb9978da852f326b14e70a5]]></title>
<description><![CDATA[[Inductor][XPU] Fix combo kernel no-bench carve-out gate for XPU (#18…]]></description>
<link>https://tsecurity.de/de/3594828/downloads/trunk2637e3bdb177c149dfb9978da852f326b14e70a5/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594828/downloads/trunk2637e3bdb177c149dfb9978da852f326b14e70a5/</guid>
<pubDate>Sat, 13 Jun 2026 03:16:52 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[Inductor][XPU] Fix combo kernel no-bench carve-out gate for XPU (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/19afbb4e2e81cc5702fa8cc34c48e1879b98a5aa: [c++20] Simplify waiting using std::latch (#187194)]]></title>
<description><![CDATA[This commit updates ParallelNative to use a single std::latch (C++20) in place of an atomic + mutex + condvar. #176662.
Pull Request resolved: #187194
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3594793/downloads/trunk19afbb4e2e81cc5702fa8cc34c48e1879b98a5aa-c-20-simplify-waiting-using-stdlatch-187194/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594793/downloads/trunk19afbb4e2e81cc5702fa8cc34c48e1879b98a5aa-c-20-simplify-waiting-using-stdlatch-187194/</guid>
<pubDate>Sat, 13 Jun 2026 02:46:38 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This commit updates ParallelNative to use a single std::latch (C++20) in place of an atomic + mutex + condvar. <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4031326243" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/176662" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/176662/hovercard" href="https://github.com/pytorch/pytorch/issues/176662">#176662</a>.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4651445435" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187194" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187194/hovercard" href="https://github.com/pytorch/pytorch/pull/187194">#187194</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/187140: Add _single c10d::Backend methods and migrate backends to them (#187140)]]></title>
<description><![CDATA[Summary:
Introduce the torchcomms _single collective names on the C++ c10d::Backend and migrate the in-tree backends to define them, while keeping the old names fully working for backward compatibility.
Backend now declares all_gather_single, all_gather_single_coalesced, reduce_scatter_single, re...]]></description>
<link>https://tsecurity.de/de/3594766/downloads/ciflowtrunk187140-add-single-c10dbackend-methods-and-migrate-backends-to-them-187140/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594766/downloads/ciflowtrunk187140-add-single-c10dbackend-methods-and-migrate-backends-to-them-187140/</guid>
<pubDate>Sat, 13 Jun 2026 02:16:42 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:</p>
<p>Introduce the torchcomms <code>_single</code> collective names on the C++ <code>c10d::Backend</code> and migrate the in-tree backends to define them, while keeping the old names fully working for backward compatibility.</p>
<p><code>Backend</code> now declares <code>all_gather_single</code>, <code>all_gather_single_coalesced</code>, <code>reduce_scatter_single</code>, <code>reduce_scatter_single_coalesced</code>, and <code>all_to_all_single</code> alongside the existing <code>_allgather_base</code>, <code>allgather_into_tensor_coalesced</code>, <code>_reduce_scatter_base</code>, <code>reduce_scatter_tensor_coalesced</code>, and <code>alltoall_base</code>, which are kept as overridable, forwarding aliases.</p>
<p>Backward compatibility is preserved in both directions: each new method and its old-name alias forward to each other, so a <code>Backend</code> subclass may override EITHER name and a caller may invoke EITHER name. The old aliases simply forward to the canonical <code>_single</code> method; each canonical method opens with the <code>C10D_BACKEND_FORWARDING_GUARD()</code> macro, which installs a function-local thread-local re-entry flag using <code>__func__</code> for the message. If a backend overrides neither name the mutual forwarding re-enters the canonical method, the flag trips, and it reports "does not support " instead of recursing forever. As a result, existing callers of the old names (first-party and out-of-tree) and existing out-of-tree backends that override the old names both keep working unchanged. Removing the old overrides outright would have broken out-of-tree backends on the dispatch path, and dropping the old caller-facing names would have broken direct callers; the bidirectional forwarding avoids both.</p>
<p>The old names are intentionally NOT marked <code>C10_DEPRECATED_MESSAGE</code> in this change. PyTorch's open-source build compiles the bundled <code>third_party/torch-xpu-ops</code> submodule with <code>-Werror -Wdeprecated</code>, and its XPU collective op registration (<code>xccl/Register.cpp</code>) still calls the old <code>Backend</code> names, so a compile-time deprecation turns into a hard build error there. The deprecation will be reintroduced in a follow-up once those callers (and any other out-of-tree backends that are built in-tree) are migrated to the <code>_single</code> names and the submodule pin is bumped.</p>
<p>The in-tree backends are migrated to define the new names: <code>ProcessGroupNCCL</code>, <code>ProcessGroupGloo</code>, <code>ProcessGroupMPI</code>, <code>ProcessGroupUCC</code>, <code>ProcessGroupWrapper</code>, <code>FakeProcessGroup</code>, <code>NCCLXStub</code>, and the <code>torch_openreg</code> (<code>ProcessGroupOCCL</code>) test backend. The dispatcher op implementations in <code>Ops.cpp</code>, <code>ProcessGroupWrapper</code>'s forwarders, and the <code>Backend</code> pybind bindings in <code>init.cpp</code> use the new names.</p>
<p>The underlying aten / c10d dispatcher op names (<code>c10d::_allgather_base_</code>, <code>c10d::alltoall_base_</code>, etc.), their schema strings, the <code>IMPL_*</code> macro names in <code>Ops.cpp</code>, and the <code>OpType</code> enum values are intentionally left unchanged for backward compatibility.</p>
<p>Suggested review order: <code>Backend.hpp</code> (the new/old method pairs and the bidirectional forwarding + <code>ForwardingGuard</code>), then the per-backend subclass renames, then the <code>Ops.cpp</code> / <code>ProcessGroupWrapper</code> / <code>init.cpp</code> call sites.</p>
<p>This diff was authored with the assistance of an AI coding agent (Claude).</p>
<p>Test Plan:<br>
This revision only removes the compile-time deprecation annotations (<code>C10_DEPRECATED_MESSAGE</code>), their <code>-Wdeprecated-declarations</code> suppressions, and the now-unused <code>&lt;c10/util/Deprecated.h&gt;</code> include. The <code>_single</code> rename and the bidirectional forwarding/recursion guard are unchanged, so there is no runtime or codegen change -- dropping the annotations only removes deprecation warnings.</p>
<p>Verified by open-source CI on the exported commit: the full libtorch build across platforms, the <code>linux-noble-xpu-n-py3.10 / build-osdc</code> job (which compiles the bundled <code>torch-xpu-ops</code> <code>xccl/Register.cpp</code> with <code>-Werror -Wdeprecated</code> and previously failed on <code>-Werror=deprecated-declarations</code>), and the <code>lintrunner-clang</code> format jobs.</p>
<p>Reviewed By: kapilsh</p>
<p>Differential Revision: D108364288</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/096c3b356f9f273e9701557a3b295c0fad202de5: Fix typos in comments, docstrings, and strings across torch (#187076)]]></title>
<description><![CDATA[Correct "than"/"then" confusions, "it's"/"its" mistakes, and several other
small spelling errors ("no less then", "loose"/"lose", "coorelate",
"sparsr", "an unique" -> "a unique") in comments, docstrings, and error
messages. No functional changes.
Test Plan: Comment/string-only edits; no code beh...]]></description>
<link>https://tsecurity.de/de/3594758/downloads/trunk096c3b356f9f273e9701557a3b295c0fad202de5-fix-typos-in-comments-docstrings-and-strings-across-torch-187076/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594758/downloads/trunk096c3b356f9f273e9701557a3b295c0fad202de5-fix-typos-in-comments-docstrings-and-strings-across-torch-187076/</guid>
<pubDate>Sat, 13 Jun 2026 02:16:32 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Correct "than"/"then" confusions, "it's"/"its" mistakes, and several other<br>
small spelling errors ("no less then", "loose"/"lose", "coorelate",<br>
"sparsr", "an unique" -&gt; "a unique") in comments, docstrings, and error<br>
messages. No functional changes.</p>
<p>Test Plan: Comment/string-only edits; no code behavior affected.</p>
<p>Authored with assistance from an AI assistant.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4643639306" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187076" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187076/hovercard" href="https://github.com/pytorch/pytorch/pull/187076">#187076</a><br>
Approved by: <a href="https://github.com/d4l3k">https://github.com/d4l3k</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781302646: Validate num_heads before division in MultiheadAttention (#186376)]]></title>
<description><![CDATA[Fix: Division-by-Zero in MultiheadAttentionImpl::reset() When num_heads == 0
Summary
Fixes #186237
Fixed a division-by-zero in MultiheadAttentionImpl::reset() triggered when num_heads is 0.
Previously, head_dim was computed before any validation of num_heads:
head_dim = options.embed_dim() / opti...]]></description>
<link>https://tsecurity.de/de/3594643/downloads/viablestrict1781302646-validate-numheads-before-division-in-multiheadattention-186376/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594643/downloads/viablestrict1781302646-validate-numheads-before-division-in-multiheadattention-186376/</guid>
<pubDate>Sat, 13 Jun 2026 00:21:19 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Fix: Division-by-Zero in <code>MultiheadAttentionImpl::reset()</code> When <code>num_heads == 0</code></h1>
<h2>Summary</h2>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4589681276" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186237" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186237/hovercard" href="https://github.com/pytorch/pytorch/issues/186237">#186237</a><br>
Fixed a division-by-zero in <code>MultiheadAttentionImpl::reset()</code> triggered when <code>num_heads</code> is <code>0</code>.</p>
<p>Previously, <code>head_dim</code> was computed before any validation of <code>num_heads</code>:</p>
<div class="highlight highlight-source-c++ notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="head_dim = options.embed_dim() / options.num_heads();"><pre>head_dim = options.embed_dim() / options.num_heads();</pre></div>
<p>Constructing <code>MultiheadAttention</code> with <code>num_heads = 0</code> could invoke undefined behavior (integer division by zero / floating-point exception) before any validation logic was reached.</p>
<p>This PR adds a <code>TORCH_CHECK</code> to validate that <code>num_heads &gt; 0</code> <strong>before</strong> performing the division, replacing the silent UB with a clear, actionable error message.</p>
<h2>Changes</h2>
<ul>
<li>Added a <code>TORCH_CHECK(options.num_heads() &gt; 0, ...)</code> guard in <code>MultiheadAttentionImpl::reset()</code> prior to the <code>head_dim</code> computation.</li>
</ul>
<h2>Test Plan</h2>
<p>Added a C++ API test verifying that constructing:</p>
<div class="highlight highlight-source-c++ notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="MultiheadAttention(MultiheadAttentionOptions(0, 0))"><pre><span class="pl-en">MultiheadAttention</span>(MultiheadAttentionOptions(<span class="pl-c1">0</span>, <span class="pl-c1">0</span>))</pre></div>
<p>throws a <code>c10::Error</code>.</p>
<p><strong>Before this change:</strong> the test triggered a floating-point exception due to division by zero.</p>
<p><strong>After this change:</strong> the test passes and the invalid configuration is reported via a proper error.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4597743345" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186376" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186376/hovercard" href="https://github.com/pytorch/pytorch/pull/186376">#186376</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/7659b78f44ac1098acb26e733e76162b02a6c4d1: [MPS] Migrate sigmoid_backward from MPSGraph to Metal (#187151)]]></title>
<description><![CDATA[Replace the MPSGraph-based sigmoid_backward_out_mps with a native Metal kernel registered through the shared sigmoid_backward_stub.
Complex types are handled via the conjugated derivative (grad * conj((1 - output) * output)), matching CUDA's behavior.
Microbenchmark (us per call, A/B from the sam...]]></description>
<link>https://tsecurity.de/de/3594590/downloads/trunk7659b78f44ac1098acb26e733e76162b02a6c4d1-mps-migrate-sigmoidbackward-from-mpsgraph-to-metal-187151/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594590/downloads/trunk7659b78f44ac1098acb26e733e76162b02a6c4d1-mps-migrate-sigmoidbackward-from-mpsgraph-to-metal-187151/</guid>
<pubDate>Fri, 12 Jun 2026 23:37:15 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Replace the MPSGraph-based <code>sigmoid_backward_out_mps</code> with a native Metal kernel registered through the shared <code>sigmoid_backward_stub</code>.</p>
<p>Complex types are handled via the conjugated derivative (<code>grad * conj((1 - output) * output)</code>), matching CUDA's behavior.</p>
<p>Microbenchmark (us per call, A/B from the same script with only the impl<br>
swapped):</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="shape|dtype                      mpsgraph     metal   speedup
1024|float32                        12.81      7.43     1.72x
1024|float16                        14.82      7.10     2.09x
1024|bfloat16                       12.11      7.15     1.69x
1024x1024|float16                   18.56     14.67     1.26x
1024x1024|float32                   45.02     40.06     1.12x
8x1024x1024|float32                382.21    376.67     1.01x
32x1024x1024|float32              1487.84   1493.75     1.00x"><pre class="notranslate"><code>shape|dtype                      mpsgraph     metal   speedup
1024|float32                        12.81      7.43     1.72x
1024|float16                        14.82      7.10     2.09x
1024|bfloat16                       12.11      7.15     1.69x
1024x1024|float16                   18.56     14.67     1.26x
1024x1024|float32                   45.02     40.06     1.12x
8x1024x1024|float32                382.21    376.67     1.01x
32x1024x1024|float32              1487.84   1493.75     1.00x
</code></pre></div>
<p>Metal wins on dispatch overhead at small sizes; parity at large sizes where<br>
the op is memory-bandwidth bound, as expected for ~3 FLOPs/element.</p>
<p>Co-Authored-By: Claude Opus 4.7 <a href="mailto:noreply@anthropic.com">noreply@anthropic.com</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4646622207" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187151" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187151/hovercard" href="https://github.com/pytorch/pytorch/pull/187151">#187151</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781292498: [DTensor] Preserve symbolic local layouts without DDE guards (#187026)]]></title>
<description><![CDATA[Compiled DTensor paths can see symbolic local layout metadata that is
semantically valid but not syntactically identical to the metadata saved during
forward propagation.
For to_local() backward, AOTAutograd can produce a local gradient stride like
(Max(1, u3), 1) while the saved DTensor metadata...]]></description>
<link>https://tsecurity.de/de/3594417/downloads/viablestrict1781292498-dtensor-preserve-symbolic-local-layouts-without-dde-guards-187026/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3594417/downloads/viablestrict1781292498-dtensor-preserve-symbolic-local-layouts-without-dde-guards-187026/</guid>
<pubDate>Fri, 12 Jun 2026 21:36:06 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Compiled DTensor paths can see symbolic local layout metadata that is<br>
semantically valid but not syntactically identical to the metadata saved during<br>
forward propagation.</p>
<p>For <code>to_local()</code> backward, AOTAutograd can produce a local gradient stride like<br>
<code>(Max(1, u3), 1)</code> while the saved DTensor metadata uses <code>(u1, 1)</code>. The previous<br>
backward path recomputed the global gradient stride before deciding whether it<br>
could reuse the original DTensor spec. That forced<br>
<code>compute_global_tensor_info()</code> to evaluate symbolic stride relations and could<br>
raise a data-dependent guard for the default same-placement backward path.</p>
<p>Reuse the original DTensor spec only for the default same-placement backward<br>
when the local gradient stride, saved forward local stride, and DTensor spec<br>
stride are compatible. Exact/provable stride equality is accepted directly. For<br>
contiguous symbolic stride forms such as <code>Max(1, u*)</code>, use the existing<br>
<code>check_contiguous_sizes_strides(..., false_if_dde=True)</code> helper so equivalent<br>
contiguous layouts are recognized without requiring a brittle exact symbolic<br>
match. If neither exact nor contiguous equivalence can be proven, emit<br>
<code>torch._check</code> assertions for the required stride equalities before taking the<br>
symbolic shortcut.</p>
<p>If the placement changes or the physical local stride cannot justify the<br>
original spec layout, keep the existing recomputation path and build a fresh<br>
spec from the observed gradient stride. This avoids the symbolic guard failure<br>
without lying about memory layout. In particular, uneven channels-last shards<br>
must keep using the recomputation path so autograd can repair the physical local<br>
gradient layout correctly.</p>
<p>The same class of issue also appears in <code>aten.t</code> sharding propagation. The<br>
single-dim strategy can enumerate candidate placements that move a symbolic<br>
<code>_StridedShard</code> split factor onto the other tensor dimension. When that<br>
candidate is not provably shardable, strategy expansion should reject it instead<br>
of evaluating a Python bool on an unbacked expression such as <code>8 &lt; 2*u0</code>.<br>
Register transpose with <code>allow_unbacked_sharding=False</code> so unproven candidates<br>
are pruned while statically valid candidates, including <code>_StridedShard(0, u0)</code><br>
propagating to <code>_StridedShard(1, u0)</code>, are still kept.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4639021322" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187025" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/187025/hovercard" href="https://github.com/pytorch/pytorch/issues/187025">#187025</a></p>
<p>This PR was authored with assistance from an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python -m pytest test/distributed/tensor/test_op_strategy.py::TestCostModel::test_t_prunes_unproven_unbacked_strided_shard_candidates test/distributed/tensor/test_op_strategy.py::TestCostModel::test_mm_strategies test/distributed/tensor/test_dtensor_compile.py::TestDTensorCompile::test_to_local_backward_unbacked_symbolic_stride test/distributed/tensor/test_tensor_ops.py::TestNewEmptyStridedUneven::test_backward_channels_last -q -s"><pre>python -m pytest test/distributed/tensor/test_op_strategy.py::TestCostModel::test_t_prunes_unproven_unbacked_strided_shard_candidates test/distributed/tensor/test_op_strategy.py::TestCostModel::test_mm_strategies test/distributed/tensor/test_dtensor_compile.py::TestDTensorCompile::test_to_local_backward_unbacked_symbolic_stride test/distributed/tensor/test_tensor_ops.py::TestNewEmptyStridedUneven::test_backward_channels_last -q -s</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4639032064" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187026" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187026/hovercard" href="https://github.com/pytorch/pytorch/pull/187026">#187026</a><br>
Approved by: <a href="https://github.com/pianpwk">https://github.com/pianpwk</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5545ef3cd14633e943ce244e838eedfb6aff1235: [MPS][BE] Migrate upsample nearest forward to Metal kernels (#186989)]]></title>
<description><![CDATA[Route the 1D/2D nearest and nearest-exact forward ops to Metal kernels instead of MPSGraph, joining the linear/bilinear/bicubic/3D forward paths that were already Metal. This also fixes a latent MPSGraph correctness bug.
Unify the host-side dispatch so a single templated launcher drives every Met...]]></description>
<link>https://tsecurity.de/de/3593709/downloads/trunk5545ef3cd14633e943ce244e838eedfb6aff1235-mpsbe-migrate-upsample-nearest-forward-to-metal-kernels-186989/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3593709/downloads/trunk5545ef3cd14633e943ce244e838eedfb6aff1235-mpsbe-migrate-upsample-nearest-forward-to-metal-kernels-186989/</guid>
<pubDate>Fri, 12 Jun 2026 16:16:55 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Route the 1D/2D nearest and nearest-exact forward ops to Metal kernels instead of MPSGraph, joining the linear/bilinear/bicubic/3D forward paths that were already Metal. This also fixes a latent MPSGraph correctness bug.</p>
<p>Unify the host-side dispatch so a single templated launcher drives every Metal upsample kernel and UpsampleParams supplies the geometry for all ranks, collapsing the previously duplicated per-rank dispatch overloads.</p>
<p>Backward is intentionally left on MPSGraph: it has no correctness bug and its fused gradient kernel is several times faster than a naive Metal gather.</p>
<p>Co-authored-by: Claude <a href="mailto:noreply@anthropic.com">noreply@anthropic.com</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4635444107" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186989" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186989/hovercard" href="https://github.com/pytorch/pytorch/pull/186989">#186989</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a>, <a href="https://github.com/kurtamohler">https://github.com/kurtamohler</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/7120d05eddfb5563e592a89f83bcdee7baa4911c: [MPS] median and nanmedian to metal (#187060)]]></title>
<description><![CDATA[Fixes #187017
median and nanmedian to metal. Perf:
Global reduction (torch.median(x) / torch.nanmedian(x))



config
dtype
before (us)
after (us)
speedup




global 1e6
f32
376.1
104.9
3.59x


global 1e7
f32
3,403.7
725.0
4.69x


global 1e6
bf16
416.3
51.5
8.08x


global 1e7
bf16
3,961.4
238.6
16...]]></description>
<link>https://tsecurity.de/de/3593708/downloads/trunk7120d05eddfb5563e592a89f83bcdee7baa4911c-mps-median-and-nanmedian-to-metal-187060/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3593708/downloads/trunk7120d05eddfb5563e592a89f83bcdee7baa4911c-mps-median-and-nanmedian-to-metal-187060/</guid>
<pubDate>Fri, 12 Jun 2026 16:16:54 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4637702988" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187017" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/187017/hovercard" href="https://github.com/pytorch/pytorch/issues/187017">#187017</a><br>
median and nanmedian to metal. Perf:</p>
<h3>Global reduction (<code>torch.median(x)</code> / <code>torch.nanmedian(x)</code>)</h3>
<table>
<thead>
<tr>
<th>config</th>
<th>dtype</th>
<th align="right">before (us)</th>
<th align="right">after (us)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>global 1e6</td>
<td>f32</td>
<td align="right">376.1</td>
<td align="right">104.9</td>
<td align="right">3.59x</td>
</tr>
<tr>
<td>global 1e7</td>
<td>f32</td>
<td align="right">3,403.7</td>
<td align="right">725.0</td>
<td align="right">4.69x</td>
</tr>
<tr>
<td>global 1e6</td>
<td>bf16</td>
<td align="right">416.3</td>
<td align="right">51.5</td>
<td align="right">8.08x</td>
</tr>
<tr>
<td>global 1e7</td>
<td>bf16</td>
<td align="right">3,961.4</td>
<td align="right">238.6</td>
<td align="right">16.61x</td>
</tr>
<tr>
<td>global 1e6</td>
<td>i32</td>
<td align="right">427.4</td>
<td align="right">114.5</td>
<td align="right">3.73x</td>
</tr>
<tr>
<td>global 1e7</td>
<td>i32</td>
<td align="right">3,994.5</td>
<td align="right">647.5</td>
<td align="right">6.17x</td>
</tr>
<tr>
<td>global 1e6</td>
<td>i64</td>
<td align="right">1,114.7</td>
<td align="right">206.9</td>
<td align="right">5.39x</td>
</tr>
<tr>
<td>global 1e7</td>
<td>i64</td>
<td align="right">16,138.2</td>
<td align="right">2,511.3</td>
<td align="right">6.43x</td>
</tr>
<tr>
<td>nan10% 1e7 nanmed glob</td>
<td>f32</td>
<td align="right">18,441.6</td>
<td align="right">660.2</td>
<td align="right">27.93x</td>
</tr>
<tr>
<td>nan10% 1e7 nanmed glob</td>
<td>bf16</td>
<td align="right">OOM</td>
<td align="right">239.8</td>
<td align="right">inf (OOM before)</td>
</tr>
</tbody>
</table>
<h3>Reduction along a dim (<code>torch.median(x, dim)</code>)</h3>
<table>
<thead>
<tr>
<th>config</th>
<th>dtype</th>
<th align="right">before (us)</th>
<th align="right">after (us)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>4096x4096 dim=1</td>
<td>f32</td>
<td align="right">10,186.3</td>
<td align="right">8,234.8</td>
<td align="right">1.24x</td>
</tr>
<tr>
<td>4096x4096 dim=0</td>
<td>f32</td>
<td align="right">18,222.7</td>
<td align="right">10,010.9</td>
<td align="right">1.82x</td>
</tr>
<tr>
<td>64x1e6 dim=1</td>
<td>f32</td>
<td align="right">131,392.2</td>
<td align="right">44,488.6</td>
<td align="right">2.95x</td>
</tr>
<tr>
<td>1e6x64 dim=1</td>
<td>f32</td>
<td align="right">799,542.7</td>
<td align="right">53,340.6</td>
<td align="right">14.99x</td>
</tr>
<tr>
<td>1024x65536 dim=1</td>
<td>f32</td>
<td align="right">165,552.8</td>
<td align="right">36,321.2</td>
<td align="right">4.56x</td>
</tr>
<tr>
<td>512x131072 dim=1</td>
<td>f32</td>
<td align="right">192,594.4</td>
<td align="right">50,550.8</td>
<td align="right">3.81x</td>
</tr>
<tr>
<td>256^3 dim=2</td>
<td>f32</td>
<td align="right">85,459.8</td>
<td align="right">8,062.8</td>
<td align="right">10.60x</td>
</tr>
<tr>
<td>256^3 dim=0</td>
<td>f32</td>
<td align="right">99,853.3</td>
<td align="right">7,937.4</td>
<td align="right">12.58x</td>
</tr>
<tr>
<td>128x1024 dim=1</td>
<td>f32</td>
<td align="right">263.4</td>
<td align="right">65.8</td>
<td align="right">4.00x</td>
</tr>
<tr>
<td>2x3x4x5x6 dim=1</td>
<td>f32</td>
<td align="right">210.5</td>
<td align="right">25.9</td>
<td align="right">8.13x</td>
</tr>
<tr>
<td>4096x4096 dim=1</td>
<td>bf16</td>
<td align="right">8,259.1</td>
<td align="right">4,460.5</td>
<td align="right">1.85x</td>
</tr>
<tr>
<td>4096x4096 dim=0</td>
<td>bf16</td>
<td align="right">14,593.3</td>
<td align="right">5,263.7</td>
<td align="right">2.77x</td>
</tr>
<tr>
<td>64x1e6 dim=1</td>
<td>bf16</td>
<td align="right">71,337.7</td>
<td align="right">16,727.9</td>
<td align="right">4.26x</td>
</tr>
<tr>
<td>1e6x64 dim=1</td>
<td>bf16</td>
<td align="right">718,870.6</td>
<td align="right">63,828.1</td>
<td align="right">11.26x</td>
</tr>
<tr>
<td>1024x65536 dim=1</td>
<td>bf16</td>
<td align="right">99,392.7</td>
<td align="right">31,771.3</td>
<td align="right">3.13x</td>
</tr>
<tr>
<td>512x131072 dim=1</td>
<td>bf16</td>
<td align="right">69,957.3</td>
<td align="right">26,466.0</td>
<td align="right">2.64x</td>
</tr>
<tr>
<td>256^3 dim=2</td>
<td>bf16</td>
<td align="right">61,144.0</td>
<td align="right">8,276.0</td>
<td align="right">7.39x</td>
</tr>
<tr>
<td>256^3 dim=0</td>
<td>bf16</td>
<td align="right">73,448.6</td>
<td align="right">8,124.6</td>
<td align="right">9.04x</td>
</tr>
<tr>
<td>128x1024 dim=1</td>
<td>bf16</td>
<td align="right">192.0</td>
<td align="right">72.1</td>
<td align="right">2.66x</td>
</tr>
<tr>
<td>2x3x4x5x6 dim=1</td>
<td>bf16</td>
<td align="right">214.8</td>
<td align="right">28.5</td>
<td align="right">7.54x</td>
</tr>
<tr>
<td>4096x4096 dim=1</td>
<td>i32</td>
<td align="right">7,865.9</td>
<td align="right">3,397.0</td>
<td align="right">2.32x</td>
</tr>
<tr>
<td>4096x4096 dim=0</td>
<td>i32</td>
<td align="right">14,520.8</td>
<td align="right">4,709.3</td>
<td align="right">3.08x</td>
</tr>
<tr>
<td>64x1e6 dim=1</td>
<td>i32</td>
<td align="right">125,680.6</td>
<td align="right">36,122.3</td>
<td align="right">3.48x</td>
</tr>
<tr>
<td>1e6x64 dim=1</td>
<td>i32</td>
<td align="right">722,882.5</td>
<td align="right">10,845.9</td>
<td align="right">66.65x</td>
</tr>
<tr>
<td>1024x65536 dim=1</td>
<td>i32</td>
<td align="right">159,814.7</td>
<td align="right">32,887.0</td>
<td align="right">4.86x</td>
</tr>
<tr>
<td>512x131072 dim=1</td>
<td>i32</td>
<td align="right">147,411.6</td>
<td align="right">40,269.4</td>
<td align="right">3.66x</td>
</tr>
<tr>
<td>256^3 dim=2</td>
<td>i32</td>
<td align="right">58,854.0</td>
<td align="right">1,665.9</td>
<td align="right">35.33x</td>
</tr>
<tr>
<td>256^3 dim=0</td>
<td>i32</td>
<td align="right">71,037.4</td>
<td align="right">2,494.1</td>
<td align="right">28.48x</td>
</tr>
<tr>
<td>128x1024 dim=1</td>
<td>i32</td>
<td align="right">188.2</td>
<td align="right">23.9</td>
<td align="right">7.88x</td>
</tr>
<tr>
<td>2x3x4x5x6 dim=1</td>
<td>i32</td>
<td align="right">210.3</td>
<td align="right">9.8</td>
<td align="right">21.49x</td>
</tr>
<tr>
<td>4096x4096 dim=1</td>
<td>i64</td>
<td align="right">15,432.2</td>
<td align="right">6,173.9</td>
<td align="right">2.50x</td>
</tr>
<tr>
<td>4096x4096 dim=0</td>
<td>i64</td>
<td align="right">24,024.7</td>
<td align="right">7,869.9</td>
<td align="right">3.05x</td>
</tr>
<tr>
<td>64x1e6 dim=1</td>
<td>i64</td>
<td align="right">384,666.6</td>
<td align="right">76,788.9</td>
<td align="right">5.01x</td>
</tr>
<tr>
<td>1e6x64 dim=1</td>
<td>i64</td>
<td align="right">1,401,332.0</td>
<td align="right">17,755.0</td>
<td align="right">78.93x</td>
</tr>
<tr>
<td>1024x65536 dim=1</td>
<td>i64</td>
<td align="right">306,227.5</td>
<td align="right">49,172.7</td>
<td align="right">6.23x</td>
</tr>
<tr>
<td>512x131072 dim=1</td>
<td>i64</td>
<td align="right">383,091.9</td>
<td align="right">61,487.8</td>
<td align="right">6.23x</td>
</tr>
<tr>
<td>256^3 dim=2</td>
<td>i64</td>
<td align="right">109,296.3</td>
<td align="right">2,737.0</td>
<td align="right">39.93x</td>
</tr>
<tr>
<td>256^3 dim=0</td>
<td>i64</td>
<td align="right">125,894.7</td>
<td align="right">3,872.0</td>
<td align="right">32.51x</td>
</tr>
<tr>
<td>128x1024 dim=1</td>
<td>i64</td>
<td align="right">317.5</td>
<td align="right">34.5</td>
<td align="right">9.21x</td>
</tr>
<tr>
<td>2x3x4x5x6 dim=1</td>
<td>i64</td>
<td align="right">397.5</td>
<td align="right">12.6</td>
<td align="right">31.57x</td>
</tr>
<tr>
<td>512^3 dim=1</td>
<td>f32</td>
<td align="right">654,396.9</td>
<td align="right">65,719.6</td>
<td align="right">9.96x</td>
</tr>
<tr>
<td>512^3 dim=1</td>
<td>bf16</td>
<td align="right">594,299.7</td>
<td align="right">73,546.9</td>
<td align="right">8.08x</td>
</tr>
<tr>
<td>512^3 dim=1</td>
<td>i32</td>
<td align="right">333,311.3</td>
<td align="right">25,239.7</td>
<td align="right">13.21x</td>
</tr>
<tr>
<td>512^3 dim=1</td>
<td>i64</td>
<td align="right">OOM</td>
<td align="right">36,270.9</td>
<td align="right">inf (OOM before)</td>
</tr>
<tr>
<td>1024x1024x512 dim=1</td>
<td>f32</td>
<td align="right">OOM</td>
<td align="right">278,978.0</td>
<td align="right">inf (OOM before)</td>
</tr>
<tr>
<td>1024x1024x512 dim=1</td>
<td>bf16</td>
<td align="right">OOM</td>
<td align="right">311,786.9</td>
<td align="right">inf (OOM before)</td>
</tr>
</tbody>
</table>
<h3>Strided / sliced layouts</h3>
<table>
<thead>
<tr>
<th>config</th>
<th>dtype</th>
<th align="right">before (us)</th>
<th align="right">after (us)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>4096^2.T dim=0</td>
<td>f32</td>
<td align="right">12,089.2</td>
<td align="right">8,482.2</td>
<td align="right">1.43x</td>
</tr>
<tr>
<td>4096^2.T dim=1</td>
<td>f32</td>
<td align="right">16,644.8</td>
<td align="right">10,942.0</td>
<td align="right">1.52x</td>
</tr>
<tr>
<td>4096^2[::2] dim=1</td>
<td>f32</td>
<td align="right">5,158.8</td>
<td align="right">4,526.9</td>
<td align="right">1.14x</td>
</tr>
<tr>
<td>4096^2[:,1k:3k] dim=0</td>
<td>f32</td>
<td align="right">9,151.6</td>
<td align="right">5,053.1</td>
<td align="right">1.81x</td>
</tr>
<tr>
<td>4096^2.T dim=0</td>
<td>bf16</td>
<td align="right">14,767.4</td>
<td align="right">4,472.3</td>
<td align="right">3.30x</td>
</tr>
<tr>
<td>4096^2.T dim=1</td>
<td>bf16</td>
<td align="right">11,008.9</td>
<td align="right">5,410.5</td>
<td align="right">2.03x</td>
</tr>
<tr>
<td>4096^2[::2] dim=1</td>
<td>bf16</td>
<td align="right">4,127.8</td>
<td align="right">2,515.7</td>
<td align="right">1.64x</td>
</tr>
<tr>
<td>4096^2[:,1k:3k] dim=0</td>
<td>bf16</td>
<td align="right">7,348.2</td>
<td align="right">2,566.0</td>
<td align="right">2.86x</td>
</tr>
<tr>
<td>4096^2.T dim=0</td>
<td>i32</td>
<td align="right">9,637.6</td>
<td align="right">3,409.0</td>
<td align="right">2.83x</td>
</tr>
<tr>
<td>4096^2.T dim=1</td>
<td>i32</td>
<td align="right">12,800.6</td>
<td align="right">4,815.9</td>
<td align="right">2.66x</td>
</tr>
<tr>
<td>4096^2[::2] dim=1</td>
<td>i32</td>
<td align="right">3,874.0</td>
<td align="right">1,711.5</td>
<td align="right">2.26x</td>
</tr>
<tr>
<td>4096^2[:,1k:3k] dim=0</td>
<td>i32</td>
<td align="right">7,214.9</td>
<td align="right">2,360.5</td>
<td align="right">3.06x</td>
</tr>
<tr>
<td>4096^2.T dim=0</td>
<td>i64</td>
<td align="right">17,010.4</td>
<td align="right">6,161.7</td>
<td align="right">2.76x</td>
</tr>
<tr>
<td>4096^2.T dim=1</td>
<td>i64</td>
<td align="right">22,469.8</td>
<td align="right">6,447.1</td>
<td align="right">3.49x</td>
</tr>
<tr>
<td>4096^2[::2] dim=1</td>
<td>i64</td>
<td align="right">7,788.3</td>
<td align="right">3,129.9</td>
<td align="right">2.49x</td>
</tr>
<tr>
<td>4096^2[:,1k:3k] dim=0</td>
<td>i64</td>
<td align="right">11,912.6</td>
<td align="right">3,966.6</td>
<td align="right">3.00x</td>
</tr>
</tbody>
</table>
<h3>nanmedian with NaNs</h3>
<table>
<thead>
<tr>
<th>config</th>
<th>dtype</th>
<th align="right">before (us)</th>
<th align="right">after (us)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>nan10% 4096^2 nanmed d1</td>
<td>f32</td>
<td align="right">10,439.2</td>
<td align="right">9,199.2</td>
<td align="right">1.13x</td>
</tr>
<tr>
<td>nan10% 4096^2 nanmed d0</td>
<td>f32</td>
<td align="right">18,530.9</td>
<td align="right">10,941.3</td>
<td align="right">1.69x</td>
</tr>
<tr>
<td>clean 4096^2 nanmed d1</td>
<td>f32</td>
<td align="right">10,470.4</td>
<td align="right">8,214.9</td>
<td align="right">1.27x</td>
</tr>
<tr>
<td>nan10% 4096^2 nanmed d1</td>
<td>bf16</td>
<td align="right">18,870.5</td>
<td align="right">4,368.2</td>
<td align="right">4.32x</td>
</tr>
<tr>
<td>nan10% 4096^2 nanmed d0</td>
<td>bf16</td>
<td align="right">14,820.6</td>
<td align="right">5,037.6</td>
<td align="right">2.94x</td>
</tr>
<tr>
<td>clean 4096^2 nanmed d1</td>
<td>bf16</td>
<td align="right">18,981.2</td>
<td align="right">4,390.5</td>
<td align="right">4.32x</td>
</tr>
<tr>
<td>nan10% 512^3 nanmed d1</td>
<td>f32</td>
<td align="right">415,150.8</td>
<td align="right">72,076.8</td>
<td align="right">5.76x</td>
</tr>
<tr>
<td>nan10% 512^3 nanmed d1</td>
<td>bf16</td>
<td align="right">633,870.6</td>
<td align="right">87,350.0</td>
<td align="right">7.26x</td>
</tr>
</tbody>
</table>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4642356307" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187060" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187060/hovercard" href="https://github.com/pytorch/pytorch/pull/187060">#187060</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/7d28bf8e2bf24163375b6904056caffa5a58242d]]></title>
<description><![CDATA[Add 2.13 to Release Compatibility Matrix and CUDA Support Matrix (#18…]]></description>
<link>https://tsecurity.de/de/3593533/downloads/trunk7d28bf8e2bf24163375b6904056caffa5a58242d/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3593533/downloads/trunk7d28bf8e2bf24163375b6904056caffa5a58242d/</guid>
<pubDate>Fri, 12 Jun 2026 15:08:47 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Add 2.13 to Release Compatibility Matrix and CUDA Support Matrix (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/415662b62d95f2f48455cbc0f7f7943ed7e8a7f6: [dynamo, 3.15] Update test error message (#187102)]]></title>
<description><![CDATA[Pull Request resolved: #187102
Approved by: https://github.com/guilhermeleobas]]></description>
<link>https://tsecurity.de/de/3592866/downloads/trunk415662b62d95f2f48455cbc0f7f7943ed7e8a7f6-dynamo-315-update-test-error-message-187102/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592866/downloads/trunk415662b62d95f2f48455cbc0f7f7943ed7e8a7f6-dynamo-315-update-test-error-message-187102/</guid>
<pubDate>Fri, 12 Jun 2026 10:31:57 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4644880150" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187102" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187102/hovercard" href="https://github.com/pytorch/pytorch/pull/187102">#187102</a><br>
Approved by: <a href="https://github.com/guilhermeleobas">https://github.com/guilhermeleobas</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/79a16845cb3bab008e8eb8177a449f1c430833b0]]></title>
<description><![CDATA[c10d: add one-sided window interfaces to Backend and ProcessGroup (#1…]]></description>
<link>https://tsecurity.de/de/3592491/downloads/trunk79a16845cb3bab008e8eb8177a449f1c430833b0/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592491/downloads/trunk79a16845cb3bab008e8eb8177a449f1c430833b0/</guid>
<pubDate>Fri, 12 Jun 2026 07:31:52 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>c10d: add one-sided window interfaces to Backend and ProcessGroup (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="171281708" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1/hovercard" href="https://github.com/pytorch/pytorch/issues/1">#1</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/e4ebddd3b67f49080147c1f2c4487148b199ea15: [vllm hash update] update the pinned vllm hash (#187114)]]></title>
<description><![CDATA[This PR is auto-generated nightly by this action.
Update the pinned vllm hash.
Pull Request resolved: #187114
Approved by: https://github.com/pytorchbot]]></description>
<link>https://tsecurity.de/de/3592417/downloads/trunke4ebddd3b67f49080147c1f2c4487148b199ea15-vllm-hash-update-update-the-pinned-vllm-hash-187114/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592417/downloads/trunke4ebddd3b67f49080147c1f2c4487148b199ea15-vllm-hash-update-update-the-pinned-vllm-hash-187114/</guid>
<pubDate>Fri, 12 Jun 2026 07:01:36 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This PR is auto-generated nightly by <a href="https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml">this action</a>.<br>
Update the pinned vllm hash.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4645354028" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187114" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187114/hovercard" href="https://github.com/pytorch/pytorch/pull/187114">#187114</a><br>
Approved by: <a href="https://github.com/pytorchbot">https://github.com/pytorchbot</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/ba71580d47143e418ee8882c54bce1d330d7242d: [BE] Make spmd_type a CI rather than CD dependency (#187067)]]></title>
<description><![CDATA[Alas, it could not be added to requirements-ci.txt, as it depends on torch and we don't want any pre-installed torch wheels inside CI docker builds
Pull Request resolved: #187067
Approved by: https://github.com/pianpwk, https://github.com/atalman, https://github.com/fegin]]></description>
<link>https://tsecurity.de/de/3592402/downloads/trunkba71580d47143e418ee8882c54bce1d330d7242d-be-make-spmdtype-a-ci-rather-than-cd-dependency-187067/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592402/downloads/trunkba71580d47143e418ee8882c54bce1d330d7242d-be-make-spmdtype-a-ci-rather-than-cd-dependency-187067/</guid>
<pubDate>Fri, 12 Jun 2026 06:46:42 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Alas, it could not be added to <code>requirements-ci.txt</code>, as it depends on torch and we don't want any pre-installed torch wheels inside CI docker builds</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4643132874" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187067" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187067/hovercard" href="https://github.com/pytorch/pytorch/pull/187067">#187067</a><br>
Approved by: <a href="https://github.com/pianpwk">https://github.com/pianpwk</a>, <a href="https://github.com/atalman">https://github.com/atalman</a>, <a href="https://github.com/fegin">https://github.com/fegin</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/664ed2f6a67ca8e2769b41c8079e916979bb0daa]]></title>
<description><![CDATA[xpu: fix CUTLASS cpp_wrapper compilation with correct code cache (#18…]]></description>
<link>https://tsecurity.de/de/3592273/downloads/trunk664ed2f6a67ca8e2769b41c8079e916979bb0daa/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592273/downloads/trunk664ed2f6a67ca8e2769b41c8079e916979bb0daa/</guid>
<pubDate>Fri, 12 Jun 2026 05:01:41 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>xpu: fix CUTLASS cpp_wrapper compilation with correct code cache (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/e0a79513f30ad45c20c5e37465b31682beacaa3c: cache _is_spmd_types_available (#187071)]]></title>
<description><![CDATA[avoid importlib overhead if any
Pull Request resolved: #187071
Approved by: https://github.com/aditvenk, https://github.com/fegin]]></description>
<link>https://tsecurity.de/de/3592271/downloads/trunke0a79513f30ad45c20c5e37465b31682beacaa3c-cache-isspmdtypesavailable-187071/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592271/downloads/trunke0a79513f30ad45c20c5e37465b31682beacaa3c-cache-isspmdtypesavailable-187071/</guid>
<pubDate>Fri, 12 Jun 2026 05:01:38 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>avoid importlib overhead if any<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4643281118" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187071" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187071/hovercard" href="https://github.com/pytorch/pytorch/pull/187071">#187071</a><br>
Approved by: <a href="https://github.com/aditvenk">https://github.com/aditvenk</a>, <a href="https://github.com/fegin">https://github.com/fegin</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/xpu/186797: [XPU] Fix test_autotune_gemm_choice_validation for mm_plus_mm, add TODO]]></title>
<description><![CDATA[Per Claude code review: removing the Triton mm_plus_mm template on XPU
broke test_autotune_gemm_choice_validation which asserts TritonTemplateCaller
is always present when max_autotune=True. Guard the assertion for the
mm_plus_mm+XPU case.
Also add TODO(#184490) in mm_plus_mm.py to mark the XPU g...]]></description>
<link>https://tsecurity.de/de/3592270/downloads/ciflowxpu186797-xpu-fix-testautotunegemmchoicevalidation-for-mmplusmm-add-todo/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592270/downloads/ciflowxpu186797-xpu-fix-testautotunegemmchoicevalidation-for-mmplusmm-add-todo/</guid>
<pubDate>Fri, 12 Jun 2026 05:01:37 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Per Claude code review: removing the Triton mm_plus_mm template on XPU<br>
broke test_autotune_gemm_choice_validation which asserts TritonTemplateCaller<br>
is always present when max_autotune=True. Guard the assertion for the<br>
mm_plus_mm+XPU case.</p>
<p>Also add TODO(<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4483674345" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184490" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/184490/hovercard" href="https://github.com/pytorch/pytorch/issues/184490">#184490</a>) in mm_plus_mm.py to mark the XPU guard as temporary<br>
— it should be lifted once the underlying Triton codegen accuracy bug is<br>
fixed so XPU can benefit from the fused template perf.</p>
<p>Co-authored-by: Copilot <a href="mailto:223556219+Copilot@users.noreply.github.com">223556219+Copilot@users.noreply.github.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/xpu/186087: [XPU] Keep max_pool2d_with_indices_backward fallback for performance]]></title>
<description><![CDATA[Per review from jianyizh and guangyey: removing the is_xpu fallback and
re-enabling the scatter_add decomposition causes atomic contention
performance regression with overlapping pooling windows. Revert the
decomposition removal and update the comment to document the performance
rationale (not pr...]]></description>
<link>https://tsecurity.de/de/3592269/downloads/ciflowxpu186087-xpu-keep-maxpool2dwithindicesbackward-fallback-for-performance/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592269/downloads/ciflowxpu186087-xpu-keep-maxpool2dwithindicesbackward-fallback-for-performance/</guid>
<pubDate>Fri, 12 Jun 2026 05:01:36 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Per review from jianyizh and guangyey: removing the is_xpu fallback and<br>
re-enabling the scatter_add decomposition causes atomic contention<br>
performance regression with overlapping pooling windows. Revert the<br>
decomposition removal and update the comment to document the performance<br>
rationale (not precision — the kernel fix from torch-xpu-ops <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="275014369" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/3765" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/3765/hovercard" href="https://github.com/pytorch/pytorch/pull/3765">#3765</a> is<br>
landed and pinned).</p>
<p>The fallback stays for performance. Tests reverted accordingly:</p>
<ul>
<li>test_torchinductor_codegen_dynamic_shapes: keep TestFailure entries<br>
(no Triton kernel generated while fallback is active)</li>
<li>test_torchinductor_opinfo: keep inductor_one_sample skip</li>
</ul>
<p>Co-authored-by: Copilot <a href="mailto:223556219+Copilot@users.noreply.github.com">223556219+Copilot@users.noreply.github.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/36f2cea3c3440f40ac205d70ed4e0a0ac5d5ce90: [XPU] Unskip test_float16_reduction_with_int_output (#185199)]]></title>
<description><![CDATA[Summary
Removes the @skipIfXpu decorator from CudaReproTests::test_float16_reduction_with_int_output in test/inductor/test_cuda_repro.py. The skip referenced intel/torch-xpu-ops#3006, where Triton codegen for argmax over float16 inputs was emitting a spurious .to(tl.float16) downcast on the int64...]]></description>
<link>https://tsecurity.de/de/3592257/downloads/trunk36f2cea3c3440f40ac205d70ed4e0a0ac5d5ce90-xpu-unskip-testfloat16reductionwithintoutput-185199/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592257/downloads/trunk36f2cea3c3440f40ac205d70ed4e0a0ac5d5ce90-xpu-unskip-testfloat16reductionwithintoutput-185199/</guid>
<pubDate>Fri, 12 Jun 2026 04:46:27 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Removes the <code>@skipIfXpu</code> decorator from <code>CudaReproTests::test_float16_reduction_with_int_output</code> in <code>test/inductor/test_cuda_repro.py</code>. The skip referenced <a href="https://github.com/intel/torch-xpu-ops/issues/3006" data-hovercard-type="issue" data-hovercard-url="/intel/torch-xpu-ops/issues/3006/hovercard">intel/torch-xpu-ops#3006</a>, where Triton codegen for <code>argmax</code> over float16 inputs was emitting a spurious <code>.to(tl.float16)</code> downcast on the int64 index result. That bug was fixed upstream by <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3898474425" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/174321" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/174321/hovercard" href="https://github.com/pytorch/pytorch/pull/174321">#174321</a> (<code>[inductor] Fix incorrect Triton reduction output dtype</code>), which is already in <code>main</code>, so the XPU skip is no longer needed.</p>
<h2>Verification</h2>
<p>Verified on Intel Data Center GPU Max 1550 (PVC) with <code>torch==2.13.0.dev20260525+xpu</code>:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="source /home/gta/intel/oneapi/setvars.sh
export LD_PRELOAD=/home/gta/intel/oneapi/ccl/2022.0/lib/libccl.so.1
export PYTORCH_TEST_WITH_SLOW=1
pytest -v test/inductor/test_cuda_repro.py \
  -k test_float16_reduction_with_int_output -rs"><pre><span class="pl-c1">source</span> /home/gta/intel/oneapi/setvars.sh
<span class="pl-k">export</span> LD_PRELOAD=/home/gta/intel/oneapi/ccl/2022.0/lib/libccl.so.1
<span class="pl-k">export</span> PYTORCH_TEST_WITH_SLOW=1
pytest -v test/inductor/test_cuda_repro.py \
  -k test_float16_reduction_with_int_output -rs</pre></div>
<p>Result:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="test/inductor/test_cuda_repro.py::CudaReproTests::test_float16_reduction_with_int_output PASSED [16.3236s] [100%]
====================== 1 passed, 110 deselected in 26.62s ======================"><pre class="notranslate"><code>test/inductor/test_cuda_repro.py::CudaReproTests::test_float16_reduction_with_int_output PASSED [16.3236s] [100%]
====================== 1 passed, 110 deselected in 26.62s ======================
</code></pre></div>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4038031747" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3006" data-hovercard-type="issue" data-hovercard-url="/intel/torch-xpu-ops/issues/3006/hovercard" href="https://github.com/intel/torch-xpu-ops/issues/3006">intel/torch-xpu-ops#3006</a>.</p>
<p>Authored by Claude (opencode).</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4522666321" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185199" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185199/hovercard" href="https://github.com/pytorch/pytorch/pull/185199">#185199</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/187083: Move backend-specific c10d files into per-backend subfolders (#187083)]]></title>
<description><![CDATA[Summary:
Pull Request resolved: #187083
Reorganizes torch/csrc/distributed/c10d by moving non-public, backend-specific implementation files and the TCPStore backend files into per-backend subfolders, while leaving the public-facing classes at the top level (the ProcessGroupGloo/NCCL/MPI/UCC backe...]]></description>
<link>https://tsecurity.de/de/3592206/downloads/ciflowtrunk187083-move-backend-specific-c10d-files-into-per-backend-subfolders-187083/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592206/downloads/ciflowtrunk187083-move-backend-specific-c10d-files-into-per-backend-subfolders-187083/</guid>
<pubDate>Fri, 12 Jun 2026 04:01:29 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4644176746" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187083" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187083/hovercard" href="https://github.com/pytorch/pytorch/pull/187083">#187083</a></p>
<p>Reorganizes <code>torch/csrc/distributed/c10d</code> by moving non-public, backend-specific implementation files and the TCPStore backend files into per-backend subfolders, while leaving the public-facing classes at the top level (the <code>ProcessGroupGloo</code>/<code>NCCL</code>/<code>MPI</code>/<code>UCC</code> backends and the <code>Store</code>/<code>TCPStore</code>/<code>FileStore</code>/<code>HashStore</code>/<code>PrefixStore</code> classes all stay put).</p>
<p>The moves are: <code>store/</code> gets <code>TCPStoreBackend.{cpp,hpp}</code> and <code>TCPStoreLibUvBackend.cpp</code>; <code>gloo/</code> gets <code>ProcessGroupGlooCuda.cpp</code>, <code>ProcessGroupGlooDetail.hpp</code>, and <code>GlooDeviceFactory.{cpp,hpp}</code>; <code>ucc/</code> gets <code>UCCTracing.{cpp,hpp}</code> and <code>UCCUtils.{cpp,hpp}</code>; <code>nccl/</code> gets <code>NCCLXStub.hpp</code>.</p>
<p><code>NCCLUtils.{cpp,hpp}</code> was deliberately kept at the top level even though it is backend-specific: it is included by several call sites outside <code>caffe2</code> (in <code>gen_ai</code>, <code>ads_mkl</code>, and <code>fbgemm_gpu</code>), so relocating it would be a wider, riskier change better done on its own. As a result the new <code>nccl/</code> folder currently holds only <code>NCCLXStub.hpp</code>.</p>
<p>All include sites were updated, covering both the canonical <code>torch/csrc/distributed/c10d/...</code> include form and the legacy short <code>c10d/...</code> form (used by <code>fb/GlooDeviceFactory.cpp</code>). Build wiring was updated in <code>build_variables.bzl</code> -- the canonical source list consumed by CMake (via <code>append_filelist</code> in <code>cmake/Codegen.cmake</code>), OSS Bazel, and OSS Buck -- and in the internal <code>fb/fbcode/target_definitions.bzl</code> for <code>ProcessGroupGlooCuda.cpp</code>. Headers are picked up by recursive globs, so no header-list edits were needed.</p>
<p>This is a pure file move: contents are unchanged apart from the relocated <code>#include</code> paths, so correctness is established by a clean build rather than by behavioral tests.</p>
<p>Authored with the assistance of an AI coding assistant (Claude Code).</p>
<p>Test Plan:<br>
Confirmed no references to the old paths remain anywhere in <code>fbcode</code>, then ran the fbcode lint and build tooling:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="arc f
arc lint
arc lint --take AUTODEPS --apply-patches
buck2 build fbcode//caffe2:_libtorch fbcode//caffe2:_libtorch_cuda"><pre class="notranslate"><code>arc f
arc lint
arc lint --take AUTODEPS --apply-patches
buck2 build fbcode//caffe2:_libtorch fbcode//caffe2:_libtorch_cuda
</code></pre></div>
<p><code>arc f</code> and <code>arc lint</code> reported no issues; AUTODEPS produced no dependency changes (the moves stayed within existing Buck targets); both the CPU (<code>_libtorch</code>) and CUDA (<code>_libtorch_cuda</code>) libraries built successfully (exit 0).</p>
<p>Reviewed By: kapilsh</p>
<p>Differential Revision: D108332288</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/183837: [Dynamic Shapes] Rebind unbacked symbols to derived expressions]]></title>
<description><![CDATA[rebind_unbacked already records equivalences when a retraced binding site maps
an unbacked symbol to another symbol or to a constant. The same invariant applies
when the new value is a derived symbolic expression such as (u1 + 1) // 2: the
old symbol still has a concrete binding relationship and ...]]></description>
<link>https://tsecurity.de/de/3592185/downloads/ciflowtrunk183837-dynamic-shapes-rebind-unbacked-symbols-to-derived-expressions/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592185/downloads/ciflowtrunk183837-dynamic-shapes-rebind-unbacked-symbols-to-derived-expressions/</guid>
<pubDate>Fri, 12 Jun 2026 03:46:33 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>rebind_unbacked</code> already records equivalences when a retraced binding site maps<br>
an unbacked symbol to another symbol or to a constant. The same invariant applies<br>
when the new value is a derived symbolic expression such as <code>(u1 + 1) // 2</code>: the<br>
old symbol still has a concrete binding relationship and should be eliminated in<br>
favor of that expression.</p>
<p>The previous assertion assumed any non-symbol replacement with free symbols was<br>
invalid. That is too strong for legitimate derived unbacked shapes and makes the<br>
binding logic reject a value it can represent with the existing ShapeEnv<br>
replacement mechanism. Record the replacement with <code>_eliminate_unbacked</code>, which<br>
is the existing path for replacing an unbacked symbol by a non-symbol expression.</p>
<p>This intentionally does not restore the reverted broad HOP fake-trace<br>
suppression. The FlexAttention/HOP reproducer for <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4444345402" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183677" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/183677/hovercard" href="https://github.com/pytorch/pytorch/issues/183677">#183677</a> now passes on current<br>
main without that suppression, and the broad suppression was the source of the<br>
internal cond, AOTInductor, Executorch, and FlexAttention regressions.</p>
<p>This change was authored with assistance from an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/test_dynamic_shapes.py TestUnbacked.test_rebind_unbacked_to_symbolic_expression -v"><pre>python test/test_dynamic_shapes.py TestUnbacked.test_rebind_unbacked_to_symbolic_expression -v</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/dynamo/test_misc.py MiscTests.test_cond_runtime_assert_generation -v"><pre>python test/dynamo/test_misc.py MiscTests.test_cond_runtime_assert_generation -v</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/functorch/test_control_flow.py TestControlFlowTraced.test_cond_functionalized_nested -v"><pre>python test/functorch/test_control_flow.py TestControlFlowTraced.test_cond_functionalized_nested -v</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_aot_inductor.py -k cond_non_tensor_predicates_dynamic_True_cpu -v"><pre>python test/inductor/test_aot_inductor.py -k cond_non_tensor_predicates_dynamic_True_cpu -v</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python - &lt;&lt;'PY'
import torch
from torch.nn.attention.flex_attention import create_block_mask, flex_attention

if not torch.cuda.is_available():
    raise RuntimeError('CUDA unavailable')

device = 'cuda'
dtype = torch.float16

def score_mod(score, batch, head, q_idx, kv_idx):
    return score

compiled_create_block_mask = torch.compile(create_block_mask, dynamic=True, fullgraph=True)

def create_dynamic_block_mask(q_batch, kv_batch):
    q_len = q_batch.size(0)
    kv_len = kv_batch.size(0)
    def mask_mod(batch, head, q_idx, kv_idx):
        q_group = q_batch[q_idx]
        kv_group = kv_batch[kv_idx]
        return (q_group == kv_group) &amp; (q_group != -1) &amp; (kv_group != -1)
    return compiled_create_block_mask(mask_mod, B=None, H=None, Q_LEN=q_len, KV_LEN=kv_len, device=device, BLOCK_SIZE=128)

groups = torch.zeros(128, dtype=torch.int64, device=device)
block_mask = create_dynamic_block_mask(groups, groups)
q = torch.randn(1, 1, 128, 64, device=device, dtype=dtype, requires_grad=True)
k = torch.randn(1, 1, 128, 64, device=device, dtype=dtype, requires_grad=True)
v = torch.randn(1, 1, 128, 64, device=device, dtype=dtype, requires_grad=True)
compiled_flex_attention = torch.compile(flex_attention, fullgraph=True, dynamic=True, backend='aot_eager')
out = compiled_flex_attention(q, k, v, score_mod=score_mod, block_mask=block_mask)
out.sum().backward()
torch.cuda.synchronize()
print('ok', out.shape)
PY"><pre>python - <span class="pl-s"><span class="pl-k">&lt;&lt;</span>'<span class="pl-k">PY</span>'</span>
<span class="pl-s">import torch</span>
<span class="pl-s">from torch.nn.attention.flex_attention import create_block_mask, flex_attention</span>
<span class="pl-s"></span>
<span class="pl-s">if not torch.cuda.is_available():</span>
<span class="pl-s">    raise RuntimeError('CUDA unavailable')</span>
<span class="pl-s"></span>
<span class="pl-s">device = 'cuda'</span>
<span class="pl-s">dtype = torch.float16</span>
<span class="pl-s"></span>
<span class="pl-s">def score_mod(score, batch, head, q_idx, kv_idx):</span>
<span class="pl-s">    return score</span>
<span class="pl-s"></span>
<span class="pl-s">compiled_create_block_mask = torch.compile(create_block_mask, dynamic=True, fullgraph=True)</span>
<span class="pl-s"></span>
<span class="pl-s">def create_dynamic_block_mask(q_batch, kv_batch):</span>
<span class="pl-s">    q_len = q_batch.size(0)</span>
<span class="pl-s">    kv_len = kv_batch.size(0)</span>
<span class="pl-s">    def mask_mod(batch, head, q_idx, kv_idx):</span>
<span class="pl-s">        q_group = q_batch[q_idx]</span>
<span class="pl-s">        kv_group = kv_batch[kv_idx]</span>
<span class="pl-s">        return (q_group == kv_group) &amp; (q_group != -1) &amp; (kv_group != -1)</span>
<span class="pl-s">    return compiled_create_block_mask(mask_mod, B=None, H=None, Q_LEN=q_len, KV_LEN=kv_len, device=device, BLOCK_SIZE=128)</span>
<span class="pl-s"></span>
<span class="pl-s">groups = torch.zeros(128, dtype=torch.int64, device=device)</span>
<span class="pl-s">block_mask = create_dynamic_block_mask(groups, groups)</span>
<span class="pl-s">q = torch.randn(1, 1, 128, 64, device=device, dtype=dtype, requires_grad=True)</span>
<span class="pl-s">k = torch.randn(1, 1, 128, 64, device=device, dtype=dtype, requires_grad=True)</span>
<span class="pl-s">v = torch.randn(1, 1, 128, 64, device=device, dtype=dtype, requires_grad=True)</span>
<span class="pl-s">compiled_flex_attention = torch.compile(flex_attention, fullgraph=True, dynamic=True, backend='aot_eager')</span>
<span class="pl-s">out = compiled_flex_attention(q, k, v, score_mod=score_mod, block_mask=block_mask)</span>
<span class="pl-s">out.sum().backward()</span>
<span class="pl-s">torch.cuda.synchronize()</span>
<span class="pl-s">print('ok', out.shape)</span>
<span class="pl-s"><span class="pl-k">PY</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>stack-info: PR: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851792" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183837" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183837/hovercard" href="https://github.com/pytorch/pytorch/pull/183837">#183837</a>, branch: sanketpurandare/stack/11</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/a4097e577fe5d1e21dfe2fa8c36af3fdf8854e34: Revert "[BE] Make spmd_type a CI rather than CD dependency (#187067)"]]></title>
<description><![CDATA[This reverts commit d4c98cd.
Reverted #187067 on behalf of https://github.com/pytorch-auto-revert due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (comment)]]></description>
<link>https://tsecurity.de/de/3592173/downloads/trunka4097e577fe5d1e21dfe2fa8c36af3fdf8854e34-revert-be-make-spmdtype-a-ci-rather-than-cd-dependency-187067/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592173/downloads/trunka4097e577fe5d1e21dfe2fa8c36af3fdf8854e34-revert-be-make-spmdtype-a-ci-rather-than-cd-dependency-187067/</guid>
<pubDate>Fri, 12 Jun 2026 03:16:34 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/d4c98cdffd561ed6c1d55473271e19a091063a37/hovercard" href="https://github.com/pytorch/pytorch/commit/d4c98cdffd561ed6c1d55473271e19a091063a37"><tt>d4c98cd</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4643132874" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/187067" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187067/hovercard" href="https://github.com/pytorch/pytorch/pull/187067">#187067</a> on behalf of <a href="https://github.com/pytorch-auto-revert">https://github.com/pytorch-auto-revert</a> due to Reverted automatically by pytorch's autorevert, to avoid this behaviour add the tag autorevert: disable (<a href="https://github.com/pytorch/pytorch/pull/187067#issuecomment-4686401401" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/187067/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5ffde693e13e101c8a4f5ea685dfbaef0c7e7466]]></title>
<description><![CDATA[[c10] Make basic_string_view inherit from std::basic_string_view (#18…]]></description>
<link>https://tsecurity.de/de/3592135/downloads/trunk5ffde693e13e101c8a4f5ea685dfbaef0c7e7466/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592135/downloads/trunk5ffde693e13e101c8a4f5ea685dfbaef0c7e7466/</guid>
<pubDate>Fri, 12 Jun 2026 02:46:00 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[c10] Make basic_string_view inherit from std::basic_string_view (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/1807281cecd9a6506a8156e27d02dd50440e7235: [Inductor] Fix TMA compat to check returned options, not class (#185983)]]></title>
<description><![CDATA[Summary:
[Inductor] Fix TMA compat to check returned options, not class
Summary
Fix a bug in the TMA compatibility check that prevented backends from falling back to BlockPtrOptions, causing unwanted scalar-pointer fallbacks and PE timeouts.
Problem
The TMA compatibility check at triton.py:3595 u...]]></description>
<link>https://tsecurity.de/de/3592069/downloads/trunk1807281cecd9a6506a8156e27d02dd50440e7235-inductor-fix-tma-compat-to-check-returned-options-not-class-185983/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3592069/downloads/trunk1807281cecd9a6506a8156e27d02dd50440e7235-inductor-fix-tma-compat-to-check-returned-options-not-class-185983/</guid>
<pubDate>Fri, 12 Jun 2026 01:45:56 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:<br>
[Inductor] Fix TMA compat to check returned options, not class</p>
<p><strong>Summary</strong><br>
Fix a bug in the TMA compatibility check that prevented backends from falling back to BlockPtrOptions, causing unwanted scalar-pointer fallbacks and PE timeouts.</p>
<p><strong>Problem</strong><br>
The TMA compatibility check at triton.py:3595 used <code>issubclass(options_class, TensorDescriptorOptions)</code> on the <em>queried</em> class instead of the <em>returned</em> options. When backends like MTIA override <code>TensorDescriptorOptions.create</code> to fall back to <code>BlockPtrOptions</code> (for shapes incompatible with hardware layout), the check incorrectly fired anyway, rejecting block params and triggering Triton's scalar-pointer fallback.</p>
<p><strong>Solution</strong><br>
Changed to <code>isinstance(options, TensorDescriptorOptions)</code> to check the actual returned type. Now when a backend's <code>create</code> returns <code>BlockPtrOptions</code>, the TMA check is correctly skipped, allowing the block pointer to flow to <code>codegen_block_ptr</code> and backends to register autotune metadata (e.g., MTIA's <code>inner_blocks</code> for bump-to-minimum heuristic).</p>
<p><strong>Impact</strong></p>
<ul>
<li>Fixes PE timeouts caused by unwanted scalar-pointer fallbacks</li>
<li>Enables backends to properly use their custom autotune heuristics</li>
<li>Small change (1 line in 2 files) with targeted fix</li>
</ul>
<p>Test Plan:</p>
<ul>
<li>Ran <code>buck2 test fbcode//mode/opt fbcode//mtia/compiler/graph_compiler/inductor/test:test_inductor_kernels -- --regex 'test_bump_block_size'</code></li>
<li>Before fix: XBLOCK=2 for int8/float16/float32 (heuristic skipped because <code>inner_blocks</code> empty)</li>
<li>After fix: XBLOCK bumped to 32/16/8 per dtype, matching the 32-byte hardware floor</li>
<li><strong>TODO: Paste full test output showing the XBLOCK values before/after</strong></li>
</ul>
<p>Differential Revision: D106090810</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4574397097" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185983" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185983/hovercard" href="https://github.com/pytorch/pytorch/pull/185983">#185983</a><br>
Approved by: <a href="https://github.com/blaine-rister">https://github.com/blaine-rister</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781213548: Add PyTorch QuACK GEMM epilogue adapter e.g FlexGemm (#186483)]]></title>
<description><![CDATA[Summary
Groundwork PR

Adds the hop
Adds the quack impl
Only supports pointwise
no aux buffers
All this to make follow up prs easier to review

sanity check

Notes for reviewer
API;
# Do be made public later
 from torch._higher_order_ops import flex_gemm

 out = flex_gemm(
     gemm_op,          ...]]></description>
<link>https://tsecurity.de/de/3591898/downloads/viablestrict1781213548-add-pytorch-quack-gemm-epilogue-adapter-eg-flexgemm-186483/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3591898/downloads/viablestrict1781213548-add-pytorch-quack-gemm-epilogue-adapter-eg-flexgemm-186483/</guid>
<pubDate>Thu, 11 Jun 2026 23:46:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Groundwork PR</p>
<ol>
<li>Adds the hop</li>
<li>Adds the quack impl</li>
<li>Only supports pointwise</li>
<li>no aux buffers</li>
<li>All this to make follow up prs easier to review</li>
</ol>
<p>sanity check<br>
<a target="_blank" rel="noopener noreferrer" href="https://private-user-images.githubusercontent.com/32754868/605346325-eafd7127-4584-4c1b-9349-668239f77727.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODEyMTQ2OTAsIm5iZiI6MTc4MTIxNDM5MCwicGF0aCI6Ii8zMjc1NDg2OC82MDUzNDYzMjUtZWFmZDcxMjctNDU4NC00YzFiLTkzNDktNjY4MjM5Zjc3NzI3LnBuZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNjA2MTElMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjYwNjExVDIxNDYzMFomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPWZmOTczOTlhZjIxNTNkMGI2ZDhiM2NhMjEwY2JkYzU1MWIyMmFiNTFjZDRkNzAzMTc2YmI5OTkxNDNlZGZmZTUmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0JnJlc3BvbnNlLWNvbnRlbnQtdHlwZT1pbWFnZSUyRnBuZyJ9.lMiiMggM2OhhBZSovIks6u2xVELAEmDnPU04PLDOX28"><img width="2000" height="349" alt="image" src="https://private-user-images.githubusercontent.com/32754868/605346325-eafd7127-4584-4c1b-9349-668239f77727.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODEyMTQ2OTAsIm5iZiI6MTc4MTIxNDM5MCwicGF0aCI6Ii8zMjc1NDg2OC82MDUzNDYzMjUtZWFmZDcxMjctNDU4NC00YzFiLTkzNDktNjY4MjM5Zjc3NzI3LnBuZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNjA2MTElMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjYwNjExVDIxNDYzMFomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPWZmOTczOTlhZjIxNTNkMGI2ZDhiM2NhMjEwY2JkYzU1MWIyMmFiNTFjZDRkNzAzMTc2YmI5OTkxNDNlZGZmZTUmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0JnJlc3BvbnNlLWNvbnRlbnQtdHlwZT1pbWFnZSUyRnBuZyJ9.lMiiMggM2OhhBZSovIks6u2xVELAEmDnPU04PLDOX28" content-type-secured-asset="image/png"></a></p>
<h3>Notes for reviewer</h3>
<p>API;</p>
<div class="highlight highlight-source-python notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="# Do be made public later
 from torch._higher_order_ops import flex_gemm

 out = flex_gemm(
     gemm_op,          # torch.mm / torch.addmm / aten op overload expanded to a wider set scaled_mm, grouped friends
     gemm_args,        # tuple of tensor/scalar operands orignal args ot he base func
     epilogue_fn,      # Python fn over accumulator/result
     *,
     gemm_kwargs=None, # non-tensor kwargs only
     kernel_options={},  # dict of options, will be backend, fastmath, tune
 )
"><pre><span class="pl-c"># Do be made public later</span>
 <span class="pl-k">from</span> <span class="pl-s1">torch</span>.<span class="pl-s1">_higher_order_ops</span> <span class="pl-k">import</span> <span class="pl-s1">flex_gemm</span>

 <span class="pl-s1">out</span> <span class="pl-c1">=</span> <span class="pl-en">flex_gemm</span>(
     <span class="pl-s1">gemm_op</span>,          <span class="pl-c"># torch.mm / torch.addmm / aten op overload expanded to a wider set scaled_mm, grouped friends</span>
     <span class="pl-s1">gemm_args</span>,        <span class="pl-c"># tuple of tensor/scalar operands orignal args ot he base func</span>
     <span class="pl-s1">epilogue_fn</span>,      <span class="pl-c"># Python fn over accumulator/result</span>
     <span class="pl-c1">*</span><span class="pl-s1"></span>,
     <span class="pl-s1">gemm_kwargs</span><span class="pl-c1">=</span><span class="pl-c1">None</span>, <span class="pl-c"># non-tensor kwargs only</span>
     <span class="pl-s1">kernel_options</span><span class="pl-c1">=</span>{},  <span class="pl-c"># dict of options, will be backend, fastmath, tune</span>
 )</pre></div>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4604765393" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186483" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186483/hovercard" href="https://github.com/pytorch/pytorch/pull/186483">#186483</a><br>
Approved by: <a href="https://github.com/mlazos">https://github.com/mlazos</a>, <a href="https://github.com/eellison">https://github.com/eellison</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4633136494" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186944" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186944/hovercard" href="https://github.com/pytorch/pytorch/pull/186944">#186944</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/906979aa9de162759f996c8d859d5ec82d1faa79: Fix QuACK vendoring source and patch split (#186944)]]></title>
<description><![CDATA[Summary
Pointed quack at my upstream fork, pointing back to main and using patchsets as intended
Pull Request resolved: #186944
Approved by: https://github.com/slayton58]]></description>
<link>https://tsecurity.de/de/3591263/downloads/trunk906979aa9de162759f996c8d859d5ec82d1faa79-fix-quack-vendoring-source-and-patch-split-186944/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3591263/downloads/trunk906979aa9de162759f996c8d859d5ec82d1faa79-fix-quack-vendoring-source-and-patch-split-186944/</guid>
<pubDate>Thu, 11 Jun 2026 18:32:05 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Pointed quack at my upstream fork, pointing back to main and using patchsets as intended</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4633136494" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186944" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186944/hovercard" href="https://github.com/pytorch/pytorch/pull/186944">#186944</a><br>
Approved by: <a href="https://github.com/slayton58">https://github.com/slayton58</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/19791183fec14fa4a6dc3a82004ed29cd4dc6704]]></title>
<description><![CDATA[[CUDA graphs] Annotate kernels across graphs captured in sequence (#1…]]></description>
<link>https://tsecurity.de/de/3591262/downloads/trunk19791183fec14fa4a6dc3a82004ed29cd4dc6704/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3591262/downloads/trunk19791183fec14fa4a6dc3a82004ed29cd4dc6704/</guid>
<pubDate>Thu, 11 Jun 2026 18:32:04 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[CUDA graphs] Annotate kernels across graphs captured in sequence (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="171281708" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1/hovercard" href="https://github.com/pytorch/pytorch/issues/1">#1</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5935675e1b4bd686068b4f7d7fc664fb38bb0119]]></title>
<description><![CDATA[Revert "Fix Dynamo set operations over pre-existing generators (#1860…]]></description>
<link>https://tsecurity.de/de/3590114/downloads/trunk5935675e1b4bd686068b4f7d7fc664fb38bb0119/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3590114/downloads/trunk5935675e1b4bd686068b4f7d7fc664fb38bb0119/</guid>
<pubDate>Thu, 11 Jun 2026 12:17:02 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Revert "Fix Dynamo set operations over pre-existing generators (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="237453334" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1860" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1860/hovercard" href="https://github.com/pytorch/pytorch/issues/1860">#1860</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/719055fbee1551db2b6efcb2d089e23c305fb8f4: Fall back to eager _scaled_mm_v2 for swizzled scale layouts (#186384)]]></title>
<description><![CDATA[TorchInductor _scaled_mm_v2 lowering had no template or extern path for non-trivial swizzle patterns and asserted during compilation, breaking blockwise MXFP8/NVFP4 compile tests on Blackwell.
For swizzled scale layouts, defer to the eager _scaled_mm_v2 op before mm_args. The trivial-swizzle path...]]></description>
<link>https://tsecurity.de/de/3589698/downloads/trunk719055fbee1551db2b6efcb2d089e23c305fb8f4-fall-back-to-eager-scaledmmv2-for-swizzled-scale-layouts-186384/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589698/downloads/trunk719055fbee1551db2b6efcb2d089e23c305fb8f4-fall-back-to-eager-scaledmmv2-for-swizzled-scale-layouts-186384/</guid>
<pubDate>Thu, 11 Jun 2026 09:02:33 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>TorchInductor <code>_scaled_mm_v2</code> lowering had no template or extern path for non-trivial swizzle patterns and asserted during compilation, breaking blockwise MXFP8/NVFP4 compile tests on Blackwell.</p>
<p>For swizzled scale layouts, defer to the eager <code>_scaled_mm_v2</code> op before mm_args. The trivial-swizzle path is unchanged.</p>
<p>Authored with Claude Opus 4.8.</p>
<p>rel: <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/9eb8bcb2f55f6c5c7e65d8a098127cba1be75b5c/hovercard" href="https://github.com/pytorch/pytorch/commit/9eb8bcb2f55f6c5c7e65d8a098127cba1be75b5c"><tt>9eb8bcb</tt></a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4598254984" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186384" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186384/hovercard" href="https://github.com/pytorch/pytorch/pull/186384">#186384</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/8537afcf2575f43ceecf660c16912c733e81bd8c: Fix Dynamo random.shuffle and random.sample (#186499)]]></title>
<description><![CDATA[Dynamo could not trace random.shuffle or random.sample, on either an
explicit random.Random object or the module-level helpers. The methods
live in the skip-listed random module, so calls raised "Attempted to
call function marked as skipped". This blocked the CPython
test_dict-DictTest.test_liter...]]></description>
<link>https://tsecurity.de/de/3589651/downloads/trunk8537afcf2575f43ceecf660c16912c733e81bd8c-fix-dynamo-randomshuffle-and-randomsample-186499/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589651/downloads/trunk8537afcf2575f43ceecf660c16912c733e81bd8c-fix-dynamo-randomshuffle-and-randomsample-186499/</guid>
<pubDate>Thu, 11 Jun 2026 08:32:00 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Dynamo could not trace random.shuffle or random.sample, on either an<br>
explicit random.Random object or the module-level helpers. The methods<br>
live in the skip-listed random module, so calls raised "Attempted to<br>
call function marked as skipped". This blocked the CPython<br>
test_dict-DictTest.test_literal_constructor gate, whose only obstacle<br>
was using random.sample/random.shuffle to build its test data.</p>
<p>Authored with assistance from Claude Code.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4605454728" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186499" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186499/hovercard" href="https://github.com/pytorch/pytorch/pull/186499">#186499</a><br>
Approved by: <a href="https://github.com/rtimpe">https://github.com/rtimpe</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4575081030" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185998" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185998/hovercard" href="https://github.com/pytorch/pytorch/pull/185998">#185998</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4575081561" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185999" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185999/hovercard" href="https://github.com/pytorch/pytorch/pull/185999">#185999</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4576971253" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186042" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186042/hovercard" href="https://github.com/pytorch/pytorch/pull/186042">#186042</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/646f5c9c27d4293b87fbef024c59209f26c4e4b6: Fix Dynamo defaultdict shallow copy (copy.copy / __copy__) (#186758)]]></title>
<description><![CDATA[copy.copy(d) on a defaultdict resolves type(d).copy and calls it.
Under Dynamo this became UserDefinedClassVariable(defaultdict).call_method(
"copy"), which fell through to the bare dict type via
SourcelessBuilder(dict). Since dict has no copy, tracing failed with
"Dynamo does not know how to tra...]]></description>
<link>https://tsecurity.de/de/3589650/downloads/trunk646f5c9c27d4293b87fbef024c59209f26c4e4b6-fix-dynamo-defaultdict-shallow-copy-copycopy-copy-186758/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589650/downloads/trunk646f5c9c27d4293b87fbef024c59209f26c4e4b6-fix-dynamo-defaultdict-shallow-copy-copycopy-copy-186758/</guid>
<pubDate>Thu, 11 Jun 2026 08:31:58 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>copy.copy(d) on a defaultdict resolves type(d).<strong>copy</strong> and calls it.<br>
Under Dynamo this became UserDefinedClassVariable(defaultdict).call_method(<br>
"<strong>copy</strong>"), which fell through to the bare dict type via<br>
SourcelessBuilder(dict). Since dict has no <strong>copy</strong>, tracing failed with<br>
"Dynamo does not know how to trace method <strong>copy</strong>".</p>
<p>The fix handles <strong>copy</strong> in two places. DefaultDictVariable.call_method now<br>
treats "<strong>copy</strong>" alongside "copy", returning a fresh defaultdict that<br>
preserves default_factory and performs a shallow copy of the base dict.<br>
UserDefinedClassVariable.call_method gains a narrow branch that dispatches<br>
defaultdict.<strong>copy</strong>(inst) (one positional arg, no kwargs) to the instance,<br>
mirroring the existing OrderedDict.move_to_end pattern.</p>
<p>Test Plan:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="CUDA_VISIBLE_DEVICES= PYTORCH_TESTING_DEVICE_ONLY_FOR=cpu PYTORCH_TEST_WITH_DYNAMO=1 \
  pixi run -w pytorch -e pytorch313 python -m pytest test/cpython/v3_13/test_defaultdict.py -q"><pre class="notranslate"><code>CUDA_VISIBLE_DEVICES= PYTORCH_TESTING_DEVICE_ONLY_FOR=cpu PYTORCH_TEST_WITH_DYNAMO=1 \
  pixi run -w pytorch -e pytorch313 python -m pytest test/cpython/v3_13/test_defaultdict.py -q
</code></pre></div>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="CUDA_VISIBLE_DEVICES= PYTORCH_TESTING_DEVICE_ONLY_FOR=cpu \
  pixi run -w pytorch -e pytorch313 python -m pytest test/dynamo/test_dicts.py -q"><pre class="notranslate"><code>CUDA_VISIBLE_DEVICES= PYTORCH_TESTING_DEVICE_ONLY_FOR=cpu \
  pixi run -w pytorch -e pytorch313 python -m pytest test/dynamo/test_dicts.py -q
</code></pre></div>
<p>Authored with the help of an AI assistant.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4618559271" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186758" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186758/hovercard" href="https://github.com/pytorch/pytorch/pull/186758">#186758</a><br>
Approved by: <a href="https://github.com/rtimpe">https://github.com/rtimpe</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4575081030" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185998" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185998/hovercard" href="https://github.com/pytorch/pytorch/pull/185998">#185998</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4575081561" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185999" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185999/hovercard" href="https://github.com/pytorch/pytorch/pull/185999">#185999</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4576971253" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186042" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186042/hovercard" href="https://github.com/pytorch/pytorch/pull/186042">#186042</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4605454728" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186499" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186499/hovercard" href="https://github.com/pytorch/pytorch/pull/186499">#186499</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5e3ec847f80ed0f0a3936094608793a61bf1380c: [vllm hash update] update the pinned vllm hash (#186858)]]></title>
<description><![CDATA[This PR is auto-generated nightly by this action.
Update the pinned vllm hash.
Pull Request resolved: #186858
Approved by: https://github.com/pytorchbot]]></description>
<link>https://tsecurity.de/de/3589545/downloads/trunk5e3ec847f80ed0f0a3936094608793a61bf1380c-vllm-hash-update-update-the-pinned-vllm-hash-186858/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589545/downloads/trunk5e3ec847f80ed0f0a3936094608793a61bf1380c-vllm-hash-update-update-the-pinned-vllm-hash-186858/</guid>
<pubDate>Thu, 11 Jun 2026 07:21:31 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This PR is auto-generated nightly by <a href="https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml">this action</a>.<br>
Update the pinned vllm hash.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4626919533" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186858" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186858/hovercard" href="https://github.com/pytorch/pytorch/pull/186858">#186858</a><br>
Approved by: <a href="https://github.com/pytorchbot">https://github.com/pytorchbot</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/020bf1f03e7684e962dced71f6b7e5f6d5b30245]]></title>
<description><![CDATA[[inductor] fix remainder for fp16/bf16 tensors and scalar inputs (#18…]]></description>
<link>https://tsecurity.de/de/3589544/downloads/trunk020bf1f03e7684e962dced71f6b7e5f6d5b30245/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589544/downloads/trunk020bf1f03e7684e962dced71f6b7e5f6d5b30245/</guid>
<pubDate>Thu, 11 Jun 2026 07:21:22 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[inductor] fix remainder for fp16/bf16 tensors and scalar inputs (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/053eaef2c546cc3dfe715cc995a70d2b20e6a075: Fix Windows cpp wrapper int array FileCheck (#185145)]]></title>
<description><![CDATA[The Windows CPU cpp-wrapper repro in #162366 still had a test expectation that assumed generated C++ int64 array literals always use the L suffix. MSVC-generated code uses LL for these literals, and the test already had the same platform split elsewhere through target_assert_size_stride_str. That...]]></description>
<link>https://tsecurity.de/de/3589543/downloads/trunk053eaef2c546cc3dfe715cc995a70d2b20e6a075-fix-windows-cpp-wrapper-int-array-filecheck-185145/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589543/downloads/trunk053eaef2c546cc3dfe715cc995a70d2b20e6a075-fix-windows-cpp-wrapper-int-array-filecheck-185145/</guid>
<pubDate>Thu, 11 Jun 2026 07:21:10 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The Windows CPU cpp-wrapper repro in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3392761828" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/162366" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/162366/hovercard" href="https://github.com/pytorch/pytorch/issues/162366">#162366</a> still had a test expectation that assumed generated C++ int64 array literals always use the <code>L</code> suffix. MSVC-generated code uses <code>LL</code> for these literals, and the test already had the same platform split elsewhere through <code>target_assert_size_stride_str</code>. That made <code>CPUReproTests.test_require_stride_order_non_owning</code> fail on Windows even when the generated allocation and stride-order behavior was correct.</p>
<p>Factor the cpp-wrapper int-array formatting into <code>cpp_int_array_str()</code> and reuse it in both the existing assert-size-stride helper and the repro test. This keeps the expectation tied to the platform-specific generated C++ spelling instead of hard-coding Linux formatting in the Windows test path.</p>
<p>The attached convolution_backward C-shim failures from the issue are already addressed on current main by <code>1b421fae114 [inductor] Fix MSVC const pointer emission in cpp wrapper temporary arrays (#179846)</code>, so this change is limited to the remaining Windows-sensitive FileCheck expectation.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3392761828" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/162366" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/162366/hovercard" href="https://github.com/pytorch/pytorch/issues/162366">#162366</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>TORCHINDUCTOR_CPP_WRAPPER=1 python - &lt;&lt;'PY'<br>
import runpy<br>
import sys<br>
sys.modules['torchvision'] = None<br>
sys.path.insert(0, 'test/inductor')<br>
sys.argv = ['test/inductor/test_cpu_repro.py', 'CPUReproTests.test_require_stride_order_non_owning']<br>
runpy.run_path('test/inductor/test_cpu_repro.py', run_name='<strong>main</strong>')<br>
PY</li>
<li>python - &lt;&lt;'PY'<br>
import sys<br>
sys.modules['torchvision'] = None<br>
sys.path.insert(0, 'test/inductor')<br>
import test_torchinductor<br>
orig = sys.platform<br>
try:<br>
sys.platform = 'win32'<br>
assert test_torchinductor.cpp_int_array_str([2, 3, 4, 4]) == '{2LL, 3LL, 4LL, 4LL}'<br>
sys.platform = 'linux'<br>
assert test_torchinductor.cpp_int_array_str([2, 3, 4, 4]) == '{2L, 3L, 4L, 4L}'<br>
finally:<br>
sys.platform = orig<br>
print('cpp_int_array_str platform formatting OK')<br>
PY</li>
<li>TORCHINDUCTOR_CPP_WRAPPER=1 python - &lt;&lt;'PY'<br>
import runpy<br>
import sys<br>
sys.modules['torchvision'] = None<br>
sys.path.insert(0, 'test/inductor')<br>
sys.argv = ['test/inductor/test_torchinductor.py', 'CpuTests.test_conv_backward_cpu']<br>
runpy.run_path('test/inductor/test_torchinductor.py', run_name='<strong>main</strong>')<br>
PY</li>
<li>TORCHINDUCTOR_CPP_WRAPPER=1 python - &lt;&lt;'PY'<br>
import runpy<br>
import sys<br>
sys.modules['torchvision'] = None<br>
sys.path.insert(0, 'test/inductor')<br>
sys.argv = ['test/inductor/test_torchinductor.py', 'CpuTests.test_conv2d_backward_channels_last_cpu']<br>
runpy.run_path('test/inductor/test_torchinductor.py', run_name='<strong>main</strong>')<br>
PY</li>
<li>python - &lt;&lt;'PY'<br>
import runpy<br>
import sys<br>
sys.modules['torchvision'] = None<br>
sys.path.insert(0, 'test/inductor')<br>
sys.argv = ['test/inductor/test_cpu_repro.py', 'CPUReproTests.test_require_stride_order_non_owning']<br>
runpy.run_path('test/inductor/test_cpu_repro.py', run_name='<strong>main</strong>')<br>
PY</li>
<li>lintrunner -a</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4518056793" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185145" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185145/hovercard" href="https://github.com/pytorch/pytorch/pull/185145">#185145</a><br>
Approved by: <a href="https://github.com/desertfire">https://github.com/desertfire</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5c2969a1182d020d3387c03c6de90d488cfaad1f: Fix clang-tidy warnings in CUDACachingAllocator.cpp (#186308)]]></title>
<description><![CDATA[Pull Request resolved: #186308
Approved by: https://github.com/malfet, https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3589533/downloads/trunk5c2969a1182d020d3387c03c6de90d488cfaad1f-fix-clang-tidy-warnings-in-cudacachingallocatorcpp-186308/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589533/downloads/trunk5c2969a1182d020d3387c03c6de90d488cfaad1f-fix-clang-tidy-warnings-in-cudacachingallocatorcpp-186308/</guid>
<pubDate>Thu, 11 Jun 2026 07:02:21 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4593604870" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186308" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186308/hovercard" href="https://github.com/pytorch/pytorch/pull/186308">#186308</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a>, <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/6428e60868d8a93ec1df4475ac09c46b53acc4ea: [dtensor] migrating matrix_ops to single dim strategies (#186667)]]></title>
<description><![CDATA[**Summary: ** Adds 11 single dim strategies, 2 of which were not registered before.
Test Cases

pytest test/distributed/tensor/test_matrix_ops.py
pytest test/distributed/tensor/test_op_strategy.py -k test_redistribute_cost_latency
pytest test/distributed/tensor/test_op_strategy.py -k test_redistr...]]></description>
<link>https://tsecurity.de/de/3589441/downloads/trunk6428e60868d8a93ec1df4475ac09c46b53acc4ea-dtensor-migrating-matrixops-to-single-dim-strategies-186667/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589441/downloads/trunk6428e60868d8a93ec1df4475ac09c46b53acc4ea-dtensor-migrating-matrixops-to-single-dim-strategies-186667/</guid>
<pubDate>Thu, 11 Jun 2026 05:31:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>**Summary: ** Adds 11 single dim strategies, 2 of which were not registered before.</p>
<p><strong>Test Cases</strong></p>
<ol>
<li>pytest test/distributed/tensor/test_matrix_ops.py</li>
<li>pytest test/distributed/tensor/test_op_strategy.py -k test_redistribute_cost_latency</li>
<li>pytest test/distributed/tensor/test_op_strategy.py -k test_redistribute_cost_mesh_2d</li>
<li>pytest test/distributed/tensor/test_op_strategy.py -k test_bmm_strategies</li>
</ol>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4617233070" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186667" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186667/hovercard" href="https://github.com/pytorch/pytorch/pull/186667">#186667</a><br>
Approved by: <a href="https://github.com/pianpwk">https://github.com/pianpwk</a>, <a href="https://github.com/weifengpy">https://github.com/weifengpy</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/fbcec9b828dfc73354428c4168d51736e566f192: [ptd_triage_bot] preventing duplicated ptd triage bot comments (#186966)]]></title>
<description><![CDATA[Pull Request resolved: #186966
Approved by: https://github.com/aditvenk]]></description>
<link>https://tsecurity.de/de/3589415/downloads/trunkfbcec9b828dfc73354428c4168d51736e566f192-ptdtriagebot-preventing-duplicated-ptd-triage-bot-comments-186966/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589415/downloads/trunkfbcec9b828dfc73354428c4168d51736e566f192-ptdtriagebot-preventing-duplicated-ptd-triage-bot-comments-186966/</guid>
<pubDate>Thu, 11 Jun 2026 05:01:27 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4634593997" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186966" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186966/hovercard" href="https://github.com/pytorch/pytorch/pull/186966">#186966</a><br>
Approved by: <a href="https://github.com/aditvenk">https://github.com/aditvenk</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/ff2350bab01e85090bf019269498ab6b0e953e3d: [MPS] Handle empty indexes in index_add (#186990)]]></title>
<description><![CDATA[Fixes #186972
Pull Request resolved: #186990
Approved by: https://github.com/Isalia20, https://github.com/izaitsevfb, https://github.com/atalman]]></description>
<link>https://tsecurity.de/de/3589333/downloads/trunkff2350bab01e85090bf019269498ab6b0e953e3d-mps-handle-empty-indexes-in-indexadd-186990/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589333/downloads/trunkff2350bab01e85090bf019269498ab6b0e953e3d-mps-handle-empty-indexes-in-indexadd-186990/</guid>
<pubDate>Thu, 11 Jun 2026 03:46:01 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4634899750" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186972" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186972/hovercard" href="https://github.com/pytorch/pytorch/issues/186972">#186972</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4635544631" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186990" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186990/hovercard" href="https://github.com/pytorch/pytorch/pull/186990">#186990</a><br>
Approved by: <a href="https://github.com/Isalia20">https://github.com/Isalia20</a>, <a href="https://github.com/izaitsevfb">https://github.com/izaitsevfb</a>, <a href="https://github.com/atalman">https://github.com/atalman</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/c20d62a6125f1fe187b4d17a9b8a96a84984ecf3: [Profiler] Fix a set of profiler tests (#186970)]]></title>
<description><![CDATA[A couple changes for pytorch/kineto#1429 and pytorch/kineto#1430:

Add setUpModule to make sure Kineto is created with CUDA support. #186036 recently changed the behavior to not initialize Kineto with GPU tracing if only CPU activities are requested, which is more correct behavior. However, this ...]]></description>
<link>https://tsecurity.de/de/3589331/downloads/trunkc20d62a6125f1fe187b4d17a9b8a96a84984ecf3-profiler-fix-a-set-of-profiler-tests-186970/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589331/downloads/trunkc20d62a6125f1fe187b4d17a9b8a96a84984ecf3-profiler-fix-a-set-of-profiler-tests-186970/</guid>
<pubDate>Thu, 11 Jun 2026 03:45:59 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>A couple changes for <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4630695392" data-permission-text="Title is private" data-url="https://github.com/pytorch/kineto/issues/1429" data-hovercard-type="issue" data-hovercard-url="/pytorch/kineto/issues/1429/hovercard" href="https://github.com/pytorch/kineto/issues/1429">pytorch/kineto#1429</a> and <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4630712695" data-permission-text="Title is private" data-url="https://github.com/pytorch/kineto/issues/1430" data-hovercard-type="issue" data-hovercard-url="/pytorch/kineto/issues/1430/hovercard" href="https://github.com/pytorch/kineto/issues/1430">pytorch/kineto#1430</a>:</p>
<ol>
<li>Add <code>setUpModule</code> to make sure Kineto is created with CUDA support. <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4576764274" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186036" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186036/hovercard" href="https://github.com/pytorch/pytorch/pull/186036">#186036</a> recently changed the behavior to not initialize Kineto with GPU tracing if only CPU activities are requested, which is more correct behavior. However, this changes the behavior of the following code:</li>
</ol>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="with profile(activities=[ProfilerActivity.CPU]) as prof:
   ...

with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
   ..."><pre class="notranslate"><code>with profile(activities=[ProfilerActivity.CPU]) as prof:
   ...

with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
   ...
</code></pre></div>
<p>Now Kineto is armed to only do CPU tracing. The issue is that Kineto is only registered once, so even though you added ProfilerActivity.CUDA in the second pass, the profiler still latches on to the CPU-only session, leading to no device-side activities being recorded. We need to find a more complete solution but to unblock the tests we can hack around it by forcing Kineto to instantiate with GPU-tracing before running any tests.</p>
<ol start="2">
<li>Clear KinetoStepTracker to make sure repeat runs with different activity types don't pollute <code>test_kineto_profiler_multiple_steppers</code>.</li>
<li>Make sure <code>test_forked_process</code> explicitly runs with CPU activities only.</li>
</ol>
<p>There are a few more failures due to test pollution, but I haven't tracked those down yet.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4634710412" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186970" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186970/hovercard" href="https://github.com/pytorch/pytorch/pull/186970">#186970</a><br>
Approved by: <a href="https://github.com/scotts">https://github.com/scotts</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/ce9267a4585256d1465a00a6a7687ce76f3784a5: Fetch tags in unified manywheel build job so release tags are detected]]></title>
<description><![CDATA[Release candidate builds triggered by a tag (e.g. v2.13.0-rc1) were
producing wheels that pinned triton to a dev-style version such as
triton==3.7.1+git5d6048aa instead of the release triton==3.7.1, which
then fails to install in the test job because that version does not
exist on the test index....]]></description>
<link>https://tsecurity.de/de/3589327/downloads/trunkce9267a4585256d1465a00a6a7687ce76f3784a5-fetch-tags-in-unified-manywheel-build-job-so-release-tags-are-detected/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589327/downloads/trunkce9267a4585256d1465a00a6a7687ce76f3784a5-fetch-tags-in-unified-manywheel-build-job-so-release-tags-are-detected/</guid>
<pubDate>Thu, 11 Jun 2026 03:31:31 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Release candidate builds triggered by a tag (e.g. v2.13.0-rc1) were<br>
producing wheels that pinned triton to a dev-style version such as<br>
triton==3.7.1+git5d6048aa instead of the release triton==3.7.1, which<br>
then fails to install in the test job because that version does not<br>
exist on the test index.</p>
<p>Root cause: the unified inline manywheel build job (cpu/cpu-aarch64/<br>
cuda/cuda-aarch64) checks out the raw commit via <code>ref: github.sha</code> with<br>
<code>fetch-depth: 2</code>. Checking out a bare SHA does not fetch the tag ref, so<br>
<code>tagged_version()</code> in .ci/pytorch/binary_populate_env.sh<br>
(<code>git describe --tags --exact</code>) fails and the build falls through to the<br>
<code>&lt;version&gt;.dev&lt;DATE&gt;</code> default. binary_populate_env.sh then takes its<br>
<code>.*dev.*</code> branch and appends <code>+git&lt;triton-shorthash&gt;</code> to the triton<br>
requirement that gets baked into the wheel metadata.</p>
<p>Fix: add <code>fetch-tags: true</code> to the inline build job checkout in the<br>
linux binary build template so the tag pointing at the checked-out<br>
commit is fetched, letting <code>git describe --tags --exact</code> succeed and the<br>
build resolve to the release version. The ROCm/XPU/s390x builds use the<br>
reusable _binary-build-linux.yml whose checkout has no pinned ref (it<br>
uses the default tag ref for a tag push) and are unaffected.</p>
<p>This is the same class of bug fixed for the older checkout-pytorch path<br>
in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4271932621" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/180508" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/180508/hovercard" href="https://github.com/pytorch/pytorch/pull/180508">#180508</a>; this covers the inlined unified build job introduced since.</p>
<p>Test Plan:</p>
<p>Regenerated the workflows from the template and confirmed only the two<br>
linux manywheel workflows change, each gaining <code>fetch-tags: true</code> on its<br>
four unified build jobs:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python3 .github/scripts/generate_ci_workflows.py
git diff --stat .github/workflows
grep -c 'fetch-tags: true' \
  .github/workflows/generated-linux-binary-manywheel-nightly.yml \
  .github/workflows/generated-linux-aarch64-binary-manywheel-nightly.yml"><pre class="notranslate"><code>python3 .github/scripts/generate_ci_workflows.py
git diff --stat .github/workflows
grep -c 'fetch-tags: true' \
  .github/workflows/generated-linux-binary-manywheel-nightly.yml \
  .github/workflows/generated-linux-aarch64-binary-manywheel-nightly.yml
</code></pre></div>
<p>Validated the regenerated YAML parses and passes actionlint (the only<br>
reported shellcheck warnings are pre-existing on the base files):</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python3 -c "import yaml; yaml.safe_load(open(f)) for f in (...)"
lintrunner --take ACTIONLINT \
  .github/workflows/generated-linux-binary-manywheel-nightly.yml \
  .github/workflows/generated-linux-aarch64-binary-manywheel-nightly.yml'><pre class="notranslate"><code>python3 -c "import yaml; yaml.safe_load(open(f)) for f in (...)"
lintrunner --take ACTIONLINT \
  .github/workflows/generated-linux-binary-manywheel-nightly.yml \
  .github/workflows/generated-linux-aarch64-binary-manywheel-nightly.yml
</code></pre></div>
<p>This PR was authored with the assistance of an AI coding assistant.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/8ff6c00bd7c492840847f9205cfdf936fe1584c8: Use C++20 std::numbers in c10 MathConstants (#186877)]]></title>
<description><![CDATA[Pull Request resolved: #186877
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3589313/downloads/trunk8ff6c00bd7c492840847f9205cfdf936fe1584c8-use-c-20-stdnumbers-in-c10-mathconstants-186877/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589313/downloads/trunk8ff6c00bd7c492840847f9205cfdf936fe1584c8-use-c-20-stdnumbers-in-c10-mathconstants-186877/</guid>
<pubDate>Thu, 11 Jun 2026 03:16:43 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4627310339" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186877" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186877/hovercard" href="https://github.com/pytorch/pytorch/pull/186877">#186877</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/9973ec93082df39bffc41a1d0e26d368e4defef0: [Operators] Implement nb_power / nb_inplace_power in Dynamo (#186296)]]></title>
<description><![CDATA[Wires operator.pow and operator.ipow through ternary_op /
ternary_iop in BuiltinVariable, removing them from the old
_handle_op_in_graph table. Adds nb_power_impl (with reverse)
and nb_inplace_power_impl to VariableTracker, ConstantVariable,
TensorVariable, SymNodeVariable, and UserDefinedObjectV...]]></description>
<link>https://tsecurity.de/de/3589280/downloads/trunk9973ec93082df39bffc41a1d0e26d368e4defef0-operators-implement-nbpower-nbinplacepower-in-dynamo-186296/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589280/downloads/trunk9973ec93082df39bffc41a1d0e26d368e4defef0-operators-implement-nbpower-nbinplacepower-in-dynamo-186296/</guid>
<pubDate>Thu, 11 Jun 2026 02:31:33 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Wires <code>operator.pow</code> and <code>operator.ipow</code> through <code>ternary_op</code> /<br>
<code>ternary_iop</code> in <code>BuiltinVariable</code>, removing them from the old<br>
<code>_handle_op_in_graph</code> table. Adds <code>nb_power_impl</code> (with <code>reverse</code>)<br>
and <code>nb_inplace_power_impl</code> to <code>VariableTracker</code>, <code>ConstantVariable</code>,<br>
<code>TensorVariable</code>, <code>SymNodeVariable</code>, and <code>UserDefinedObjectVariable</code>.</p>
<p>Power is ternary, so the modulus-position slot (CPython's z-slot) gets<br>
its own <code>nb_power_z_impl</code> method to avoid conflating it with the<br>
forward/reflected roles that <code>reverse: bool</code> captures. <code>ternary_op</code><br>
dispatches the z-slot through <code>nb_power_z_impl(self=z, tx, v, w)</code>.</p>
<p>Co-authored-by: Claude Sonnet</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4593270467" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186296" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186296/hovercard" href="https://github.com/pytorch/pytorch/pull/186296">#186296</a><br>
Approved by: <a href="https://github.com/rtimpe">https://github.com/rtimpe</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/63f903c3d6b04c7cb1433d1d67e2b8e21c055bc7: [cuDNN][SDPA] d=256 support for cuDNN SDPA (#185553)]]></title>
<description><![CDATA[authored with codex, requires FE upgrade as well
Pull Request resolved: #185553
Approved by: https://github.com/Skylion007]]></description>
<link>https://tsecurity.de/de/3589278/downloads/trunk63f903c3d6b04c7cb1433d1d67e2b8e21c055bc7-cudnnsdpa-d256-support-for-cudnn-sdpa-185553/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589278/downloads/trunk63f903c3d6b04c7cb1433d1d67e2b8e21c055bc7-cudnnsdpa-d256-support-for-cudnn-sdpa-185553/</guid>
<pubDate>Thu, 11 Jun 2026 02:31:30 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>authored with codex, requires FE upgrade as well</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4543147964" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185553" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185553/hovercard" href="https://github.com/pytorch/pytorch/pull/185553">#185553</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/186104: [Inductor] Allow output-input aliasing in torch.cond branches (#186104)]]></title>
<description><![CDATA[Summary:
Allow output-input aliasing in torch.cond branches.
torch.cond branches are mutually exclusive — only one branch executes at runtime. When the true branch returns an operand directly (e.g., identity/skip path), the output aliases the input. This is safe because the non-taken branch never...]]></description>
<link>https://tsecurity.de/de/3589236/downloads/ciflowtrunk186104-inductor-allow-output-input-aliasing-in-torchcond-branches-186104/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589236/downloads/ciflowtrunk186104-inductor-allow-output-input-aliasing-in-torchcond-branches-186104/</guid>
<pubDate>Thu, 11 Jun 2026 02:16:41 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:</p>
<p>Allow output-input aliasing in <code>torch.cond</code> branches.</p>
<p>torch.cond branches are mutually exclusive — only one branch executes at runtime. When the true branch returns an operand directly (e.g., identity/skip path), the output aliases the input. This is safe because the non-taken branch never runs, so there is no concurrent access to aliased buffers.</p>
<p>Previously, two checks rejected this pattern:</p>
<ol>
<li><strong>Dynamo</strong> (<code>higher_order_ops.py</code>): <code>CondHigherOrderVariable.supports_aliasing = False</code> rejected aliasing during tracing with a graph break.</li>
<li><strong>Functionalization</strong> (<code>cond.py</code>): <code>_check_alias_and_mutation</code> rejected aliasing during AOT functionalization.</li>
</ol>
<p>Note: Dynamo already allowed aliasing in inference (<code>supports_aliasing = not torch.is_grad_enabled()</code>), but <code>torch.export</code> traces with grad enabled even for inference models, so the check was ineffective for the AOTI export path.</p>
<p>Fix:</p>
<ul>
<li>Set <code>supports_aliasing = True</code> unconditionally on <code>CondHigherOrderVariable</code>. Aliasing is safe for cond regardless of grad context because branches are exclusive — even with autograd, only the taken branch's gradient flows.</li>
<li>Skip the alias check in the functionalization impl (keep mutation check).</li>
</ul>
<p>Motivation: PrismNet's <code>_clone_layer_outputs</code> uses <code>lo.clone()</code> to break output-input aliasing required by torch.cond. Removing this constraint saves ~0.046ms per inference (2 DHEN layers at B=2048).</p>
<p>Adds <code>test_cond_output_input_aliasing</code> regression test.</p>
<p>Test Plan:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="buck2 test fbcode//mode/opt -c 'fbcode.nvcc_arch=a100,h100' fbcode//caffe2/test/inductor:control_flow -- --exact 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_input_mutation_still_rejected (caffe2.test.inductor.test_control_flow.CondTests)'"><pre class="notranslate"><code>buck2 test fbcode//mode/opt -c 'fbcode.nvcc_arch=a100,h100' fbcode//caffe2/test/inductor:control_flow -- --exact 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_input_mutation_still_rejected (caffe2.test.inductor.test_control_flow.CondTests)'
</code></pre></div>
<p>Test session: <a href="https://www.internalfb.com/intern/testinfra/testrun/10414574313882076" rel="nofollow">https://www.internalfb.com/intern/testinfra/testrun/10414574313882076</a></p>
<p>PrismNet repro — all 11 tests pass, including <code>CondOutputInputAlias: without clone</code> which previously failed with "Replace return input with return input.clone()". Confirms <code>lo.clone()</code> workaround in <code>_clone_layer_outputs</code> can be removed.</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="buck2 run fbcode//mode/opt fbcode//aps_models/ads/gmp/models/ads_mtml_prism_net_dedicated_model/experimental/benchmarks:repro_aoti_inductor_bugs"><pre class="notranslate"><code>buck2 run fbcode//mode/opt fbcode//aps_models/ads/gmp/models/ads_mtml_prism_net_dedicated_model/experimental/benchmarks:repro_aoti_inductor_bugs
</code></pre></div>
<p>Reviewed By: kqfu, atalman</p>
<p>Differential Revision: D106106380</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/torchtitan/186104: [Inductor] Allow output-input aliasing in torch.cond branches (#186104)]]></title>
<description><![CDATA[Summary:
Allow output-input aliasing in torch.cond branches.
torch.cond branches are mutually exclusive — only one branch executes at runtime. When the true branch returns an operand directly (e.g., identity/skip path), the output aliases the input. This is safe because the non-taken branch never...]]></description>
<link>https://tsecurity.de/de/3589235/downloads/ciflowtorchtitan186104-inductor-allow-output-input-aliasing-in-torchcond-branches-186104/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589235/downloads/ciflowtorchtitan186104-inductor-allow-output-input-aliasing-in-torchcond-branches-186104/</guid>
<pubDate>Thu, 11 Jun 2026 02:16:40 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:</p>
<p>Allow output-input aliasing in <code>torch.cond</code> branches.</p>
<p>torch.cond branches are mutually exclusive — only one branch executes at runtime. When the true branch returns an operand directly (e.g., identity/skip path), the output aliases the input. This is safe because the non-taken branch never runs, so there is no concurrent access to aliased buffers.</p>
<p>Previously, two checks rejected this pattern:</p>
<ol>
<li><strong>Dynamo</strong> (<code>higher_order_ops.py</code>): <code>CondHigherOrderVariable.supports_aliasing = False</code> rejected aliasing during tracing with a graph break.</li>
<li><strong>Functionalization</strong> (<code>cond.py</code>): <code>_check_alias_and_mutation</code> rejected aliasing during AOT functionalization.</li>
</ol>
<p>Note: Dynamo already allowed aliasing in inference (<code>supports_aliasing = not torch.is_grad_enabled()</code>), but <code>torch.export</code> traces with grad enabled even for inference models, so the check was ineffective for the AOTI export path.</p>
<p>Fix:</p>
<ul>
<li>Set <code>supports_aliasing = True</code> unconditionally on <code>CondHigherOrderVariable</code>. Aliasing is safe for cond regardless of grad context because branches are exclusive — even with autograd, only the taken branch's gradient flows.</li>
<li>Skip the alias check in the functionalization impl (keep mutation check).</li>
</ul>
<p>Motivation: PrismNet's <code>_clone_layer_outputs</code> uses <code>lo.clone()</code> to break output-input aliasing required by torch.cond. Removing this constraint saves ~0.046ms per inference (2 DHEN layers at B=2048).</p>
<p>Adds <code>test_cond_output_input_aliasing</code> regression test.</p>
<p>Test Plan:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="buck2 test fbcode//mode/opt -c 'fbcode.nvcc_arch=a100,h100' fbcode//caffe2/test/inductor:control_flow -- --exact 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_input_mutation_still_rejected (caffe2.test.inductor.test_control_flow.CondTests)'"><pre class="notranslate"><code>buck2 test fbcode//mode/opt -c 'fbcode.nvcc_arch=a100,h100' fbcode//caffe2/test/inductor:control_flow -- --exact 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_multi_output_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cpu (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_output_input_aliasing_both_branches_device_cuda (caffe2.test.inductor.test_control_flow.CondTests)' 'fbcode//caffe2/test/inductor:control_flow - test_cond_input_mutation_still_rejected (caffe2.test.inductor.test_control_flow.CondTests)'
</code></pre></div>
<p>Test session: <a href="https://www.internalfb.com/intern/testinfra/testrun/10414574313882076" rel="nofollow">https://www.internalfb.com/intern/testinfra/testrun/10414574313882076</a></p>
<p>PrismNet repro — all 11 tests pass, including <code>CondOutputInputAlias: without clone</code> which previously failed with "Replace return input with return input.clone()". Confirms <code>lo.clone()</code> workaround in <code>_clone_layer_outputs</code> can be removed.</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="buck2 run fbcode//mode/opt fbcode//aps_models/ads/gmp/models/ads_mtml_prism_net_dedicated_model/experimental/benchmarks:repro_aoti_inductor_bugs"><pre class="notranslate"><code>buck2 run fbcode//mode/opt fbcode//aps_models/ads/gmp/models/ads_mtml_prism_net_dedicated_model/experimental/benchmarks:repro_aoti_inductor_bugs
</code></pre></div>
<p>Reviewed By: kqfu, atalman</p>
<p>Differential Revision: D106106380</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781128641]]></title>
<description><![CDATA[[ROCM][inductor][UT] Enable Inductor-Triton debug asserts on ROCm (#1…]]></description>
<link>https://tsecurity.de/de/3589040/downloads/viablestrict1781128641/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3589040/downloads/viablestrict1781128641/</guid>
<pubDate>Thu, 11 Jun 2026 00:01:52 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[ROCM][inductor][UT] Enable Inductor-Triton debug asserts on ROCm (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="171281708" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1/hovercard" href="https://github.com/pytorch/pytorch/issues/1">#1</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781125794: Fix bool mask IoU comparison for Detectron2 outputs (#186441)]]></title>
<description><![CDATA[CPU Inductor freezing can produce tiny numeric differences in Detectron2 mask logits. Those differences are thresholded into boolean pred_masks, so exact bool equality is too strict for these segmentation outputs. The benchmark harness already has an IoU-based bool mask comparator for this class ...]]></description>
<link>https://tsecurity.de/de/3588910/downloads/viablestrict1781125794-fix-bool-mask-iou-comparison-for-detectron2-outputs-186441/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588910/downloads/viablestrict1781125794-fix-bool-mask-iou-comparison-for-detectron2-outputs-186441/</guid>
<pubDate>Wed, 10 Jun 2026 23:16:48 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>CPU Inductor freezing can produce tiny numeric differences in Detectron2 mask logits. Those differences are thresholded into boolean pred_masks, so exact bool equality is too strict for these segmentation outputs. The benchmark harness already has an IoU-based bool mask comparator for this class of model, but torch._dynamo.utils.same() dropped use_iou_for_bool and iou_threshold when recursing through Detectron2 Instances objects. As a result, Instances.pred_masks still used exact equality and the affected Mask R-CNN models remained fail_accuracy under freezing.</p>
<p>Propagate the bool-mask IoU options through the object and numpy recursive same() paths, enable the existing IoU comparator for the two affected Detectron2 Mask R-CNN models, and update the CPU freezing expected-accuracy files to pass.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2314668187" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/127073" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/127073/hovercard" href="https://github.com/pytorch/pytorch/issues/127073">#127073</a></p>
<p>Generated by my agent</p>
<p>Test Plan:</p>
<p>python test/dynamo/test_utils.py -k TestUtils.test_same_iou_for_bool_propagates_to_instances</p>
<p>python -m unittest benchmarks.dynamo.test.TestDynamoBenchmark.test_detectron2_maskrcnn_uses_iou_for_bool_masks</p>
<p>Local Detectron2 benchmark workaround for detectron2_maskrcnn_r_101_fpn and detectron2_maskrcnn_r_50_c4 with --accuracy --inference --device cpu --inductor --freezing reported pass</p>
<p>git diff --check</p>
<p>lintrunner -a (fails on unrelated existing Pyrefly error in torch/autograd/graph.py for torch._C._remove_obj_from_tls)</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4601453837" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186441" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186441/hovercard" href="https://github.com/pytorch/pytorch/pull/186441">#186441</a><br>
Approved by: <a href="https://github.com/eellison">https://github.com/eellison</a>, <a href="https://github.com/mlazos">https://github.com/mlazos</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[v2.13.0-rc1: [release 2.13] Apply Release only changes to 2.13 branch (#186959)]]></title>
<description><![CDATA[[release 2.13] Apply Release only changes to 2.13 branch

Release-only changes for the release/2.13 branch cut, produced by
running scripts/release/apply-release-changes.sh. The script repoints
reusable workflows and composite actions from @main to @release/2.13,
rewrites templates to release/2.1...]]></description>
<link>https://tsecurity.de/de/3588809/downloads/v2130-rc1-release-213-apply-release-only-changes-to-213-branch-186959/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588809/downloads/v2130-rc1-release-213-apply-release-only-changes-to-213-branch-186959/</guid>
<pubDate>Wed, 10 Jun 2026 22:16:55 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<ul>
<li>[release 2.13] Apply Release only changes to 2.13 branch</li>
</ul>
<p>Release-only changes for the release/2.13 branch cut, produced by<br>
running scripts/release/apply-release-changes.sh. The script repoints<br>
reusable workflows and composite actions from <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/main/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/main">@main</a> to @release/2.13,<br>
rewrites templates to release/2.13 (with checkout_pr_head=False so PRs<br>
build the merge base rather than the PR head), pins the XLA checkout to<br>
the r2.13 branch, pins the disabled/unstable jobs and disabled-tests S3<br>
JSON blobs to fixed versionIds, sets RELEASE_VERSION_TAG=2.13 for the<br>
workflow-regeneration lint check, and drops the pull_request-specific<br>
checkout ref from the linux binary build/test workflows.</p>
<p>Only files that are tracked in release/2.13 are included. The 2.12 PR<br>
additionally touched .ci/manywheel/build_cuda.sh and a few workflow<br>
files (torchbench/nitpicker/quantization-periodic) that are not tracked<br>
in this checkout, so they are intentionally omitted here.</p>
<p>Test Plan:<br>
Ran the release script and linter:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="DRY_RUN=disabled ./scripts/release/apply-release-changes.sh
lintrunner -a"><pre class="notranslate"><code>DRY_RUN=disabled ./scripts/release/apply-release-changes.sh
lintrunner -a
</code></pre></div>
<p>lintrunner reported only pre-existing ACTIONLINT shellcheck warnings on<br>
generated workflow lines unrelated to the release-version edits, and made<br>
no changes to the staged files. Verified every staged hunk is a<br>
release-only edit (<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/main/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/main">@main</a> -&gt; @release/2.13, main -&gt; release/2.13, XLA<br>
r2.13 pin, S3 versionId pins, RELEASE_VERSION_TAG=2.13, and the<br>
pull_request ref removal) with no submodule or unrelated changes.</p>
<p>This PR was authored with the assistance of Claude Code.</p>
<ul>
<li>[release 2.13] Pin Linux manywheel builder docker images</li>
</ul>
<p>Release builds should use a fixed, reproducible build toolchain instead<br>
of the floating builder image tags that main tracks (e.g.<br>
pytorch/manylinux2_28-builder:cuda12.6). For 2.12 the binary build<br>
workflows resolved the image dynamically via calculate-docker-image; for<br>
2.13 we freeze that resolved image as a literal pin.</p>
<p>The pin is applied in the generator (generate_binary_build_matrix.py)<br>
rather than in the generated YAML directly, so re-running<br>
.github/regenerate.sh (and the lint job that asserts the generated files<br>
are up to date) reproduces the pinned tags. wheel_container_image_tag_prefix()<br>
appends the pin only for the linux manywheel OSes (linux, linux-aarch64,<br>
linux-s390x), since only those builds run inside these containers;<br>
windows and macos keep the plain tag prefix.</p>
<p>The pin suffix is the .ci/docker tree hash, f38ba0b10220982e39441d29d203d803a2b56c92<br>
(git rev-parse HEAD:.ci/docker), which is exactly the tag that<br>
.github/actions/binary-docker-build (and the s390x equivalent) publish<br>
as ${prefix}-${CI_FOLDER_SHA}. It matches the image already published on<br>
Docker Hub, e.g.<br>
pytorch/manylinux2_28_aarch64-builder:cpu-aarch64-f38ba0b10220982e39441d29d203d803a2b56c92</p>
<p>Test Plan:<br>
Regenerated the workflows in release mode and verified the result:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="RELEASE_VERSION_TAG=2.13 python3 .github/scripts/generate_ci_workflows.py"><pre class="notranslate"><code>RELEASE_VERSION_TAG=2.13 python3 .github/scripts/generate_ci_workflows.py
</code></pre></div>
<ul>
<li>Every <code>image:</code> and matrix <code>docker_image_tag_prefix</code> in the three linux<br>
manywheel generated workflows now carries the <code>-f38ba0b...</code> suffix<br>
(cpu, cpu-aarch64, cuda12.6/13.0/13.2, rocm7.1/7.2, xpu, cpu-s390x);<br>
no linux builder image remains floating.</li>
<li>Windows and macos generated workflows are unchanged (plain <code>cpu</code>,<br>
<code>cuda12.6</code>, ...), confirming the pin is linux-only.</li>
<li>Re-running the generator a second time produced byte-identical output<br>
(md5sum unchanged), so the regeneration/lint up-to-date check is<br>
stable.</li>
</ul>
<p>Note: the local linters that shell out to <code>uv</code> could not run in this<br>
environment; the change was format-checked manually.</p>
<p>This PR was authored with the assistance of Claude Code.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781115490: Remove deprecated torch.cholesky (#186817)]]></title>
<description><![CDATA[The time has come to remove deprecated linear algebra related functions. This PR removes torch.cholesky.
Fresh replacement for #70979 based on current upstream/main to avoid stale/conflicting old PR state.
Pull Request resolved: #186817
Approved by: https://github.com/zou3519]]></description>
<link>https://tsecurity.de/de/3588637/downloads/viablestrict1781115490-remove-deprecated-torchcholesky-186817/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588637/downloads/viablestrict1781115490-remove-deprecated-torchcholesky-186817/</guid>
<pubDate>Wed, 10 Jun 2026 20:31:37 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The time has come to remove deprecated linear algebra related functions. This PR removes torch.cholesky.</p>
<p>Fresh replacement for <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="1096205245" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/70979" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/70979/hovercard" href="https://github.com/pytorch/pytorch/pull/70979">#70979</a> based on current upstream/main to avoid stale/conflicting old PR state.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4624434760" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186817" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186817/hovercard" href="https://github.com/pytorch/pytorch/pull/186817">#186817</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781110191: Optimize mutable custom op version bump dispatch (#186175)]]></title>
<description><![CDATA[Mutable torch.library custom ops install an ADInplaceOrView shim to bump
the version counter for mutated inputs. The shim already precomputed which
schema arguments were mutable, but it still called utils.fill_defaults on
every invocation. That walked every schema argument before bumping only the...]]></description>
<link>https://tsecurity.de/de/3588430/downloads/viablestrict1781110191-optimize-mutable-custom-op-version-bump-dispatch-186175/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588430/downloads/viablestrict1781110191-optimize-mutable-custom-op-version-bump-dispatch-186175/</guid>
<pubDate>Wed, 10 Jun 2026 19:02:41 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Mutable torch.library custom ops install an ADInplaceOrView shim to bump<br>
the version counter for mutated inputs. The shim already precomputed which<br>
schema arguments were mutable, but it still called utils.fill_defaults on<br>
every invocation. That walked every schema argument before bumping only the<br>
mutated values, so mutable custom op overhead grew with total argument count.</p>
<p>Use the dispatcher-provided args and kwargs directly in the hot path, and skip<br>
omitted mutated arguments instead of materializing or incrementing their default<br>
values. Default values for mutated arguments are expected to be non-Tensors or<br>
fresh values, so there is no useful version counter to bump when the caller<br>
omitted them.</p>
<p>This substantially reduces the overhead reported in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2629589745" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/139494" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/139494/hovercard" href="https://github.com/pytorch/pytorch/issues/139494">#139494</a>, but does not claim<br>
to fully close the remaining gap between mutable and non-mutating custom ops.</p>
<p>The alternative was to optimize utils.fill_defaults itself, but that helper is<br>
general schema normalization and the mutable version-bump path only needs the<br>
small subset of mutated arguments that were actually provided.</p>
<p>Generated by my agent</p>
<p>Benchmark Results:<br>
Before, issue benchmark on CUDA H100, 1000 calls per do_bench function:</p>
<ul>
<li>mutate2 = 18.27797737121582</li>
<li>no_mutate2 = 13.457284654889788</li>
<li>mutate = 66.9056625366211</li>
<li>no_mutate = 17.066380310058594</li>
</ul>
<p>After:</p>
<ul>
<li>mutate2 = 16.13766403198242</li>
<li>no_mutate2 = 13.513686997549874</li>
<li>mutate = 21.02779197692871</li>
<li>no_mutate = 17.362232971191407</li>
</ul>
<p>Test Plan:</p>
<ul>
<li>python test/test_custom_ops.py TestCustomOpAPI.test_mutated_version_bump_does_not_fill_all_defaults TestCustomOpAPI.test_mutated_optional_arg_default_none TestCustomOpAPI.test_mutated TestCustomOpAPI.test_custom_op_out_tag</li>
<li>lintrunner -a<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4585264042" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186175" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186175/hovercard" href="https://github.com/pytorch/pytorch/pull/186175">#186175</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a></li>
</ul>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/16491f83cd79d8f05f68f1cb909b34a62140d5ed: Fix bool mask IoU comparison for Detectron2 outputs (#186441)]]></title>
<description><![CDATA[CPU Inductor freezing can produce tiny numeric differences in Detectron2 mask logits. Those differences are thresholded into boolean pred_masks, so exact bool equality is too strict for these segmentation outputs. The benchmark harness already has an IoU-based bool mask comparator for this class ...]]></description>
<link>https://tsecurity.de/de/3588429/downloads/trunk16491f83cd79d8f05f68f1cb909b34a62140d5ed-fix-bool-mask-iou-comparison-for-detectron2-outputs-186441/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588429/downloads/trunk16491f83cd79d8f05f68f1cb909b34a62140d5ed-fix-bool-mask-iou-comparison-for-detectron2-outputs-186441/</guid>
<pubDate>Wed, 10 Jun 2026 19:02:40 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>CPU Inductor freezing can produce tiny numeric differences in Detectron2 mask logits. Those differences are thresholded into boolean pred_masks, so exact bool equality is too strict for these segmentation outputs. The benchmark harness already has an IoU-based bool mask comparator for this class of model, but torch._dynamo.utils.same() dropped use_iou_for_bool and iou_threshold when recursing through Detectron2 Instances objects. As a result, Instances.pred_masks still used exact equality and the affected Mask R-CNN models remained fail_accuracy under freezing.</p>
<p>Propagate the bool-mask IoU options through the object and numpy recursive same() paths, enable the existing IoU comparator for the two affected Detectron2 Mask R-CNN models, and update the CPU freezing expected-accuracy files to pass.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2314668187" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/127073" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/127073/hovercard" href="https://github.com/pytorch/pytorch/issues/127073">#127073</a></p>
<p>Generated by my agent</p>
<p>Test Plan:</p>
<p>python test/dynamo/test_utils.py -k TestUtils.test_same_iou_for_bool_propagates_to_instances</p>
<p>python -m unittest benchmarks.dynamo.test.TestDynamoBenchmark.test_detectron2_maskrcnn_uses_iou_for_bool_masks</p>
<p>Local Detectron2 benchmark workaround for detectron2_maskrcnn_r_101_fpn and detectron2_maskrcnn_r_50_c4 with --accuracy --inference --device cpu --inductor --freezing reported pass</p>
<p>git diff --check</p>
<p>lintrunner -a (fails on unrelated existing Pyrefly error in torch/autograd/graph.py for torch._C._remove_obj_from_tls)</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4601453837" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186441" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186441/hovercard" href="https://github.com/pytorch/pytorch/pull/186441">#186441</a><br>
Approved by: <a href="https://github.com/eellison">https://github.com/eellison</a>, <a href="https://github.com/mlazos">https://github.com/mlazos</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781106114: [TESTING] Append rules to sample skips and xfails (#185905)]]></title>
<description><![CDATA[Since not long ago, an attempt of running pytest test/functorch/test_vmap.py -v ends up in an error:
  File "/opt/pytorch/pytorch/test/functorch/test_vmap.py", line 6579, in 
    instantiate_device_type_tests(TestVmapOperatorsOpInfo, globals(), only_for=only_for)
  File "/usr/local/lib/python3.12...]]></description>
<link>https://tsecurity.de/de/3588213/downloads/viablestrict1781106114-testing-append-rules-to-sample-skips-and-xfails-185905/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588213/downloads/viablestrict1781106114-testing-append-rules-to-sample-skips-and-xfails-185905/</guid>
<pubDate>Wed, 10 Jun 2026 17:46:36 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Since not long ago, an attempt of running <code>pytest test/functorch/test_vmap.py -v</code> ends up in an error:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='  File "/opt/pytorch/pytorch/test/functorch/test_vmap.py", line 6579, in &lt;module&gt;
    instantiate_device_type_tests(TestVmapOperatorsOpInfo, globals(), only_for=only_for)
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/common_device_type.py", line 1074, in instantiate_device_type_tests
    device_type_test_class.instantiate_test(
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/common_device_type.py", line 634, in instantiate_test
    instantiate_test_helper(
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/common_device_type.py", line 540, in instantiate_test_helper
    test = decorator(test)
           ^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/opinfo/core.py", line 1713, in __call__
    raise RuntimeError("Multiple sets of sample_skips_and_xfails defined")
RuntimeError: Multiple sets of sample_skips_and_xfails defined
================================================================================== short test summary info ===================================================================================
ERROR ../opt/pytorch/pytorch/test/functorch/test_vmap.py - RuntimeError: Multiple sets of sample_skips_and_xfails defined
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
====================================================================================== 1 error in 3.91s ======================================================================================'><pre class="notranslate"><code>  File "/opt/pytorch/pytorch/test/functorch/test_vmap.py", line 6579, in &lt;module&gt;
    instantiate_device_type_tests(TestVmapOperatorsOpInfo, globals(), only_for=only_for)
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/common_device_type.py", line 1074, in instantiate_device_type_tests
    device_type_test_class.instantiate_test(
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/common_device_type.py", line 634, in instantiate_test
    instantiate_test_helper(
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/common_device_type.py", line 540, in instantiate_test_helper
    test = decorator(test)
           ^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/testing/_internal/opinfo/core.py", line 1713, in __call__
    raise RuntimeError("Multiple sets of sample_skips_and_xfails defined")
RuntimeError: Multiple sets of sample_skips_and_xfails defined
================================================================================== short test summary info ===================================================================================
ERROR ../opt/pytorch/pytorch/test/functorch/test_vmap.py - RuntimeError: Multiple sets of sample_skips_and_xfails defined
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
====================================================================================== 1 error in 3.91s ======================================================================================
</code></pre></div>
<p>due to <code>test_vmap_exhaustive</code> having multiple decorators <code>skipIf</code> and <code>xfailIf</code> from <code>test/functorch/common_utils.py</code>. Both of these rely on <code>sample_skips_and_xfails</code>:<br>
<a href="https://github.com/pytorch/pytorch/blob/d3ce23d75ee5e488787aafb12c281b5142d91e75/test/functorch/common_utils.py#L481-L497">https://github.com/pytorch/pytorch/blob/d3ce23d75ee5e488787aafb12c281b5142d91e75/test/functorch/common_utils.py#L481-L497</a></p>
<p><a href="https://github.com/pytorch/pytorch/blob/d3ce23d75ee5e488787aafb12c281b5142d91e75/test/functorch/common_utils.py#L510-L526">https://github.com/pytorch/pytorch/blob/d3ce23d75ee5e488787aafb12c281b5142d91e75/test/functorch/common_utils.py#L510-L526</a></p>
<p>Though, using a mix of these on a certain test end up in the <code>RuntimeError</code>:<br>
<a href="https://github.com/pytorch/pytorch/blob/d3ce23d75ee5e488787aafb12c281b5142d91e75/torch/testing/_internal/opinfo/core.py#L1713">https://github.com/pytorch/pytorch/blob/d3ce23d75ee5e488787aafb12c281b5142d91e75/torch/testing/_internal/opinfo/core.py#L1713</a></p>
<p>This PR proposes appending the rules, as this simplifies the usage of <code>sample_skips_and_xfails</code>.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4503579291" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184894" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/184894/hovercard" href="https://github.com/pytorch/pytorch/issues/184894">#184894</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4569750760" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185905" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185905/hovercard" href="https://github.com/pytorch/pytorch/pull/185905">#185905</a><br>
Approved by: <a href="https://github.com/benjaminglass1">https://github.com/benjaminglass1</a>, <a href="https://github.com/eqy">https://github.com/eqy</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/474b708affd7beacd45145a5b3d011a9d59c6b5c: [ROCm] Guard NCCL one-sided API behind device support (#186888)]]></title>
<description><![CDATA[Summary
This updates NCCL_HAS_ONE_SIDED_API to require NCCL_HAS_SYMMEM_DEVICE_SUPPORT, matching the existing guard pattern used for NCCL device-side support.
NCCL_HAS_SYMMEM_DEVICE_SUPPORT is not enabled for ROCm builds because nccl_device.h is excluded there. Without this additional guard, ROCm ...]]></description>
<link>https://tsecurity.de/de/3588166/downloads/trunk474b708affd7beacd45145a5b3d011a9d59c6b5c-rocm-guard-nccl-one-sided-api-behind-device-support-186888/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588166/downloads/trunk474b708affd7beacd45145a5b3d011a9d59c6b5c-rocm-guard-nccl-one-sided-api-behind-device-support-186888/</guid>
<pubDate>Wed, 10 Jun 2026 17:32:16 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>This updates <code>NCCL_HAS_ONE_SIDED_API</code> to require <code>NCCL_HAS_SYMMEM_DEVICE_SUPPORT</code>, matching the existing guard pattern used for NCCL device-side support.</p>
<p><code>NCCL_HAS_SYMMEM_DEVICE_SUPPORT</code> is not enabled for ROCm builds because <code>nccl_device.h</code> is excluded there. Without this additional guard, ROCm release branches can still compile one-sided NCCL paths that reference device-side support such as <code>NCCLDevCommManager</code>, causing build failures.</p>
<p>A related guard already exists from <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4623145442" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186794" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186794/hovercard" href="https://github.com/pytorch/pytorch/pull/186794">#186794</a>; this extends the same protection to the one-sided API path after it caused issues on another ROCm branch.</p>
<h2>Testing</h2>
<ul>
<li>Verified the merge conflict is resolved.</li>
<li>Verified <code>nccl_dev_cap.hpp</code> has no conflict markers.</li>
<li>Verified <code>git diff --check</code> passes for the changed file.</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4627843541" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186888" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186888/hovercard" href="https://github.com/pytorch/pytorch/pull/186888">#186888</a><br>
Approved by: <a href="https://github.com/jeffdaily">https://github.com/jeffdaily</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/1190095994b3b3913e3d2e60d63219cf786ba3b4: [FakeTensor] Add hinted symbolic storage size metadata (#183839)]]></title>
<description><![CDATA[Trace tooling serializes FakeTensor storage metadata to JSON. When the
storage size is a hinted SymInt, the trace can have both useful pieces of
information:

the symbolic storage expression, which preserves provenance for downstream
trace consumers;
the optimization hint, which gives diagnostic/...]]></description>
<link>https://tsecurity.de/de/3588131/downloads/trunk1190095994b3b3913e3d2e60d63219cf786ba3b4-faketensor-add-hinted-symbolic-storage-size-metadata-183839/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3588131/downloads/trunk1190095994b3b3913e3d2e60d63219cf786ba3b4-faketensor-add-hinted-symbolic-storage-size-metadata-183839/</guid>
<pubDate>Wed, 10 Jun 2026 17:19:16 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Trace tooling serializes FakeTensor storage metadata to JSON. When the<br>
storage size is a hinted SymInt, the trace can have both useful pieces of<br>
information:</p>
<ul>
<li>the symbolic storage expression, which preserves provenance for downstream<br>
trace consumers;</li>
<li>the optimization hint, which gives diagnostic/policy tooling a concrete<br>
expected extent when one was explicitly provided.</li>
</ul>
<p>Keep the existing <code>size</code> field symbolic for symbolic storage sizes and add a<br>
separate <code>size_hint</code> field only when every free symbol in the storage-size<br>
expression has an explicit optimization hint override. This preserves the<br>
existing symbolic trace contract while exposing concrete policy metadata<br>
without specializing tensor shapes or changing runtime semantics.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450833740" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183835" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/183835/hovercard" href="https://github.com/pytorch/pytorch/issues/183835">#183835</a></p>
<p>Test Plan:</p>
<p>python test/test_fake_tensor.py FakeTensorTest.test_meta_storage_trace_uses_hint_for_symbolic_size -q</p>
<p>lintrunner --config=.lintrunner.toml torch/_subclasses/meta_utils.py test/test_fake_tensor.py</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851888" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183839" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183839/hovercard" href="https://github.com/pytorch/pytorch/pull/183839">#183839</a><br>
Approved by: <a href="https://github.com/ezyang">https://github.com/ezyang</a>, <a href="https://github.com/laithsakka">https://github.com/laithsakka</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781077762]]></title>
<description><![CDATA[[DTensor] Fix group_norm scalar adjuster crash when weight=None (#184…]]></description>
<link>https://tsecurity.de/de/3586856/downloads/viablestrict1781077762/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586856/downloads/viablestrict1781077762/</guid>
<pubDate>Wed, 10 Jun 2026 10:07:50 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[DTensor] Fix group_norm scalar adjuster crash when weight=None (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="186193628" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/184/hovercard" href="https://github.com/pytorch/pytorch/issues/184">#184</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/98b772f8ab86176b8e0b572d88f4e09510c1b587: [Inductor] Handle hinted and fallback unbacked symbols (#183840)]]></title>
<description><![CDATA[Stacked PRs:

#183838
#183839
->#183840


[Inductor] Handle hinted and fallback unbacked symbols
Inductor has several codegen paths that need to reason about unbacked symbolic extents without turning policy decisions into semantic guards. Layout constraint lowering needs to recognize common symbo...]]></description>
<link>https://tsecurity.de/de/3586854/downloads/trunk98b772f8ab86176b8e0b572d88f4e09510c1b587-inductor-handle-hinted-and-fallback-unbacked-symbols-183840/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586854/downloads/trunk98b772f8ab86176b8e0b572d88f4e09510c1b587-inductor-handle-hinted-and-fallback-unbacked-symbols-183840/</guid>
<pubDate>Wed, 10 Jun 2026 10:07:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Stacked PRs:</p>
<ul>
<li><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851844" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183838" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183838/hovercard" href="https://github.com/pytorch/pytorch/pull/183838">#183838</a></li>
<li><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851888" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183839" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183839/hovercard" href="https://github.com/pytorch/pytorch/pull/183839">#183839</a></li>
<li><strong>-&gt;</strong><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851914" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183840" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183840/hovercard" href="https://github.com/pytorch/pytorch/pull/183840">#183840</a></li>
</ul>
<hr>
<h3>[Inductor] Handle hinted and fallback unbacked symbols</h3>
<p>Inductor has several codegen paths that need to reason about unbacked symbolic extents without turning policy decisions into semantic guards. Layout constraint lowering needs to recognize common symbolic stride orderings, while fallback kernels can repropagate outputs whose unbacked symbols are graph-owned and must be bound for wrapper codegen.</p>
<p>For stride ordering, use a stride-specific symbolic greater-or-equal proof with ordinary guarded comparisons plus divisibility reasoning. This lets Inductor prove layouts such as u0 * 256 &gt;= 256 without relying on concrete hints, while still rejecting unproved unbacked layouts. When require_strides proves the current layout already satisfies the requested order, freeze the current layout directly instead of forcing guarding_hints_or_throw() or bailing out for every unbacked stride.</p>
<p>For fallback outputs, temporarily re-enable fresh unbacked symbol tracking while rerunning the fallback fake kernel. That call is the binding site for output size and stride symbols that later wrapper code references, so those symbols should go through the normal pending-symbol and compute_unbacked_bindings() path instead of being created under ignore_fresh_unbacked_symbols() and rediscovered afterward.</p>
<p>The C++ wrapper path also has to treat input unbacked symbols as already declared before emitting output bindings. Otherwise a fallback output binding can redeclare a symbol such as u0 and can emit Python-only <strong>floordiv</strong> syntax for DivideByKey paths. Emit C++ integer division for that path and reuse the existing unbacked symbol declaration.</p>
<p>Finally, do not make the post-copy stride-order sanity check prove data-dependent unbacked stride inequalities. The copy has already been materialized with the requested stride order; requiring a symbolic proof there rejects valid layouts whose unbacked size may be zero or one.</p>
<p>These changes preserve symbolic semantics: hints are not used to create semantic layout guards, fallback output extents are bound through normal ShapeEnv tracking at the fallback output binding site, and stride/order changes copy data or freeze already-valid layouts rather than manufacturing semantic guards.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450833733" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183834" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/183834/hovercard" href="https://github.com/pytorch/pytorch/issues/183834">#183834</a></p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4530789485" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185341" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/185341/hovercard" href="https://github.com/pytorch/pytorch/issues/185341">#185341</a></p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCPU.test_stride_order_uses_unbacked_optimization_hint_cpu TestUnbackedSymintsCPU.test_stride_ordered_uses_symbolic_divisibility_cpu TestUnbackedSymintsCPU.test_stride_ordered_rejects_unproved_unbacked_layout_cpu -q"><pre>python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCPU.test_stride_order_uses_unbacked_optimization_hint_cpu TestUnbackedSymintsCPU.test_stride_ordered_uses_symbolic_divisibility_cpu TestUnbackedSymintsCPU.test_stride_ordered_rejects_unproved_unbacked_layout_cpu -q</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCUDA.test_standalone_compile_reuses_fallback_unbacked_binding_cuda TestUnbackedSymintsCUDA.test_sdfpa_unbacked_strides_cuda -q"><pre>python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCUDA.test_standalone_compile_reuses_fallback_unbacked_binding_cuda TestUnbackedSymintsCUDA.test_sdfpa_unbacked_strides_cuda -q</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCPU.test_stride_ordered_uses_symbolic_divisibility_cpu TestUnbackedSymintsCPU.test_stride_ordered_rejects_unproved_unbacked_layout_cpu TestUnbackedSymintsCUDA.test_standalone_compile_reuses_fallback_unbacked_binding_cuda TestUnbackedSymintsCUDA.test_sdfpa_unbacked_strides_cuda -q"><pre>python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCPU.test_stride_ordered_uses_symbolic_divisibility_cpu TestUnbackedSymintsCPU.test_stride_ordered_rejects_unproved_unbacked_layout_cpu TestUnbackedSymintsCUDA.test_standalone_compile_reuses_fallback_unbacked_binding_cuda TestUnbackedSymintsCUDA.test_sdfpa_unbacked_strides_cuda -q</pre></div>
<p>TorchTitan graph_trainer H100 integration: aot_fx_trace_deepseek_v3_sdpa_full_inductor_ep_overlap_moe_seq</p>
<p>This PR was authored with the assistance of an AI assistant.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851914" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183840" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183840/hovercard" href="https://github.com/pytorch/pytorch/pull/183840">#183840</a><br>
Approved by: <a href="https://github.com/laithsakka">https://github.com/laithsakka</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/cef26d1e97fcb9dd61b4471f9bd7fa9a32bd42b9: [MPS] faster norms (#186076)]]></title>
<description><![CDATA[Routes norm(1/2, dim=-1) through the efficient sum_reduction_inner kernel via abs/square load + sqrt finalize, instead of the slow per-output-threadgroup norm kernel with per-element pow.
Example:
import torch
import torch.nn.functional as F

x = torch.randn(4096, 4096, device="mps", dtype=torch....]]></description>
<link>https://tsecurity.de/de/3586787/downloads/trunkcef26d1e97fcb9dd61b4471f9bd7fa9a32bd42b9-mps-faster-norms-186076/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586787/downloads/trunkcef26d1e97fcb9dd61b4471f9bd7fa9a32bd42b9-mps-faster-norms-186076/</guid>
<pubDate>Wed, 10 Jun 2026 09:31:51 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Routes <code>norm(1/2, dim=-1)</code> through the efficient <code>sum_reduction_inner</code> kernel via abs/square load + sqrt finalize, instead of the slow per-output-threadgroup norm kernel with per-element <code>pow</code>.<br>
Example:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='import torch
import torch.nn.functional as F

x = torch.randn(4096, 4096, device="mps", dtype=torch.bfloat16)

x.norm(2, dim=-1)          # L2 -&gt; norm_l2_reduction_inner kernel
x.norm(1, dim=-1)          # L1 -&gt; norm_l1_reduction_inner kernel
F.normalize(x, dim=-1)'><pre class="notranslate"><code>import torch
import torch.nn.functional as F

x = torch.randn(4096, 4096, device="mps", dtype=torch.bfloat16)

x.norm(2, dim=-1)          # L2 -&gt; norm_l2_reduction_inner kernel
x.norm(1, dim=-1)          # L1 -&gt; norm_l1_reduction_inner kernel
F.normalize(x, dim=-1)
</code></pre></div>
<p>Perf:</p>
<h3>fp32</h3>
<table>
<thead>
<tr>
<th>shape</th>
<th align="right">old (µs)</th>
<th align="right">new (µs)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>(32,768)</td>
<td align="right">6.9</td>
<td align="right">3.8</td>
<td align="right"><strong>1.8x</strong></td>
</tr>
<tr>
<td>(128,1024)</td>
<td align="right">17.1</td>
<td align="right">4.2</td>
<td align="right"><strong>4.1x</strong></td>
</tr>
<tr>
<td>(512,4096)</td>
<td align="right">171.9</td>
<td align="right">20.9</td>
<td align="right"><strong>8.2x</strong></td>
</tr>
<tr>
<td>(2048,4096)</td>
<td align="right">676.8</td>
<td align="right">114.5</td>
<td align="right"><strong>5.9x</strong></td>
</tr>
<tr>
<td>(4096,4096)</td>
<td align="right">1348.2</td>
<td align="right">238.4</td>
<td align="right"><strong>5.7x</strong></td>
</tr>
<tr>
<td>(16384,1024)</td>
<td align="right">1601.8</td>
<td align="right">236.2</td>
<td align="right"><strong>6.8x</strong></td>
</tr>
<tr>
<td>(1024,8192)</td>
<td align="right">666.3</td>
<td align="right">112.0</td>
<td align="right"><strong>5.9x</strong></td>
</tr>
<tr>
<td>(65536,256)</td>
<td align="right">1707.3</td>
<td align="right">229.9</td>
<td align="right"><strong>7.4x</strong></td>
</tr>
<tr>
<td>(8,32,768)</td>
<td align="right">27.5</td>
<td align="right">4.2</td>
<td align="right"><strong>6.6x</strong></td>
</tr>
<tr>
<td>(32,128,128)</td>
<td align="right">73.1</td>
<td align="right">7.4</td>
<td align="right"><strong>9.9x</strong></td>
</tr>
</tbody>
</table>
<h3>bf16</h3>
<table>
<thead>
<tr>
<th>shape</th>
<th align="right">old (µs)</th>
<th align="right">new (µs)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>(32,768)</td>
<td align="right">6.9</td>
<td align="right">3.6</td>
<td align="right"><strong>1.9x</strong></td>
</tr>
<tr>
<td>(128,1024)</td>
<td align="right">17.0</td>
<td align="right">3.7</td>
<td align="right"><strong>4.6x</strong></td>
</tr>
<tr>
<td>(512,4096)</td>
<td align="right">177.5</td>
<td align="right">10.3</td>
<td align="right"><strong>17.2x</strong></td>
</tr>
<tr>
<td>(2048,4096)</td>
<td align="right">715.7</td>
<td align="right">40.0</td>
<td align="right"><strong>17.9x</strong></td>
</tr>
<tr>
<td>(4096,4096)</td>
<td align="right">1402.8</td>
<td align="right">117.6</td>
<td align="right"><strong>11.9x</strong></td>
</tr>
<tr>
<td>(16384,1024)</td>
<td align="right">1650.8</td>
<td align="right">112.5</td>
<td align="right"><strong>14.7x</strong></td>
</tr>
<tr>
<td>(1024,8192)</td>
<td align="right">678.0</td>
<td align="right">42.4</td>
<td align="right"><strong>16.0x</strong></td>
</tr>
<tr>
<td>(65536,256)</td>
<td align="right">1790.2</td>
<td align="right">113.2</td>
<td align="right"><strong>15.8x</strong></td>
</tr>
<tr>
<td>(8,32,768)</td>
<td align="right">27.4</td>
<td align="right">3.9</td>
<td align="right"><strong>7.0x</strong></td>
</tr>
<tr>
<td>(32,128,128)</td>
<td align="right">74.5</td>
<td align="right">7.5</td>
<td align="right"><strong>10.0x</strong></td>
</tr>
</tbody>
</table>
<h3>fp16</h3>
<table>
<thead>
<tr>
<th>shape</th>
<th align="right">old (µs)</th>
<th align="right">new (µs)</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>(32,768)</td>
<td align="right">6.8</td>
<td align="right">3.8</td>
<td align="right"><strong>1.8x</strong></td>
</tr>
<tr>
<td>(128,1024)</td>
<td align="right">17.0</td>
<td align="right">3.9</td>
<td align="right"><strong>4.3x</strong></td>
</tr>
<tr>
<td>(512,4096)</td>
<td align="right">172.6</td>
<td align="right">10.2</td>
<td align="right"><strong>16.9x</strong></td>
</tr>
<tr>
<td>(2048,4096)</td>
<td align="right">678.6</td>
<td align="right">40.4</td>
<td align="right"><strong>16.8x</strong></td>
</tr>
<tr>
<td>(4096,4096)</td>
<td align="right">1352.9</td>
<td align="right">112.8</td>
<td align="right"><strong>12.0x</strong></td>
</tr>
<tr>
<td>(16384,1024)</td>
<td align="right">1604.2</td>
<td align="right">107.6</td>
<td align="right"><strong>14.9x</strong></td>
</tr>
<tr>
<td>(1024,8192)</td>
<td align="right">670.2</td>
<td align="right">42.5</td>
<td align="right"><strong>15.8x</strong></td>
</tr>
<tr>
<td>(65536,256)</td>
<td align="right">1711.6</td>
<td align="right">108.1</td>
<td align="right"><strong>15.8x</strong></td>
</tr>
<tr>
<td>(8,32,768)</td>
<td align="right">27.5</td>
<td align="right">4.0</td>
<td align="right"><strong>6.8x</strong></td>
</tr>
<tr>
<td>(32,128,128)</td>
<td align="right">73.2</td>
<td align="right">7.7</td>
<td align="right"><strong>9.5x</strong></td>
</tr>
</tbody>
</table>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4579919157" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186076" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186076/hovercard" href="https://github.com/pytorch/pytorch/pull/186076">#186076</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[KI-Dienstleister finden: Ein AI-as-a-Service-Ratgeber]]></title>
<description><![CDATA[Wer nicht die Zeit oder das Geld hat, um eigene KI-Instanzen aufzusetzen, könnte mit einem AI-as-a-Service-Angebot glücklich werden. Insofern die Voraussetzungen stimmen. 
					Foto: VesnaArt | shutterstock.com




Künstliche Intelligenz (KI) ist weiter auf dem Vormarsch. Laut Gartner werden mehr...]]></description>
<link>https://tsecurity.de/de/3586471/it-security-nachrichten/ki-dienstleister-finden-ein-ai-as-a-service-ratgeber/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586471/it-security-nachrichten/ki-dienstleister-finden-ein-ai-as-a-service-ratgeber/</guid>
<pubDate>Wed, 10 Jun 2026 06:53:20 +0200</pubDate>
<category>📰 IT Security Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<div>
<div class="grid grid--cols-10@md grid--cols-8@lg article-column">
					  <div class="col-12 col-10@md col-6@lg col-start-3@lg">
						<div class="article-column__content">
<section class="wp-block-bigbite-multi-title"><div class="container"></div></section>



<div class="extendedBlock-wrapper block-coreImage"><figure class="wp-block-image size-large"><img loading="lazy" decoding="async" alt="Wer nicht die Zeit oder das Geld hat, um eigene KI-Instanzen aufzusetzen, könnte mit einem AI-as-a-Service-Angebot glücklich werden. Insofern die Voraussetzungen stimmen. " title="Wer nicht die Zeit oder das Geld hat, um eigene KI-Instanzen aufzusetzen, könnte mit einem AI-as-a-Service-Angebot glücklich werden. Insofern die Voraussetzungen stimmen. " src="https://images.computerwoche.de/bdb/3391899/840x473.jpg" width="840" height="473"><figcaption class="wp-element-caption"><p class="foundryImageCaption">Wer nicht die Zeit oder das Geld hat, um eigene KI-Instanzen aufzusetzen, könnte mit einem AI-as-a-Service-Angebot glücklich werden. Insofern die Voraussetzungen stimmen. </p></figcaption></figure><p class="imageCredit">
					Foto: VesnaArt | shutterstock.com</p></div>




<p><a href="https://www.computerwoche.de/k/kuenstliche-intelligenz-artifical-intelligence,3544" target="_blank" class="idgGlossaryLink">Künstliche Intelligenz</a> (KI) ist weiter auf dem Vormarsch. Laut Gartner werden mehr als 80 Prozent aller Unternehmen bis zum Jahr 2026 in irgendeiner Form Generative-AI-APIs oder -Apps <a href="https://www.computerwoche.de/article/2828107/bis-2026-nutzen-80-prozent-der-unternehmen-genai.html" title="nutzen" target="_blank">nutzen</a>. Diejenigen, die dazu gehören, müssen sich überlegen, wie sie ihre KI-Instanzen optimal trainieren und einsetzen – On-Premises oder <a href="https://www.computerwoche.de/article/2824042/generative-ai-aus-der-cloud-wird-teuer.html" title="in der Cloud" target="_blank">in der Cloud</a>.</p>



<p>Eine KI zu trainieren, erfordert spezielle Hardware, die im Vergleich zu Standard-Servern deutlich teurer ist. Die Kosten beginnen im mittleren sechsstelligen Bereich und können in die Millionen gehen. Allerdings kann das teure Equipment nicht für andere Zwecke (etwa Datenbanken) genutzt respektive wiederverwendet werden. Abgesehen von den Anschaffungs- und Wartungskosten stellt das <a href="https://www.computerwoche.de/article/2823313/14-ki-algorithmen-die-sie-kennen-sollten.html" title="KI-Training" target="_blank">KI-Training</a> auch den schwierigsten und prozessintensivsten Part eines solchen Unterfangens dar. Je nach zugrundeliegendem Datensatz kann das Wochen oder auch Monate in Anspruch nehmen – die Sie unter Umständen nicht warten können oder wollen. Im Grunde haben Sie also zwei Möglichkeiten:</p>



<ul class="wp-block-list">
<li><p>Sie kaufen Hardware und setzen Ihre KI im DIY-Verfahren auf.</p></li>



<li><p>Sie wenden sich an einen KI-Dienstleister.</p></li>
</ul>



<p>AI as a Service – kurz AIaaS – ist der neueste Schrei auf dem As-a-Service-Markt und speziell auf Initiativen in Zusammenhang mit künstlicher Intelligenz ausgerichtet. <a href="https://www.computerwoche.de/article/2812429/9-alternativen-zu-aws-azure-und-google-cloud.html" title="Cloud Services Provider" target="_blank">Cloud Services Provider</a> – aber auch kleinere Anbieter – haben solche Dienstleistungsangebote im Portfolio. </p>



<p>Im Folgenden werfen wir einen Blick darauf, wie AIaaS funktioniert, wo die Vor- und Nachteile des Modells liegen und wie die Schlüsselkriterien für entsprechende Plattformen aussehen. Abschließend stellen wir Ihnen die fünf wichtigsten KI-Dienstleister – und ihre Angebote – vor. </p>



<h2 class="wp-block-heading">Wie funktioniert AI as a Service?</h2>



<p>KI als Service bietet Anwenderunternehmen Cloud-basierten Zugang zu AI-Funktionalitäten, um diese in ihre Projekte oder Applikationen zu integrieren, ohne dafür eine eigene Infrastruktur aufbauen und pflegen zu müssen. Vorgefertigte respektive -trainierte KI-Modelle für grundlegende Anwendungsfälle wie <a href="https://www.computerwoche.de/article/2807033/was-ist-ein-chatbot.html" title="Chatbots" target="_blank">Chatbots</a> gewährleisten dabei, dass den Kunden zeitintensive Grundlagenarbeit erspart bleibt.</p>



<p>“AI as a Service ermöglicht grundsätzlich, Anwendungen im Unternehmen schneller entwickeln und bereitzustellen”, konstatiert Gartner-Analyst Chirag Dekate. Dem Experten zufolge gibt es im Bereich AIaaS drei Einstiegspunkte:</p>



<ul class="wp-block-list">
<li><p>die Anwendungsebene,</p></li>



<li><p>die Ebene der Modellentwicklung und</p></li>



<li><p>die Ebene der benutzerdefinierten Modellentwicklung.</p></li>
</ul>



<p>Die Offerings der KI-Dienstleister beinhalten darüber hinaus oft Data Preparation für <a href="https://www.computerwoche.de/article/2790671/unbequeme-ki-wahrheiten.html" title="unstrukturierte Daten" target="_blank">unstrukturierte Daten</a> und die Möglichkeit, vom Anwender bereitgestellte Modelle zu trainieren. Vorgefertigte Modelle stehen im Regelfall für diverse Tasks zur Verfügung – unter anderem: </p>



<ul class="wp-block-list">
<li><p>Bilderkennung,</p></li>



<li><p>Datenanalyse,</p></li>



<li><p><a title="Natural Language Processing" href="https://www.computerwoche.de/article/2799474/was-ist-natural-language-processing.html" target="_blank">Natural Language Processing</a> und</p></li>



<li><p><a title="Predictive Analytics" href="https://www.computerwoche.de/article/2797839/die-besten-predictive-analytics-tools.html" target="_blank">Predictive Analytics</a>.</p></li>
</ul>



<p>Der Zugriff auf die AI Services erfolgt dabei entweder über <a href="https://www.computerwoche.de/article/2790525/was-sie-ueber-application-programming-interfaces-wissen-muessen.html" title="APIs" target="_blank">APIs</a> oder User Interfaces. Das ermöglicht, KI-Funktionen (oft) mit minimalem Programmieraufwand in eigene Applikationen und/oder Plattformen zu integrieren.</p>



<h2 class="wp-block-heading">KI als Service – Kosten und Anforderungen</h2>



<p>Die meisten AI-as-a-Service-Anbieter nutzen ein <a href="https://www.computerwoche.de/article/2804386/11-gruende-gegen-die-cloud.html" title="Pay-as-you-go-Modell" target="_blank">Pay-as-you-go-Modell</a> – entweder auf Nutzungs- oder Flatrate-Basis. Das ist im Vergleich zu <a href="https://www.computerwoche.de/article/2768407/was-ist-infrastructure-as-a-service-die-moderne-data-center-plattform.html" target="_blank" class="idgGlossaryLink">IaaS</a>– oder <a href="https://www.computerwoche.de/article/2645453/was-sie-ueber-die-cloud-wissen-muessen.html" target="_blank" class="idgGlossaryLink">PaaS</a>-Szenarien deutlich kostenintensiver. So ruft beispielsweise Nvidia für seinen <a href="https://www.computerwoche.de/article/2826057/supercomputer-service-mit-genai-fokus.html" title="Supercomputer-Service DGX Cloud" target="_blank">Supercomputer-Service DGX Cloud</a> monatlich pauschal 37.000 Dollar auf.</p>



<p>Das schlagende Argument für AIaaS sind die Kosten für die Hardware, die nötig ist, um KI On-Premises zu realisieren. Nvidias <a href="https://www.nvidia.com/de-de/data-center/dgx-h100/" title="DGX Pod H100-Server" target="_blank" rel="noopener">DGX Pod H100-Server</a> ist beispielsweise ab 500.000 Dollar zu haben, das wesentlich größere Modell <a href="https://www.nvidia.com/de-de/data-center/dgx-superpod/" title="SuperPOD" target="_blank" rel="noopener">SuperPOD</a> hat einen Einstiegspreis von 7.000.000 Dollar. Unternehmen, die das Preisgefüge von x86-Servern gewöhnt sind, dürften bei diesen Zahlen schnell einem Preisschock erliegen. Speziell für KMUs ist AI as a Service deshalb des Öfteren die einzige realistische Option.</p>



<p>Laut Mike Gualtieri, Vice President und Chefanalyst bei Forrester Research, eignet sich KI als Service auch deshalb besonders gut, um zu experimentieren: “Ihre eigene Infrastruktur für 30 Use Cases hochzufahren, wenn vielleicht nur ein paar davon funktionieren, ist nicht wirtschaftlich. Dennoch brauchen Sie die Infrastruktur, um zu testen.”</p>



<p>Neben den Kosten für die Hardware stellt sich außerdem die Frage der Verfügbarkeit: So hat Nvidia beispielsweise mit einem enormen Auftrags-Backlog zu kämpfen. Wer heute einen On-Premises SuperPOD kaufen möchte, muss bis zu acht Monaten warten. Selbst wenn Sie also kaufen möchten und das nötige Kleingeld dafür haben, ist die einzige Option für ein zeitnahes Deployment ein KI-Dienstleister.</p>



<h2 class="wp-block-heading">AI as a Service – Vor- und Nachteile</h2>



<p>Die Vorteile von AI as a Service im Überblick:</p>



<ul class="wp-block-list">
<li><p><strong>Niedrigere Einstiegsschwelle:</strong> AIaaS ist wesentlich kostenintensiver als herkömmliche <a title="Cloud Services" href="https://www.computerwoche.de/article/2806737/die-30-wichtigsten-saas-anbieter.html" target="_blank">Cloud Services</a>. Dennoch ist KI als Dienstleistung deutlich günstiger als die Hardware selbst anzuschaffen. Vom Modelltraining ganz zu schweigen.</p></li>



<li><p><strong>Kürzere Time-to-Market:</strong> KI-Hardware on Premises zu installieren und zu konfigurieren, ist ebenfalls kostspielig und zeitaufwändig. Zudem setzt das voraus, dass Sie auch <a title="die richtigen Experten" href="https://www.computerwoche.de/article/2828187/die-11-gefragtesten-ki-berufe.html" target="_blank">die richtigen Experten</a> an Bord haben, die entsprechende Lösungen bereitstellen und supporten können. Das Hardware-Management einem Serviceanbieter zu überlassen, spart Zeit und ermöglicht den Anwenderunternehmen, sich auf Kerngeschäft und -kompetenzen zu fokussieren.</p></li>



<li><p><strong>Zugang zu State-of-the-Art-Technologie:</strong> KI <a title="kann" href="https://www.computerwoche.de/article/2830668/unternehmen-nicht-bereit-fuer-ki-erfolg.html" target="_blank">kann</a> Wettbewerbsvorteile erschließen. Weil die Anbieter Kunden binden wollen, sind sie darauf bedacht, sich beständig zu verbessern und ihre Angebote weiterzuentwickeln.</p></li>



<li><p><strong>Skalierbarkeit:</strong> Da es sich um Cloud-basierte Dienstleistungen handelt, ist eines der wichtigsten Verkaufsargumente, dass sich AIaaS-Lösungen entsprechend den Bedürfnissen des Anwenderunternehmens anpassen lassen.</p></li>



<li><p><strong>Zugang zu KI-Expertise:</strong> KI steckt immer noch in den Kinderschuhen und die Anzahl der IT-Experten, die die Hardware konfigurieren und managen können, ist überschaubar – außer bei den <a title="großen Cloud Services Providern" href="https://www.computerwoche.de/article/2816695/welche-cloud-ist-fuer-sie-die-richtige.html" target="_blank">großen Cloud Services Providern</a>.</p></li>
</ul>



<p>Wo Licht ist, ist auch Schatten – da bildet AI as a Service keine Ausnahme:</p>



<ul class="wp-block-list">
<li><p><strong>Vendor Lock-in.</strong> Wenn Sie sich einmal für eine Plattform entschieden haben, kann der Umstieg auf einen anderen Anbieter problematisch sein.</p></li>



<li><p><strong>Begrenzte Anpassungsmöglichkeiten.</strong> Vorgefertigte Modelle sind für allgemeine Zwecke gut geeignet, können aber nicht auf spezifische Anforderungen abgestimmt werden. Sie sind deshalb möglicherweise gezwungen, Ihre eigenen Modelle zu erstellen und zu verarbeiten.</p></li>



<li><p><strong>Security- und Datenschutzbedenken.</strong> Wenn es um Unternehmensdaten geht, ist sorgfältig abzuwägen, ob <a title="Drittanbieter miteinbezogen werden" href="https://www.csoonline.com/de/a/5-wege-mit-drittanbietern-unterzugehen,3674508" target="_blank">Drittanbieter miteinbezogen werden</a>.</p></li>
</ul>



<h2 class="wp-block-heading">AIaaS-Plattformen – Schlüsselkriterien</h2>



<p>Bei der Auswahl einer AIaaS-Plattform sollten Sie (unter anderem) verstärkt auf die folgenden Aspekte achten:</p>



<ul class="wp-block-list">
<li><p><strong>Unterstützte Workloads:</strong> Laut Forrester-Chefanalyst Gaultieri ist das wichtigste Kriterium bei der Auswahl eines AIaaS-Anbieters, ob er alle KI-Phasen unterstützt: <a title="Datenvorbereitung" href="https://www.computerwoche.de/article/2803194/so-bereiten-sie-daten-fuer-ki-gestuetzte-prognosen-vor.html" target="_blank">Datenvorbereitung</a>, Modelltraining und Inferencing. Gerade ersteres wird in der KI-Debatte oft unter den Tisch gekehrt, ist jedoch unumgänglich, weil KI-Instanzen in der Regel auf unstrukturierte Daten zugreifen.</p></li>



<li><p><strong>Regionale Infrastruktur:</strong> Geht es nach Gartner-Analyst Dikate, sollten Anwender in erster Linie darauf achten, dass AIaaS-Anbieter über ausreichende Skalierungskapazitäten in ihrer Region verfügen. Nicht alle Cloud-Anbieter verfügen über global verteilte Ressourcen.</p></li>



<li><p><strong>Erfahrung sticht:</strong> Suchen Sie gezielt nach Anbietern mit Erfahrung in Ihrer Branche oder mit Projekten, die ähnliche Herausforderungen aufwerfen. Fragen Sie nach Fallstudien, Kundenreferenzen und <a title="Erfahrungsberichten" href="https://www.computerwoche.de/article/2832966/wie-ibm-von-adobes-genai-tool-profitiert.html" target="_blank">Erfahrungsberichten</a>.</p></li>



<li><p><strong>KI-Spezifizierung:</strong> Bilderkennung ist etwas völlig anderes als Intrusion Detection oder ein Chatbot. Nicht jeder AIaaS-Anbieter ist auf sämtliche KI-Arten spezialisiert. Sie sollten deshalb sicherstellen, dass die Spezialisierung Ihres KI-Dienstleisters möglichst Ihren Bedürfnissen entspricht.</p></li>



<li><p><strong>Daten- und Compliance-Kompatibilität:</strong> Stellen Sie sicher, dass die Anbieterplattform sowohl Ihr Datenformat als auch -volumen unterstützt. Wenn es sich um stark regulierte Daten handelt, sollten Sie sichergehen, dass der Anbieter entsprechend zertifiziert ist, um diese zu verarbeiten.</p></li>



<li><p><strong>Skalierbarkeit:</strong> KI-Dienstleister verfügen möglicherweise nicht mehr über die nötige Kapazität, wenn Ihr Bedarf weiter wächst. Natürlich sind Zukunftsprognosen schwierig, insbesondere bei sich rasant verändernden Branchen wie KI. Dennoch ist es ratsam, eine zufriedenstellende Zukunftsperspektive einzuholen.</p></li>



<li><p><strong>Modell-Updates und -Wartung:</strong> KI-Modelle funktionieren fast nie nur einmal und dann nie wieder. Sie müssen regelmäßig und routinemäßig auf den aktuellen Stand gebracht werden. Studieren Sie deshalb die Richtlinien des Anbieters mit Blick darauf, wie das Modell gespeichert und aktualisiert werden kann und wie es sich “aus dem System” On-Premises nehmen lässt.</p></li>



<li><p><strong>Workload Management Software:</strong> Schließlich sollten Sie sicherstellen, dass der Anbieter in der Lage ist, einen Workload neu zu starten, wenn während der Verarbeitung ein Problem auftritt. Gaultieri erklärt, warum: “Stellen Sie sich vor, Sie erstellen ein <a title="Large Language Model" href="https://www.computerwoche.de/article/2823883/was-sind-llms.html" target="_blank">Large Language Model</a>, lassen es eine Woche lang laufen und dann geht etwas schief. Wenn es sich um einen mehrwöchigen Workload handelt, möchten Sie nicht von vorne beginnen. Features wie Checkpointing sind deshalb ratsam.”</p></li>
</ul>



<h2 class="wp-block-heading">Die 5 wichtigsten KI-Dienstleister</h2>



<p>AI as a Service ist kein kleines Unterfangen. Während <a href="https://www.computerwoche.de/article/2830844/microsoft-openai-und-nvidia-dominieren-genai-markt.html" title="neuere Anbieter" target="_blank">neuere Anbieter</a> wie Nvidia, OpenAI und auch einige <a href="https://www.computerwoche.de/article/2818139/5-tipps-fuer-die-zusammenarbeit-mit-msps.html" title="Managed-Services-Anbieter" target="_blank">Managed-Services-Anbieter</a> auf den Plan treten, sind die wichtigsten Akteure im Bereich KI-Dienstleistungen in erster Linie die Cloud-Hyperscaler. Sie verfügen über die Ressourcen, um AIaaS im Enterprise-Maßstab zu supporten.</p>



<p><strong><a href="https://aws.amazon.com/de/machine-learning/ai-services/" title="Amazon Web Services" target="_blank" rel="noopener">Amazon Web Services</a></strong></p>



<p>Amazon Web Services (<a href="https://www.computerwoche.de/article/2732199/amazon-web-services-viel-cloud-fuer-wenig-geld.html" target="_blank" class="idgGlossaryLink">AWS</a>) hat eine breite Palette von KI-Dienstleistungen im Angebot, angefangen bei vorgefertigten, sofort einsetzbaren Services, die den Start eines KI-Projekts erleichtern und den Bedarf an erfahrenen Datenwissenschaftlern und <a href="https://www.computerwoche.de/article/2825328/generative-ai-praxistipps-von-entwicklern.html" title="KI-Entwicklern" target="_blank">KI-Entwicklern</a> minimieren können. Zu diesen Services gehören beispielsweise:</p>



<ul class="wp-block-list">
<li><p>Amazon Translate (Echtzeit-Übersetzungen),</p></li>



<li><p>Amazon Rekognition (Bild- und Videoanalyse),</p></li>



<li><p>Amazon Polly (Text-to-Speech) und</p></li>



<li><p>Amazon Transcribe (Speech-to-Text).</p></li>
</ul>



<p>Zu den Tools im Bereich Managed Infrastructure gehören:</p>



<ul class="wp-block-list">
<li><p>Amazon SageMaker (Machine-Learning-Modelle erstellen, trainieren und bereitstellen),</p></li>



<li><p>Amazon Machine Learning Drag-and-Drop-Tools und -Templates (ML-Modelle einfacher bereitstellen),</p></li>



<li><p>Amazon Comprehend (Natural Language Processing),</p></li>



<li><p>Amazon Forecast (akkurate Zeitreihenvorhersagen) und</p></li>



<li><p>Amazon Personalize (personalisierte Produkt- und Content-Vorschläge).</p></li>
</ul>



<p>Im Bereich <a href="https://www.computerwoche.de/article/2832232/wie-genai-geschaeftswert-schafft.html" title="generative KI" target="_blank">generative KI</a> bietet <a href="https://www.computerwoche.de/article/2732199/amazon-web-services-viel-cloud-fuer-wenig-geld.html" target="_blank" class="idgGlossaryLink">AWS</a>:</p>



<ul class="wp-block-list">
<li><p>Amazon Lex (KI-Chatbots erstellen),</p></li>



<li><p>Amazon CodeGuru (Code analysieren und optimieren) sowie</p></li>



<li><p>Amazon Kendra (intelligente Suchen).</p></li>
</ul>



<p><strong><a href="https://azure.microsoft.com/de-de/products/ai-services/" title="Microsoft" target="_blank" rel="noopener">Microsoft</a></strong></p>



<p>Microsofts Azure KI Services richten sich an Entwickler und Datenwissenschaftler und basieren auf Anwendungen wie SQL Server, Office und Dynamics. Dabei haben die Redmonder künstliche Intelligenz in verschiedene Business-Apps integriert – sowohl in der Cloud als auch On-Premises.</p>



<p>Bekanntermaßen ist Microsoft eine enge Partnerschaft mit ChatGPT-Entwickler OpenAI eingegangen. Entsprechend sind viele KI-Anwendungen auf dem Azure-Marktplatz zu finden. Darüber hinaus steht ein OpenAI Service mit vortrainierten Large Language Models wie GPT-3.5, Codex und DALL-E 2 zur Verfügung.</p>



<p>Zu den vorgefertigten AI Services im Angebot von Microsoft zählen etwa:</p>



<ul class="wp-block-list">
<li><p>Spracherkennung,</p></li>



<li><p>Textanalyse,</p></li>



<li><p>Übersetzung,</p></li>



<li><p>Bildverarbeitung und</p></li>



<li><p>ML-Modell-Deployment.</p></li>
</ul>



<p><strong><a href="https://ai.google/" title="Google Cloud" target="_blank" rel="noopener">Google Cloud</a></strong></p>



<p>Der KI-Service von Google ist auf Data Analytics fokussiert und bietet Tools wie BigQuery und AI Platform sowie den Service AutoML, der Benutzern mit eingeschränkten Coding-Fähigkeiten ermöglicht, automatisch Modelle zu erstellen. Google Cloud bietet zudem die Plattform <a href="https://www.computerwoche.de/article/2821909/google-cloud-gmail-docs-und-co-erhalten-ki-features.html" title="Vertex AI" target="_blank">Vertex AI</a>, um KI-Workflows zu optimieren und Entwicklung und Bereitstellung zu vereinfachen. Dazu gehört auch eine breite Palette von Services mit vorgefertigten Lösungen, benutzerdefiniertem Modelltraining und Generative-AI-Werkzeugen. Mit der Vertex AI Workbench stellt Google außerdem eine kollaborative KI-Projektumgebung für Datenwissenschaftler und Entwickler zur Verfügung.</p>



<p>Die vorgefertigten KI-Lösungen von Google Cloud umfassen:</p>



<ul class="wp-block-list">
<li><p>Dialogflow (Plattform, um Chatbots und virtuelle Assistenten zu entwickeln),</p></li>



<li><p>Natural Language API (Sentiment-Textanalysen, Entity-Extraktion etc.),</p></li>



<li><p>Vision AI (Objekterkennung in Bildern und Videos),</p></li>



<li><p>Translation API (maschinelle Übersetzungen in verschiedenen Sprachen) sowie</p></li>



<li><p>Speech-to-Text und Text-to-Speech (Konversionen zwischen gesprochener Sprache und Text).</p></li>
</ul>



<p>Was <a title="Generative AI" href="https://www.computerwoche.de/article/2821922/was-ist-generative-ai.html" target="_blank">Generative AI</a> angeht, bietet Vertex AI Search and Conversation eine Tool-Sammlung, die speziell darauf konzipiert ist, GenAI-Applikationen wie Suchmaschinen und <a class="idgGlossaryLink" href="https://www.computerwoche.de/article/2753417/was-unternehmen-ueber-chatbots-wissen-muessen.html" target="_blank">Chatbots</a> zu entwickeln. Diese Suite enthält mehr als 130 vortrainierte Foundational-LLMs wie PaLM und Imagen, um Texte und/oder Bilder zu generieren. Darüber hinaus hat Google auch den KI-Assisstenten Gemini im Programm, der in verschiedenen Versionen <a href="https://ai.google.dev/gemini-api/docs/models?hl=de" target="_blank" rel="noreferrer noopener">erhältlich ist</a>.</p>



<p><strong><a href="https://www.ibm.com/watsonx" title="IBM" target="_blank" rel="noopener">IBM</a></strong></p>



<p>IBMs <a href="https://www.computerwoche.de/article/2825363/wie-sich-ibm-gegen-chatgpt-und-bard-behaupten-will.html" title="Watsonx" target="_blank">Watsonx</a>, ist ein umfassendes KI-Tool- und -Dienstleistungsangebot, das für seinen Fokus auf die Automatisierung komplexer Geschäftsprozesse und seine branchenspezifischen Lösungen, insbesondere für das Gesundheits- und Finanzwesen, bekannt ist. Watsonx.ai Studio ist das Herzstück dieser Plattform, auf der Sie KI-Modelle trainieren, validieren, abstimmen und einsetzen können – sowohl für maschinelles Lernen als auch für generative KI. Ein Data Lakehouse gewährleistet dabei ein sicheres und skalierbares Speichersystem für Ihre Daten (sowohl strukturierte als auch unstrukturierte).</p>



<p>IBMs AI Toolkit ist eine Sammlung vorgefertigter Tools und Konnektoren, die die Integration von KI in Ihre bestehenden Workflows erleichtern. So lassen sich Tasks automatisieren, Erkenntnisse aus Daten gewinnen und intelligente Applikationen erstellen. Watsonx enthält auch eine Reihe vortrainierter KI-Modelle, die Sie direkt und ohne Training einsetzen können. Diese Modelle decken verschiedene Tasks ab – beispielsweise:</p>



<ul class="wp-block-list">
<li><p>Natural Language Processing,</p></li>



<li><p><a title="Computer Vision" href="https://www.computerwoche.de/article/2799318/was-ist-computer-vision.html" target="_blank">Computer Vision</a> und</p></li>



<li><p>Spracherkennung.</p></li>
</ul>



<p><strong><a href="https://www.oracle.com/artificial-intelligence/ai-services/" title="Oracle" target="_blank" rel="noopener">Oracle</a></strong></p>



<p>Oracle liegt bislang weit hinter den Cloud-Hyperscalern zurück, bringt jedoch einige Vorteile mit, die Sie auf dem Schirm haben sollten – in erster Linie, weil es ein Riese in Sachen Business-Anwendungen und -Datenbanken ist. Alle On-Premises installierten Anwendungen können in die Cloud verlagert werden, um eine hybride Konfiguration zu verwirklichen. Das erleichtert es erheblich, Ihre lokalen Daten zu Data-Preparation- und Schulungszwecken in die Cloud zu verlagern. Oracle hat sehr stark in GPU-Technologie investiert, die derzeit das wichtigste Mittel zur KI-Datenverarbeitung darstellt. Wenn Sie also KI-Anwendungen auf Nvidia-Technologie laufen lassen wollen, können Sie sich an Oracle wenden. Ein weiterer gewichtiger Vorteil: Die KI-Dienstleistungen von Oracle sind mit die günstigsten.</p>



<p>Die Oracle Cloud Infrastructure (OCI) AI Services decken ein breit gefächertes Portfolio von Tools und Services ab, um Unternehmen mit diversen KI-Funktionen zu versorgen. Ähnlich wie im Fall von IBMs Watsonx handelt es sich nicht um einen einzigen Service, sondern um eine Sammlung von Funktionen, die unterschiedliche Anforderungen erfüllen – darunter:</p>



<ul class="wp-block-list">
<li><p>Betrugserkennung und -prävention,</p></li>



<li><p>Spracherkennung sowie </p></li>



<li><p>Sprach- und Textanalyse.</p></li>
</ul>



<p>Oracles <a href="https://www.computerwoche.de/article/2831669/oracle-tunt-seine-cloud-mit-genai-services.html" title="Generative-AI-Services" target="_blank">Generative-AI-Services</a> unterstützt LLMs wie Cohere und Llama 2 und ermöglicht Anwendungsfälle wie:</p>



<ul class="wp-block-list">
<li><p>Schreib-Assistenten,</p></li>



<li><p>Textzusammenfassungen, </p></li>



<li><p>Chatbots oder</p></li>



<li><p>Code-Generierung.</p></li>
</ul>



<p>Die <a href="https://www.computerwoche.de/article/2752649/was-sie-ueber-maschinelles-lernen-wissen-muessen.html" target="_blank" class="idgGlossaryLink">Machine Learning</a> Services von Oracle bieten Tools für Datenwissenschaftler, um ML-Modelle zu erstellen, zu trainieren und zu managen. Dabei werden populäre <a href="https://www.computerwoche.de/k/linux-open-source,3472" target="_blank" class="idgGlossaryLink">Open-Source</a>-Frameworks wie <a href="https://www.computerwoche.de/article/2789265/darum-geht-s-beim-fuehrenden-ki-framework.html" title="TensorFlow" target="_blank">TensorFlow</a> und <a href="https://www.computerwoche.de/article/2794453/5-gruende-fuer-das-deep-learning-framework.html" title="PyTorch" target="_blank">PyTorch</a> unterstützt. Mit OCI Data Science ist es schließlich möglich, virtuelle Maschinen mit vorkonfigurierten Umgebungen für Data-Science-Aufgaben bereitzustellen – inklusive Jupyter-Notebooks und Zugang zu beliebten Bibliotheken, die die Datenexploration und den Modellentwicklungs-Workflow vereinfachen. (fm)

</p>



<p><strong>Dieser Artikel ist <a href="https://www.computerworld.com/article/1612061/aiaas-artificial-intelligence-as-a-service-buyers-guide.html" target="_blank">im Original</a> bei unserer Schwesterpublikation Computerworld.com erschienen.</strong></p>
</div></div></div>
</div>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/183840: [Inductor] Handle hinted and fallback unbacked symbols]]></title>
<description><![CDATA[Inductor has several codegen paths that need to reason about unbacked symbolic
extents without turning policy decisions into semantic guards. Layout constraint
lowering needs to recognize common symbolic stride orderings, while fallback
kernels can repropagate outputs whose unbacked symbols are g...]]></description>
<link>https://tsecurity.de/de/3586309/downloads/ciflowtrunk183840-inductor-handle-hinted-and-fallback-unbacked-symbols/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586309/downloads/ciflowtrunk183840-inductor-handle-hinted-and-fallback-unbacked-symbols/</guid>
<pubDate>Wed, 10 Jun 2026 04:01:29 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Inductor has several codegen paths that need to reason about unbacked symbolic<br>
extents without turning policy decisions into semantic guards. Layout constraint<br>
lowering needs to recognize common symbolic stride orderings, while fallback<br>
kernels can repropagate outputs whose unbacked symbols are graph-owned and must<br>
be bound for wrapper codegen.</p>
<p>For stride ordering, use a stride-specific symbolic greater-or-equal proof with<br>
ordinary guarded comparisons plus divisibility reasoning. This lets Inductor<br>
prove layouts such as <code>u0 * 256 &gt;= 256</code> without relying on concrete hints, while<br>
still rejecting unproved unbacked layouts. When <code>require_strides()</code> proves the<br>
current layout already satisfies the requested order, freeze the current layout<br>
directly instead of forcing <code>guarding_hints_or_throw()</code> or bailing out for every<br>
unbacked stride.</p>
<p>For fallback outputs, temporarily re-enable fresh unbacked symbol tracking while<br>
rerunning the fallback fake kernel. That call is the binding site for output<br>
size and stride symbols that later wrapper code references, so those symbols go<br>
through the normal pending-symbol and <code>compute_unbacked_bindings()</code> path instead<br>
of being created under <code>ignore_fresh_unbacked_symbols()</code> and rediscovered<br>
afterward.</p>
<p>The C++ wrapper path also has to treat input unbacked symbols as already<br>
declared before emitting output bindings. Otherwise a fallback output binding<br>
can redeclare a symbol such as <code>u0</code> and can emit Python-only <code>__floordiv__</code><br>
syntax for <code>DivideByKey</code> paths. Emit C++ integer division for that path and<br>
reuse the existing unbacked symbol declaration.</p>
<p>Finally, do not make the post-copy stride-order sanity check prove<br>
data-dependent unbacked stride inequalities. The copy has already been<br>
materialized with the requested stride order; requiring a symbolic proof there<br>
rejects valid layouts whose unbacked size may be zero or one.</p>
<p>These changes preserve symbolic semantics: hints are not used to create semantic<br>
layout guards, fallback output extents are bound through normal ShapeEnv<br>
tracking at the fallback output binding site, and stride/order changes copy data<br>
or freeze already-valid layouts rather than manufacturing semantic guards.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450833733" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183834" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/183834/hovercard" href="https://github.com/pytorch/pytorch/issues/183834">#183834</a></p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4530789485" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185341" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/185341/hovercard" href="https://github.com/pytorch/pytorch/issues/185341">#185341</a></p>
<p>This PR was authored with the assistance of an AI assistant.</p>
<p>Test Plan:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCPU.test_stride_order_uses_unbacked_optimization_hint_cpu TestUnbackedSymintsCPU.test_stride_ordered_uses_symbolic_divisibility_cpu TestUnbackedSymintsCPU.test_stride_ordered_rejects_unproved_unbacked_layout_cpu -q"><pre>python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCPU.test_stride_order_uses_unbacked_optimization_hint_cpu TestUnbackedSymintsCPU.test_stride_ordered_uses_symbolic_divisibility_cpu TestUnbackedSymintsCPU.test_stride_ordered_rejects_unproved_unbacked_layout_cpu -q</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCUDA.test_standalone_compile_reuses_fallback_unbacked_binding_cuda TestUnbackedSymintsCUDA.test_sdfpa_unbacked_strides_cuda -q"><pre>python test/inductor/test_unbacked_symints.py TestUnbackedSymintsCUDA.test_standalone_compile_reuses_fallback_unbacked_binding_cuda TestUnbackedSymintsCUDA.test_sdfpa_unbacked_strides_cuda -q</pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content='python -m pytest test/inductor/test_unbacked_symints.py -q -s -k "stride_order"'><pre>python -m pytest test/inductor/test_unbacked_symints.py -q -s -k <span class="pl-s"><span class="pl-pds">"</span>stride_order<span class="pl-pds">"</span></span></pre></div>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="lintrunner -a"><pre>lintrunner -a</pre></div>
<p>stack-info: PR: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4450851914" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183840" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183840/hovercard" href="https://github.com/pytorch/pytorch/pull/183840">#183840</a>, branch: sanketpurandare/stack/14</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/2a3d5d1fb4a5c90d8d34770cd876858739533d93]]></title>
<description><![CDATA[[DTensor] Fix group_norm scalar adjuster crash when weight=None (#184…]]></description>
<link>https://tsecurity.de/de/3586304/downloads/trunk2a3d5d1fb4a5c90d8d34770cd876858739533d93/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586304/downloads/trunk2a3d5d1fb4a5c90d8d34770cd876858739533d93/</guid>
<pubDate>Wed, 10 Jun 2026 03:46:47 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[DTensor] Fix group_norm scalar adjuster crash when weight=None (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="186193628" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/184/hovercard" href="https://github.com/pytorch/pytorch/issues/184">#184</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/186854: Simplify endianness check using std::endian]]></title>
<description><![CDATA[std::endian is available since C++20 #176662]]></description>
<link>https://tsecurity.de/de/3586245/downloads/ciflowtrunk186854-simplify-endianness-check-using-stdendian/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586245/downloads/ciflowtrunk186854-simplify-endianness-check-using-stdendian/</guid>
<pubDate>Wed, 10 Jun 2026 02:31:27 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>std::endian</code> is available since C++20 <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4031326243" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/176662" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/176662/hovercard" href="https://github.com/pytorch/pytorch/issues/176662">#176662</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/vllm/186854: Simplify endianness check using std::endian]]></title>
<description><![CDATA[std::endian is available since C++20 #176662]]></description>
<link>https://tsecurity.de/de/3586244/downloads/ciflowvllm186854-simplify-endianness-check-using-stdendian/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586244/downloads/ciflowvllm186854-simplify-endianness-check-using-stdendian/</guid>
<pubDate>Wed, 10 Jun 2026 02:31:26 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>std::endian</code> is available since C++20 <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4031326243" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/176662" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/176662/hovercard" href="https://github.com/pytorch/pytorch/issues/176662">#176662</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781048736: Validate delta type in nn.HuberLoss constructor (#184012)]]></title>
<description><![CDATA[Fixes #133796
nn.HuberLoss(delta=1.+0.j) silently accepts a complex value even though the error message for the underlying huber_loss() function states that delta must be a float. Other invalid types like strings are also accepted by the constructor but fail later with confusing errors.
This adds...]]></description>
<link>https://tsecurity.de/de/3586196/downloads/viablestrict1781048736-validate-delta-type-in-nnhuberloss-constructor-184012/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586196/downloads/viablestrict1781048736-validate-delta-type-in-nnhuberloss-constructor-184012/</guid>
<pubDate>Wed, 10 Jun 2026 01:46:31 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2471790867" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/133796" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/133796/hovercard" href="https://github.com/pytorch/pytorch/issues/133796">#133796</a></p>
<p><code>nn.HuberLoss(delta=1.+0.j)</code> silently accepts a complex value even though the error message for the underlying <code>huber_loss()</code> function states that delta must be a float. Other invalid types like strings are also accepted by the constructor but fail later with confusing errors.</p>
<p>This adds an explicit type check in <code>HuberLoss.__init__</code> that rejects non-numeric delta values (e.g. complex, str) with a clear <code>TypeError</code>. Ints and bools are still accepted for backward compatibility since <code>int</code> is a reasonable numeric type and <code>bool</code> is a subclass of <code>int</code>.</p>
<p><strong>Changes:</strong></p>
<ul>
<li><code>torch/nn/modules/loss.py</code>: Add <code>isinstance(delta, (float, int))</code> check in <code>__init__</code></li>
<li><code>torch/testing/_internal/common_modules.py</code>: Add <code>ErrorModuleInput</code> test for complex delta</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4459600093" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184012" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184012/hovercard" href="https://github.com/pytorch/pytorch/pull/184012">#184012</a><br>
Approved by: <a href="https://github.com/jbschlosser">https://github.com/jbschlosser</a>, <a href="https://github.com/zou3519">https://github.com/zou3519</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781045526: [MPS] Metal cumsum cumprod kernels (#185609)]]></title>
<description><![CDATA[Fixes #154881 and #184844
PR is a bit of a big one so I'll try to summarize below:
Each row is an example tensor (scan dim + dtype) and the kernel it lands on.  1M = a long axis, the specialized long-scan kernels need axis >= 65536 with  scan the block sums -> re-read + scan), with 32 x stride sh...]]></description>
<link>https://tsecurity.de/de/3586145/downloads/viablestrict1781045526-mps-metal-cumsum-cumprod-kernels-185609/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3586145/downloads/viablestrict1781045526-mps-metal-cumsum-cumprod-kernels-185609/</guid>
<pubDate>Wed, 10 Jun 2026 01:01:31 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3111311551" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/154881" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/154881/hovercard" href="https://github.com/pytorch/pytorch/issues/154881">#154881</a> and <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4499614146" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184844" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/184844/hovercard" href="https://github.com/pytorch/pytorch/issues/184844">#184844</a></p>
<p>PR is a bit of a big one so I'll try to summarize below:<br>
Each row is an example tensor (scan dim + dtype) and the kernel it lands on. <code> 1M</code> = a long axis, the specialized long-scan kernels need axis <code>&gt;= 65536</code> with <code>&lt;= 1024</code> scans, everything else takes the generic fallback.</p>
<table>
<thead>
<tr>
<th>example shape</th>
<th>kernel</th>
<th>what it does</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>[8, 1M]</code> dim=-1, float/half/bf16</td>
<td><code>scan_contig_decoupled</code></td>
<td>each 1M row is cut into 4096-elem tiles, <strong>one threadgroup per tile</strong> (256 threads, ~256 tiles/row), all running at once; a decoupled look-back hands each tile's carry to the next via small atomics. <strong>"single pass" = every element read once + written once</strong> (many threadgroups, <em>not</em> one; data never re-read).</td>
</tr>
<tr>
<td><code>[8, 1M]</code> dim=-1, int/long</td>
<td><code>scan_block_reduce</code> + <code>scan_block_carry</code></td>
<td>int can't use the float look-back, due to not having a sentinel value like float(NaN) to signal the "word" not being ready. Therefore we need <strong>2 passes over the data</strong>: pass 1 splits the row into blocks (one threadgroup each) and reduces each to a sum; pass 2 re-reads + scans each block, seeded by the running total of preceding block sums.</td>
</tr>
<tr>
<td><code>[1M, 2]</code> dim=0, float</td>
<td><code>scan_vec_decoupled</code></td>
<td>the contig kernel, but each "element" is the 2-wide row read as one coalesced <code>float2</code>; one threadgroup per tile, look-back per component, 1 read + 1 write.</td>
</tr>
<tr>
<td><code>[1M, 8]</code> dim=0, float</td>
<td><code>scan_strided_col_decoupled</code></td>
<td><strong>one simdgroup (32 lanes) per column</strong> (16 strided reads/lane); 8 simdgroups per threadgroup claim adjacent columns so the strided reads coalesce; look-back per column, 1 read + 1 write.</td>
</tr>
<tr>
<td><code>[1M, 8]</code> dim=0, int/long</td>
<td><code>scan_strided_block_*</code></td>
<td>int version of the column scan: <strong>3 passes</strong> (reduce blocks -&gt; scan the block sums -&gt; re-read + scan), with 32 x stride shared-memory tiles.</td>
</tr>
<tr>
<td><code>[1M, 1]</code> dim=0</td>
<td>reshape -&gt; <code>scan_contig_decoupled</code></td>
<td>trailing size-1 dim =&gt; stride 1 =&gt; identical to the contiguous case (row 1).</td>
</tr>
<tr>
<td><code>[8192, 16]</code> dim=-1</td>
<td><code>scan_tiny_innermost</code></td>
<td>rows too short to justify a threadgroup each, so <strong>one threadgroup scans ~128 rows</strong> (packs ~2048 elems into shared memory), serial-scanning each.</td>
</tr>
<tr>
<td><code>[128, 4096]</code> dim=any; any <code>&gt;4D</code>; <code>logcumsumexp</code>; complex</td>
<td><code>scan_innermost_dim</code> / <code>scan_outer_dim</code></td>
<td><strong>one threadgroup per scan line</strong>, whole line scanned in threadgroup memory; the simple baseline, handles any dim/op.</td>
</tr>
</tbody>
</table>
<h2>tiny - scan_tiny_innermost</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>262144x16</td>
<td align="right">-1</td>
<td align="right">80.0</td>
<td align="right">413.7</td>
<td align="right">5.17x</td>
</tr>
<tr>
<td>bf16</td>
<td>262144x32</td>
<td align="right">-1</td>
<td align="right">188.8</td>
<td align="right">415.6</td>
<td align="right">2.20x</td>
</tr>
<tr>
<td>bf16</td>
<td>262144x8</td>
<td align="right">-1</td>
<td align="right">41.0</td>
<td align="right">412.9</td>
<td align="right">10.08x</td>
</tr>
<tr>
<td>fp32</td>
<td>262144x16</td>
<td align="right">-1</td>
<td align="right">116.2</td>
<td align="right">347.6</td>
<td align="right">2.99x</td>
</tr>
<tr>
<td>fp32</td>
<td>262144x32</td>
<td align="right">-1</td>
<td align="right">260.2</td>
<td align="right">358.3</td>
<td align="right">1.38x</td>
</tr>
<tr>
<td>fp32</td>
<td>262144x8</td>
<td align="right">-1</td>
<td align="right">41.6</td>
<td align="right">345.5</td>
<td align="right">8.30x</td>
</tr>
<tr>
<td>i32</td>
<td>262144x16</td>
<td align="right">-1</td>
<td align="right">178.9</td>
<td align="right">540.9</td>
<td align="right">3.02x</td>
</tr>
<tr>
<td>i32</td>
<td>262144x32</td>
<td align="right">-1</td>
<td align="right">466.7</td>
<td align="right">755.4</td>
<td align="right">1.62x</td>
</tr>
<tr>
<td>i32</td>
<td>262144x8</td>
<td align="right">-1</td>
<td align="right">55.8</td>
<td align="right">435.8</td>
<td align="right">7.81x</td>
</tr>
<tr>
<td>i64</td>
<td>262144x16</td>
<td align="right">-1</td>
<td align="right">246.4</td>
<td align="right">1136.0</td>
<td align="right">4.61x</td>
</tr>
<tr>
<td>i64</td>
<td>262144x32</td>
<td align="right">-1</td>
<td align="right">555.5</td>
<td align="right">1143.0</td>
<td align="right">2.06x</td>
</tr>
<tr>
<td>i64</td>
<td>262144x8</td>
<td align="right">-1</td>
<td align="right">110.0</td>
<td align="right">1136.1</td>
<td align="right">10.32x</td>
</tr>
</tbody>
</table>
<h2>contig - decoupled (float) / multiblock 3-pass (int)</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>512x65536</td>
<td align="right">-1</td>
<td align="right">496.2</td>
<td align="right">769.0</td>
<td align="right">1.55x</td>
</tr>
<tr>
<td>bf16</td>
<td>64x1048576</td>
<td align="right">-1</td>
<td align="right">988.4</td>
<td align="right">1528.1</td>
<td align="right">1.55x</td>
</tr>
<tr>
<td>bf16</td>
<td>8x1048576</td>
<td align="right">-1</td>
<td align="right">118.2</td>
<td align="right">153.2</td>
<td align="right">1.30x</td>
</tr>
<tr>
<td>fp32</td>
<td>512x65536</td>
<td align="right">-1</td>
<td align="right">984.1</td>
<td align="right">1530.0</td>
<td align="right">1.55x</td>
</tr>
<tr>
<td>fp32</td>
<td>64x1048576</td>
<td align="right">-1</td>
<td align="right">2006.3</td>
<td align="right">3075.1</td>
<td align="right">1.53x</td>
</tr>
<tr>
<td>fp32</td>
<td>8x1048576</td>
<td align="right">-1</td>
<td align="right">252.7</td>
<td align="right">374.9</td>
<td align="right">1.48x</td>
</tr>
<tr>
<td>i32</td>
<td>512x65536</td>
<td align="right">-1</td>
<td align="right">2007.4</td>
<td align="right">3133.6</td>
<td align="right">1.56x</td>
</tr>
<tr>
<td>i32</td>
<td>64x1048576</td>
<td align="right">-1</td>
<td align="right">4000.3</td>
<td align="right">6216.6</td>
<td align="right">1.55x</td>
</tr>
<tr>
<td>i32</td>
<td>8x1048576</td>
<td align="right">-1</td>
<td align="right">502.0</td>
<td align="right">769.2</td>
<td align="right">1.53x</td>
</tr>
<tr>
<td>i64</td>
<td>512x65536</td>
<td align="right">-1</td>
<td align="right">2850.1</td>
<td align="right">3116.5</td>
<td align="right">1.09x</td>
</tr>
<tr>
<td>i64</td>
<td>64x1048576</td>
<td align="right">-1</td>
<td align="right">5688.6</td>
<td align="right">6211.4</td>
<td align="right">1.09x</td>
</tr>
<tr>
<td>i64</td>
<td>8x1048576</td>
<td align="right">-1</td>
<td align="right">731.9</td>
<td align="right">851.9</td>
<td align="right">1.16x</td>
</tr>
</tbody>
</table>
<h2>inner_fb - scan_innermost_dim (generic)</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>1100x65536</td>
<td align="right">-1</td>
<td align="right">1066.5</td>
<td align="right">1632.6</td>
<td align="right">1.53x</td>
</tr>
<tr>
<td>bf16</td>
<td>256x8192</td>
<td align="right">-1</td>
<td align="right">24.0</td>
<td align="right">35.9</td>
<td align="right">1.49x</td>
</tr>
<tr>
<td>bf16</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">248.8</td>
<td align="right">251.4</td>
<td align="right">1.01x</td>
</tr>
<tr>
<td>fp32</td>
<td>1100x65536</td>
<td align="right">-1</td>
<td align="right">2091.5</td>
<td align="right">3284.4</td>
<td align="right">1.57x</td>
</tr>
<tr>
<td>fp32</td>
<td>256x8192</td>
<td align="right">-1</td>
<td align="right">36.0</td>
<td align="right">56.5</td>
<td align="right">1.57x</td>
</tr>
<tr>
<td>fp32</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">487.4</td>
<td align="right">517.9</td>
<td align="right">1.06x</td>
</tr>
<tr>
<td>i32</td>
<td>1100x65536</td>
<td align="right">-1</td>
<td align="right">3306.3</td>
<td align="right">6697.6</td>
<td align="right">2.03x</td>
</tr>
<tr>
<td>i32</td>
<td>256x8192</td>
<td align="right">-1</td>
<td align="right">62.2</td>
<td align="right">145.6</td>
<td align="right">2.34x</td>
</tr>
<tr>
<td>i32</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">768.1</td>
<td align="right">1323.8</td>
<td align="right">1.72x</td>
</tr>
<tr>
<td>i64</td>
<td>1100x65536</td>
<td align="right">-1</td>
<td align="right">4183.8</td>
<td align="right">6638.6</td>
<td align="right">1.59x</td>
</tr>
<tr>
<td>i64</td>
<td>256x8192</td>
<td align="right">-1</td>
<td align="right">114.0</td>
<td align="right">243.2</td>
<td align="right">2.13x</td>
</tr>
<tr>
<td>i64</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">972.1</td>
<td align="right">1009.6</td>
<td align="right">1.04x</td>
</tr>
</tbody>
</table>
<h2>vec - vec_decoupled (float) / strided_multiblock (int)</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">41.8</td>
<td align="right">40.4</td>
<td align="right">0.97x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">50.3</td>
<td align="right">66.7</td>
<td align="right">1.33x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">90.3</td>
<td align="right">150.0</td>
<td align="right">1.66x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">152.1</td>
<td align="right">176.3</td>
<td align="right">1.16x</td>
</tr>
</tbody>
</table>
<h2>col - strided_col decoupled (float) / strided_multiblock (int)</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">909.2</td>
<td align="right">1809.8</td>
<td align="right">1.99x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">136.0</td>
<td align="right">218.4</td>
<td align="right">1.61x</td>
</tr>
<tr>
<td>bf16</td>
<td>8192x256</td>
<td align="right">0</td>
<td align="right">49.9</td>
<td align="right">88.7</td>
<td align="right">1.78x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">1295.2</td>
<td align="right">3441.8</td>
<td align="right">2.66x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">262.4</td>
<td align="right">492.9</td>
<td align="right">1.88x</td>
</tr>
<tr>
<td>fp32</td>
<td>8192x256</td>
<td align="right">0</td>
<td align="right">52.2</td>
<td align="right">90.2</td>
<td align="right">1.73x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">2360.9</td>
<td align="right">5134.7</td>
<td align="right">2.17x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">541.3</td>
<td align="right">949.8</td>
<td align="right">1.75x</td>
</tr>
<tr>
<td>i32</td>
<td>8192x256</td>
<td align="right">0</td>
<td align="right">141.2</td>
<td align="right">173.9</td>
<td align="right">1.23x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">3031.4</td>
<td align="right">5004.6</td>
<td align="right">1.65x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">777.9</td>
<td align="right">1124.8</td>
<td align="right">1.45x</td>
</tr>
<tr>
<td>i64</td>
<td>8192x256</td>
<td align="right">0</td>
<td align="right">177.9</td>
<td align="right">317.8</td>
<td align="right">1.79x</td>
</tr>
</tbody>
</table>
<h2>outer_fb - scan_outer_dim (generic, axis&lt;1024)</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>512x16384</td>
<td align="right">0</td>
<td align="right">127.3</td>
<td align="right">186.6</td>
<td align="right">1.47x</td>
</tr>
<tr>
<td>bf16</td>
<td>512x256</td>
<td align="right">0</td>
<td align="right">15.2</td>
<td align="right">12.6</td>
<td align="right">0.83x</td>
</tr>
<tr>
<td>fp32</td>
<td>512x16384</td>
<td align="right">0</td>
<td align="right">256.5</td>
<td align="right">303.0</td>
<td align="right">1.18x</td>
</tr>
<tr>
<td>fp32</td>
<td>512x256</td>
<td align="right">0</td>
<td align="right">15.0</td>
<td align="right">12.3</td>
<td align="right">0.82x</td>
</tr>
<tr>
<td>i32</td>
<td>512x16384</td>
<td align="right">0</td>
<td align="right">404.2</td>
<td align="right">733.1</td>
<td align="right">1.81x</td>
</tr>
<tr>
<td>i32</td>
<td>512x256</td>
<td align="right">0</td>
<td align="right">14.1</td>
<td align="right">15.6</td>
<td align="right">1.11x</td>
</tr>
<tr>
<td>i64</td>
<td>512x16384</td>
<td align="right">0</td>
<td align="right">504.4</td>
<td align="right">739.7</td>
<td align="right">1.47x</td>
</tr>
<tr>
<td>i64</td>
<td>512x256</td>
<td align="right">0</td>
<td align="right">14.5</td>
<td align="right">27.6</td>
<td align="right">1.90x</td>
</tr>
</tbody>
</table>
<h2>col|outer - strided_col (float) / scan_outer_dim (int, cols&gt;=64)</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>1024x4096</td>
<td align="right">0</td>
<td align="right">86.1</td>
<td align="right">106.6</td>
<td align="right">1.24x</td>
</tr>
<tr>
<td>bf16</td>
<td>2048x8192</td>
<td align="right">0</td>
<td align="right">386.0</td>
<td align="right">414.5</td>
<td align="right">1.07x</td>
</tr>
<tr>
<td>fp32</td>
<td>1024x4096</td>
<td align="right">0</td>
<td align="right">138.0</td>
<td align="right">134.9</td>
<td align="right">0.98x</td>
</tr>
<tr>
<td>fp32</td>
<td>2048x8192</td>
<td align="right">0</td>
<td align="right">641.0</td>
<td align="right">810.9</td>
<td align="right">1.27x</td>
</tr>
<tr>
<td>i32</td>
<td>1024x4096</td>
<td align="right">0</td>
<td align="right">195.6</td>
<td align="right">339.4</td>
<td align="right">1.74x</td>
</tr>
<tr>
<td>i32</td>
<td>2048x8192</td>
<td align="right">0</td>
<td align="right">830.3</td>
<td align="right">1710.4</td>
<td align="right">2.06x</td>
</tr>
<tr>
<td>i64</td>
<td>1024x4096</td>
<td align="right">0</td>
<td align="right">253.4</td>
<td align="right">323.1</td>
<td align="right">1.27x</td>
</tr>
<tr>
<td>i64</td>
<td>2048x8192</td>
<td align="right">0</td>
<td align="right">1020.7</td>
<td align="right">1741.3</td>
<td align="right">1.71x</td>
</tr>
</tbody>
</table>
<h2>reshape - trailing size-1 dims -&gt; contiguous scan</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>8x65536x1</td>
<td align="right">1</td>
<td align="right">10.5</td>
<td align="right">12.8</td>
<td align="right">1.22x</td>
</tr>
<tr>
<td>fp32</td>
<td>8x65536x1</td>
<td align="right">1</td>
<td align="right">13.9</td>
<td align="right">16.3</td>
<td align="right">1.17x</td>
</tr>
<tr>
<td>i32</td>
<td>8x65536x1</td>
<td align="right">1</td>
<td align="right">22.4</td>
<td align="right">28.8</td>
<td align="right">1.29x</td>
</tr>
<tr>
<td>i64</td>
<td>8x65536x1</td>
<td align="right">1</td>
<td align="right">31.1</td>
<td align="right">37.1</td>
<td align="right">1.19x</td>
</tr>
</tbody>
</table>
<h2>mid3d - middle-dim outer scan</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>16x70000x5</td>
<td align="right">1</td>
<td align="right">79.5</td>
<td align="right">158.1</td>
<td align="right">1.99x</td>
</tr>
<tr>
<td>bf16</td>
<td>64x4096x8</td>
<td align="right">1</td>
<td align="right">34.0</td>
<td align="right">49.7</td>
<td align="right">1.46x</td>
</tr>
<tr>
<td>fp32</td>
<td>16x70000x5</td>
<td align="right">1</td>
<td align="right">173.6</td>
<td align="right">381.4</td>
<td align="right">2.20x</td>
</tr>
<tr>
<td>fp32</td>
<td>64x4096x8</td>
<td align="right">1</td>
<td align="right">43.5</td>
<td align="right">69.4</td>
<td align="right">1.59x</td>
</tr>
<tr>
<td>i32</td>
<td>16x70000x5</td>
<td align="right">1</td>
<td align="right">504.3</td>
<td align="right">857.8</td>
<td align="right">1.70x</td>
</tr>
<tr>
<td>i32</td>
<td>64x4096x8</td>
<td align="right">1</td>
<td align="right">105.7</td>
<td align="right">155.2</td>
<td align="right">1.47x</td>
</tr>
<tr>
<td>i64</td>
<td>16x70000x5</td>
<td align="right">1</td>
<td align="right">533.6</td>
<td align="right">1069.0</td>
<td align="right">2.00x</td>
</tr>
<tr>
<td>i64</td>
<td>64x4096x8</td>
<td align="right">1</td>
<td align="right">180.6</td>
<td align="right">308.4</td>
<td align="right">1.71x</td>
</tr>
</tbody>
</table>
<h2>nc_in - non-contiguous (transposed) input</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">303.8</td>
<td align="right">391.9</td>
<td align="right">1.29x</td>
</tr>
<tr>
<td>fp32</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">540.8</td>
<td align="right">673.8</td>
<td align="right">1.25x</td>
</tr>
<tr>
<td>i32</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">840.2</td>
<td align="right">1496.5</td>
<td align="right">1.78x</td>
</tr>
<tr>
<td>i64</td>
<td>4096x4096</td>
<td align="right">-1</td>
<td align="right">1019.0</td>
<td align="right">1338.5</td>
<td align="right">1.31x</td>
</tr>
</tbody>
</table>
<h2>model_seq - [batch, seq]</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>32x2048</td>
<td align="right">-1</td>
<td align="right">3.1</td>
<td align="right">11.7</td>
<td align="right">3.79x</td>
</tr>
<tr>
<td>fp32</td>
<td>32x2048</td>
<td align="right">-1</td>
<td align="right">2.8</td>
<td align="right">11.8</td>
<td align="right">4.17x</td>
</tr>
<tr>
<td>i32</td>
<td>32x2048</td>
<td align="right">-1</td>
<td align="right">3.7</td>
<td align="right">14.9</td>
<td align="right">4.08x</td>
</tr>
<tr>
<td>i64</td>
<td>32x2048</td>
<td align="right">-1</td>
<td align="right">3.7</td>
<td align="right">23.7</td>
<td align="right">6.37x</td>
</tr>
</tbody>
</table>
<h2>model_bhs - [batch, heads, seq]</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>16x16x4096</td>
<td align="right">-1</td>
<td align="right">14.3</td>
<td align="right">21.6</td>
<td align="right">1.51x</td>
</tr>
<tr>
<td>fp32</td>
<td>16x16x4096</td>
<td align="right">-1</td>
<td align="right">18.9</td>
<td align="right">33.4</td>
<td align="right">1.77x</td>
</tr>
<tr>
<td>i32</td>
<td>16x16x4096</td>
<td align="right">-1</td>
<td align="right">28.6</td>
<td align="right">58.7</td>
<td align="right">2.05x</td>
</tr>
<tr>
<td>i64</td>
<td>16x16x4096</td>
<td align="right">-1</td>
<td align="right">37.0</td>
<td align="right">133.2</td>
<td align="right">3.60x</td>
</tr>
</tbody>
</table>
<h2>model_mamba - [d_inner, seq] outer</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>8192x1024</td>
<td align="right">0</td>
<td align="right">188.2</td>
<td align="right">340.8</td>
<td align="right">1.81x</td>
</tr>
<tr>
<td>fp32</td>
<td>8192x1024</td>
<td align="right">0</td>
<td align="right">305.2</td>
<td align="right">497.1</td>
<td align="right">1.63x</td>
</tr>
<tr>
<td>i32</td>
<td>8192x1024</td>
<td align="right">0</td>
<td align="right">655.1</td>
<td align="right">905.2</td>
<td align="right">1.38x</td>
</tr>
<tr>
<td>i64</td>
<td>8192x1024</td>
<td align="right">0</td>
<td align="right">833.7</td>
<td align="right">1174.8</td>
<td align="right">1.41x</td>
</tr>
</tbody>
</table>
<h2>sweep_1Mx - [1M, n] dim0 inner-stride sweep</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">41.6</td>
<td align="right">40.4</td>
<td align="right">0.97x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x3</td>
<td align="right">0</td>
<td align="right">67.6</td>
<td align="right">64.3</td>
<td align="right">0.95x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x4</td>
<td align="right">0</td>
<td align="right">80.3</td>
<td align="right">95.9</td>
<td align="right">1.19x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">136.5</td>
<td align="right">216.3</td>
<td align="right">1.58x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x16</td>
<td align="right">0</td>
<td align="right">297.5</td>
<td align="right">613.7</td>
<td align="right">2.06x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x24</td>
<td align="right">0</td>
<td align="right">502.2</td>
<td align="right">961.4</td>
<td align="right">1.91x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">904.1</td>
<td align="right">1544.7</td>
<td align="right">1.71x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x48</td>
<td align="right">0</td>
<td align="right">1239.9</td>
<td align="right">2239.2</td>
<td align="right">1.81x</td>
</tr>
<tr>
<td>bf16</td>
<td>1048576x64</td>
<td align="right">0</td>
<td align="right">1847.0</td>
<td align="right">4507.1</td>
<td align="right">2.44x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">50.3</td>
<td align="right">66.8</td>
<td align="right">1.33x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x3</td>
<td align="right">0</td>
<td align="right">90.7</td>
<td align="right">94.2</td>
<td align="right">1.04x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x4</td>
<td align="right">0</td>
<td align="right">138.0</td>
<td align="right">168.1</td>
<td align="right">1.22x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">263.3</td>
<td align="right">491.8</td>
<td align="right">1.87x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x16</td>
<td align="right">0</td>
<td align="right">643.8</td>
<td align="right">1133.0</td>
<td align="right">1.76x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x24</td>
<td align="right">0</td>
<td align="right">963.0</td>
<td align="right">1813.9</td>
<td align="right">1.88x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">1297.6</td>
<td align="right">3365.6</td>
<td align="right">2.59x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x48</td>
<td align="right">0</td>
<td align="right">1889.4</td>
<td align="right">4937.8</td>
<td align="right">2.61x</td>
</tr>
<tr>
<td>fp32</td>
<td>1048576x64</td>
<td align="right">0</td>
<td align="right">2588.0</td>
<td align="right">7967.3</td>
<td align="right">3.08x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">93.2</td>
<td align="right">145.6</td>
<td align="right">1.56x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x3</td>
<td align="right">0</td>
<td align="right">167.5</td>
<td align="right">232.7</td>
<td align="right">1.39x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x4</td>
<td align="right">0</td>
<td align="right">225.0</td>
<td align="right">356.8</td>
<td align="right">1.59x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">541.2</td>
<td align="right">933.5</td>
<td align="right">1.72x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x16</td>
<td align="right">0</td>
<td align="right">1162.5</td>
<td align="right">2373.7</td>
<td align="right">2.04x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x24</td>
<td align="right">0</td>
<td align="right">1969.4</td>
<td align="right">3806.6</td>
<td align="right">1.93x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">2367.8</td>
<td align="right">5066.6</td>
<td align="right">2.14x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x48</td>
<td align="right">0</td>
<td align="right">4325.3</td>
<td align="right">7599.3</td>
<td align="right">1.76x</td>
</tr>
<tr>
<td>i32</td>
<td>1048576x64</td>
<td align="right">0</td>
<td align="right">4637.8</td>
<td align="right">11633.8</td>
<td align="right">2.51x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x2</td>
<td align="right">0</td>
<td align="right">152.2</td>
<td align="right">178.1</td>
<td align="right">1.17x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x3</td>
<td align="right">0</td>
<td align="right">276.1</td>
<td align="right">443.2</td>
<td align="right">1.61x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x4</td>
<td align="right">0</td>
<td align="right">375.2</td>
<td align="right">549.3</td>
<td align="right">1.46x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x8</td>
<td align="right">0</td>
<td align="right">777.5</td>
<td align="right">1143.2</td>
<td align="right">1.47x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x16</td>
<td align="right">0</td>
<td align="right">1480.6</td>
<td align="right">2349.6</td>
<td align="right">1.59x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x24</td>
<td align="right">0</td>
<td align="right">2360.8</td>
<td align="right">3733.4</td>
<td align="right">1.58x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x32</td>
<td align="right">0</td>
<td align="right">3029.8</td>
<td align="right">4991.4</td>
<td align="right">1.65x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x48</td>
<td align="right">0</td>
<td align="right">4791.1</td>
<td align="right">7845.6</td>
<td align="right">1.64x</td>
</tr>
<tr>
<td>i64</td>
<td>1048576x64</td>
<td align="right">0</td>
<td align="right">6117.7</td>
<td align="right">10750.9</td>
<td align="right">1.76x</td>
</tr>
</tbody>
</table>
<h2>sweep_4x256k - [4, 256k, n] dim1 inner-stride sweep</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>4x262144x2</td>
<td align="right">1</td>
<td align="right">40.1</td>
<td align="right">37.3</td>
<td align="right">0.93x</td>
</tr>
<tr>
<td>bf16</td>
<td>4x262144x4</td>
<td align="right">1</td>
<td align="right">69.7</td>
<td align="right">88.5</td>
<td align="right">1.27x</td>
</tr>
<tr>
<td>bf16</td>
<td>4x262144x8</td>
<td align="right">1</td>
<td align="right">125.9</td>
<td align="right">219.2</td>
<td align="right">1.74x</td>
</tr>
<tr>
<td>bf16</td>
<td>4x262144x16</td>
<td align="right">1</td>
<td align="right">288.5</td>
<td align="right">618.1</td>
<td align="right">2.14x</td>
</tr>
<tr>
<td>bf16</td>
<td>4x262144x32</td>
<td align="right">1</td>
<td align="right">821.8</td>
<td align="right">1627.7</td>
<td align="right">1.98x</td>
</tr>
<tr>
<td>fp32</td>
<td>4x262144x2</td>
<td align="right">1</td>
<td align="right">50.3</td>
<td align="right">64.1</td>
<td align="right">1.27x</td>
</tr>
<tr>
<td>fp32</td>
<td>4x262144x4</td>
<td align="right">1</td>
<td align="right">121.7</td>
<td align="right">170.8</td>
<td align="right">1.40x</td>
</tr>
<tr>
<td>fp32</td>
<td>4x262144x8</td>
<td align="right">1</td>
<td align="right">257.7</td>
<td align="right">486.7</td>
<td align="right">1.89x</td>
</tr>
<tr>
<td>fp32</td>
<td>4x262144x16</td>
<td align="right">1</td>
<td align="right">572.3</td>
<td align="right">1171.6</td>
<td align="right">2.05x</td>
</tr>
<tr>
<td>fp32</td>
<td>4x262144x32</td>
<td align="right">1</td>
<td align="right">1239.0</td>
<td align="right">3542.3</td>
<td align="right">2.86x</td>
</tr>
<tr>
<td>i32</td>
<td>4x262144x2</td>
<td align="right">1</td>
<td align="right">91.7</td>
<td align="right">148.0</td>
<td align="right">1.61x</td>
</tr>
<tr>
<td>i32</td>
<td>4x262144x4</td>
<td align="right">1</td>
<td align="right">224.7</td>
<td align="right">363.0</td>
<td align="right">1.62x</td>
</tr>
<tr>
<td>i32</td>
<td>4x262144x8</td>
<td align="right">1</td>
<td align="right">543.2</td>
<td align="right">936.7</td>
<td align="right">1.72x</td>
</tr>
<tr>
<td>i32</td>
<td>4x262144x16</td>
<td align="right">1</td>
<td align="right">1192.4</td>
<td align="right">2314.3</td>
<td align="right">1.94x</td>
</tr>
<tr>
<td>i32</td>
<td>4x262144x32</td>
<td align="right">1</td>
<td align="right">2413.2</td>
<td align="right">5172.1</td>
<td align="right">2.14x</td>
</tr>
<tr>
<td>i64</td>
<td>4x262144x2</td>
<td align="right">1</td>
<td align="right">152.2</td>
<td align="right">163.8</td>
<td align="right">1.08x</td>
</tr>
<tr>
<td>i64</td>
<td>4x262144x4</td>
<td align="right">1</td>
<td align="right">372.6</td>
<td align="right">524.1</td>
<td align="right">1.41x</td>
</tr>
<tr>
<td>i64</td>
<td>4x262144x8</td>
<td align="right">1</td>
<td align="right">774.6</td>
<td align="right">1132.0</td>
<td align="right">1.46x</td>
</tr>
<tr>
<td>i64</td>
<td>4x262144x16</td>
<td align="right">1</td>
<td align="right">1482.6</td>
<td align="right">2396.9</td>
<td align="right">1.62x</td>
</tr>
<tr>
<td>i64</td>
<td>4x262144x32</td>
<td align="right">1</td>
<td align="right">3047.8</td>
<td align="right">5210.7</td>
<td align="right">1.71x</td>
</tr>
</tbody>
</table>
<h2>sweep_large - [axis, n] dim0 large-stride sweep</h2>
<table>
<thead>
<tr>
<th>dtype</th>
<th>shape</th>
<th align="right">dim</th>
<th align="right">Metal us</th>
<th align="right">MPSGraph us</th>
<th align="right">speedup</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16</td>
<td>262144x64</td>
<td align="right">0</td>
<td align="right">424.3</td>
<td align="right">704.8</td>
<td align="right">1.66x</td>
</tr>
<tr>
<td>bf16</td>
<td>262144x128</td>
<td align="right">0</td>
<td align="right">852.3</td>
<td align="right">1472.3</td>
<td align="right">1.73x</td>
</tr>
<tr>
<td>bf16</td>
<td>262144x256</td>
<td align="right">0</td>
<td align="right">1736.2</td>
<td align="right">3562.1</td>
<td align="right">2.05x</td>
</tr>
<tr>
<td>bf16</td>
<td>131072x512</td>
<td align="right">0</td>
<td align="right">1602.9</td>
<td align="right">3035.3</td>
<td align="right">1.89x</td>
</tr>
<tr>
<td>bf16</td>
<td>131072x1024</td>
<td align="right">0</td>
<td align="right">3149.5</td>
<td align="right">7540.6</td>
<td align="right">2.39x</td>
</tr>
<tr>
<td>fp32</td>
<td>262144x64</td>
<td align="right">0</td>
<td align="right">627.6</td>
<td align="right">1141.2</td>
<td align="right">1.82x</td>
</tr>
<tr>
<td>fp32</td>
<td>262144x128</td>
<td align="right">0</td>
<td align="right">1265.6</td>
<td align="right">3779.8</td>
<td align="right">2.99x</td>
</tr>
<tr>
<td>fp32</td>
<td>262144x256</td>
<td align="right">0</td>
<td align="right">2476.1</td>
<td align="right">7887.5</td>
<td align="right">3.19x</td>
</tr>
<tr>
<td>fp32</td>
<td>131072x512</td>
<td align="right">0</td>
<td align="right">2450.2</td>
<td align="right">7993.3</td>
<td align="right">3.26x</td>
</tr>
<tr>
<td>fp32</td>
<td>131072x1024</td>
<td align="right">0</td>
<td align="right">5103.4</td>
<td align="right">16029.0</td>
<td align="right">3.14x</td>
</tr>
<tr>
<td>i32</td>
<td>262144x64</td>
<td align="right">0</td>
<td align="right">1244.2</td>
<td align="right">2363.8</td>
<td align="right">1.90x</td>
</tr>
<tr>
<td>i32</td>
<td>262144x128</td>
<td align="right">0</td>
<td align="right">2360.0</td>
<td align="right">5240.1</td>
<td align="right">2.22x</td>
</tr>
<tr>
<td>i32</td>
<td>262144x256</td>
<td align="right">0</td>
<td align="right">4567.1</td>
<td align="right">11069.9</td>
<td align="right">2.42x</td>
</tr>
<tr>
<td>i32</td>
<td>131072x512</td>
<td align="right">0</td>
<td align="right">4612.6</td>
<td align="right">10710.7</td>
<td align="right">2.32x</td>
</tr>
<tr>
<td>i32</td>
<td>131072x1024</td>
<td align="right">0</td>
<td align="right">9130.5</td>
<td align="right">22414.9</td>
<td align="right">2.45x</td>
</tr>
<tr>
<td>i64</td>
<td>262144x64</td>
<td align="right">0</td>
<td align="right">1597.2</td>
<td align="right">2493.7</td>
<td align="right">1.56x</td>
</tr>
<tr>
<td>i64</td>
<td>262144x128</td>
<td align="right">0</td>
<td align="right">3048.3</td>
<td align="right">5219.8</td>
<td align="right">1.71x</td>
</tr>
<tr>
<td>i64</td>
<td>262144x256</td>
<td align="right">0</td>
<td align="right">6004.4</td>
<td align="right">10937.7</td>
<td align="right">1.82x</td>
</tr>
<tr>
<td>i64</td>
<td>131072x512</td>
<td align="right">0</td>
<td align="right">6041.5</td>
<td align="right">10788.0</td>
<td align="right">1.79x</td>
</tr>
<tr>
<td>i64</td>
<td>131072x1024</td>
<td align="right">0</td>
<td align="right">12158.0</td>
<td align="right">22534.8</td>
<td align="right">1.85x</td>
</tr>
</tbody>
</table>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4546723122" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185609" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185609/hovercard" href="https://github.com/pytorch/pytorch/pull/185609">#185609</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/498536e07a1a5173db6dfb83aacb2f0b8da27044]]></title>
<description><![CDATA[Repro for Inductor tl.broadcast_to(False) Triton compile crash (#1866…]]></description>
<link>https://tsecurity.de/de/3585947/downloads/trunk498536e07a1a5173db6dfb83aacb2f0b8da27044/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3585947/downloads/trunk498536e07a1a5173db6dfb83aacb2f0b8da27044/</guid>
<pubDate>Tue, 09 Jun 2026 22:31:23 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Repro for Inductor tl.broadcast_to(False) Triton compile crash (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="237601643" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1866" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1866/hovercard" href="https://github.com/pytorch/pytorch/issues/1866">#1866</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/02776472c5f02856cb4ed9f67250e584c5d1126e: Revert "[inductor] Guard FakeTensorUpdater lowering metadata (#184023)"]]></title>
<description><![CDATA[This reverts commit e48af70.
Reverted #184023 on behalf of https://github.com/eellison due to causing failures internally (comment)]]></description>
<link>https://tsecurity.de/de/3585747/downloads/trunk02776472c5f02856cb4ed9f67250e584c5d1126e-revert-inductor-guard-faketensorupdater-lowering-metadata-184023/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3585747/downloads/trunk02776472c5f02856cb4ed9f67250e584c5d1126e-revert-inductor-guard-faketensorupdater-lowering-metadata-184023/</guid>
<pubDate>Tue, 09 Jun 2026 21:01:57 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/e48af70dc4264c343f76009ef00289adb590dc70/hovercard" href="https://github.com/pytorch/pytorch/commit/e48af70dc4264c343f76009ef00289adb590dc70"><tt>e48af70</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4460315368" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184023" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184023/hovercard" href="https://github.com/pytorch/pytorch/pull/184023">#184023</a> on behalf of <a href="https://github.com/eellison">https://github.com/eellison</a> due to causing failures internally (<a href="https://github.com/pytorch/pytorch/pull/184023#issuecomment-4663005489" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184023/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/060670aadfb5995aaf62aa31f2d7c23db24bd791: Preserve FX graph cache guard provenance (#184193)]]></title>
<description><![CDATA[Summary
Preserve source locations for Inductor guards that are saved in the FX graph cache, so a cache-hit recompile reason still points at the original guard producer instead of the cache replay site.
Fixes #141595
Generated by my agent
Root Cause
Without the FX graph cache, ShapeEnv records eac...]]></description>
<link>https://tsecurity.de/de/3585574/downloads/trunk060670aadfb5995aaf62aa31f2d7c23db24bd791-preserve-fx-graph-cache-guard-provenance-184193/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3585574/downloads/trunk060670aadfb5995aaf62aa31f2d7c23db24bd791-preserve-fx-graph-cache-guard-provenance-184193/</guid>
<pubDate>Tue, 09 Jun 2026 20:15:53 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Preserve source locations for Inductor guards that are saved in the FX graph cache, so a cache-hit recompile reason still points at the original guard producer instead of the cache replay site.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2695664819" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/141595" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/141595/hovercard" href="https://github.com/pytorch/pytorch/issues/141595">#141595</a><br>
Generated by my agent</p>
<h2>Root Cause</h2>
<p>Without the FX graph cache, <code>ShapeEnv</code> records each <code>ShapeGuard</code> with the source location captured when Inductor creates the guard, for example <code>can_use_32bit_indexing</code>.</p>
<p>On an FX graph cache hit, Inductor only had <code>CompiledFxGraph.guards_expr</code>, a single combined Python expression string. Re-evaluating that string installs equivalent guards in the current <code>ShapeEnv</code>, but the new guards get the replay callsite as their source location. That is why recompile logs pointed at <code>_lookup_graph</code>, <code>autograd_cache.py</code>, or <code>&lt;string&gt;:1</code> instead of the original Inductor guard.</p>
<h2>Review Guide</h2>
<ul>
<li><code>torch/fx/experimental/symbolic_shapes.py</code>: threads source locations alongside each produced guard expression, adds <code>ShapeGuardExpression</code>, and replays cached guards while temporarily using the stored source location for any <code>ShapeGuard</code> created by evaluation.</li>
<li><code>torch/_inductor/output_code.py</code>: adds optional cached metadata fields for the per-guard expression/source list and the SymInt argument count used when saving it.</li>
<li><code>torch/_inductor/codecache.py</code>: saves the new metadata, replays it on normal FX graph cache hits, and falls back to the existing combined-expression path for old cache entries, unsafe-skip mode, or AOTAutograd loads whose argument list differs.</li>
<li>Tests cover both the small <code>ShapeEnv</code> round trip and the original CUDA/flex-attention cache-hit failure mode.</li>
</ul>
<h2>Test Plan</h2>
<ul>
<li><code>PYTHONPATH=. python test/test_dynamic_shapes.py TestGuardsExpressions.test_guards_expression_source_info</code></li>
<li><code>PYTHONPATH=. python test/inductor/test_codecache.py TestFxGraphCache.test_cache_hit_inductor_guard_preserves_source</code></li>
<li><code>PYTHONPATH=. python test/dynamo/test_aot_autograd_cache.py AOTAutogradCacheTests.test_output_views_input_dynamic</code></li>
<li><code>lintrunner -a</code></li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4468971201" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184193" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184193/hovercard" href="https://github.com/pytorch/pytorch/pull/184193">#184193</a><br>
Approved by: <a href="https://github.com/oulgen">https://github.com/oulgen</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/34631b2af67b6ba64a6204c3cdb330fd401190fa: Fix CPU Inductor FX cache thread specialization (#184444)]]></title>
<description><![CDATA[Include the resolved CPU runtime thread count in the FX graph cache key when C++ codegen may specialize on it, including no-input CPU factory graphs, so cached serial kernels are not reused under a different thread count. Making the resolved thread count part of the FX graph cache key fixes all r...]]></description>
<link>https://tsecurity.de/de/3585573/downloads/trunk34631b2af67b6ba64a6204c3cdb330fd401190fa-fix-cpu-inductor-fx-cache-thread-specialization-184444/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3585573/downloads/trunk34631b2af67b6ba64a6204c3cdb330fd401190fa-fix-cpu-inductor-fx-cache-thread-specialization-184444/</guid>
<pubDate>Tue, 09 Jun 2026 20:15:52 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Include the resolved CPU runtime thread count in the FX graph cache key when C++ codegen may specialize on it, including no-input CPU factory graphs, so cached serial kernels are not reused under a different thread count. Making the resolved thread count part of the FX graph cache key fixes all reported cases where a CPU graph compiled under one <code>torch.set_num_threads()</code> value was incorrectly served from cache after the thread count changed.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="1928645113" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/110611" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/110611/hovercard" href="https://github.com/pytorch/pytorch/issues/110611">#110611</a><br>
Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3607621605" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/167463" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/167463/hovercard" href="https://github.com/pytorch/pytorch/issues/167463">#167463</a><br>
Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3327318404" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/160812" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/160812/hovercard" href="https://github.com/pytorch/pytorch/issues/160812">#160812</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4481664338" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184444" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184444/hovercard" href="https://github.com/pytorch/pytorch/pull/184444">#184444</a><br>
Approved by: <a href="https://github.com/oulgen">https://github.com/oulgen</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/521f9d784733cb2f4571804b0f6e4687fb7a91be: [inductor] Unify threads-per-wave (warp_size) extraction (#181112)]]></title>
<description><![CDATA[Route threads-per-wave through a single source of truth
(DeviceProperties.warp_size, from torch.cuda.get_device_properties) and
thread it through the config-generation helpers, replacing _NUM_THREADS_PER_WARP,
cc_warp_size, and scattered warp_size or 32 / torch.version.hip branches.
Behavior chan...]]></description>
<link>https://tsecurity.de/de/3585442/downloads/trunk521f9d784733cb2f4571804b0f6e4687fb7a91be-inductor-unify-threads-per-wave-warpsize-extraction-181112/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3585442/downloads/trunk521f9d784733cb2f4571804b0f6e4687fb7a91be-inductor-unify-threads-per-wave-warpsize-extraction-181112/</guid>
<pubDate>Tue, 09 Jun 2026 19:46:55 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Route threads-per-wave through a single source of truth<br>
(DeviceProperties.warp_size, from torch.cuda.get_device_properties) and<br>
thread it through the config-generation helpers, replacing _NUM_THREADS_PER_WARP,<br>
cc_warp_size, and scattered <code>warp_size or 32</code> / <code>torch.version.hip</code> branches.</p>
<p>Behavior change: _num_warps halves num_warps only when warp_size == 64<br>
(AMD CDNA/gfx9), not on every HIP device; AMD RDNA (wave32) now follows<br>
the NVIDIA path.</p>
<p>Authored with: Claude</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4309751213" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/181112" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/181112/hovercard" href="https://github.com/pytorch/pytorch/pull/181112">#181112</a><br>
Approved by: <a href="https://github.com/jansel">https://github.com/jansel</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/8182b8f1f6c7aaebd2da18786f4667dea7542f02: [BE] Make `IMPSAllocator` inherit form `c10::DeviceAllocator` (#186748)]]></title>
<description><![CDATA[Instead of c10::Allocator to make it more compatible with torch.accelerator APIs, that route all torch.accelerator.memory_* cals thru  at::getDeviceAllocator, which does
dynamic_cast(c10::GetAllocator(device)).
Test Plan:
python -c "import torch;x = torch.rand(1024, 1024, device='mps'); print(tor...]]></description>
<link>https://tsecurity.de/de/3585314/downloads/trunk8182b8f1f6c7aaebd2da18786f4667dea7542f02-be-make-impsallocator-inherit-form-c10deviceallocator-186748/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3585314/downloads/trunk8182b8f1f6c7aaebd2da18786f4667dea7542f02-be-make-impsallocator-inherit-form-c10deviceallocator-186748/</guid>
<pubDate>Tue, 09 Jun 2026 19:16:48 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Instead of <code>c10::Allocator</code> to make it more compatible with <code>torch.accelerator</code> APIs, that route all <code>torch.accelerator.memory_*</code> cals thru  <code>at::getDeviceAllocator</code>, which does<br>
<code>dynamic_cast&lt;DeviceAllocator*&gt;(c10::GetAllocator(device))</code>.</p>
<p>Test Plan:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="python -c &quot;import torch;x = torch.rand(1024, 1024, device='mps'); print(torch.mps.current_allocated_memory()); del x; torch.accelerator.empty_cache(); print(torch.accelerator.memory_allocated())&quot;"><pre class="notranslate"><code>python -c "import torch;x = torch.rand(1024, 1024, device='mps'); print(torch.mps.current_allocated_memory()); del x; torch.accelerator.empty_cache(); print(torch.accelerator.memory_allocated())"
</code></pre></div>
<p>Should allow next iterations of <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4475263372" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184326" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184326/hovercard" href="https://github.com/pytorch/pytorch/pull/184326">#184326</a> to land successfully on MacOS</p>
<p>This change was authored with the assistance of Claude (Claude Code).<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4617738254" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186748" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186748/hovercard" href="https://github.com/pytorch/pytorch/pull/186748">#186748</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1781016133: [TMA] Fix NameError in combo kernel when TMA enabled (#185514)]]></title>
<description><![CDATA[Fixes: #185513
Root cause: uniquify_block_sizes() in triton_combo_kernel.py only handled str and DeferredLine objects in the body buffer, but TMA load expressions are stored as DelayReplaceLine objects (a different subclass of DeferredLineBase). These were passed through unmodified, leaving R0_BL...]]></description>
<link>https://tsecurity.de/de/3584869/downloads/viablestrict1781016133-tma-fix-nameerror-in-combo-kernel-when-tma-enabled-185514/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3584869/downloads/viablestrict1781016133-tma-fix-nameerror-in-combo-kernel-when-tma-enabled-185514/</guid>
<pubDate>Tue, 09 Jun 2026 17:01:38 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4540768497" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185513" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/185513/hovercard" href="https://github.com/pytorch/pytorch/issues/185513">#185513</a></p>
<p>Root cause: <code>uniquify_block_sizes()</code> in <code>triton_combo_kernel.py</code> only handled <code>str</code> and <code>DeferredLine</code> objects in the body buffer, but TMA load expressions are stored as <code>DelayReplaceLine</code> objects (a different subclass of <code>DeferredLineBase</code>). These were passed through unmodified, leaving <code>R0_BLOCK</code> unrenamed.</p>
<p>Fix: Changed the isinstance check from <code>DeferredLine</code> to <code>DeferredLineBase</code> and use <code>_new_line()</code> to create the replacement — this handles all deferred line types generically (DeferredLine, DelayReplaceLine, and any future subclasses).</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4540773265" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185514" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185514/hovercard" href="https://github.com/pytorch/pytorch/pull/185514">#185514</a><br>
Approved by: <a href="https://github.com/kshitij12345">https://github.com/kshitij12345</a>, <a href="https://github.com/jansel">https://github.com/jansel</a>, <a href="https://github.com/mlazos">https://github.com/mlazos</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780997283: Add oneDNN LSTM primitive support for XPU inference (#185531)]]></title>
<description><![CDATA[Summary
Use dnnl::lstm_forward primitive on XPU for LSTM inference, replacing the per-timestep fused cell approach. The oneDNN primitive processes the entire sequence internally, eliminating O(T) kernel launches.
Changes

New file aten/src/ATen/native/mkldnn/xpu/RNN.cpp: Implements lstm_onednn_xp...]]></description>
<link>https://tsecurity.de/de/3583981/downloads/viablestrict1780997283-add-onednn-lstm-primitive-support-for-xpu-inference-185531/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583981/downloads/viablestrict1780997283-add-onednn-lstm-primitive-support-for-xpu-inference-185531/</guid>
<pubDate>Tue, 09 Jun 2026 11:31:20 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Use <code>dnnl::lstm_forward</code> primitive on XPU for LSTM inference, replacing the per-timestep fused cell approach. The oneDNN primitive processes the entire sequence internally, eliminating O(T) kernel launches.</p>
<h2>Changes</h2>
<ul>
<li><strong>New file <code>aten/src/ATen/native/mkldnn/xpu/RNN.cpp</code></strong>: Implements <code>lstm_onednn_xpu</code> using <code>dnnl::lstm_forward</code>, registered via <code>REGISTER_XPU_DISPATCH(lstm_mkldnn_stub, ...)</code>. Handles weight layout transformation (PyTorch <code>[4*H, I]</code> → oneDNN <code>ldgoi</code>), weight reorder, and scratchpad management.</li>
<li><strong><code>aten/src/ATen/native/RNN.cpp</code></strong>:
<ul>
<li>Extended <code>use_mkldnn()</code> to return true for XPU in inference mode (float/bf16/fp16)</li>
<li>Added packed sequence unwrap: when <code>batch_sizes</code> are uniform (common in batch=1 inference), reshape to regular 3D tensor and route to oneDNN</li>
<li>Added XPU-specific <code>LSTMCell</code> path with pre-computed input gates as fallback for training</li>
</ul>
</li>
</ul>
<h2>Correctness</h2>
<ul>
<li>oneDNN and PyTorch use identical gate order: <code>i, f, c̃, o</code></li>
<li>oneDNN formula: <code>gates = W·x + U·h + B</code> where <code>B = b_ih + b_hh</code> (we sum before passing)</li>
<li>Verified max diff &lt; 1e-6 vs CPU reference across multiple configurations</li>
</ul>
<h2>Performance (Intel PVC, Kokoro TTS, bidirectional LSTM, hidden=256)</h2>
<table>
<thead>
<tr>
<th>Path</th>
<th>LSTM latency</th>
<th>E2E Kokoro</th>
</tr>
</thead>
<tbody>
<tr>
<td>Original (per-step mm + fused cell)</td>
<td>~34 ms</td>
<td>1.06s</td>
</tr>
<tr>
<td>oneDNN LSTM primitive</td>
<td>~5 ms</td>
<td><strong>0.635s (-40%)</strong></td>
</tr>
</tbody>
</table>
<h2>Dependencies</h2>
<ul>
<li>Depends on <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4538986478" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3770" data-hovercard-type="pull_request" data-hovercard-url="/intel/torch-xpu-ops/pull/3770/hovercard" href="https://github.com/intel/torch-xpu-ops/pull/3770">intel/torch-xpu-ops#3770</a> for the fused cell fallback path (bias fix)</li>
<li>Only affects XPU. CUDA and CPU paths unchanged.</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4541871110" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185531" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185531/hovercard" href="https://github.com/pytorch/pytorch/pull/185531">#185531</a><br>
Approved by: <a href="https://github.com/EikanWang">https://github.com/EikanWang</a>, <a href="https://github.com/guangyey">https://github.com/guangyey</a>, <a href="https://github.com/atalman">https://github.com/atalman</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[NVIDIA cuTile Python Tutorial: Building Tiled GPU Kernels for Vector Addition, Matrix Addition, and Matrix Multiplication in Colab]]></title>
<description><![CDATA[In this tutorial, we implement a hands-on workflow for NVIDIA cuTile Python, a tile-based GPU programming interface for CUDA-style kernels in Python. We prepare a Colab-friendly environment and check GPU, driver, CUDA, and cuTile availability before running kernels. We then build tiled vector add...]]></description>
<link>https://tsecurity.de/de/3583882/ai-nachrichten/nvidia-cutile-python-tutorial-building-tiled-gpu-kernels-for-vector-addition-matrix-addition-and-matrix-multiplication-in-colab/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583882/ai-nachrichten/nvidia-cutile-python-tutorial-building-tiled-gpu-kernels-for-vector-addition-matrix-addition-and-matrix-multiplication-in-colab/</guid>
<pubDate>Tue, 09 Jun 2026 10:48:44 +0200</pubDate>
<category>🔧 AI Nachrichten</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>In this tutorial, we implement a hands-on workflow for NVIDIA cuTile Python, a tile-based GPU programming interface for CUDA-style kernels in Python. We prepare a Colab-friendly environment and check GPU, driver, CUDA, and cuTile availability before running kernels. We then build tiled vector addition, matrix addition, and matrix multiplication, keeping a PyTorch fallback so the notebook stays executable. We validate correctness against PyTorch and benchmark median runtimes at every stage.</p>
<p>The post <a href="https://www.marktechpost.com/2026/06/09/nvidia-cutile-python-tutorial-building-tiled-gpu-kernels-for-vector-addition-matrix-addition-and-matrix-multiplication-in-colab/">NVIDIA cuTile Python Tutorial: Building Tiled GPU Kernels for Vector Addition, Matrix Addition, and Matrix Multiplication in Colab</a> appeared first on <a href="https://www.marktechpost.com/">MarkTechPost</a>.</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/2ee4eba06fb726736fd1ac93f07a3d86af25d9a3: [TMA] Fix NameError in combo kernel when TMA enabled (#185514)]]></title>
<description><![CDATA[Fixes: #185513
Root cause: uniquify_block_sizes() in triton_combo_kernel.py only handled str and DeferredLine objects in the body buffer, but TMA load expressions are stored as DelayReplaceLine objects (a different subclass of DeferredLineBase). These were passed through unmodified, leaving R0_BL...]]></description>
<link>https://tsecurity.de/de/3583865/downloads/trunk2ee4eba06fb726736fd1ac93f07a3d86af25d9a3-tma-fix-nameerror-in-combo-kernel-when-tma-enabled-185514/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583865/downloads/trunk2ee4eba06fb726736fd1ac93f07a3d86af25d9a3-tma-fix-nameerror-in-combo-kernel-when-tma-enabled-185514/</guid>
<pubDate>Tue, 09 Jun 2026 10:46:36 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4540768497" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185513" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/185513/hovercard" href="https://github.com/pytorch/pytorch/issues/185513">#185513</a></p>
<p>Root cause: <code>uniquify_block_sizes()</code> in <code>triton_combo_kernel.py</code> only handled <code>str</code> and <code>DeferredLine</code> objects in the body buffer, but TMA load expressions are stored as <code>DelayReplaceLine</code> objects (a different subclass of <code>DeferredLineBase</code>). These were passed through unmodified, leaving <code>R0_BLOCK</code> unrenamed.</p>
<p>Fix: Changed the isinstance check from <code>DeferredLine</code> to <code>DeferredLineBase</code> and use <code>_new_line()</code> to create the replacement — this handles all deferred line types generically (DeferredLine, DelayReplaceLine, and any future subclasses).</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4540773265" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185514" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185514/hovercard" href="https://github.com/pytorch/pytorch/pull/185514">#185514</a><br>
Approved by: <a href="https://github.com/kshitij12345">https://github.com/kshitij12345</a>, <a href="https://github.com/jansel">https://github.com/jansel</a>, <a href="https://github.com/mlazos">https://github.com/mlazos</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/8ebaef469416972efa51750db0f833ff028221c0: Add XPU dispatch for _fused_adagrad_ operator (#185577)]]></title>
<description><![CDATA[Register XPU dispatch key in native_functions.yaml for both scalar and tensor lr overloads of fused_adagrad
Add 'xpu' to Adagrad's supports_fused_on tuple in common_optimizers.py

Fixes: #185477
Pull Request resolved: #185577
Approved by: https://github.com/guangyey, https://github.com/EikanWang,...]]></description>
<link>https://tsecurity.de/de/3583715/downloads/trunk8ebaef469416972efa51750db0f833ff028221c0-add-xpu-dispatch-for-fusedadagrad-operator-185577/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583715/downloads/trunk8ebaef469416972efa51750db0f833ff028221c0-add-xpu-dispatch-for-fusedadagrad-operator-185577/</guid>
<pubDate>Tue, 09 Jun 2026 09:31:24 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<ul>
<li>Register XPU dispatch key in native_functions.yaml for both scalar and tensor lr overloads of <em>fused_adagrad</em></li>
<li>Add 'xpu' to Adagrad's supports_fused_on tuple in common_optimizers.py</li>
</ul>
<p>Fixes: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4538434732" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185477" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/185477/hovercard" href="https://github.com/pytorch/pytorch/issues/185477">#185477</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4544979916" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185577" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185577/hovercard" href="https://github.com/pytorch/pytorch/pull/185577">#185577</a><br>
Approved by: <a href="https://github.com/guangyey">https://github.com/guangyey</a>, <a href="https://github.com/EikanWang">https://github.com/EikanWang</a>, <a href="https://github.com/janeyx99">https://github.com/janeyx99</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/d5262ac6a97040903a946589ce5db0f1497ec8b6]]></title>
<description><![CDATA[[inductor] BatchLinearLHSFusion: also match torch._C._nn.linear (#186…]]></description>
<link>https://tsecurity.de/de/3583583/downloads/trunkd5262ac6a97040903a946589ce5db0f1497ec8b6/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583583/downloads/trunkd5262ac6a97040903a946589ce5db0f1497ec8b6/</guid>
<pubDate>Tue, 09 Jun 2026 07:46:00 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>[inductor] BatchLinearLHSFusion: also match torch._C._nn.linear (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="186298929" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186/hovercard" href="https://github.com/pytorch/pytorch/issues/186">#186</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/c933d423a9953e8a2ecc2c1de19e3552de3ce959: Cleaning up the 32/64bit offset checks for im2col/col2im (#186109)]]></title>
<description><![CDATA[These ops each pick between 32- and 64-bit index kernels and pack their stride/size args in the chosen width. This factors the shared machinery into OperationUtils.h. A duplicated geometry formula for window position computation in im2col/col2im is moved to im2col_shape_check.h.
This PR is a foll...]]></description>
<link>https://tsecurity.de/de/3583564/downloads/trunkc933d423a9953e8a2ecc2c1de19e3552de3ce959-cleaning-up-the-3264bit-offset-checks-for-im2colcol2im-186109/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583564/downloads/trunkc933d423a9953e8a2ecc2c1de19e3552de3ce959-cleaning-up-the-3264bit-offset-checks-for-im2colcol2im-186109/</guid>
<pubDate>Tue, 09 Jun 2026 07:31:25 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>These ops each pick between 32- and 64-bit index kernels and pack their stride/size args in the chosen width. This factors the shared machinery into OperationUtils.h. A duplicated geometry formula for window position computation in im2col/col2im is moved to im2col_shape_check.h.</p>
<p>This PR is a follow-up to discussion in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4566749414" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185860" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185860/hovercard" href="https://github.com/pytorch/pytorch/pull/185860">#185860</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4582399139" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186109" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186109/hovercard" href="https://github.com/pytorch/pytorch/pull/186109">#186109</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>
<p>Co-authored-by: Aaron Gokaslan <a href="mailto:aaronGokaslan@gmail.com">aaronGokaslan@gmail.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/1100802ab6ce07f48d6ad543fa87318d10bbfb49: Make topk deterministic under deterministic algorithms (#186653)]]></title>
<description><![CDATA[When torch.use_deterministic_algorithms(True) is enabled, make topk resolve tied values with stable index tie-breaking.
CPU topk now compares by value and then index in deterministic mode. CUDA keeps the existing top-k selection path and uses stable sorting for the selected top-k outputs so thres...]]></description>
<link>https://tsecurity.de/de/3583508/downloads/trunk1100802ab6ce07f48d6ad543fa87318d10bbfb49-make-topk-deterministic-under-deterministic-algorithms-186653/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583508/downloads/trunk1100802ab6ce07f48d6ad543fa87318d10bbfb49-make-topk-deterministic-under-deterministic-algorithms-186653/</guid>
<pubDate>Tue, 09 Jun 2026 06:46:51 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>When torch.use_deterministic_algorithms(True) is enabled, make topk resolve tied values with stable index tie-breaking.</p>
<p>CPU topk now compares by value and then index in deterministic mode. CUDA keeps the existing top-k selection path and uses stable sorting for the selected top-k outputs so threshold ties preserve input-index order without falling back to a full sort on NVIDIA. ROCm uses stable sort in deterministic mode because its current top-k gather reserves tie slots through atomics.</p>
<p>This gives deterministic mode a topk behavior analogous to scatter_add: fast default kernels remain available, while deterministic mode uses stable algorithms when requested.</p>
<p>Test Plan:</p>
<ul>
<li>
<p>Rebuilt PyTorch from source with CUDA enabled.</p>
</li>
<li>
<p>python -m py_compile test/test_sort_and_select.py torch/<strong>init</strong>.py torch/_torch_docs.py</p>
</li>
<li>
<p>python test/test_sort_and_select.py -k topk_deterministic</p>
</li>
<li>
<p>python test/test_sort_and_select.py -k topk</p>
</li>
<li>
<p>git diff --check HEAD^</p>
</li>
<li>
<p>lintrunner --config=.lintrunner.toml --skip=PYREFLY aten/src/ATen/native/TopKImpl.h aten/src/ATen/native/cuda/TensorTopK.cpp test/test_sort_and_select.py torch/<strong>init</strong>.py torch/_torch_docs.py</p>
</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4616776557" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186653" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186653/hovercard" href="https://github.com/pytorch/pytorch/pull/186653">#186653</a><br>
Approved by: <a href="https://github.com/drisspg">https://github.com/drisspg</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/d288c4204390786fb6ecadcba8a45c395b91da6e: Add oneDNN LSTM primitive support for XPU inference (#185531)]]></title>
<description><![CDATA[Summary
Use dnnl::lstm_forward primitive on XPU for LSTM inference, replacing the per-timestep fused cell approach. The oneDNN primitive processes the entire sequence internally, eliminating O(T) kernel launches.
Changes

New file aten/src/ATen/native/mkldnn/xpu/RNN.cpp: Implements lstm_onednn_xp...]]></description>
<link>https://tsecurity.de/de/3583507/downloads/trunkd288c4204390786fb6ecadcba8a45c395b91da6e-add-onednn-lstm-primitive-support-for-xpu-inference-185531/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583507/downloads/trunkd288c4204390786fb6ecadcba8a45c395b91da6e-add-onednn-lstm-primitive-support-for-xpu-inference-185531/</guid>
<pubDate>Tue, 09 Jun 2026 06:46:50 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Use <code>dnnl::lstm_forward</code> primitive on XPU for LSTM inference, replacing the per-timestep fused cell approach. The oneDNN primitive processes the entire sequence internally, eliminating O(T) kernel launches.</p>
<h2>Changes</h2>
<ul>
<li><strong>New file <code>aten/src/ATen/native/mkldnn/xpu/RNN.cpp</code></strong>: Implements <code>lstm_onednn_xpu</code> using <code>dnnl::lstm_forward</code>, registered via <code>REGISTER_XPU_DISPATCH(lstm_mkldnn_stub, ...)</code>. Handles weight layout transformation (PyTorch <code>[4*H, I]</code> → oneDNN <code>ldgoi</code>), weight reorder, and scratchpad management.</li>
<li><strong><code>aten/src/ATen/native/RNN.cpp</code></strong>:
<ul>
<li>Extended <code>use_mkldnn()</code> to return true for XPU in inference mode (float/bf16/fp16)</li>
<li>Added packed sequence unwrap: when <code>batch_sizes</code> are uniform (common in batch=1 inference), reshape to regular 3D tensor and route to oneDNN</li>
<li>Added XPU-specific <code>LSTMCell</code> path with pre-computed input gates as fallback for training</li>
</ul>
</li>
</ul>
<h2>Correctness</h2>
<ul>
<li>oneDNN and PyTorch use identical gate order: <code>i, f, c̃, o</code></li>
<li>oneDNN formula: <code>gates = W·x + U·h + B</code> where <code>B = b_ih + b_hh</code> (we sum before passing)</li>
<li>Verified max diff &lt; 1e-6 vs CPU reference across multiple configurations</li>
</ul>
<h2>Performance (Intel PVC, Kokoro TTS, bidirectional LSTM, hidden=256)</h2>
<table>
<thead>
<tr>
<th>Path</th>
<th>LSTM latency</th>
<th>E2E Kokoro</th>
</tr>
</thead>
<tbody>
<tr>
<td>Original (per-step mm + fused cell)</td>
<td>~34 ms</td>
<td>1.06s</td>
</tr>
<tr>
<td>oneDNN LSTM primitive</td>
<td>~5 ms</td>
<td><strong>0.635s (-40%)</strong></td>
</tr>
</tbody>
</table>
<h2>Dependencies</h2>
<ul>
<li>Depends on <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4538986478" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3770" data-hovercard-type="pull_request" data-hovercard-url="/intel/torch-xpu-ops/pull/3770/hovercard" href="https://github.com/intel/torch-xpu-ops/pull/3770">intel/torch-xpu-ops#3770</a> for the fused cell fallback path (bias fix)</li>
<li>Only affects XPU. CUDA and CPU paths unchanged.</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4541871110" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185531" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185531/hovercard" href="https://github.com/pytorch/pytorch/pull/185531">#185531</a><br>
Approved by: <a href="https://github.com/EikanWang">https://github.com/EikanWang</a>, <a href="https://github.com/guangyey">https://github.com/guangyey</a>, <a href="https://github.com/atalman">https://github.com/atalman</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/ebf43868c89e7f0cfb2cce8764f993b7f010ee57: [dynamo, 3.15] Update IMPORT_NAME generation (#186402)]]></title>
<description><![CDATA[In 3.15 the lower 2 bits of the IMPORT_NAME oparg are used to indicate lazy vs eager imports.  We need to shift the
index left (this was causing SyntaxError: 'lazy import' is only allowed at module level previously in generated
IMPORT_NAME instructions) and set 2nd lowest bit to indicate eager im...]]></description>
<link>https://tsecurity.de/de/3583402/downloads/trunkebf43868c89e7f0cfb2cce8764f993b7f010ee57-dynamo-315-update-importname-generation-186402/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583402/downloads/trunkebf43868c89e7f0cfb2cce8764f993b7f010ee57-dynamo-315-update-importname-generation-186402/</guid>
<pubDate>Tue, 09 Jun 2026 05:16:27 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>In 3.15 the lower 2 bits of the IMPORT_NAME oparg are used to indicate lazy vs eager imports.  We need to shift the<br>
index left (this was causing <code>SyntaxError: 'lazy import' is only allowed at module level</code> previously in generated<br>
IMPORT_NAME instructions) and set 2nd lowest bit to indicate eager imports.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4599600869" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186402" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186402/hovercard" href="https://github.com/pytorch/pytorch/pull/186402">#186402</a><br>
Approved by: <a href="https://github.com/guilhermeleobas">https://github.com/guilhermeleobas</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/9922478dffa606fd798cc2346a227d4867e8b6ee: UCC/test: Undo migration to reduce_scatter_single (#186666)]]></title>
<description><![CDATA[_reduce_scatter_base was changed to reduce_scatter_single in #186135. ProcessGroupUCC inherits from c10::Backend which does not have a reduce_scatter_single member function which causes this test to fail. Other ProcessGroup* inherit from c10::Backend as well but don't show a test failure because ...]]></description>
<link>https://tsecurity.de/de/3583374/downloads/trunk9922478dffa606fd798cc2346a227d4867e8b6ee-ucctest-undo-migration-to-reducescattersingle-186666/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583374/downloads/trunk9922478dffa606fd798cc2346a227d4867e8b6ee-ucctest-undo-migration-to-reducescattersingle-186666/</guid>
<pubDate>Tue, 09 Jun 2026 04:46:20 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>_reduce_scatter_base</code> was changed to <code>reduce_scatter_single</code> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4583758676" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186135" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186135/hovercard" href="https://github.com/pytorch/pytorch/pull/186135">#186135</a>. <code>ProcessGroupUCC</code> inherits from <code>c10::Backend</code> which does not have a <code>reduce_scatter_single</code> member function which causes this test to fail. Other <code>ProcessGroup*</code> inherit from <code>c10::Backend</code> as well but don't show a test failure because all the other tests call the distributed APIs via a <code>ProcessGroup</code> object unlike the UCC test which uses the <code>Backend</code> object. Confirmed with <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/d4l3k/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/d4l3k">@d4l3k</a> that reverting is okay for now.</p>
<p>This failure made it into the repo because the CI does not have support for the UCC backend.</p>
<p>Fixes: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4617070743" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186664" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186664/hovercard" href="https://github.com/pytorch/pytorch/issues/186664">#186664</a><br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4617165908" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186666" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186666/hovercard" href="https://github.com/pytorch/pytorch/pull/186666">#186666</a><br>
Approved by: <a href="https://github.com/d4l3k">https://github.com/d4l3k</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/torchtitan/186752: [Inductor] Host-side TMA descriptors for Blackwell mm templates]]></title>
<description><![CDATA[Add a {{tma_descriptor()}} Jinja hook that lets mm templates declare TMA descriptors declaratively. When config.triton.enable_host_side_tma is True, the hook registers the input's pointer arg in host_tma_descriptor_args so the launcher (from PR #185825) replaces it with a TensorDescriptor at laun...]]></description>
<link>https://tsecurity.de/de/3583297/downloads/ciflowtorchtitan186752-inductor-host-side-tma-descriptors-for-blackwell-mm-templates/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583297/downloads/ciflowtorchtitan186752-inductor-host-side-tma-descriptors-for-blackwell-mm-templates/</guid>
<pubDate>Tue, 09 Jun 2026 03:46:27 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Add a <code>{{tma_descriptor()}}</code> Jinja hook that lets mm templates declare TMA descriptors declaratively. When <code>config.triton.enable_host_side_tma</code> is True, the hook registers the input's pointer arg in <code>host_tma_descriptor_args</code> so the launcher (from PR <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4564769337" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185825" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185825/hovercard" href="https://github.com/pytorch/pytorch/pull/185825">#185825</a>) replaces it with a <code>TensorDescriptor</code> at launch time. When False, the hook emits device-side <code>tl.make_tensor_descriptor()</code> code.</p>
<p>This builds on the pointwise host-side TMA infrastructure from <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4564769337" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185825" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185825/hovercard" href="https://github.com/pytorch/pytorch/pull/185825">#185825</a> and extends it to the template codegen path in <code>select_algorithm.py</code>. The Blackwell warp-specialized persistent TMA template (<code>triton_blackwell_ws_persistent_tma_mm.py.jinja</code>) is updated to use the hook, eliminating the inline <code>tl.make_tensor_descriptor()</code> calls and supporting both host-side and device-side paths from a single template.</p>
<p>Key changes:</p>
<ul>
<li><code>config.py</code>: <code>enable_host_side_tma</code> flag</li>
<li><code>select_algorithm.py</code>: <code>tma_descriptor()</code> hook with <code>dim_order</code> support for transposed inputs, signature upgrade to <code>tensordesc&lt;&gt;</code> types, <code>host_tma_descriptor_args</code> export to <code>inductor_meta</code></li>
<li><code>template_heuristics/triton.py</code>: <code>HOST_SIDE_TMA</code> config in Blackwell mixin</li>
<li><code>mm.py</code>: renamed template to <code>blackwell_ws_persistent_tma</code> (agnostic to host/device)</li>
<li><code>static_triton_launcher.py</code>: <code>tensordesc&lt;&gt;</code> type fallback to dynamic launcher</li>
<li><code>codegen/triton.py</code>, <code>triton_heuristics.py</code>: support dict format in <code>host_tma_descriptor_args</code> for <code>dim_order</code></li>
</ul>
<p>Authored with Claude.</p>
<p>Test Plan:</p>
<p>Tested with tritonbench on B200 with <code>TRITON_USE_META_WS=1</code>:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="CUDA_VISIBLE_DEVICES=2 TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 TRITON_USE_META_WS=1 TRITON_PRINT_AUTOTUNING=1 python run.py --op gemm --only aten_matmul,pt2_matmul_maxautotune_device_side_tma_only,pt2_matmul_maxautotune_host_side_tma_only --shapes 4096x4096x4096 --metrics tflops,latency,accuracy"><pre class="notranslate"><code>CUDA_VISIBLE_DEVICES=2 TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 TRITON_USE_META_WS=1 TRITON_PRINT_AUTOTUNING=1 python run.py --op gemm --only aten_matmul,pt2_matmul_maxautotune_device_side_tma_only,pt2_matmul_maxautotune_host_side_tma_only --shapes 4096x4096x4096 --metrics tflops,latency,accuracy
</code></pre></div>
<p>Results (4096x4096x4096 bf16, B200):</p>
<ul>
<li>cuBLAS: 0.107ms</li>
<li>device-side TMA: 0.193ms, accuracy=1</li>
<li>host-side TMA: 0.192ms, accuracy=1</li>
</ul>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/trunk/186653: Make topk deterministic under deterministic algorithms]]></title>
<description><![CDATA[When torch.use_deterministic_algorithms(True) is enabled, make topk resolve tied values with stable index tie-breaking.
CPU topk now compares by value and then index in deterministic mode. CUDA keeps the existing top-k selection path and uses stable sorting for the selected top-k outputs so thres...]]></description>
<link>https://tsecurity.de/de/3583265/downloads/ciflowtrunk186653-make-topk-deterministic-under-deterministic-algorithms/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583265/downloads/ciflowtrunk186653-make-topk-deterministic-under-deterministic-algorithms/</guid>
<pubDate>Tue, 09 Jun 2026 03:00:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>When torch.use_deterministic_algorithms(True) is enabled, make topk resolve tied values with stable index tie-breaking.</p>
<p>CPU topk now compares by value and then index in deterministic mode. CUDA keeps the existing top-k selection path and uses stable sorting for the selected top-k outputs so threshold ties preserve input-index order without falling back to a full sort on NVIDIA. ROCm uses stable sort in deterministic mode because its current top-k gather reserves tie slots through atomics.</p>
<p>This gives deterministic mode a topk behavior analogous to scatter_add: fast default kernels remain available, while deterministic mode uses stable algorithms when requested.</p>
<p>Test Plan:</p>
<ul>
<li>
<p>Rebuilt PyTorch from source with CUDA enabled.</p>
</li>
<li>
<p>python -m py_compile test/test_sort_and_select.py torch/<strong>init</strong>.py torch/_torch_docs.py</p>
</li>
<li>
<p>python test/test_sort_and_select.py -k topk_deterministic</p>
</li>
<li>
<p>python test/test_sort_and_select.py -k topk</p>
</li>
<li>
<p>git diff --check HEAD^</p>
</li>
<li>
<p>lintrunner --config=.lintrunner.toml --skip=PYREFLY aten/src/ATen/native/TopKImpl.h aten/src/ATen/native/cuda/TensorTopK.cpp test/test_sort_and_select.py torch/<strong>init</strong>.py torch/_torch_docs.py</p>
</li>
</ul>
<p>ghstack-source-id: c9b62071c22c09c660e34d957311c6d0234300c8</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4616776557" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186653" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186653/hovercard" href="https://github.com/pytorch/pytorch/pull/186653">#186653</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/torchtitan/186653: Make topk deterministic under deterministic algorithms]]></title>
<description><![CDATA[When torch.use_deterministic_algorithms(True) is enabled, make topk resolve tied values with stable index tie-breaking.
CPU topk now compares by value and then index in deterministic mode. CUDA keeps the existing top-k selection path and uses stable sorting for the selected top-k outputs so thres...]]></description>
<link>https://tsecurity.de/de/3583217/downloads/ciflowtorchtitan186653-make-topk-deterministic-under-deterministic-algorithms/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583217/downloads/ciflowtorchtitan186653-make-topk-deterministic-under-deterministic-algorithms/</guid>
<pubDate>Tue, 09 Jun 2026 02:31:30 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>When torch.use_deterministic_algorithms(True) is enabled, make topk resolve tied values with stable index tie-breaking.</p>
<p>CPU topk now compares by value and then index in deterministic mode. CUDA keeps the existing top-k selection path and uses stable sorting for the selected top-k outputs so threshold ties preserve input-index order without falling back to a full sort on NVIDIA. ROCm uses stable sort in deterministic mode because its current top-k gather reserves tie slots through atomics.</p>
<p>This gives deterministic mode a topk behavior analogous to scatter_add: fast default kernels remain available, while deterministic mode uses stable algorithms when requested.</p>
<p>Test Plan:</p>
<ul>
<li>
<p>Rebuilt PyTorch from source with CUDA enabled.</p>
</li>
<li>
<p>python -m py_compile test/test_sort_and_select.py torch/<strong>init</strong>.py torch/_torch_docs.py</p>
</li>
<li>
<p>python test/test_sort_and_select.py -k topk_deterministic</p>
</li>
<li>
<p>python test/test_sort_and_select.py -k topk</p>
</li>
<li>
<p>git diff --check HEAD^</p>
</li>
<li>
<p>lintrunner --config=.lintrunner.toml --skip=PYREFLY aten/src/ATen/native/TopKImpl.h aten/src/ATen/native/cuda/TensorTopK.cpp test/test_sort_and_select.py torch/<strong>init</strong>.py torch/_torch_docs.py</p>
</li>
</ul>
<p>ghstack-source-id: c9b62071c22c09c660e34d957311c6d0234300c8</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4616776557" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186653" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186653/hovercard" href="https://github.com/pytorch/pytorch/pull/186653">#186653</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780959143: Revert "Fix custom op autograd gradient count errors (#186226)"]]></title>
<description><![CDATA[This reverts commit 211af95.
Reverted #186226 on behalf of https://github.com/atalman due to reverted internally (comment)]]></description>
<link>https://tsecurity.de/de/3583149/downloads/viablestrict1780959143-revert-fix-custom-op-autograd-gradient-count-errors-186226/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583149/downloads/viablestrict1780959143-revert-fix-custom-op-autograd-gradient-count-errors-186226/</guid>
<pubDate>Tue, 09 Jun 2026 01:16:29 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/211af9577bfc98328aae3d1bd00ad49bc1253075/hovercard" href="https://github.com/pytorch/pytorch/commit/211af9577bfc98328aae3d1bd00ad49bc1253075"><tt>211af95</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4588171410" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186226" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186226/hovercard" href="https://github.com/pytorch/pytorch/pull/186226">#186226</a> on behalf of <a href="https://github.com/atalman">https://github.com/atalman</a> due to reverted internally (<a href="https://github.com/pytorch/pytorch/pull/186226#issuecomment-4652201344" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186226/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780953904: Skip Dynamo tracing for pad_packed_sequence (#185143)]]></title>
<description><![CDATA[Dynamo already treats pack_padded_sequence as unsupported and graph-breaks around it, but pad_packed_sequence was not in the same trace-rule map. In default graph-break mode this lets the packed/RNN segment run eagerly and then resumes tracing into pad_packed_sequence, where fake tensor evaluatio...]]></description>
<link>https://tsecurity.de/de/3583004/downloads/viablestrict1780953904-skip-dynamo-tracing-for-padpackedsequence-185143/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3583004/downloads/viablestrict1780953904-skip-dynamo-tracing-for-padpackedsequence-185143/</guid>
<pubDate>Mon, 08 Jun 2026 23:46:48 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Dynamo already treats pack_padded_sequence as unsupported and graph-breaks around it, but pad_packed_sequence was not in the same trace-rule map. In default graph-break mode this lets the packed/RNN segment run eagerly and then resumes tracing into pad_packed_sequence, where fake tensor evaluation reaches _VF._pad_packed_sequence and crashes with "data is not allocated".</p>
<p>Add pad_packed_sequence to the Dynamo skip list so packed sequence unpacking stays in the eager unsupported segment. This preserves fullgraph behavior as an explicit Unsupported instead of adding partial PackedSequence tracing support, which would require broader fake/meta coverage for the packed-sequence operators.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3393287445" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/162374" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/162374/hovercard" href="https://github.com/pytorch/pytorch/issues/162374">#162374</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>python test/dynamo/test_repros.py -k "pad_packed_sequence"</li>
<li>python test/dynamo/test_trace_rules.py</li>
<li>lintrunner -a</li>
<li>exact issue script with default torch.compile(model)</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4517893972" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185143" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185143/hovercard" href="https://github.com/pytorch/pytorch/pull/185143">#185143</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780951453: [dynamo] Box resume frame values when del can run (#185561)]]></title>
<description><![CDATA[Dynamo resume functions pass restored stack values and locals as normal
positional arguments. CPython keeps those argument references alive for the
whole resume call, so a DELETE_FAST in the resumed bytecode can remove the
local name without actually releasing the tensor. That makes compiled grap...]]></description>
<link>https://tsecurity.de/de/3582856/downloads/viablestrict1780951453-dynamo-box-resume-frame-values-when-del-can-run-185561/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3582856/downloads/viablestrict1780951453-dynamo-box-resume-frame-values-when-del-can-run-185561/</guid>
<pubDate>Mon, 08 Jun 2026 22:46:24 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Dynamo resume functions pass restored stack values and locals as normal<br>
positional arguments. CPython keeps those argument references alive for the<br>
whole resume call, so a <code>DELETE_FAST</code> in the resumed bytecode can remove the<br>
local name without actually releasing the tensor. That makes compiled graph<br>
breaks hold deleted intermediates longer than eager execution.</p>
<p>Detect resume functions that may execute <code>DELETE_FAST</code> after the resume target<br>
and use a boxed frame-values argument for those calls. The resume prologue now<br>
loads each saved stack/local value from the list, stores locals where needed,<br>
and clears each list slot immediately. Resume functions without a reachable<br>
<code>DELETE_FAST</code> keep the previous positional calling convention.</p>
<p>Nested resume calls need to preserve the child's calling convention, so nested<br>
boxed callees receive the frame-values list as one argument while non-boxed<br>
callees keep the existing list extension behavior.</p>
<p>This fixes the actionable resume-frame/intermediate lifetime issue and the<br>
issue-shaped list-cleared input path. Bare direct temporary inputs remain owned<br>
by the outer <code>torch.compile</code> wrapper's <code>*args</code> tuple until the compiled call<br>
returns; a pure Python <code>def wrapper(*args): return fn(*args)</code> has the same<br>
lifetime behavior, so this patch does not try to make <code>del x</code> release such<br>
bare inputs before return.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3068590996" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/153701" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/153701/hovercard" href="https://github.com/pytorch/pytorch/issues/153701">#153701</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>python test/dynamo/test_subgraphs.py SubGraphTests.test_nested_resume_del_releases_tensor SubGraphTests.test_resume_del_releases_tensor SubGraphTests.test_issue_shape_list_clear_and_intermediate_del_release_tensors SubGraphTests.test_del_compiled_only_local_before_graph_break</li>
<li>python test/dynamo/test_subgraphs.py</li>
<li>python test/dynamo/test_nested_graph_breaks.py -k test_step_graph_break_frame_values_not_corrupted</li>
<li>python test/dynamo/test_repros.py ReproTests.test_weakref_reconstruct ReproTests.test_weakref_del ReproTests.test_weakref_callback</li>
<li>lintrunner -a</li>
</ul>
<p>Benchmark Results:<br>
Tiny CPU graph-break runtime microbenchmark with <code>DELETE_FAST</code>, backend="eager",<br>
7 samples of 3000 cached calls each:</p>
<ul>
<li>Before: samples_us [26.349, 26.217, 26.076, 26.089, 26.544, 26.578, 26.099], median 26.217 us</li>
<li>After: samples_us [28.897, 29.262, 30.618, 29.686, 31.57, 32.353, 33.83], median 30.618 us (+16.8%)</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4543708479" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185561" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185561/hovercard" href="https://github.com/pytorch/pytorch/pull/185561">#185561</a><br>
Approved by: <a href="https://github.com/IvanKobzarev">https://github.com/IvanKobzarev</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780943741: Add highlevel C++ torch::stable::Generator (#186423)]]></title>
<description><![CDATA[Written with Claude and reviewed by me. We handle Generators the way we handle Tensors.
@desertfire had added the shims for these before but there was no way from C++ to retrieve an at::Generator in an ABI stable way (it was always passed as null). This PR changes that by letting generators be pa...]]></description>
<link>https://tsecurity.de/de/3582528/downloads/viablestrict1780943741-add-highlevel-c-torchstablegenerator-186423/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3582528/downloads/viablestrict1780943741-add-highlevel-c-torchstablegenerator-186423/</guid>
<pubDate>Mon, 08 Jun 2026 20:46:28 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Written with Claude and reviewed by me. We handle Generators the way we handle Tensors.</p>
<p><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/desertfire/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/desertfire">@desertfire</a> had added the shims for these before but there was no way from C++ to retrieve an at::Generator in an ABI stable way (it was always passed as null). This PR changes that by letting generators be passed through the dispatcher from Python to a stable kernel + adds memory management similar to Tensor.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4600389438" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186423" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186423/hovercard" href="https://github.com/pytorch/pytorch/pull/186423">#186423</a><br>
Approved by: <a href="https://github.com/albanD">https://github.com/albanD</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780940612: [inductor] Fix combo kernel benchmark device argument (#184868)]]></title>
<description><![CDATA[Summary:
Quote the generated combo-kernel benchmark device argument so the benchmark script passes a Python string such as 'cuda' instead of referencing an undefined variable.
Review:
@desertfire @karthickai, would you mind taking a look when you have a chance? Your guidance on the Inductor combo...]]></description>
<link>https://tsecurity.de/de/3582316/downloads/viablestrict1780940612-inductor-fix-combo-kernel-benchmark-device-argument-184868/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3582316/downloads/viablestrict1780940612-inductor-fix-combo-kernel-benchmark-device-argument-184868/</guid>
<pubDate>Mon, 08 Jun 2026 19:46:55 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:<br>
Quote the generated combo-kernel benchmark <code>device</code> argument so the benchmark script passes a Python string such as <code>'cuda'</code> instead of referencing an undefined variable.</p>
<p>Review:<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/desertfire/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/desertfire">@desertfire</a> <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/karthickai/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/karthickai">@karthickai</a>, would you mind taking a look when you have a chance? Your guidance on the Inductor combo-kernel benchmark generation here would be greatly appreciated.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4501131349" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184868" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184868/hovercard" href="https://github.com/pytorch/pytorch/pull/184868">#184868</a><br>
Approved by: <a href="https://github.com/karthickai">https://github.com/karthickai</a>, <a href="https://github.com/mlazos">https://github.com/mlazos</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/5626fdca7af8043cf2095aaec6a12e7853031131: Fix FakeTensor embedding with meta indices (#185060)]]></title>
<description><![CDATA[FakeTensor currently lets aten.embedding.default fall through the generic
meta-kernel wrapping path. The embedding meta kernel can infer the correct
metadata for mixed weight/index devices, but the generic FakeTensor device
propagation runs afterward and rejects a valid CPU weight plus meta indic...]]></description>
<link>https://tsecurity.de/de/3582150/downloads/trunk5626fdca7af8043cf2095aaec6a12e7853031131-fix-faketensor-embedding-with-meta-indices-185060/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3582150/downloads/trunk5626fdca7af8043cf2095aaec6a12e7853031131-fix-faketensor-embedding-with-meta-indices-185060/</guid>
<pubDate>Mon, 08 Jun 2026 18:46:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>FakeTensor currently lets aten.embedding.default fall through the generic<br>
meta-kernel wrapping path. The embedding meta kernel can infer the correct<br>
metadata for mixed weight/index devices, but the generic FakeTensor device<br>
propagation runs afterward and rejects a valid CPU weight plus meta indices<br>
case as an unhandled mixed-device operation.</p>
<p>Register a narrow FakeTensor implementation for aten.embedding.default,<br>
matching the existing aten._embedding_bag.default pattern. The implementation<br>
calls the existing embedding meta registration under FakeTensorMode, so the<br>
output metadata follows embedding semantics: indices provide the shape, while<br>
weight provides dtype and device. This avoids broad changes to generic device<br>
propagation, such as the stale PR's aten.where special case, because embedding<br>
has an op-specific output-device rule.</p>
<p>Tests cover direct FakeTensor execution for CPU weight with meta indices and<br>
the inverse meta-weight case, plus a torch.compile backend="eager" regression<br>
matching the issue.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3571348336" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/166644" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/166644/hovercard" href="https://github.com/pytorch/pytorch/issues/166644">#166644</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>python test/test_fake_tensor.py -k embedding_meta_indices</li>
<li>git diff --cached --check</li>
<li>lintrunner -a<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4511105934" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185060" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185060/hovercard" href="https://github.com/pytorch/pytorch/pull/185060">#185060</a><br>
Approved by: <a href="https://github.com/ezyang">https://github.com/ezyang</a></li>
</ul>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/536abdd19620b1afa53085c8b1a92219e235f925: Skip Dynamo tracing for pad_packed_sequence (#185143)]]></title>
<description><![CDATA[Dynamo already treats pack_padded_sequence as unsupported and graph-breaks around it, but pad_packed_sequence was not in the same trace-rule map. In default graph-break mode this lets the packed/RNN segment run eagerly and then resumes tracing into pad_packed_sequence, where fake tensor evaluatio...]]></description>
<link>https://tsecurity.de/de/3582149/downloads/trunk536abdd19620b1afa53085c8b1a92219e235f925-skip-dynamo-tracing-for-padpackedsequence-185143/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3582149/downloads/trunk536abdd19620b1afa53085c8b1a92219e235f925-skip-dynamo-tracing-for-padpackedsequence-185143/</guid>
<pubDate>Mon, 08 Jun 2026 18:46:34 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Dynamo already treats pack_padded_sequence as unsupported and graph-breaks around it, but pad_packed_sequence was not in the same trace-rule map. In default graph-break mode this lets the packed/RNN segment run eagerly and then resumes tracing into pad_packed_sequence, where fake tensor evaluation reaches _VF._pad_packed_sequence and crashes with "data is not allocated".</p>
<p>Add pad_packed_sequence to the Dynamo skip list so packed sequence unpacking stays in the eager unsupported segment. This preserves fullgraph behavior as an explicit Unsupported instead of adding partial PackedSequence tracing support, which would require broader fake/meta coverage for the packed-sequence operators.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3393287445" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/162374" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/162374/hovercard" href="https://github.com/pytorch/pytorch/issues/162374">#162374</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>python test/dynamo/test_repros.py -k "pad_packed_sequence"</li>
<li>python test/dynamo/test_trace_rules.py</li>
<li>lintrunner -a</li>
<li>exact issue script with default torch.compile(model)</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4517893972" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185143" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185143/hovercard" href="https://github.com/pytorch/pytorch/pull/185143">#185143</a><br>
Approved by: <a href="https://github.com/zou3519">https://github.com/zou3519</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/2297675f716888f44713606fcc3622415babaac1]]></title>
<description><![CDATA[Revert "inductor: skip TF32 padding for non-contiguous fp32 mm (#1840…]]></description>
<link>https://tsecurity.de/de/3582064/downloads/trunk2297675f716888f44713606fcc3622415babaac1/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3582064/downloads/trunk2297675f716888f44713606fcc3622415babaac1/</guid>
<pubDate>Mon, 08 Jun 2026 18:16:53 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Revert "inductor: skip TF32 padding for non-contiguous fp32 mm (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="236793455" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/1840" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/1840/hovercard" href="https://github.com/pytorch/pytorch/issues/1840">#1840</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780930769: Fix typos in comments and docstrings (#186234)]]></title>
<description><![CDATA[Correct a batch of spelling mistakes in source comments, docstrings, and
log messages across torch. These are documentation-only changes with no
effect on behavior.
Authored with Claude (typo_terminator2).
Pull Request resolved: #186234
Approved by: https://github.com/aorenste, https://github.com...]]></description>
<link>https://tsecurity.de/de/3581871/downloads/viablestrict1780930769-fix-typos-in-comments-and-docstrings-186234/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3581871/downloads/viablestrict1780930769-fix-typos-in-comments-and-docstrings-186234/</guid>
<pubDate>Mon, 08 Jun 2026 17:16:43 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Correct a batch of spelling mistakes in source comments, docstrings, and<br>
log messages across torch. These are documentation-only changes with no<br>
effect on behavior.</p>
<p>Authored with Claude (typo_terminator2).</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4589255212" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186234" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186234/hovercard" href="https://github.com/pytorch/pytorch/pull/186234">#186234</a><br>
Approved by: <a href="https://github.com/aorenste">https://github.com/aorenste</a>, <a href="https://github.com/zou3519">https://github.com/zou3519</a>, <a href="https://github.com/cyyever">https://github.com/cyyever</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/84dc008f1b15477177143dab753cb1f80d527d36: Revert "[inductor] Fix debug sync in GPU cpp wrapper (#184217)"]]></title>
<description><![CDATA[This reverts commit 96a92ea.
Reverted #184217 on behalf of https://github.com/atalman due to new tests test_debug_sync_graph and test_debug_sync_kernel are failing internally (comment)]]></description>
<link>https://tsecurity.de/de/3581720/downloads/trunk84dc008f1b15477177143dab753cb1f80d527d36-revert-inductor-fix-debug-sync-in-gpu-cpp-wrapper-184217/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3581720/downloads/trunk84dc008f1b15477177143dab753cb1f80d527d36-revert-inductor-fix-debug-sync-in-gpu-cpp-wrapper-184217/</guid>
<pubDate>Mon, 08 Jun 2026 16:16:32 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This reverts commit <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/pytorch/pytorch/commit/96a92eac680c796da1e946f24df1281b75b294b8/hovercard" href="https://github.com/pytorch/pytorch/commit/96a92eac680c796da1e946f24df1281b75b294b8"><tt>96a92ea</tt></a>.</p>
<p>Reverted <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4470161892" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184217" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184217/hovercard" href="https://github.com/pytorch/pytorch/pull/184217">#184217</a> on behalf of <a href="https://github.com/atalman">https://github.com/atalman</a> due to new tests test_debug_sync_graph and test_debug_sync_kernel are failing internally (<a href="https://github.com/pytorch/pytorch/pull/184217#issuecomment-4649731922" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184217/hovercard">comment</a>)</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/9d12814e74bea2317c097218cca8c01b9cc16fe2: [inductor] Fix combo kernel benchmark device argument (#184868)]]></title>
<description><![CDATA[Summary:
Quote the generated combo-kernel benchmark device argument so the benchmark script passes a Python string such as 'cuda' instead of referencing an undefined variable.
Review:
@desertfire @karthickai, would you mind taking a look when you have a chance? Your guidance on the Inductor combo...]]></description>
<link>https://tsecurity.de/de/3581506/downloads/trunk9d12814e74bea2317c097218cca8c01b9cc16fe2-inductor-fix-combo-kernel-benchmark-device-argument-184868/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3581506/downloads/trunk9d12814e74bea2317c097218cca8c01b9cc16fe2-inductor-fix-combo-kernel-benchmark-device-argument-184868/</guid>
<pubDate>Mon, 08 Jun 2026 15:01:26 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Summary:<br>
Quote the generated combo-kernel benchmark <code>device</code> argument so the benchmark script passes a Python string such as <code>'cuda'</code> instead of referencing an undefined variable.</p>
<p>Review:<br>
<a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/desertfire/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/desertfire">@desertfire</a> <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/karthickai/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/karthickai">@karthickai</a>, would you mind taking a look when you have a chance? Your guidance on the Inductor combo-kernel benchmark generation here would be greatly appreciated.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4501131349" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184868" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184868/hovercard" href="https://github.com/pytorch/pytorch/pull/184868">#184868</a><br>
Approved by: <a href="https://github.com/karthickai">https://github.com/karthickai</a>, <a href="https://github.com/mlazos">https://github.com/mlazos</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780912404: [XPU][Test] Migrate 6 UT test suites for Intel GPU (#174370)]]></title>
<description><![CDATA[Description
Fixes #114850, we will port dynamo, fsdp tests to Intel GPU
We could enable Intel GPU with following methods and try the best to keep the original code styles:
Changes

Get device type with from accelerator and get_devtype helper method
Replace the requires cuda statement with require...]]></description>
<link>https://tsecurity.de/de/3581050/downloads/viablestrict1780912404-xputest-migrate-6-ut-test-suites-for-intel-gpu-174370/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3581050/downloads/viablestrict1780912404-xputest-migrate-6-ut-test-suites-for-intel-gpu-174370/</guid>
<pubDate>Mon, 08 Jun 2026 12:16:46 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Description</h1>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2018027027" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/114850" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/114850/hovercard" href="https://github.com/pytorch/pytorch/issues/114850">#114850</a>, we will port dynamo, fsdp tests to Intel GPU<br>
We could enable Intel GPU with following methods and try the best to keep the original code styles:</p>
<h1>Changes</h1>
<ol>
<li>Get device type with from accelerator and get_devtype helper method</li>
<li>Replace the requires cuda statement with requires_gpu.</li>
<li>Replace the cuda() with to(device_type).</li>
<li>Add is_xpu check into the test logic.</li>
</ol>
<h1>Notify</h1>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3900673626" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/174370" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/174370/hovercard" href="https://github.com/pytorch/pytorch/pull/174370">#174370</a><br>
Approved by: <a href="https://github.com/jansel">https://github.com/jansel</a>, <a href="https://github.com/guangyey">https://github.com/guangyey</a></p>
<p>Co-authored-by: Benedykt Bela <a href="mailto:benedykt.bela@intel.com">benedykt.bela@intel.com</a><br>
Co-authored-by: Wang, Chuanqi <a href="mailto:chuanqi.wang@intel.com">chuanqi.wang@intel.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/25b4b9d3abccb22d469499fe09ddb0a7f1a67147: [xpu] Update xpu.txt pin for hardtanh_backward XPU opmath fix (#186327)]]></title>
<description><![CDATA[Summary
Update third_party/xpu.txt to pick up the kernel fix from intel/torch-xpu-ops#3873
(commit a227858) which corrects HardtanhBackwardFunctor to use opmath_t
instead of scalar_t for boundary value storage.
Root Cause
The XPU HardtanhBackwardFunctor stored min_val_/max_val_ as scalar_t (bf16)...]]></description>
<link>https://tsecurity.de/de/3580746/downloads/trunk25b4b9d3abccb22d469499fe09ddb0a7f1a67147-xpu-update-xputxt-pin-for-hardtanhbackward-xpu-opmath-fix-186327/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580746/downloads/trunk25b4b9d3abccb22d469499fe09ddb0a7f1a67147-xpu-update-xputxt-pin-for-hardtanhbackward-xpu-opmath-fix-186327/</guid>
<pubDate>Mon, 08 Jun 2026 10:06:41 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p>Update <code>third_party/xpu.txt</code> to pick up the kernel fix from <strong><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594526940" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3873" data-hovercard-type="pull_request" data-hovercard-url="/intel/torch-xpu-ops/pull/3873/hovercard" href="https://github.com/intel/torch-xpu-ops/pull/3873">intel/torch-xpu-ops#3873</a></strong><br>
(commit <a href="https://github.com/intel/torch-xpu-ops/commit/a2278580c37bc699662f23aa74c7c31a433284cf"><code>a227858</code></a>) which corrects <code>HardtanhBackwardFunctor</code> to use <code>opmath_t</code><br>
instead of <code>scalar_t</code> for boundary value storage.</p>
<h2>Root Cause</h2>
<p>The XPU <code>HardtanhBackwardFunctor</code> stored <code>min_val_</code>/<code>max_val_</code> as <code>scalar_t</code> (bf16),<br>
truncating the caller's correctly-computed fp32 values. E.g., float 0.7 → bf16 0.69921875.<br>
The eager kernel used wrong bounds, while the inductor-decompiled path computed correct bounds<br>
from constant folding. This caused eager vs. inductor numerical disagreement.</p>
<h2>Fix (in torch-xpu-ops)</h2>
<p>Changed <code>HardtanhBackwardFunctor</code> to store <code>min_val_</code>/<code>max_val_</code> as <code>opmath_t</code> (fp32)<br>
instead of <code>scalar_t</code>, preserving full precision of boundary values.</p>
<p>See: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594526940" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3873" data-hovercard-type="pull_request" data-hovercard-url="/intel/torch-xpu-ops/pull/3873/hovercard" href="https://github.com/intel/torch-xpu-ops/pull/3873">intel/torch-xpu-ops#3873</a></p>
<h2>Related</h2>
<ul>
<li><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4593757328" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186313" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186313/hovercard" href="https://github.com/pytorch/pytorch/issues/186313">#186313</a> — issue tracking the hardtanh_backward XPU numerical discrepancy</li>
<li><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4593754123" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186312" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186312/hovercard" href="https://github.com/pytorch/pytorch/issues/186312">#186312</a> — bfloat16 autocast fix (torch-xpu-ops dependency for this pin update)</li>
<li><a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594239157" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186325" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186325/hovercard" href="https://github.com/pytorch/pytorch/pull/186325">#186325</a> — conv_large dtype fix (separate XPU test fix)</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594340173" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186327" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186327/hovercard" href="https://github.com/pytorch/pytorch/pull/186327">#186327</a><br>
Approved by: <a href="https://github.com/guangyey">https://github.com/guangyey</a>, <a href="https://github.com/EikanWang">https://github.com/EikanWang</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780898144: Support non-blocking D2H pinned copy for XPU/MTIA (#186224)]]></title>
<description><![CDATA[Motivation
Fix intel/torch-xpu-ops#3861
The original test case comes from https://github.com/pytorch/pytorch/blob/9680b3aaf628b22b9162507ec972cabd1c8725bf/test/test_torch.py#L4917-L4929
Pull Request resolved: #186224
Approved by: https://github.com/janeyx99, https://github.com/gujinghui
ghstack d...]]></description>
<link>https://tsecurity.de/de/3580542/downloads/viablestrict1780898144-support-non-blocking-d2h-pinned-copy-for-xpumtia-186224/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580542/downloads/viablestrict1780898144-support-non-blocking-d2h-pinned-copy-for-xpumtia-186224/</guid>
<pubDate>Mon, 08 Jun 2026 08:01:35 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Motivation</h1>
<p>Fix <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4585493738" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3861" data-hovercard-type="issue" data-hovercard-url="/intel/torch-xpu-ops/issues/3861/hovercard" href="https://github.com/intel/torch-xpu-ops/issues/3861">intel/torch-xpu-ops#3861</a><br>
The original test case comes from <a href="https://github.com/pytorch/pytorch/blob/9680b3aaf628b22b9162507ec972cabd1c8725bf/test/test_torch.py#L4917-L4929">https://github.com/pytorch/pytorch/blob/9680b3aaf628b22b9162507ec972cabd1c8725bf/test/test_torch.py#L4917-L4929</a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4588014338" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186224" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186224/hovercard" href="https://github.com/pytorch/pytorch/pull/186224">#186224</a><br>
Approved by: <a href="https://github.com/janeyx99">https://github.com/janeyx99</a>, <a href="https://github.com/gujinghui">https://github.com/gujinghui</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4587958356" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186223" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186223/hovercard" href="https://github.com/pytorch/pytorch/pull/186223">#186223</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/1c94cbdb6208bbb4035b724bdd11c093947f42a5: [XPU][Test] Migrate 6 UT test suites for Intel GPU (#174370)]]></title>
<description><![CDATA[Description
Fixes #114850, we will port dynamo, fsdp tests to Intel GPU
We could enable Intel GPU with following methods and try the best to keep the original code styles:
Changes

Get device type with from accelerator and get_devtype helper method
Replace the requires cuda statement with require...]]></description>
<link>https://tsecurity.de/de/3580531/downloads/trunk1c94cbdb6208bbb4035b724bdd11c093947f42a5-xputest-migrate-6-ut-test-suites-for-intel-gpu-174370/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580531/downloads/trunk1c94cbdb6208bbb4035b724bdd11c093947f42a5-xputest-migrate-6-ut-test-suites-for-intel-gpu-174370/</guid>
<pubDate>Mon, 08 Jun 2026 07:46:29 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Description</h1>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2018027027" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/114850" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/114850/hovercard" href="https://github.com/pytorch/pytorch/issues/114850">#114850</a>, we will port dynamo, fsdp tests to Intel GPU<br>
We could enable Intel GPU with following methods and try the best to keep the original code styles:</p>
<h1>Changes</h1>
<ol>
<li>Get device type with from accelerator and get_devtype helper method</li>
<li>Replace the requires cuda statement with requires_gpu.</li>
<li>Replace the cuda() with to(device_type).</li>
<li>Add is_xpu check into the test logic.</li>
</ol>
<h1>Notify</h1>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3900673626" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/174370" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/174370/hovercard" href="https://github.com/pytorch/pytorch/pull/174370">#174370</a><br>
Approved by: <a href="https://github.com/jansel">https://github.com/jansel</a>, <a href="https://github.com/guangyey">https://github.com/guangyey</a></p>
<p>Co-authored-by: Benedykt Bela <a href="mailto:benedykt.bela@intel.com">benedykt.bela@intel.com</a><br>
Co-authored-by: Wang, Chuanqi <a href="mailto:chuanqi.wang@intel.com">chuanqi.wang@intel.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/e3bfefb9348d4eb4b8eb639d74406f249e3ec1bf: [vllm hash update] update the pinned vllm hash (#186165)]]></title>
<description><![CDATA[This PR is auto-generated nightly by this action.
Update the pinned vllm hash.
Pull Request resolved: #186165
Approved by: https://github.com/pytorchbot]]></description>
<link>https://tsecurity.de/de/3580502/downloads/trunke3bfefb9348d4eb4b8eb639d74406f249e3ec1bf-vllm-hash-update-update-the-pinned-vllm-hash-186165/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580502/downloads/trunke3bfefb9348d4eb4b8eb639d74406f249e3ec1bf-vllm-hash-update-update-the-pinned-vllm-hash-186165/</guid>
<pubDate>Mon, 08 Jun 2026 07:31:33 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>This PR is auto-generated nightly by <a href="https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml">this action</a>.<br>
Update the pinned vllm hash.<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4584891165" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186165" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186165/hovercard" href="https://github.com/pytorch/pytorch/pull/186165">#186165</a><br>
Approved by: <a href="https://github.com/pytorchbot">https://github.com/pytorchbot</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780891878: Preserve aten.hardtanh meta semantics for export (#185298)]]></title>
<description><![CDATA[torch.export runs aten.hardtanh through FakeTensor/Meta dispatch. Without a
Meta kernel, that path fell back to the _refs.nn.functional.hardtanh
decomposition, which intentionally follows the Python frontend and rejects
min_val > max_val. Native aten.hardtanh has legacy behavior that allows
inver...]]></description>
<link>https://tsecurity.de/de/3580408/downloads/viablestrict1780891878-preserve-atenhardtanh-meta-semantics-for-export-185298/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580408/downloads/viablestrict1780891878-preserve-atenhardtanh-meta-semantics-for-export-185298/</guid>
<pubDate>Mon, 08 Jun 2026 06:16:03 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>torch.export</code> runs <code>aten.hardtanh</code> through FakeTensor/Meta dispatch. Without a<br>
Meta kernel, that path fell back to the <code>_refs.nn.functional.hardtanh</code><br>
decomposition, which intentionally follows the Python frontend and rejects<br>
<code>min_val &gt; max_val</code>. Native <code>aten.hardtanh</code> has legacy behavior that allows<br>
inverted bounds and delegates the result to clamp, so export failed for a model<br>
that eager ATen accepted.</p>
<p>Add dedicated Meta kernels for <code>aten.hardtanh</code>, <code>aten.hardtanh.out</code>, and<br>
<code>aten.hardtanh_</code> so export preserves the ATen operator semantics without<br>
changing eager behavior or the Python frontend. The Meta implementation mirrors<br>
native validation order for bool/complex inputs, integer scalar conversion,<br>
unsigned negative limits, out device/dtype errors, scalar range checks, and<br>
resized out strides.</p>
<p>The alternative of adding validation to native <code>hardtanh</code> was rejected because<br>
it would be BC-breaking. Changing the functional/ref frontend would also blur<br>
the existing distinction between <code>torch.nn.functional.hardtanh</code> and<br>
<code>torch.ops.aten.hardtanh</code>. A dedicated Meta implementation keeps the fix scoped<br>
to the path that was diverging.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3339336747" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/161081" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/161081/hovercard" href="https://github.com/pytorch/pytorch/issues/161081">#161081</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>python test/test_meta.py -k hardtanh</li>
<li>python test/export/test_export.py -k test_export_allows_aten_hardtanh_with_inverted_bounds</li>
<li>lintrunner -a<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4528248743" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185298" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185298/hovercard" href="https://github.com/pytorch/pytorch/pull/185298">#185298</a><br>
Approved by: <a href="https://github.com/yushangdi">https://github.com/yushangdi</a></li>
</ul>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/cd4fdbbf0edbb513777fd0bf63b5690af58df89e: Support non-blocking D2H pinned copy for XPU/MTIA (#186224)]]></title>
<description><![CDATA[Motivation
Fix intel/torch-xpu-ops#3861
The original test case comes from https://github.com/pytorch/pytorch/blob/9680b3aaf628b22b9162507ec972cabd1c8725bf/test/test_torch.py#L4917-L4929
Pull Request resolved: #186224
Approved by: https://github.com/janeyx99, https://github.com/gujinghui
ghstack d...]]></description>
<link>https://tsecurity.de/de/3580294/downloads/trunkcd4fdbbf0edbb513777fd0bf63b5690af58df89e-support-non-blocking-d2h-pinned-copy-for-xpumtia-186224/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580294/downloads/trunkcd4fdbbf0edbb513777fd0bf63b5690af58df89e-support-non-blocking-d2h-pinned-copy-for-xpumtia-186224/</guid>
<pubDate>Mon, 08 Jun 2026 04:16:39 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h1>Motivation</h1>
<p>Fix <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4585493738" data-permission-text="Title is private" data-url="https://github.com/intel/torch-xpu-ops/issues/3861" data-hovercard-type="issue" data-hovercard-url="/intel/torch-xpu-ops/issues/3861/hovercard" href="https://github.com/intel/torch-xpu-ops/issues/3861">intel/torch-xpu-ops#3861</a><br>
The original test case comes from <a href="https://github.com/pytorch/pytorch/blob/9680b3aaf628b22b9162507ec972cabd1c8725bf/test/test_torch.py#L4917-L4929">https://github.com/pytorch/pytorch/blob/9680b3aaf628b22b9162507ec972cabd1c8725bf/test/test_torch.py#L4917-L4929</a></p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4588014338" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186224" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186224/hovercard" href="https://github.com/pytorch/pytorch/pull/186224">#186224</a><br>
Approved by: <a href="https://github.com/janeyx99">https://github.com/janeyx99</a>, <a href="https://github.com/gujinghui">https://github.com/gujinghui</a><br>
ghstack dependencies: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4587958356" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186223" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186223/hovercard" href="https://github.com/pytorch/pytorch/pull/186223">#186223</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[ciflow/xpu/186327: Update xpu.txt pin: pick up hardtanh_backward opmath fix]]></title>
<description><![CDATA[Pick up kernel fix from intel/torch-xpu-ops@236e7514 which corrects
HardtanhBackwardFunctor to use opmath_t (float) instead of scalar_t (bf16)
for min_val_/max_val_ storage, matching the CUDA reference implementation.
Fixes #186313
Fixes #186312]]></description>
<link>https://tsecurity.de/de/3580270/downloads/ciflowxpu186327-update-xputxt-pin-pick-up-hardtanhbackward-opmath-fix/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580270/downloads/ciflowxpu186327-update-xputxt-pin-pick-up-hardtanhbackward-opmath-fix/</guid>
<pubDate>Mon, 08 Jun 2026 03:46:09 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Pick up kernel fix from <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/intel/torch-xpu-ops/commit/236e7514/hovercard" href="https://github.com/intel/torch-xpu-ops/commit/236e7514">intel/torch-xpu-ops@<tt>236e7514</tt></a> which corrects<br>
HardtanhBackwardFunctor to use opmath_t (float) instead of scalar_t (bf16)<br>
for min_val_/max_val_ storage, matching the CUDA reference implementation.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4593757328" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186313" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186313/hovercard" href="https://github.com/pytorch/pytorch/issues/186313">#186313</a><br>
Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4593754123" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186312" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/186312/hovercard" href="https://github.com/pytorch/pytorch/issues/186312">#186312</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/bbccf436af934a5ef2a92f2681407e7ae48b242b: Update torch-xpu-ops commit pin (#186208)]]></title>
<description><![CDATA[Update the torch-xpu-ops commit to intel/torch-xpu-ops@151bb8, includes:

Add symmetric memory support on XPU device
Temp disable xpu n-1 build job cause this new distributed feature enabling

Pull Request resolved: #186208
Approved by: https://github.com/EikanWang
Co-authored-by: Wang, Chuanqi c...]]></description>
<link>https://tsecurity.de/de/3580263/downloads/trunkbbccf436af934a5ef2a92f2681407e7ae48b242b-update-torch-xpu-ops-commit-pin-186208/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580263/downloads/trunkbbccf436af934a5ef2a92f2681407e7ae48b242b-update-torch-xpu-ops-commit-pin-186208/</guid>
<pubDate>Mon, 08 Jun 2026 03:16:30 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Update the torch-xpu-ops commit to <a href="https://github.com/intel/torch-xpu-ops/commit/151bb86a992f73f4f53e41174bcb900760362d41">intel/torch-xpu-ops@151bb8</a>, includes:</p>
<ul>
<li>Add symmetric memory support on XPU device</li>
<li>Temp disable xpu n-1 build job cause this new distributed feature enabling</li>
</ul>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4587282692" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186208" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186208/hovercard" href="https://github.com/pytorch/pytorch/pull/186208">#186208</a><br>
Approved by: <a href="https://github.com/EikanWang">https://github.com/EikanWang</a></p>
<p>Co-authored-by: Wang, Chuanqi <a href="mailto:chuanqi.wang@intel.com">chuanqi.wang@intel.com</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/411c8477fa2478b2318f3823d57cf684a3a1f389: [BE][Ez]: More semi-automated edits to move return values (#186480)]]></title>
<description><![CDATA[Otherwise they are copied into the std::tuple
Pull Request resolved: #186480
Approved by: https://github.com/malfet]]></description>
<link>https://tsecurity.de/de/3580249/downloads/trunk411c8477fa2478b2318f3823d57cf684a3a1f389-beez-more-semi-automated-edits-to-move-return-values-186480/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580249/downloads/trunk411c8477fa2478b2318f3823d57cf684a3a1f389-beez-more-semi-automated-edits-to-move-return-values-186480/</guid>
<pubDate>Mon, 08 Jun 2026 03:01:32 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Otherwise they are copied into the std::tuple</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4604531002" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186480" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186480/hovercard" href="https://github.com/pytorch/pytorch/pull/186480">#186480</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/b4ea33e73a667e4d977daa561ac4b5cd99f77f0b: Validate MaxUnpool output sizes (#184706)]]></title>
<description><![CDATA[Reject non-positive inferred or explicit MaxUnpool output dimensions before tensor allocation so eager, compiled, decomposition, and native backend paths report a clear validation error.
Fixes #178483
Generated by my agent
Pull Request resolved: #184706
Approved by: https://github.com/aorenste]]></description>
<link>https://tsecurity.de/de/3580210/downloads/trunkb4ea33e73a667e4d977daa561ac4b5cd99f77f0b-validate-maxunpool-output-sizes-184706/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580210/downloads/trunkb4ea33e73a667e4d977daa561ac4b5cd99f77f0b-validate-maxunpool-output-sizes-184706/</guid>
<pubDate>Mon, 08 Jun 2026 02:31:39 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Reject non-positive inferred or explicit MaxUnpool output dimensions before tensor allocation so eager, compiled, decomposition, and native backend paths report a clear validation error.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4141078406" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/178483" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/178483/hovercard" href="https://github.com/pytorch/pytorch/issues/178483">#178483</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4494536695" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184706" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184706/hovercard" href="https://github.com/pytorch/pytorch/pull/184706">#184706</a><br>
Approved by: <a href="https://github.com/aorenste">https://github.com/aorenste</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/68c0b26bd612cea5af71f7343492f2e24f6bd0a6: [BE][Ez]: Append single char instead of str literal overload (#186477)]]></title>
<description><![CDATA[A micro optimization that uses the correct overload of std::string
Pull Request resolved: #186477
Approved by: https://github.com/malfet]]></description>
<link>https://tsecurity.de/de/3580205/downloads/trunk68c0b26bd612cea5af71f7343492f2e24f6bd0a6-beez-append-single-char-instead-of-str-literal-overload-186477/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580205/downloads/trunk68c0b26bd612cea5af71f7343492f2e24f6bd0a6-beez-append-single-char-instead-of-str-literal-overload-186477/</guid>
<pubDate>Mon, 08 Jun 2026 02:31:33 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>A micro optimization that uses the correct overload of std::string</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4604470921" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186477" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186477/hovercard" href="https://github.com/pytorch/pytorch/pull/186477">#186477</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/f75a3b132520d11656ceb1703c0ab8a423dd55fe: Preserve aten.hardtanh meta semantics for export (#185298)]]></title>
<description><![CDATA[torch.export runs aten.hardtanh through FakeTensor/Meta dispatch. Without a
Meta kernel, that path fell back to the _refs.nn.functional.hardtanh
decomposition, which intentionally follows the Python frontend and rejects
min_val > max_val. Native aten.hardtanh has legacy behavior that allows
inver...]]></description>
<link>https://tsecurity.de/de/3580202/downloads/trunkf75a3b132520d11656ceb1703c0ab8a423dd55fe-preserve-atenhardtanh-meta-semantics-for-export-185298/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580202/downloads/trunkf75a3b132520d11656ceb1703c0ab8a423dd55fe-preserve-atenhardtanh-meta-semantics-for-export-185298/</guid>
<pubDate>Mon, 08 Jun 2026 02:16:31 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p><code>torch.export</code> runs <code>aten.hardtanh</code> through FakeTensor/Meta dispatch. Without a<br>
Meta kernel, that path fell back to the <code>_refs.nn.functional.hardtanh</code><br>
decomposition, which intentionally follows the Python frontend and rejects<br>
<code>min_val &gt; max_val</code>. Native <code>aten.hardtanh</code> has legacy behavior that allows<br>
inverted bounds and delegates the result to clamp, so export failed for a model<br>
that eager ATen accepted.</p>
<p>Add dedicated Meta kernels for <code>aten.hardtanh</code>, <code>aten.hardtanh.out</code>, and<br>
<code>aten.hardtanh_</code> so export preserves the ATen operator semantics without<br>
changing eager behavior or the Python frontend. The Meta implementation mirrors<br>
native validation order for bool/complex inputs, integer scalar conversion,<br>
unsigned negative limits, out device/dtype errors, scalar range checks, and<br>
resized out strides.</p>
<p>The alternative of adding validation to native <code>hardtanh</code> was rejected because<br>
it would be BC-breaking. Changing the functional/ref frontend would also blur<br>
the existing distinction between <code>torch.nn.functional.hardtanh</code> and<br>
<code>torch.ops.aten.hardtanh</code>. A dedicated Meta implementation keeps the fix scoped<br>
to the path that was diverging.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3339336747" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/161081" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/161081/hovercard" href="https://github.com/pytorch/pytorch/issues/161081">#161081</a><br>
Generated by my agent</p>
<p>Test Plan:</p>
<ul>
<li>python test/test_meta.py -k hardtanh</li>
<li>python test/export/test_export.py -k test_export_allows_aten_hardtanh_with_inverted_bounds</li>
<li>lintrunner -a<br>
Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4528248743" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/185298" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/185298/hovercard" href="https://github.com/pytorch/pytorch/pull/185298">#185298</a><br>
Approved by: <a href="https://github.com/yushangdi">https://github.com/yushangdi</a></li>
</ul>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/75eef9d8859933986ec75ce29aa4e472aefe0a4c: [BE][Ez]: Use rvalue overload for stringstream str (#186552)]]></title>
<description><![CDATA[CPP20 introduced an rvalue overload to stringstream.str() . Without it, the string is copied out for the str() call every time. Now, we can steal the internal string buffer and return it directly through the rvalue overload. Mechanical find and replace assisted by codex.
Pull Request resolved: #1...]]></description>
<link>https://tsecurity.de/de/3580200/downloads/trunk75eef9d8859933986ec75ce29aa4e472aefe0a4c-beez-use-rvalue-overload-for-stringstream-str-186552/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580200/downloads/trunk75eef9d8859933986ec75ce29aa4e472aefe0a4c-beez-use-rvalue-overload-for-stringstream-str-186552/</guid>
<pubDate>Mon, 08 Jun 2026 02:16:28 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>CPP20 introduced an rvalue overload to stringstream.str() . Without it, the string is copied out for the str() call every time. Now, we can steal the internal string buffer and return it directly through the rvalue overload. Mechanical find and replace assisted by codex.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4607922322" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186552" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186552/hovercard" href="https://github.com/pytorch/pytorch/pull/186552">#186552</a><br>
Approved by: <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780872028: Fix torch.combinations with symbolic sizes (#186329)]]></title>
<description><![CDATA[Fixes #163759
When compiled by \�ot_eager\ with \dynamic=True, \	orch.combinations\ failed with \RuntimeError: Cannot call numel() on tensor with symbolic sizes/strides\ because the C++ implementation used \ umel()\ instead of \sym_numel().
This commit fixes it by using \sym_numel()\ and passing ...]]></description>
<link>https://tsecurity.de/de/3580113/downloads/viablestrict1780872028-fix-torchcombinations-with-symbolic-sizes-186329/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580113/downloads/viablestrict1780872028-fix-torchcombinations-with-symbolic-sizes-186329/</guid>
<pubDate>Mon, 08 Jun 2026 00:46:19 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="3449334226" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/163759" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/163759/hovercard" href="https://github.com/pytorch/pytorch/issues/163759">#163759</a></p>
<p>When compiled by \�ot_eager\ with \dynamic=True, \	orch.combinations\ failed with \RuntimeError: Cannot call numel() on tensor with symbolic sizes/strides\ because the C++ implementation used \ umel()\ instead of \sym_numel().<br>
This commit fixes it by using \sym_numel()\ and passing \c10::SymInt\ to the mask generation.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4594515153" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186329" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186329/hovercard" href="https://github.com/pytorch/pytorch/pull/186329">#186329</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[viable/strict/1780866717: Use normalized name spmd-types in wheel Requires-Dist (#186545)]]></title>
<description><![CDATA[Summary
PYTORCH_EXTRA_INSTALL_REQUIREMENTS becomes the torch wheel's Requires-Dist, so SPMD_TYPES_REQUIREMENT="spmd_types==0.2.0" (added in #180880) emitted Requires-Dist: spmd_types==0.2.0.
download.pytorch.org serves the dependency under its PEP 503-normalized name (spmd-types/), so when instal...]]></description>
<link>https://tsecurity.de/de/3580034/downloads/viablestrict1780866717-use-normalized-name-spmd-types-in-wheel-requires-dist-186545/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580034/downloads/viablestrict1780866717-use-normalized-name-spmd-types-in-wheel-requires-dist-186545/</guid>
<pubDate>Sun, 07 Jun 2026 23:17:13 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<h2>Summary</h2>
<p><code>PYTORCH_EXTRA_INSTALL_REQUIREMENTS</code> becomes the torch wheel's <code>Requires-Dist</code>, so <code>SPMD_TYPES_REQUIREMENT="spmd_types==0.2.0"</code> (added in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4297638522" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/180880" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/180880/hovercard" href="https://github.com/pytorch/pytorch/pull/180880">#180880</a>) emitted <code>Requires-Dist: spmd_types==0.2.0</code>.</p>
<p><code>download.pytorch.org</code> serves the dependency under its PEP 503-normalized name (<code>spmd-types/</code>), so when installing with <code>--index-url https://download.pytorch.org/whl/nightly/&lt;arch&gt;</code> pip looks up the normalized path. Emit the normalized requirement <code>spmd-types==0.2.0</code> so the metadata matches the index path.</p>
<p>PyPI treats <code>spmd_types</code> and <code>spmd-types</code> as equivalent, so default-index installs are unaffected.</p>
<h2>Companion change</h2>
<p>The matching mirror/index fixes on <code>download.pytorch.org</code> are in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4607353895" data-permission-text="Title is private" data-url="https://github.com/pytorch/test-infra/issues/8157" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/test-infra/pull/8157/hovercard" href="https://github.com/pytorch/test-infra/pull/8157">pytorch/test-infra#8157</a>.</p>
<h2>Note</h2>
<p>Changing <code>PYTORCH_EXTRA_INSTALL_REQUIREMENTS</code> requires triggering <code>ciflow/binaries</code> to validate the binary builds.</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4607354264" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/186545" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/186545/hovercard" href="https://github.com/pytorch/pytorch/pull/186545">#186545</a><br>
Approved by: <a href="https://github.com/Skylion007">https://github.com/Skylion007</a>, <a href="https://github.com/malfet">https://github.com/malfet</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/26296f2e0ecbe4b93e6cfe937ffca58cfa8c50f8: Preserve metadata for fused Inductor RNG nodes (#184316)]]></title>
<description><![CDATA[When replace_random fuses RNG seed and offset helper nodes, copy source metadata from the original per-RNG helper onto the fused node before recomputing tensor metadata.
Fixes #129813
Generated by my agent
Pull Request resolved: #184316
Approved by: https://github.com/choijon5]]></description>
<link>https://tsecurity.de/de/3580033/downloads/trunk26296f2e0ecbe4b93e6cfe937ffca58cfa8c50f8-preserve-metadata-for-fused-inductor-rng-nodes-184316/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580033/downloads/trunk26296f2e0ecbe4b93e6cfe937ffca58cfa8c50f8-preserve-metadata-for-fused-inductor-rng-nodes-184316/</guid>
<pubDate>Sun, 07 Jun 2026 23:17:12 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>When replace_random fuses RNG seed and offset helper nodes, copy source metadata from the original per-RNG helper onto the fused node before recomputing tensor metadata.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="2381448130" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/129813" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/129813/hovercard" href="https://github.com/pytorch/pytorch/issues/129813">#129813</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4474726965" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/184316" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/184316/hovercard" href="https://github.com/pytorch/pytorch/pull/184316">#184316</a><br>
Approved by: <a href="https://github.com/choijon5">https://github.com/choijon5</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/597d8e5fe8a7e1853d371fad8d6ea081dd833c2f: test/inductor: disable caches in max reads fusion test (#183695)]]></title>
<description><![CDATA[The max reads fusion test asserts scheduler/codegen metrics that are only populated during fresh compilation. Disable Inductor caches for this test so a warm FX graph cache cannot bypass the metric-producing path; prior abandoned context is chuanqi129#5.
Fixes #181699
Generated by my agent
Pull R...]]></description>
<link>https://tsecurity.de/de/3580032/downloads/trunk597d8e5fe8a7e1853d371fad8d6ea081dd833c2f-testinductor-disable-caches-in-max-reads-fusion-test-183695/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580032/downloads/trunk597d8e5fe8a7e1853d371fad8d6ea081dd833c2f-testinductor-disable-caches-in-max-reads-fusion-test-183695/</guid>
<pubDate>Sun, 07 Jun 2026 23:17:11 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>The max reads fusion test asserts scheduler/codegen metrics that are only populated during fresh compilation. Disable Inductor caches for this test so a warm FX graph cache cannot bypass the metric-producing path; prior abandoned context is <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4348982328" data-permission-text="Title is private" data-url="https://github.com/chuanqi129/pytorch/issues/5" data-hovercard-type="pull_request" data-hovercard-url="/chuanqi129/pytorch/pull/5/hovercard" href="https://github.com/chuanqi129/pytorch/pull/5">chuanqi129#5</a>.</p>
<p>Fixes <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4340177066" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/181699" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/181699/hovercard" href="https://github.com/pytorch/pytorch/issues/181699">#181699</a></p>
<p>Generated by my agent</p>
<p>Pull Request resolved: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4445019337" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/183695" data-hovercard-type="pull_request" data-hovercard-url="/pytorch/pytorch/pull/183695/hovercard" href="https://github.com/pytorch/pytorch/pull/183695">#183695</a><br>
Approved by: <a href="https://github.com/shunting314">https://github.com/shunting314</a></p>]]></content:encoded>
</item>
<item>
<title><![CDATA[trunk/825f10cd64ccb8a15cc70a6d76f4840b93b0cc24]]></title>
<description><![CDATA[Revert "Use oneDNN for int8 quantization AArch64 with x86 engine (#18…]]></description>
<link>https://tsecurity.de/de/3580031/downloads/trunk825f10cd64ccb8a15cc70a6d76f4840b93b0cc24/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580031/downloads/trunk825f10cd64ccb8a15cc70a6d76f4840b93b0cc24/</guid>
<pubDate>Sun, 07 Jun 2026 23:17:09 +0200</pubDate>
<category>💾 Downloads</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<p>Revert "Use oneDNN for int8 quantization AArch64 with x86 engine (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="176034421" data-permission-text="Title is private" data-url="https://github.com/pytorch/pytorch/issues/18" data-hovercard-type="issue" data-hovercard-url="/pytorch/pytorch/issues/18/hovercard" href="https://github.com/pytorch/pytorch/issues/18">#18</a>…</p>]]></content:encoded>
</item>
<item>
<title><![CDATA[Picking a distro for an RTX 5090 (Blackwell) CUDA + Python workstation... CachyOS?]]></title>
<description><![CDATA[I've been going back and forth on this for a while and figured the people here would have actual experience rather than just opinions. Posting my hardware, what I do with it, and my reasoning, happy to be argued out of it. The hardware  Laptop (TongFang barebone): Ryzen 9 9955HX, 64 GB RAM, ~3.7 ...]]></description>
<link>https://tsecurity.de/de/3580018/linux-tipps/picking-a-distro-for-an-rtx-5090-blackwell-cuda-python-workstation-cachyos/</link>
<guid isPermaLink="true">https://tsecurity.de/de/3580018/linux-tipps/picking-a-distro-for-an-rtx-5090-blackwell-cuda-python-workstation-cachyos/</guid>
<pubDate>Sun, 07 Jun 2026 23:07:46 +0200</pubDate>
<category>🐧 Linux Tipps</category>
<source url="https://tsecurity.de">tsecurity.de</source>
<content:encoded><![CDATA[<!-- SC_OFF --><div class="md"><p>I've been going back and forth on this for a while and figured the people here would have actual experience rather than just opinions. Posting my hardware, what I do with it, and my reasoning, happy to be argued out of it.</p> <h1>The hardware</h1> <ul> <li>Laptop (TongFang barebone): Ryzen 9 9955HX, 64 GB RAM, ~3.7 TB</li> <li>GPU: RTX 5090 Laptop (Blackwell, ~23 GB) + AMD Radeon 610M iGPU (hybrid)</li> <li>Dual-booting an existing Windows 11 install</li> </ul> <h1>What I actually do with it</h1> <p>Research computing. The specific science doesn't really matter for the distro choice (gravitational-wave data analysis, if you're curious), so here's the shape that does matter:</p> <ul> <li>Heavy CUDA + scientific Python: numpy/scipy, PyTorch / CuPy / JAX, the usual suspects</li> <li>Everything lives in Conda/Miniforge environments, deliberately kept off the system Python</li> <li>VS Code Remote-SSH into HPC clusters; but also heavy local dev + GPU runs</li> <li>Desktop: KDE Plasma or Gnome with Tweaks + Extensions on Wayland, 2-4 monitors with independent fractional scaling (e.g. one screen at 150%, another at 100%)</li> </ul> <h1>The constraints that actually drive the decision</h1> <ul> <li>Blackwell needs the open NVIDIA kernel modules + a recent driver (570+), so I want a reasonably fresh kernel/driver</li> <li>It's a work machine, so I want stability + a real rollback path (snapshots), not heroics</li> <li>Clean separation between system / Flatpak GUI apps / Conda science stack / vendor dev tools</li> </ul> <h1>Why I'm leaning CachyOS</h1> <p>Shortlist was</p> <ul> <li>Fedora (Plasma or KDE),</li> <li>Kubuntu (KDE) / Ubuntu (Gnome),</li> <li>openSUSE Tumbleweed,</li> <li>EndeavourOS and</li> <li>CachyOS.</li> </ul> <p>CachyOS keeps pulling me back because:</p> <ul> <li>Freshest kernel + driver, which matters for a launch-window GPU</li> <li>Btrfs bootable snapshots + an LTS fallback kernel by default</li> <li>NVIDIA handled in the installer</li> </ul> <p>The honest counterpoint I keep arguing with myself about: My compute stack is Conda binaries, which ship their own optimized BLAS/FFT, so CachyOS's x86-64-v3/v4 repo optimizations mostly benefit system-level stuff, not the science I actually run. So some of the appeal might just be vibes. Fedora KDE is the calmer alternative (fixed release, and RPM Fusion's akmods auto-signs the NVIDIA module so Secure Boot), and Tumbleweed arguably has the best out-of-the-box rollback story.</p> <p>I was also thinking about Ubuntu/Kubuntu, but I don't want a bloated setup and snap gets forced on you. On the other side it is the industry standard.</p> <h1>What I'd genuinely love input on</h1> <ol> <li>Anyone running Blackwell / RTX 50-series on Arch or CachyOS: How has the open-module + rolling-kernel combo held up? Any breakages on kernel bumps?</li> <li>Hybrid AMD iGPU + NVIDIA dGPU on Wayland: On these laptops the external outputs are often wired to the dGPU. PRIME / reverse-PRIME experiences and gotchas?</li> <li>Rolling vs fixed for a CUDA workstation: Does the freshness actually pay off, or does it just turn into babysitting the kernel/driver before every update?</li> <li>Secure Boot on the Arch family with out-of-tree NVIDIA: Worth the signing setup, or do you just disable it and move on?</li> <li>Anyone who picked CachyOS specifically for compute: did the optimized repos make a measurable difference, or is Fedora/Tumbleweed effectively the same once your real work is in Conda containers?</li> <li>Because someone mentioned Arch Linux: Shouldn't have CachyOS the same customization options? I think they just have added a bit above Arch Linux. I also like the btrfs snapshot and rollback feature. I was thinking about using EndeavourOS and add it, but then I was questioning myself why even doing the extra work to rebuild CachyOS if CachyOS is already there.</li> </ol> </div><!-- SC_ON -->   submitted by   <a href="https://www.reddit.com/user/Grelueen"> /u/Grelueen </a> <br> <span><a href="https://www.reddit.com/r/linux/comments/1tznceo/picking_a_distro_for_an_rtx_5090_blackwell_cuda/">[link]</a></span>   <span><a href="https://www.reddit.com/r/linux/comments/1tznceo/picking_a_distro_for_an_rtx_5090_blackwell_cuda/">[comments]</a></span>]]></content:encoded>
</item>
</channel>
</rss>
<!-- Generated in 0,11ms -->