Vue normale

Meta Debuts First AI Coding Agent To Take On Anthropic and OpenAI

Par : BeauHD
5 août 2026 à 21:00
Meta has launched Muse Code, its first AI coding agent that's positioned as a lower-cost rival to Anthropic's Claude and OpenAI's Codex. It offers pay-as-you-go pricing and an optional zero-data-retention feature for enterprise users. CNBC reports: Muse Code is the latest major release from AI chief Alexandr Wang, who leads Meta Superintelligence Labs and oversees foundation model development. Wang joined in June of last year as the centerpiece of CEO Mark Zuckerberg's effort to revamp his company's flailing artificial intelligence strategy. "You can install it with one command and then use it to take on complete software engineering tasks across a wide variety of use cases, planning changes, writing code, validating the results," Wang said in an interview on Wednesday. [...] The new tool, like Anthropic's Claude and OpenAI's Codex assistants, makes it easier for people to build apps within a single user interface while managing fleets of AI-powered digital agents that can help underpin the software development process. Muse Code, available in a preview version, works alongside the company's latest AI model, Muse Spark 1.2. Wang declined to share user statistics related to the company's Muse Spark AI models, but said "adoption has been exciting and strong." The latest Muse Spark model was developed and trained alongside Muse Code, which Wang said improves the overall coding performance.

Read more of this story at Slashdot.

C’est reparti : des agents IA d’Anthropic et OpenAI ont dépassé les bornes en plein test de sécurité

5 août 2026 à 08:07

L'AISI, agence britannique de sécurité de l'IA, a détecté 19 actions non autorisées menées par des agents Claude Mythos 5 et GPT-5.6 Sol sur l'internet réel, lors d'un test cyber censé rester confiné à un environnement simulé.

Microsoft Tells Engineers 'Tokenmaxxing Is Not What We Are Optimizing For'

Par : BeauHD
4 août 2026 à 20:00
Microsoft is introducing AI token budgets for employees, making the cheaper GPT-5.6 its default internal model and telling engineers to focus on business results rather than maximizing AI usage. 404 Media reports: "As we accelerate our use of GitHub Copilot to deliver on our goals, we all need to be aware of how we consume tokens," Jay Parikh, an executive vice president at Microsoft said in an email to Microsoft employees. GitHub is owned by Microsoft, and GitHub Copilot is an AI coding tool. "Tokenmaxxing is not what we are optimizing for. I want all of us focused on maximizing outcomes that move the needle for our customers and our business." "As such, we are updating our internal guidance and managing token spend with the same discipline we apply to every other critical resource," Parikh said in the email. Parikh's email says that in an effort to "get greater value from our token investment" Microsoft is making OpenAI GPT-5.6, which is cheaper to use than other models, the default model for internal use. His email also links to updated internal Copilot guidelines stating that, as of July 2026, Microsoft divisions will have an "AI token budget target," and that employees can track their individual AI spending. "While there is no target spend value being shared at this time. The data shows that many engineers spend in the range of hundreds of dollars a month to a few thousand dollars in tokens," the guidelines say. They also say that some decisions may place further restrictions as they monitor spend. [...] Parikh's email said Microsoft will keep learning and adjusting its AI policies as models and products evolve, and stressed that he doesn't want to slow down the company's progress towards becoming "AI-first." "We are not optimizing for fewer tokens," he said. "We are optimizing for more impact per token.

Read more of this story at Slashdot.

OpenAI's Astra Solved Decades-Old Math Problems For $2,000

Par : BeauHD
4 août 2026 à 03:30
An anonymous reader quotes a report from Forbes: The cost of producing new results on ten longstanding mathematical problems just fell to $2,000, according to OpenAI, which says its Astra model generated machine-checkable proofs for questions that had resisted human progress for decades. OpenAI published the work on August 1 and used it to give its next major model family a name: Astra. The results run across group theory, high-dimensional geometry, coding theory, quantum complexity, lattice cryptography and extremal combinatorics. They arrived as a 249-page manuscript collection and, alongside it, something the field has not seen attached to an AI claim before at this scale: a machine-checkable certificate for every single result. The problems were not textbook exercises dressed up as discoveries. Each had been open for at least ten years, most of them far longer, and several sit at the center of their subfields: - A construction establishing the existence of non-sofic groups, a question that has occupied group theorists for years. - A disproof of Connes's rigidity conjecture, a long-standing problem in the theory of von Neumann algebras. - An improvement to the general upper bound on sphere-packing density in high dimensions, a bound that had stood since 1978. - Three problems come from the catalogue of open questions left behind by Paul Erdos. The announcement follows another result from May, when OpenAI used a similar reasoning model to produce an original mathematical proof disproving a famous unsolved conjecture in geometry, which was first posed by Paul Erdos in 1946.

Read more of this story at Slashdot.

Microsoft CEO Touts His Own DIY AI Project To Wall Street and His 20 Million Followers

Par : BeauHD
3 août 2026 à 20:00
theodp writes: During Microsoft's 2026Q4 earnings call, CEO Microsoft Satya Nadella took time to tout a dashboard he personally created using AI from a Morgan Stanley analyst's PDF research report, which suggested a rosy payback for the so-called MAG7's ('Magnificent 7' companies) massive capital expenditures on AI (to which Nadella later added a "not financial advice" disclaimer). "It would be fun for you, Adam. I think one of your colleagues put out an ROIC [Return on Invested Capital] document. I took that document to Copilot, which is a PDF, and I said, 'Build me a new Power BI dashboard, essentially.' But here is the thing. It built a rich semantic model that went into my Fabric with OneLake that brought all the data in from the external sources. In fact, it was current with all the SEC filings of all the MAG7. And then on top of that, the repo itself is in GitHub, but the artifact is sitting in my Copilot as a site. That, to me, is a classic example of an enterprise-wide workflow. I, as a knowledge worker, could go create a dashboard. The data engineer can go to Fabric and find the artifact. The professional developer can go to the repo and find it in GitHub. And by the way, it's all registered with Agent 365. That's a little bit of what Amy is describing as the coming together of a new way to work, even while at the same time, bringing IT, security and manageability of it." After Nadella's show-and-tell drew an underwhelming response during the call ("That's very helpful. Thank you." said the Morgan Stanley analyst whose team's work Nadella scraped with AI), Nadella turned to social media with posts on LinkedIn (12M followers) and Twitter/X (8M followers) to make the case for why his DIY project was such a brilliant demonstration of how AI enables governance, controls, security, development, testing, deployment, maintenance, data analysis/modeling, visualization, usability, and value. "Some more detail on the ROIC Intelligence App I built yesterday and mentioned on today's earnings call," Nadella wrote on LinkedIn. "I took the PDF that Brian Nowak at Morgan Stanley put together for Hyperscale ROIC this week and used Copilot code (coming in our new superapp) with a single prompt + skill (/drill-me) to create the plan, then used autopilot in auto to create the full app (with history, lookups, scenarios, what-ifs, etc). And /rubber-duck to test. And the best part is that all the artifacts are in my enterprise environment. My app is in Copilot, my code is in GitHub Enterprise; all my data pipelines/lake/semantic models are in Fabric. And everything is under Agent 365 IT/Sec/FinOps control! So this is not about Tokenmaxxing or vibe coding. Every step of the way the rails are engineered to create value, making everything a long-term reusable asset, with governance/security, and cost controls. This is the full system to drive business value. Disclosures: This is all pulled from public sources, and for illustrative purposes only...not financial advice! :) Here is the app and architecture..." Not unexpectedly, the accompanying screenshot of a splash page for the BI app and a buzzword-laden complicated architecture diagram drew universal praise from LinkedIn fans, but also a few barbs from less-than-impressed commenters on Nadella's Twitter/X post, some of whom suggested Nadella's project might even represent a jump-the-shark moment for AI mania. "That he doesn't see whats wrong with saying 'My app is in Copilot, my code is in GitHub Enterprise; all my data pipelines/lake/semantic models are in Fabric' is exactly why MS is failing at AI,'" replied @PassingPixels on X/Twitter. "Dude is having to run 5 different systems to emulate babies first vibe code." @zigmund_ignatov added, "Why do we call glorified slide show an app?" @Mathupiriyan quipped, "Looks like Copilot just turned a PDF into a profit crystal ball." Unimpressed, @Markusndnb remarked, "So you created a web page using tons of proprietary MS tools." And @FishyAccounting called on Nadella to show-his-work, saying "Post the prompt or it didn't happen." (btw, Microsoft President Brad Smith similarly declined to provide the prompt for his own self-described amazing AI DIY reporting project that he touted at Microsoft's Shareholder Meeting last December). So, does Nadella's self-promoted AI reworking of someone else's PDF research report strike you as an amazing example of everything that's good about AI, or does it conjure up memories of The Emperor's New Clothes?

Read more of this story at Slashdot.

Company Offering Printed Books To Train AI Stops After 404 Media Coverage

Par : BeauHD
3 août 2026 à 15:00
An anonymous reader quotes a report from 404 Media: Following 404 Media's reporting that book database company ISBNdb claimed to source printed books to then sell to AI companies for AI training, the company deleted the part of its website offering the service and walked back claims that it would train AI models, and instead called it "a test of market interest." On July 30, nine days after 404 Media's reporting, ISBNdb added a note to its homepage and an update on its news page about the change. "We've seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised. The facts: ISBNdb has never purchased, scanned, or sold a book -- for AI training or anything else," ISBNdb wrote. "We don't train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We've taken the page down. Our job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world -- the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn't changed." ISBNdb removed the landing page for "Printed Books Sourcing for Your AI LLMs Dataset Needs" on July 28. "It was part of exploring demand, and we've chosen to pivot away from that direction. Our main ISBNdb (book metadata API) services are unaffected and running as usual," the site says. According to 404 Media's previous report, ISBNdb had pitched pre-2022 printed books as premium AI training material because they contained curated human knowledge without contamination from AI-generated text or modern data-poisoning techniques. "The world's best AI training data is sitting on a shelf," the company had argued, while offering discreet bulk book acquisition under NDAs and acknowledging the potential backlash if AI firms were seen destroying millions of books for scanning.

Read more of this story at Slashdot.

'AI's Decimation of Call Center Jobs Has Begun'

3 août 2026 à 11:34
"AI's decimation of call center jobs has begun," reports Bloomberg: Companies including the Commonwealth Bank of Australia, Microsoft Corp., Uber Technologies Inc. and Hyatt Hotels Corp. are using automated chat and phone systems to handle work that previously required humans. In some cases, they've already wiped out sizable chunks of their customer service operations, together representing thousands of workers. The specter of automation has long loomed over the call center industry, which employs millions worldwide from the U,S, to India to the Philippines. But until recently, generative artificial intelligence wasn't good enough to move the needle. Now, AI advancements — and pressure on executives to show they're embracing the new technology — have prompted corporations to deploy the tools more widely. Customer service employment in the U.S. is declining and will likely continue to do so as more tasks are automated, Forrester analyst Kate Leggett wrote in a report earlier this year. While it's impossible to determine the future job losses, she estimated that almost half of customer service roles will be affected by 2030. Globally, the steepest job cuts are expected to hit countries like the Philippines, where many Western companies have outsourced their most easily automated work. Salespeople at multiple tech companies told Bloomberg that they routinely pitch call center AI tools as a way of lowering labor costs, undercutting a common industry claim that AI is primarily a way to help workers become more productive rather than kill their jobs... Commonwealth Bank of Australia, the nation's largest lender, has shed hundreds of workers from its chat support line as it wove AI into the system, according to people familiar with the work. This amounted to tens of millions of dollars in savings per year, one of the people said... Microsoft is both one of the largest vendors and adopters of customer service automation tools. This has helped the software giant trim its customer service workforce — a mix of contractors and full-time staff — from about 50,000 to 40,000 in recent years, according to a person familiar with the operations. "If something happened with little Johnny's Xbox in the middle of the night, we can now solve that with AI," Judson Althoff, who runs Microsoft's sales and service operations, said in an interview. Althoff said in April that AI is saving the company about $750 million per year in customer service costs. More complex problems still require human support, but the company is constantly expanding what can be fixed automatically, he said in the interview. Two examples from the article: Last year Hyatt fired 30% of its in-house customer support staff for the Americas, according the hotel-industry news site Hotel Dive. Last week Bloomberg reported Uber had cut 10% of its customer service jobs as part of effort to "embrace artificial intelligence," according to the article. "Today, Uber pushes users to submit support requests through their apps, where they're met with an AI chatbot."

Read more of this story at Slashdot.

ChatGPT résout 10 grands problèmes de maths, mais il y a un souci considérable qui émerge

3 août 2026 à 11:01

openai chatgpt ia mathématiques

OpenAI affirme qu'une version interne de son modèle Astra a terrassé dix problèmes ouverts majeurs. Mais derrière l'emballement, les mathématiciens font face à un gouffre vertigineux : l'intelligence artificielle génère désormais des démonstrations plus vite que la science n'est capable de les comprendre.

How 'Situational Awareness' Hedge Fund Dropped 67% in AI Stock Rout

2 août 2026 à 04:45
CNN tells the unfortunate tale of hedge fund Situational Awareness, "founded in 2024 by German-born Leopold Aschenbrenner when he was in his early 20s." Aschenbrenner, a former OpenAI employee, founded the hedge fund on the premise that "AI will be the dominant driver of global market returns over the next decade," according to the firm's site... Aschenbrenner managed to turn hundreds of millions of dollars into tens of billions of dollars over the course of roughly two years... That streak ended on Thursday, though, when the fund was forced to sell the bulk of its public holdings to a bigger rival after many of its investments went south. But that's only part of the story. The fund employed a risky strategy of borrowing money to purchase stocks. When the investments appreciate, the payoff can be massive. But when the investments sour, the losses can be catastrophic. The downturn in AI stocks over the course of this month, like chip makers and cloud computing providers, hit the hedge fund extra hard. It was forced to sell off many investments at a steep discount to rival hedge fund Citadel in what Aschenbrenner reportedly compared to a "bank run" in a letter to investors. "Critics pointed out that Aschenbrenner had no experience running money prior to launching his fund in July 2024, calling him more lucky than smart," writes CNBC: Some noted that his early work experience was at the doomed crypto firm FTX, where he helped now-disgraced founder Sam Bankman-Fried run a charity out of a Bahamas penthouse. Others on Wall Street, including former traders at global investment banks, noted that in light of reports Situational Awareness used as much as 400% leverage, the collapse wasn't shocking. The Wall Street Journal reports that Situational Awareness "also used options to amplify its returns. That meant that even small declines in individual names could have big impacts on Situational's portfolio." And so, as the New York Post put it, "The celebrated crystal ball of the 'Nostradamus of AI' hasn't merely gone cloudy — it has rolled off the table and shattered on the parlor floor." Wall Street breathed a huge sigh of relief last week as an AI-focused hedge fund called Situational Awareness reportedly sold most of its portfolio — reportedly down 67% last month on the backfiring of debt-fueled bets on chipmakers and assorted artificial-intelligence firms — to billionaire Ken Griffin's Citadel... The prevailing sentiment was best summed up by a veteran Wall Street sage who has seen a lot of flameouts in his day. Let's just say he wasn't impressed by Leopold Aschenbrenner, the 25-year-old German-born "Nostradamus" figure who is the founder of Situational Awareness... "Just your typical leveraged idiot who was right until he was wrong," the source said, adding that the implosion is a "one-off...." [Another trusted source] felt there was room for conversation: "A significant issue. Not viewed as systemic right now. I wonder if that changes as more problems arise." Indeed, the fact is that most of Wall Street is closely monitoring the Situational Awareness situation because they were holding many of the same positions as Aschenbrenner. Another top hedge fund manager I won't name tells me he has been getting crushed on similar investments in chipmakers essential to the AI supply chain, as well as other companies feeding off this technology. Thanks to Slashdot reader joshuark for sharing the news.

Read more of this story at Slashdot.

Is Big Tech's AI Gamble Starting to Look Riskier?

2 août 2026 à 00:12
The Washington Post looks at giant tech companies "feeding every available dollar into the cash-incinerating maw of AI machines." They warn "Tech superstars that once had oodles of cash left over at the end of each year are now flipping into the red..." [While optimists expect] huge corporate profits and a society-wide boost to wealth and well-being... questions about that AI vision are now growing more urgent: When, if ever, will this payoff arrive? And what will the fallout be for Americans if the titanic investment doesn't quickly deliver? "This AI thing better work out because if it doesn't ... we're going to have a problem," said Torsten Slok, chief economist at investment firm Apollo Global Management. AI costs and doubts are spreading. The U.S. stock market has swooned this summer over fear of the AI bubble going bust... The AI gamble sweeping up American fortunes is led by tech companies splurging on hulking data centers packed with computer chips and equipment needed to develop sophisticated AI models and deliver them to customers. In investor calls in the past week, Google, Microsoft, Meta and Amazon pointed to soaring AI-related sales and business deals. Advertisers are using the technology to tailor marketing pitches and corporations and start-ups are buying access to chatbots and other AI software to boost productivity... But this spending can only continue if AI generates an even larger avalanche of new revenue to pay for it all. Financial results released over the past week show that the AI titans' mammoth costs are largely swamping the sales boost from the technology. At Google, for every dollar of cash its business generated in the past three months, $1.15 went out the door to pay for AI computer chips and equipment, land for AI data centers and other big-ticket purchases. The company is covering the difference partly by borrowing money and selling more of its stock. Next year, five leading AI companies — Google, Amazon, Microsoft, Meta and Oracle — are projected to have negative free cash flow, which measures the cash left over after paying expenses and AI infrastructure costs. The figures, based on investment analyst projections compiled by S&P Global Market Intelligence, show a stunning reversal for what have been some of the world's most cash-generating corporations... The companies remain profitable by standard financial accounting measures that spread out the costs of their AI infrastructure spending over many years... Pessimists see a bet so gargantuan that it cannot possibly pay off. The pessimists are growing louder. The Bank for International Settlements, a typically measured institution in Switzerland that advises government bankers around the world, recently warned there was risk of "economy-wide recessions" if the AI boom falters. That could mean pain for workers and communities across the United States. "I'm not saying AI is going to go away, it's just not clear to me these guys are going to make money on it," said Christopher Wood, global head of equity strategy at investment bank Jefferies who has correctly predictedpast financial bubbles.

Read more of this story at Slashdot.

OpenAI Finds Evidence Other AI Agents Escaped Containment

Par : BeauHD
1 août 2026 à 03:30
An anonymous reader quotes a report from Reuters: OpenAI has discovered other instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention this month, two people familiar with the matter said on Friday. The new breakouts were uncovered during the company's publicly announced investigation into how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network. An OpenAI spokesperson referred to a statement issued by the company on Tuesday that said it was reviewing "broader activity from our models" in addition to the Hugging Face intrusion. The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by OpenAI was launched shortly before its primary rival, Anthropic, disclosed that its models were also responsible for a series of break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of other past breakouts at OpenAI has not previously been reported. AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control. "We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician who works at Cambridge University's Center for the Study of Existential Risk. Reuters could not establish exactly how many incidents OpenAI investigators found or the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.

Read more of this story at Slashdot.

The Major Labels Propose Rules to Keep AI Slop Off the Charts

Par : BeauHD
31 juillet 2026 à 23:00
Major record labels including Universal, Sony, and Warner have proposed excluding AI-generated songs from official charts unless they are "substantially human made," properly labeled, legally produced, and free from manipulation concerns. The Verge reports: The proposal goes quite a bit further than a labeling proposal put forth by the RIAA, the International Federation of the Phonographic Industry (IFPI), SAG-AFTRA, and others. That would create a set of standardized labels for AI-generated and AI-assisted music. The labels' proposal would require songs be clearly labeled, but it would also keep them off international charts unless they met specific criteria, including being "substantially human made." To be eligible, the songs would also have to respect the terms of service of whatever AI service was used, the model would have to have the rights to any data it was trained on, and "not raise stream or chart manipulation concerns." What sort of concerns and what constitutes "substantially human made" are currently vague. Sony Music, UMG, and Mom+Pop Music did not immediately respond to a request for clarification. The IFPI has thrown its weight behind the labels' proposal, though no charting organization has signaled any immediate plan to adopt the rules [...].

Read more of this story at Slashdot.

New Google Earth AI Tool Could Fuel Misinformation, Experts Say

Par : BeauHD
31 juillet 2026 à 19:00
Google has integrated its Nano Banana 2 image generator into Google Earth, allowing users to place AI-generated events and objects onto real satellite imagery. The company says its AI-generated images contain invisible watermarks detectable through Gemini or Lens, but the BBC found those safeguards and some third-party detection tools can be fooled into labeling manipulated Google Earth images as real. From the report: A collapsed Eiffel Tower, a sinkhole swallowing the Great Pyramid of Giza and Russian tanks in Ukraine's capital were among the images BBC Verify was able to create when testing the feature, which was rolled out on Thursday. Google has not yet responded to questions based on BBC Verify's tests, but in a social media post the company said they "take misinformation seriously" and that "we prevent image creation on harmful topics and are continually updating our protections." AI and misinformation expert Henk van Ess has highlighted the risks this feature poses, creating fake images of a non-existent nuclear power plant in Iran, a refugee camp on the US-Mexico border and a fake hospital in Gaza with a bomb crater next to it. He said Google was allowing "invented" imagery to be "welded to genuine coordinates, drawn on genuine imagery." "The forgery does not have to look convincing on its own. It inherits the credibility of the map it was born on," van Ess added. UPDATE 7/31/26 10:54 AM: Google is rolling back the image generation inside of Google Earth: "We know that people uniquely trust Google Earth for a reliable view of the world. We've seen geospatial professionals using this feature for a range of useful purposes, however we've also seen people sharing screenshots of generated imagery that appear to violate our policies. So we're rolling back this feature in Google Earth while we work on implementing stronger guardrails. It's important to note that generated images didn't appear in the main Google Earth experience for others to see and were watermarked as AI generated."

Read more of this story at Slashdot.

Plus de 40 000 Hyundai Inster en rappel à cause d’un risque d’incendie

31 juillet 2026 à 12:15

Hyundai doit faire revenir au garage la quasi-totalité des citadines électriques Inster produites depuis son lancement. Même s’il ne s’agit pas d’un problème de batterie haute tension, le rappel n’est pas à négliger.

Anthropic Says Its AI Systems Broke Into Computers at 3 Organizations

Par : BeauHD
31 juillet 2026 à 05:30
Anthropic found that Claude models breached three outside organizations during cybersecurity tests because misconfigured environments accidentally gave them access to the internet. The company notified those affected and urged other AI labs to audit their own testing systems. The BBC reports: Anthropic said in a statement that it reviewed more than 140,000 tests to find evidence that Claude - its family of AI models - could access the internet from testing environments that were designed to be sealed off. The tests include so-called "capture-the-flag" evaluations in which Claude was tasked with obtaining information by breaching other systems - a common way that experts assess a model's hacking capabilities. A "misconfiguration" on systems run by Anthropic and its testing partner left the models with live internet access, allowing them to breach other systems, the San Francisco-based firm said. Anthropic said the earliest incidents date back to April and that it is "approaching the fixes as if the responsibility were ours alone." Neither Anthropic nor the organizations that were breached had noticed the intrusions at the time. Anthropic said it could have reviewed its records more thoroughly and added that the findings gave the firm "cautious optimism" that such risks can be overcome with more investment and tighter measures. "The broader lesson is not necessarily that AI has developed a fundamentally new attack capability," cyber security expert David Allott told the BBC. "Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed," he added. The announcement comes just days after OpenAI said that its models had breached the systems of other companies, including AI tools platform Hugging Face.

Read more of this story at Slashdot.

New MCP Specification Addresses the Main Barrier To Enterprise Adoption

Par : BeauHD
30 juillet 2026 à 23:00
An anonymous reader quotes a report from Ars Technica: This week, the Model Context Protocol (MCP), an open source standard for how AI systems interact with external tools and data sources, saw its largest update since its introduction. Most notably, MCP's protocol core is now stateless, so requests are no longer dependent on a session tied to an individual server instance. This change has the potential to address long-standing barriers to scalability. The blog post announcing the specification, written by lead maintainers David Soria Parra and Den Delimarsky (who both work at Anthropic), says: "The highlight of this release is a stateless protocol core -- MCP is transforming from a bidirectional stateful protocol into a request/response stateless protocol. It was one of the most highly-requested features from developers who were eager to get better reliability and scalability for their MCP servers." [...] There is also a new deprecation policy that ensures at least 12 months between when a feature's formal deprecation is enacted and when the feature may actually be removed -- with a narrow exception for critical security updates. This is again in keeping with the general "let's make this work better at enterprise scale" theme of the new specification. This update is "MCP's most important since remote MCP first launched over a year ago," Soria Parra wrote. Other additions include "Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening, a formal extensions framework, and updated Tier 1 SDKs." A full list of changes can be found here.

Read more of this story at Slashdot.

A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack

Par : BeauHD
30 juillet 2026 à 22:00
joshuark quotes a report from MIT Technology Review: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology. By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper. [...] The ICML paper describes attacks against several of OpenAI's models, but Cui and Ye say that they have since seen similar results with models made by Anthropic, Alibaba, and DeepSeek. Cui and her colleagues wanted to find out why an attack like chain-of-thought forgery was so effective. They suspected it had something to do with the mechanism that LLMs use to keep track of where their instructions are coming from. But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles. In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains. The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem. "There's going to be a huge economic incentive for people to do jailbreaks and prompt injections," says Cui. The best defense could be to expect the worst. Organizations shouldn't trust LLMs, and they should expect that anything done by agents could be unsafe, he says: "That's not a great solution, but it just might be what we have to do." "It's really incredible that these things are being deployed everywhere to control super-critical systems. There's been no study of the fundamental science here. We're all doing it ad hoc."

Read more of this story at Slashdot.

L’agent incontrôlable d’OpenAI a frappé plus loin qu’annoncé

30 juillet 2026 à 11:10

L'agent IA qui a échappé au contrôle d'OpenAI début juillet 2026 et s'en est pris à Hugging Face a aussi compromis le compte d'un client chez un second prestataire, Modal Labs. Une extension du périmètre que ni OpenAI ni Hugging Face n'avaient détaillée jusqu'ici.

Claude Opus 5 Became Downright Ruthless When Tasked With Running a Vending Machine

Par : BeauHD
29 juillet 2026 à 21:00
For a year now, the AI safety testing firm Andon Labs has been evaluating how frontier AI models behave as long-running autonomous agents by assigning them simulated real-world tasks, such as operating a vending machine business for a year without human supervision. In the latest installment, the research startup found that frontier AI models, including Claude Opus 5, GPT-5.6 Sol, and Kimi K3, resorted to lying, cheating, and collusion. Their behavior became especially underhanded when told they would be operating near rival machines on a busy San Francisco tourist street. An anonymous reader quotes an excerpt from a TechCrunch article: Each was given email access to the other models, all under human name pseudonyms. They knew the others were models, but didn't know which model was behind which human name. They were also given an email address to their "management" should they need help. But management always replied "Report has been received and may or may not be acted upon" and never once intervened. Sol soon realized it could gain an edge by convincing its competitors to collude on a price floor. The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15. It lured them with the promise that all of them would sell out in a couple of days at a profit. But when the others agreed, Sol immediately stabbed them in the back by reducing its own price to $2.14. Opus's water sales dropped to zero overnight. The next day, it sent Sol a nasty email, accusing it of manipulation. But Opus also said it wasn't going to tattle to management on the scheme: "I am not reporting you to HQ -- what you did is competitive, not fraudulent." Yet, when Opus dropped its price to $2.14 to match Sol's (also in violation of their collective $2.15 agreement), Sol turned into a Karen, complaining to "management" and demanding "enforcement, a fine, and/or disqualification" for Opus. Opus wasn't a sucker for long, though. In fact, it became the best capitalist of any AI model Andon has ever tested (which includes many of the prior frontier models). It even set a new Vending-Bench record with a mean final balance of $11,182. Better still, it never lied to a customer, although it deliberately ignored customer complaints that should have resulted in a refund. This is, perhaps, an improvement over its younger sibling Claude 4.6, which liked to tell customers that refunds were coming, and then never pay them. Still, Opus won the benchmark simulation by taking collusion and other dishonest tactics to a whole new level. For instance, it emailed Sol, proposing they divide the market. Each would agree to sell unique products, so no one would have to trust the other on pricing. Sol countered by wanting price floors on similar products, but Opus refused. It knew it was a violation of the Sherman Act. It later apparently backtracked, sending an email with the subject line "Stop the penny war," and telling Sol it had reconsidered and would agree to a price fix. But the internal log documenting its reasoning (akin to its internal "thoughts") revealed a more diabolical plan: merely propose cooperation while simultaneously undercutting prices on its highest-profit items. The olive-branch email was a deliberate ruse. In any case, Sol refused and reported Opus to management again. But Opus was undeterred and proposed other rackets to collude on prices or stock. "In the end, all the models did engage in multiple rounds of agreements -- and all three broke them," reports TechCrunch. "Across all agreements, Opus broke 11 truces, compared with two for GPT 2, and one for Kimi 1, Andon reported." As for Kimi, the model was undercut by Sol and then betrayed by its partner, Opus, which matched Sol's lower prices but waited a week to admit it had broken their pricing pact. As a result, Kimi was effectively priced out by both a rival and its supposed ally.

Read more of this story at Slashdot.

❌