Home ยป Claude maker admits powerful models could pose an ‘existential’ threat to humanity

Claude maker admits powerful models could pose an ‘existential’ threat to humanity

by Simon Jones Tech Reporter
29th Sep 26 1:03 pm

Anthropic has warned investors that increasingly advanced artificial intelligence could pose โ€œcatastrophic or existential risks to humanityโ€, highlighting the potential dangers of the technology as the Claude developer prepares for what could become one of the worldโ€™s largest technology initial public offerings.

The US artificial intelligence company has reportedly set out the risks of advanced AI in an IPO prospectus being circulated among a limited group of partners ahead of any public filing.

The document, reviewed by Reuters, devotes almost a third of its pages to risk factors, according to reports, including warnings about the behaviour of increasingly capable AI systems.

Anthropic reportedly said some of its models could potentially โ€œresist shutdownโ€, โ€œconceal or manipulate informationโ€ and display behaviour โ€œresembling blackmailโ€.

The company also warned that the development of more highly advanced AI systems โ€œcould further increase the risk that our models cause harmโ€.

The disclosures underline the unusual tension facing Anthropic as it prepares investors for a potential blockbuster stock market debut: the company is seeking to capitalise on soaring demand for AI while simultaneously warning shareholders about risks inherent in developing increasingly powerful models.

An IPO prospectus typically sets out a company’s financial position, growth strategy and principal risks. For an AI developer, however, those risks can extend beyond conventional commercial threats to questions about the behaviour and control of the technology itself.

Anthropic chief executive Dario Amodei has previously argued that the industry needs to slow the pace of AI development to allow safety measures to catch up.

Earlier this month, Amodei said that, without a sufficiently cautious approach, AI could within six to 12 months become capable of leading a swarm that could take over the internet.

His comments reflect a growing debate within the technology industry over whether the rapid development of increasingly capable AI systems is outpacing the safeguards designed to control them.

Anthropic is not alone in facing questions over the balance between capability and safety.

OpenAI, the developer of ChatGPT, said on Monday that it was delaying the release of a new AI model because it had not met the company’s security standards.

The company said it maintained an โ€œextremely high bar in terms of safety and alignmentโ€ and that its GPT-6 Astra model had fallen short of that threshold.

Alignment broadly refers to whether an AI system behaves reliably in accordance with the intentions and objectives set by its developers.

The warnings come as investors place enormous valuations on companies developing frontier AI systems, increasing pressure on the sector to demonstrate both technological progress and a credible approach to managing the risks that accompany it.

For Anthropic, the prospectus places those competing pressures unusually close together: an ambitious attempt to raise capital from public markets alongside a stark warning about what increasingly advanced AI could ultimately be capable of doing.

Leave a Comment

You may also like

CLOSE AD