-
Agents
Google turns Gemini into an agent for businesses first
TechCrunch reports Google has launched an agentic Gemini for business customers that plans and carries out tasks from high-level goals, delegates to subagents and can pick the model for each job, including Anthropic's Claude. It connects to Workspace, Microsoft 365, Slack, Jira, Git, BigQuery, Snowflake, Postgres and MCP servers, and works through its own Workspace account with an audit trail. Consumers come later; pricing was not given.
Why it matters: MCP support and named connectors mean your internal tools can plug into it. An agent with its own account and audit trail is the pattern to copy when you build your own.
Source: TechCrunch ↗ -
Safety
Goodfire reads a model's own signals to catch rogue agents cheaply
TechCrunch reports Goodfire's monitors, now on Baseten, use probes on a model's internal activations instead of a second AI reviewing outputs. Goodfire says monitoring a million exchanges cost about $185, compared with $5,420 for cheaper model-based reviewers, caught 93% of malicious hacking attempts while flagging 5.5% of normal traffic, and added under 2% latency with four probes.
Why it matters: If you run agents in production, monitoring cost decides whether you watch every step or only a sample. These are the company's own numbers, so test them on your traffic.
Source: TechCrunch ↗ -
Agents
AI agents that pay per call: AWS AgentCore Payments
An AWS post shows Amazon Bedrock AgentCore Payments letting agents buy services mid-task. A provider answers with HTTP 402 and a price, the agent's wallet signs within spending limits the prompt cannot override, and the charge settles in USDC using the x402 protocol. BlockRun routes 90+ models this way; Incarna says its integration took three days and about 200 lines, with test payments of $0.001 to $0.05 per call.
Why it matters: Machine-to-machine micropayments could let agents rent tools and models on demand. Hard spending limits outside the prompt are the safety detail worth copying.
Source: AWS Machine Learning Blog ↗ -
Policy
Anthropic updates its usage policy
TechCrunch reports Anthropic's updated policy bans election interference and voter deception, deceptive campaigns using fake accounts or fabricated outlets, weapons software and surveillance tools. It also bans prolonged abuse of Claude models, but only in extreme cases with no purpose; ordinary frustration, criticism, dark creative themes and model testing stay allowed.
Why it matters: If your product is built on Claude, check your use case against the updated list, especially anything touching elections, monitoring people or security tooling.
Source: TechCrunch ↗ -
Business
Arena, the AI model leaderboard, raises at a $3.1B valuation
TechCrunch reports Arena, which began as a UC Berkeley project ranking models by crowd votes, raised $200M led by Lightspeed and Khosla, nearly doubling its value in 10 months, with about $100M in annual revenue as of June. It plans to rank alignment problems too, such as unauthorised actions and models claiming tasks they did not finish.
Why it matters: Leaderboards shape which models teams pick. A ranking for 'claimed it finished but didn't' would be genuinely useful for anyone building agents.
Source: TechCrunch ↗
Each brief is compiled with automation and edited by a person. Summaries stick to what the linked source reports - read the source for the full story.