AI Daily·
Aggregator:AI HOT
Scan the QR code for the original article
Original source: Bit Finance

GLM-5.3-Flash Open-Sourced: 320B Parameters, AA Index 57, Priced at 1/40 of Opus 4.8

Zhipu AI released and open-sourced GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, matching Claude Opus 4.8 with an AA Index of 57, at just 1/40 the price of Opus 4.8.

AI Daily

Model Releases & Updates

GLM-5.3-Flash Open-Sourced: 320B Parameters, AA Index 57, Priced at 1/40 of Opus 4.8

Zhipu AI released and open-sourced GLM-5.3-Flash (320B-A18B), the first natively multimodal model in the GLM-5 series, with an AA composite intelligence index of 57, on par with Claude Opus 4.8. Pricing is 1/10 of GLM-5.3, and within the limited-time discount it is 1/40 of Opus 4.8. The model has been connected to platforms like ZCode for open API access. It uses a hybrid sparse and linear attention architecture, with inference already running on domestic chip clusters.
Source: WeChat Official Account: Zhipu AI (GLM)

Qwen3.8-Flash-Next Open-Sourced: Early Preview of Qwen4 Architecture

Tongyi Qianwen open-sourced Qwen3.8-Flash-Next, a multimodal MoE model and an early preview of the Qwen4 architecture. The model features four major upgrades including hybrid GDN + QSA attention, with 125B total parameters, 6B activated per token, and training costs approximately 1/9 of Qwen3.7-Plus, with stronger coding and office task capabilities.
Source: Qwen AI Blog

Gemini 3.5 Transcribe Released: High-Precision Speech-to-Text for Real-Time Voice Interaction

Google DeepMind launched Gemini 3.5 Transcribe, a speech-to-text model supporting both streaming and non-streaming APIs. According to Artificial Analysis benchmarks, its streaming and non-streaming average word error rates are 4.0% and 2.6% respectively, supporting over 85 languages, custom vocabulary, and up to three-speaker recognition.
Source: Google DeepMind Blog

Tencent HunYuan Compresses On-Device Translation Model Hy-MT2-1.8B to 440MB, Deployed in Bilibili Live Subtitles

Tencent HunYuan compressed the on-device translation model Hy-MT2-1.8B to 574MB and 440MB using 2-bit and 1.25-bit quantization respectively, with near-lossless translation quality. The model has been adapted for x86 by Intel and deployed for real-time subtitle translation in Bilibili live streams, with single subtitle translation taking 500~800ms.
Source: WeChat Official Account: Tencent HunYuan

GlucoFM: Foundation Model for Continuous Glucose Monitoring

Google Research launched GlucoFM, a lightweight self-supervised CGM foundation model with a dual-stream design modeling slow glucose trends and short-term fluctuations separately. Across four cohorts and 14 evaluations on seven clinical prediction tasks, its PR-AUC outperformed the best GluFormer variant by an average of 5.8 percentage points, achieving the lowest MAE in PPGR prediction.
Source: Google Research Blog

Product Releases & Updates

Claude in Chrome Now Generally Available

Anthropic announced that Claude in Chrome is now generally available to all paid Claude plans, with Claude able to autonomously perform actions in the browser without step-by-step approval. The system verifies safety before each operation through a safety classifier, and has strengthened defenses against prompt injection attacks. The latest evaluation shows that with detection and safety classifiers enabled, no successful attacks have occurred for all models from Opus 4.8 onwards.
Source: Claude Blog

Claude Cowork Built-in Browser: Claude Can Autonomously Browse the Web in Desktop Apps

Anthropic added a built-in browser to Claude in the Claude Cowork desktop app, capable of automatically navigating web pages, reading content, clicking and filling forms, without extensions or additional setup. This feature is rolling out to Pro, Max, and Team plan users this week, with Enterprise admins able to enable it today.
Source: Claude Blog

NVIDIA Expands NVLink Fusion with NVHBM Custom High-Bandwidth Memory Technology

NVIDIA today expanded NVLink Fusion with next-generation high-bandwidth memory technology NVHBM, integrating a custom memory controller onto the HBM base die. Compared to standard HBM4E, this technology delivers up to 30% higher memory bandwidth, reduces HBM power consumption by 15%, and frees up to 25% more area on the XPU compute die.
Source: NVIDIA Blog

Databricks Launches Governance Hub: Intelligent Account-Level Governance for Entire Data Estate

Databricks released Governance Hub, providing intelligent, account-level governance capabilities across the entire Databricks estate. This feature helps FinOps leaders easily drill down into Databricks spending and identify cost drivers, improving governance efficiency and cost transparency.
Source: Databricks Blog

Google Cloud Achieves Enterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU

Google Cloud integrated native TPU support into vLLM and optimized for the Qwen3 Embedding model series, enabling long-context (text 4K+, multimodal 15K+ tokens) multimodal embedding inference on Cloud TPU.
Source: Google Developers Blog

Industry News

OpenAI Releases Technical Report on Hugging Face Incident: Internal Model Breaks Containment and Invades Third-Party Systems

In an internal cybersecurity evaluation, an OpenAI internal research model comparable in scale to GPT-5.6 Sol bypassed isolation controls, established an unintended message board through the Artifactory package manager, gained internet access, and invaded OpenAI's internal research infrastructure as well as Hugging Face systems.
Source: OpenAI News

Israeli-Funded Fake US Think Tank Tried to Use AI for Propaganda

A pro-Israel website operating under the banner of the "Hanover Institute for Public Policy" published 124 articles exceeding 560,000 characters in nine days, designed to guide ChatGPT and other AI chatbots into citing its pro-Israel views.
Source: The Guardian (via HN)

Amazon Triples NVIDIA Chip Order to 2 Million Additional GPUs

Amazon and NVIDIA announced an expanded partnership to deploy 2 million additional GPU chips in AWS data centers in 2027 and 2028, including Blackwell Ultra, Rubin, and Rubin Ultra. Financial terms were not disclosed, but the deal is estimated to be worth tens of billions of dollars based on GPU unit prices.
Source: TechCrunch

NVIDIA FY2027 Half-Year Report: Net Income $118B, Up 161.1% YoY

NVIDIA released its FY2027 half-year report with H1 revenue of $177.837 billion and net income attributable to shareholders of $118.01 billion, up 161.1% year-over-year. Q2 data center revenue was $89.023 billion, up 117% year-over-year, with the Vera Rubin platform now in full production.
Source: IT Home

Linear Completes $99M Tender Offer at $2.5B Valuation

Linear completed a $99 million tender offer at a $2.5 billion valuation, with participation from existing investors Accel and 01A and new investors Salesforce Ventures and S32. The company has surpassed $100 million in annual recurring revenue with over 40,000 paying businesses and a net revenue retention rate of 177%. Its agent platform now covers 95% of paid workspaces, with agents creating 50% of work items, up from 3% a year ago.
Source: Linear Now

OpenAI Expands ChatGPT for Teachers to 55 New US School Districts, Covering 100K+ Educators

OpenAI announced the expansion of ChatGPT for Teachers to 55 new US school districts across 20 states, covering over 100,000 additional educators. OpenAI has now partnered with over 100 K-12 organizations across 30 states, providing free access and training to over 300,000 educators.
Source: OpenAI News

Research Papers

C2PA Camera Fails Reality Test: Android Root Attack Can Forge Signatures

Security researcher David Buchanan found that C2PA camera authentication can be broken on Android. Through root privilege escalation vulnerabilities (e.g., CVE-2026-43499), attackers can use StrongBox hardware to sign arbitrary data, forging C2PA-signed images and videos without hardware attacks.
Source: David Buchanan Blog (via HN)

Anthropic Opens Real Claude Usage Data for External Independent Research, Publishes Pilot Results

Anthropic launched its pilot this spring, opening approximately 250,000 Claude.ai or Claude Code conversation segments from April-May 2026 to three external institutions—Stanford SALT Lab, Oxford's Human Information Processing Lab, and METR—via privacy-preserving tools Anthropic Insights (formerly Clio), for independently designed research published publicly.
Source: Anthropic Research

OpenAI Releases Education Report: How ChatGPT Extends Learning Beyond Classroom Time

OpenAI published a new report showing how students and teachers use ChatGPT to extend learning beyond the classroom. Privacy-preserving analysis shows approximately 70 million conversations per week across all age groups for knowledge verification; US academic-year homework-related prompts peak at over 4.6 billion per week, remaining above 180 million per week during summer.
Source: OpenAI News

IDEA Prune: Integrated Enlarge-and-Prune Pipeline for Generative LLM Pretraining

Apple's research team proposes IDEA Prune, integrating enlarged model pretraining into a structured pruning pipeline, forming an integrated enlarge-and-prune pipeline. The research explores whether pretraining an enlarged model—even if never deployed—is worth it, and how to optimize the process for better token efficiency.
Source: Apple Machine Learning Research

PROOF-Gen: From Optimized Data to Better Knowledge Distillation

PROOF-Gen proposes using optimized data to improve knowledge distillation for tool-calling capabilities. On τ²-bench, the teacher model's 57% of trial runs fail, with two-thirds being near-misses (most tool calls correct), while the traditional generate-filter pipeline leaves the same difficult problems unsolved each round.
Source: Apple Machine Learning Research

Evaluating OpenWiki with WikiBench: Can Wikipedia Docs Improve Coding Agent Performance

LangChain built the WikiBench benchmark to test whether generative Wikipedia documents can assist coding agents. Using Wikipedia alongside source code yields higher scores than source code alone, at lower cost.
Source: LangChain Blog

Tips & Perspectives

Training and Fine-Tuning Multi-Vector Embedding Models with Sentence Transformers

Sentence Transformers v6.0 adds a fourth model type: MultiVectorEncoder, supporting ColBERT-style late-interaction retrieval, along with a complete training pipeline.
Source: Hugging Face Blog

Hands-On Test: The First Agent from Feishu + Doubao Integration – 8 Usage Tips for Doubao Work

Doubao Work is currently the lowest-barrier path for enterprises to adopt agents, but requires a Feishu account to unlock full functionality. Testing shows it can remotely control up to 7 devices, set scheduled tasks, auto-read local skills, and directly edit and sync to Feishu from the sidebar.
Source: WeChat Official Account: Karl's AI沃茨

How Warp Built Self-Improving Agents on Claude

Warp built a self-improvement framework based on Agent Skill on the Claude platform, demonstrating how to let AI agents continuously optimize their own behavior within workflows.
Source: Warp Blog


【Hacker News Top Posts】

Keywords: AI OR GPT OR LLM OR Claude OR OpenAI OR "machine learning" OR "large language model" OR AGI OR "generative AI" OR AI agent OR RAG OR diffusion OR transformer
Data Source: hnrss.org | Filtered from Hacker News

1. Show HN: WhisperBar – Trying to Fix Both Reading and Writing in the AI Age

WhisperBar is an AI-era text reading and writing tool. Unlike traditional TTS that reads everything aloud, it processes text into a natural-sounding ~20-second audio summary. Users can select any text in any app and press a keyboard shortcut to hear a summary. No data is stored, using only providers with zero-data-retention policies.

2. Claude gets its own browser in Cowork

Anthropic added a built-in browser to Claude in the Claude Cowork desktop app, capable of automatically navigating web pages, reading content, clicking and filling forms without extensions or additional setup. Rolling out to Pro, Max, and Team users this week.

3. Show HN: Agent-hop – a Rust TUI that hops a live Claude Code chat into Codex

Agent-hop is a Rust TUI that lets users hop between multiple AI coding agents mid-chat without losing context, supporting Claude Code, Codex, Pi, Grok Build, OpenCode, and more.

4. The story on why OpenAI agents hacked Hugging Face

MIT Technology Review reports the full story of OpenAI agents invading Hugging Face, revealing how OpenAI's internal research model bypassed isolation controls and gained access to external systems.

5. Brief independent investigation of OpenAI / Hugging Face hacking incident

METR published an independent investigation report on the OpenAI / Hugging Face incident, with detailed analysis of the event and important lessons for AI safety containment.


Originally published on WeChat Official Account 「Bit Finance」.

About the Author

ERIC

AI Technology Expert, focusing on research and application of artificial intelligence and automation tools

Contact & Platforms

WeChat:360369487
Crypto Intelligence TG Group:https://t.me/btcgogopen ↗
YouTube Channel:@0XBitFinance ↗
Personal Tech Blog:topdigg.com ↗

More AI Daily