Category: AI

  • SSPAI Morning Brief: Google IPv6 Traffic Surpasses 50% Milestone, Claude Opus 4.7 Token Costs Rise Significantly

    SSPAI Morning Brief: Google IPv6 Traffic Surpasses 50% Milestone, Claude Opus 4.7 Token Costs Rise Significantly

    Morning Brief

    1. China’s low-altitude authority responds to drone “takeoff difficulties”
    2. Seven e-commerce platforms fined RMB 3.597 billion over ghost kitchens
    3. Kindle for PC to shut down, fewer options for DRM removal
    4. Claude Opus 4.7 shows significantly higher token usage
    5. Google’s IPv6 traffic share surpasses 50% for the first time
    6. Digital and smart products become key drivers in trade-in consumption
    7. News Worth a Quick Look

    China’s low-altitude authority responds to drone “takeoff difficulties”

    According to Caixin, at a State Council Information Office press conference on April 17, Zheng Jian, Director of the Low-Altitude Economy Development Department under the National Development and Reform Commission, responded that in light of difficulties in obtaining approvals for drone flights, the department is working with relevant agencies to promote effective local practices such as “scan-to-fly,” aiming to improve the efficiency of flight plan approvals. This marks the first official response to widespread user complaints since tighter drone regulations were introduced.

    Under current regulations, drones must be registered with real-name identification on either the Civil Aviation Administration’s Unmanned Aircraft Management Platform (UOM) or the Public Security Drone Management Platform. Each flight must then be submitted for prior approval, and only after authorization can the drone be operated—otherwise, users may face fines or detention from public security authorities. In practice, these rules have proven difficult for many drone users to adapt to.

    The “scan-to-fly” approach mentioned in the response originates from Sichuan Province, where local authorities have experimented with easing flight demand. After completing an online filing via a mini program, users can fly within designated permissible airspace without having to visit local police stations in person. In addition to Sichuan, Shanghai introduced a similar service module in early February via its “Suishenban” government app, designating three drone flight experience zones within controlled urban airspace.

    Industry insiders note that China’s low-altitude airspace is primarily divided between military and civil aviation authorities, with the military holding final authority over approvals. As a result, local governments have limited involvement. A representative from a domestic drone company pointed out that while flight application volumes are large, approval capacity remains constrained, and suggested that granting local governments more control over suitable and partially restricted airspace, along with improving automation in approvals, could be a solution.

    Previously, on November 8, 2023, the Civil Aviation Administration released a draft of the Airspace Management Regulations, proposing the establishment of hierarchical air traffic management coordination bodies responsible for airspace governance. Industry experts believe the legislation could be finalized by 2026.


    Seven e-commerce platforms fined RMB 3.597 billion over ghost kitchens

    According to Caixin, on April 17, China’s State Administration for Market Regulation (SAMR) announced administrative penalties against seven e-commerce platforms—including Pinduoduo, Meituan, JD.com, Ele.me (now Taobao Flash Delivery), Douyin, Taobao, and Tmall—for violations related to “ghost kitchen” operations. The platforms were ordered to rectify illegal practices, suspend the onboarding of new bakery merchants for periods ranging from three to nine months, and pay combined fines and confiscations totaling RMB 3.597 billion. Additionally, legal representatives and food safety directors of the seven companies were fined a total of RMB 19.6874 million. Pinduoduo faced the largest penalty at RMB 1.522 billion, followed by Meituan, JD.com, and Ele.me at RMB 746 million, RMB 635 million, and RMB 558 million respectively; Douyin, Taobao, and Tmall were fined RMB 56.89 million, RMB 46.97 million, and RMB 31.74 million.

    “Ghost kitchens” refer to vendors that lack physical dining spaces or have already ceased operations but continue to operate on delivery platforms. Order-routing platforms allow merchants to transfer orders to other food businesses for fulfillment.

    According to SAMR, the platforms failed to properly verify the licenses of food vendors, neglected their legal obligations for qualification review, and entered into agreements with order-routing platforms while knowing—or having reason to know—that such practices infringed on consumer rights, yet failed to take necessary measures. Legal representatives and food safety directors also failed to fully perform their responsibilities.

    Additionally, SAMR’s administrative penalty document against Shanghai Xunmeng Information Technology Co., Ltd. (Pinduoduo’s operating entity) revealed that during the investigation, the company repeatedly refused to provide materials without justification, submitted false information, and even obstructed enforcement through confrontational tactics. Previous reports indicated that on December 3, multiple Pinduoduo employees clashed with regulators during the investigation, leading to the dismissal of several staff in its government relations department and triggering public labor disputes.

    Earlier, in November 2025, SAMR guided eight major online food trading platforms—including JD.com, Meituan, Pinduoduo, Douyin e-commerce, Xiaohongshu, Taobao, WeChat Shops, and Kuaishou e-commerce—to jointly sign a self-regulatory agreement on food safety management. On February 26, 2026, SAMR announced that new regulations on food safety responsibilities for online catering service operators will take effect on June 1, requiring delivery platforms to implement real-name registration for vendors and conduct on-site verification to ensure license information matches actual conditions.


    Kindle for PC to shut down, fewer options for DRM removal

    According to Good e-Reader, Amazon recently notified users via pop-up that the current Kindle for PC desktop client will be discontinued on June 30. After that date, the software will no longer function. Amazon confirmed it is developing a new Kindle app for PC, compatible only with Windows 11 and available exclusively through the Microsoft Store.

    The original Kindle for PC client was launched in 2009 and has long been used by users to download e-books locally and remove DRM (digital rights management) protections. In recent years, Amazon has repeatedly urged users to upgrade the client and restricted access for older versions in order to patch vulnerabilities and combat piracy. In 2023, Amazon had already discontinued the older Kindle for Mac, replacing it with a version distributed solely via the Mac App Store.

    Compared to standalone installers, app store–distributed versions are generally harder to bypass technically. By shifting its PC client entirely to the Microsoft Store, Amazon aims to further tighten control over Kindle and meet publishers’ requirements for preventing e-book piracy.


    Claude Opus 4.7 shows significantly higher token usage

    According to the official migration guide for Claude Opus 4.7, the model adopts a new tokenizer, which can increase token usage by up to 35% compared to previous versions.

    However, user testing suggests this may be a conservative estimate. In one test, TypeScript code saw token usage increase by 1.36×, while English technical documentation rose by 1.47×. In a debugging session involving 80 conversational turns, costs increased from approximately $6.65 to nearly $8.80. Another dataset of over 500 samples showed an average token usage increase of 38.6%. In contrast, token consumption for non-Latin scripts such as Chinese, Japanese, and Korean (CJK) remained largely unchanged.

    The new tokenizer breaks English text and code into smaller segments, reducing the number of characters per token. Anthropic claims this improves task accuracy, but many users have already complained that usage limits are now depleted even faster.


    Google’s IPv6 traffic share surpasses 50% for the first time

    According to data released by Google, on March 28, 2026, global user traffic accessing its services via IPv6 reached 50.1%, up from 46.33% during the same period last year—marking the first time the metric has exceeded the halfway point.

    However, data from other internet infrastructure organizations suggests that IPv6 has not yet fully taken the lead. Monitoring by Cloudflare indicates that IPv6 currently accounts for only 40.1% of global HTTP request sources; data from the Asia-Pacific Network Information Centre (APNIC) shows that the proportion of networks with IPv6 capability stands at around 43.13%.

    IPv6 was introduced to address the exhaustion of IPv4 addresses. IPv4 can provide only about 4.3 billion IP addresses, while IPv6, using 128-bit addressing, offers an almost limitless address space. However, because the new protocol did not bring many disruptive new features, and due to the widespread use of Network Address Translation (NAT)—which allows a large number of devices to share a single public IPv4 address—many organizations have relied on NAT to mitigate address shortages. As a result, the global adoption of IPv6 has long lagged behind expectations.

    The pace of IPv6 adoption also varies significantly across regions. Due to relatively lenient allocation mechanisms in the early days of the internet, Europe and the United States secured large amounts of IPv4 resources. In contrast, populous countries such as China and India received far fewer IPv4 addresses, prompting earlier and more aggressive deployment of IPv6. According to APNIC data, 29 countries in the Asia-Pacific region, including China, surpassed the 50% IPv6 adoption threshold in 2025.


    Digital and smart products become key drivers in trade-in consumption

    According to Xinhua News Agency, based on data from the Ministry of Commerce’s national system for home appliance trade-ins and digital and smart product purchases, as of April 16, purchases of digital and smart products reached 42.433 million units, up 31.7% year-on-year, with total sales of RMB 126.153 billion, up 36.4%. Among these, mobile phone sales accounted for more than 80%.

    Since the beginning of this year, subsidies for new purchases of digital products—such as smartphones—have been expanded and upgraded to cover a broader range of digital and smart products, with smart glasses included for the first time. This has driven rapid growth in retail sales of communication equipment. By mid-April, 16 domestic smart glasses brands had participated in the subsidy program, boosting sales volume and revenue of key enterprises by 42.4% and 46.8% year-on-year, respectively.

    It is reported that under the combined effect of home appliance trade-in programs and digital product subsidies, retail sales of communication equipment above the designated size reached RMB 284 billion from January to March, up 20.8% year-on-year—18.4 percentage points higher than the overall growth rate of total retail sales of consumer goods, ranking first among 16 product categories.


    News Worth a Quick Look

    • According to Nikkei, the three major DRAM suppliers—Samsung Electronics, SK Hynix, and Micron Technology—can currently meet only about 60% of total market demand. By mid-2026, the share of memory costs in low-end smartphone manufacturing is expected to double from 20% to nearly 40%.
    • Recently, NVIDIA CEO Jensen Huang stated on the Dwarkesh Podcast that he opposes stricter U.S. export controls on chip equipment to China, arguing that China’s abundant energy resources and manufacturing capabilities could allow it to achieve large-scale computing power through system scaling. He warned that strict export controls could inadvertently push China to build a strong, independent technology stack, potentially causing the U.S. to lose access to the world’s second-largest market and weakening the global influence of U.S. technology standards.
    • According to Microsoft’s release notes, in Windows 11 Insider Preview Build 26300.8170, the FAT32 partition size limit has been increased from 32GB to 2TB, although the change is currently only accessible via command-line operations. The 32GB limit had long been an artificial restriction imposed by Microsoft and remained unchanged for decades.
    • According to MacWorld, this year’s iPhone 18 Pro is expected to feature “Dark Cherry” as its signature color, while the foldable iPhone is rumored to measure just 4.7 mm thick when unfolded, with currently tested color options including silver-white and indigo.
  • SSPAI Morning Brief: Canva Launches AI 2.0 Productivity Platform, OpenAI Upgrades Codex with Advanced AI Agent Capabilities

    SSPAI Morning Brief: Canva Launches AI 2.0 Productivity Platform, OpenAI Upgrades Codex with Advanced AI Agent Capabilities

    Morning Brief

    1. Canva AI 2.0 launched
    2. Anthropic releases Claude Opus 4.7
    3. OpenAI upgrades Codex with multiple practical features
    4. DJI launches Osmo Pocket 4 gimbal camera
    5. Amazon introduces the thinnest Fire TV Stick HD
    6. Adobe releases Firefly AI assistant
    7. Tencent unveils Hunyuan 3D World Model 2.0
    8. Mastercard enables cross-border Apple Pay support for Chinese cardholders
    9. Apple Wallet now supports NFC transit cards via Alipay
    10. News Worth a Quick Look

    Canva AI 2.0 launched

    On April 16, Canva unveiled Canva AI 2.0, announcing its shift toward an “integrated productivity system.”

    This update introduces a new underlying architecture, with core capabilities including: conversational design, allowing users to generate fully editable designs through natural language or voice; agent orchestration, automatically coordinating multiple tools to accomplish complex tasks (such as generating full multi-channel marketing campaigns); intelligent object editing, enabling precise adjustments to individual elements without affecting the overall design; and persistent memory, which learns user workflows and automatically applies brand styles.

    In terms of intelligent workflows, Canva AI 2.0 expands its application scenarios into everyday office tasks. The new system integrates with commonly used tools such as Slack, Gmail, Zoom, and Google Drive, enabling direct extraction of key points from audio, video, or chat logs to generate documents. It also introduces automated planning with background offline operation, web-wide research capabilities, Canva Code 2.0 with support for importing and editing HTML files, and Sheets AI, which can generate structured tables with a single prompt.

    Canva AI 2.0 will begin rolling out in limited regions worldwide for early access. Source


    Anthropic releases Claude Opus 4.7

    Anthropic has released the Claude Opus 4.7 model. Based on Opus 4.6, the new version focuses on improving performance in complex software engineering tasks, reducing reliance on manual intervention. It also enhances capabilities in image analysis, instruction following, and the generation of documents and presentations, and is considered to demonstrate stronger creativity.

    However, Anthropic noted that Opus 4.7 does not push the company’s capability boundaries further, and its overall performance still falls short of the previously released Claude Mythos Preview model, which outperforms it across multiple benchmarks.

    For safety reasons, Anthropic is currently providing Claude Mythos Preview only to selected partners, including NVIDIA, JPMorgan Chase, Google, Apple, and Microsoft. Opus 4.7, as a public model, is being used to test new cybersecurity defense mechanisms and includes additional safety restrictions. The company has also introduced a Cyber Verification Program, allowing security researchers to conduct vulnerability research under specific conditions.

    Pricing remains unchanged at $5 per million input tokens and $25 per million output tokens. However, Anthropic noted that Opus 4.7 features an upgraded tokenizer, meaning the same text may consume 1.0 to 1.35 times more tokens than before; under high-reasoning settings, especially in multi-turn interactions, the model produces deeper reasoning and longer outputs. Source


    OpenAI upgrades Codex with multiple practical features

    On April 17, OpenAI announced an upgrade to Codex, further enhancing its agent-based development capabilities. The new Codex can directly operate desktop applications, execute tasks in the background, and allow multiple agents to work in parallel, making it suitable for front-end debugging, application testing, and development workflows without APIs. It also introduces an in-app browser, image generation and editing capabilities, memory features, and support for plugins such as GitLab Issues, Atlassian Rovo, and Microsoft Suite.

    According to OpenAI, the new “computer operation” capability enables agents to perform actions such as clicking and typing on the user’s computer without interfering with other applications. The built-in browser supports web browsing and page annotation, making it easier for developers to provide precise instructions. The image capabilities, powered by gpt-image-1.5, can be used for generating and iterating on product prototypes, UI designs, and game assets.

    In addition to feature expansion, Codex is beginning to introduce “memory,” which can store user preferences, historical modifications, and common workflows to improve efficiency in future tasks. These personalization features will be gradually rolled out.

    These updates have begun rolling out to users of the Codex desktop app within ChatGPT. At the initial stage, desktop operation features are only available on macOS. Support for Enterprise, Edu, as well as users in the EU and UK, will be added in future updates. Source


    DJI launches Osmo Pocket 4 gimbal camera

    On April 16, DJI introduced the new-generation pocket gimbal camera, Osmo Pocket 4.

    The device features a new 1-inch CMOS sensor combined with an ƒ/2.0 large aperture, achieving 14 stops of dynamic range and supporting 10-bit D-Log professional color mode. In terms of video capabilities, it supports up to 4K recording and 240 fps slow motion, while also adding spatial audio recording and zoom-based audio pickup. It supports direct connection to DJI microphone transmitters via OsmoAudio, enabling four-channel audio recording.

    On the system side, Osmo Pocket 4 is equipped with the Intelligent Tracking 7.0 system, supporting up to 4× distance tracking. The autofocus system has been upgraded with new “Subject Lock Tracking” and “Registered Subject Priority” modes, along with gesture control support. For imaging, the device includes optimized skin tone rendering and adjustable beauty filters. It also supports external fill lights with adjustable color temperature and brightness.

    The product is available in Standard and Creator Combo versions, priced at RMB 2,999 and RMB 3,799 respectively. The Creator Combo additionally includes a DJI Mic 3 transmitter, fill light, wide-angle lens, and mini tripod. DJI has also introduced the DJI Care Refresh service, priced at RMB 219 for one year (including two replacement claims) and RMB 349 for two years (including four replacement claims). Source

    At the same time, DJI announced plans to release a dual-camera version, Pocket 4P. Source


    Amazon introduces the thinnest Fire TV Stick HD

    On April 15, Amazon officially launched the new-generation Fire TV Stick HD streaming device, emphasizing a slimmer design and improved performance. The device supports 1080p video output, is priced at $34.99, and is now available for pre-order, with shipments expected to begin on April 29 across multiple markets.

    According to official information, the new Fire TV Stick HD features a significantly optimized design, with overall thickness reduced by about 30% compared to the previous generation. Amazon describes it as the “thinnest streaming device ever.” It can be powered directly via a TV’s USB port, or through a USB-C cable with a power adapter for TVs without USB ports.

    In terms of performance, the new device delivers an average improvement of over 30% compared to the previous HD model, with optimizations in startup speed and app loading times. It also supports Wi-Fi 6 and Bluetooth 5.3 standards. On the software side, the device runs Vega OS, based on the Linux kernel, and integrates the Alexa+ voice assistant for more natural voice interactions. The interface adopts Amazon’s updated content layout, categorizing movies, live TV, sports, and news into distinct sections.

    The device will initially launch in the United States, United Kingdom, Canada, Mexico, Japan, Australia, and New Zealand, with expansion to additional European markets planned. Source

    Adobe releases Firefly AI assistant

    On April 15, Adobe announced the Firefly AI Assistant. According to the company, this is an AI creative assistant with agent capabilities, capable of executing multi-step tasks across various Creative Cloud applications.

    Unlike traditional AI tools that rely on step-by-step instructions, Firefly AI Assistant adopts a goal-oriented task execution model. After users describe their needs in natural language, the system can automatically plan workflows and perform multi-step operations across applications such as Adobe Photoshop, Adobe Premiere Pro, Adobe Lightroom, Adobe Illustrator, and Adobe Express.

    Adobe stated that the assistant will provide a unified conversational interface to manage task context and synchronize generated results across applications. It also includes preset creative functions, such as completing image style adjustments with a single prompt, simplifying common workflows.

    Firefly AI Assistant also offers a degree of personalization, gradually adapting output styles based on user history. Additionally, it integrates Frame.io’s review features, enabling users to organize project feedback and share it with collaborators, while allowing external participants to submit actionable revision suggestions.

    Firefly AI Assistant has not yet been officially released. Adobe plans to roll out a public beta to test users in the coming weeks. Source


    Tencent unveils Hunyuan 3D World Model 2.0

    On April 16, Tencent announced the official release and open-sourcing of the Hunyuan 3D World Model 2.0 (HY-World 2.0). This model is a multimodal world model capable of automatically generating, reconstructing, and simulating 3D scenes based on text, images, and video inputs.

    According to Tencent, HY-World 2.0 supports exporting multiple 3D asset formats, including Mesh, 3D Gaussian Splatting (3DGS), and point clouds. It can also integrate with existing game development workflows, enabling rapid generation of maps and level prototypes. In terms of performance, Tencent claims improvements in scene completeness (such as generating object sides and backs) and fidelity to input content.

    Additionally, the model adopts a hybrid representation combining 3DGS and Mesh, allowing generated scenes to support realistic collision-based interactions. The upgraded HY-Pano 2.0 model introduces an end-to-end implicit learning approach, capable of converting standard images into 360-degree panoramas without requiring camera parameters.

    For training, Tencent stated that the model combines real panoramic images with synthetic data generated via Unreal Engine (UE) to enhance quality and generalization. The model and related technical materials have now been open-sourced on GitHub. Source


    Mastercard enables cross-border Apple Pay support for Chinese cardholders

    On April 16, Mastercard announced progress in its collaboration with NetsUnion Clearing Corporation, a domestic bank card clearing institution in China: Mastercard-branded bank cards issued within mainland China now support cross-border payments via Apple Pay.

    Currently, Mastercard single-branded or dual-branded credit cards issued by Bank of China, Agricultural Bank of China, China CITIC Bank, and Shanghai Pudong Development Bank, as well as Mastercard debit cards issued by China CITIC Bank, are supported for Apple Pay binding. In terms of usage, users can add their bank cards to Apple Wallet (Apple Pay) through the latest version of their bank’s app, or directly tap the “+” button in the iPhone Wallet app and select a credit card to complete the setup. Once added, users can make payments using Apple Pay on iPhone, Apple Watch, or iPad. Source


    Apple Wallet now supports NFC transit cards via Alipay

    On April 14, Alipay announced that Apple Wallet now supports adding NFC transit cards via Alipay, covering cities including Beijing, Shanghai, Nanjing, Changsha, Xiamen, Suzhou, Kunming, Qingdao, Shijiazhuang, and Tianjin. Users can tap the “+” button in Apple Wallet to add a transit card, select “Transit Card” – “Alipay,” confirm opening Alipay, and complete the setup. The corresponding city transit card will then appear in Apple Wallet. Source

    News Worth a Quick Look

    • Google Quick Share has recently encountered a series of cross-platform transfer issues.
      • Some Pixel 10 users reported that upon opening the Quick Share interface, their devices immediately disconnect from Wi-Fi, making it impossible to display available networks. According to user feedback, uninstalling related extension updates can temporarily alleviate the issue. The problem has also appeared on Google Issue Tracker, but the related entry was quickly closed, and users were directed to continue reporting through official support forums. As of now, no clear timeline for a fix has been announced. Source
      • Samsung Galaxy users have reported that when transferring photos to iPhone via Quick Share, EXIF metadata such as location information is not fully preserved. A moderator on Samsung’s forum has confirmed the issue and stated that a fix is currently in development, expected to be resolved in a future software update. Source
    • Canon has introduced new additions to its Cinema Servo lens lineup: the CN30×40 IAS J/R1 and CN30×40 IAS J/P1, equipped with RF and PL mounts respectively. While maintaining portability, both lenses achieve the longest focal length in the series, offering 1200 mm ultra-telephoto reach and up to 30× optical zoom, covering a range from 40 mm to 1200 mm. With the built-in 1.5× extender enabled, the focal length can be extended to 1800 mm, and the lenses support cameras equipped with full-frame 35 mm sensors. In terms of features, the RF mount version supports Dual Pixel CMOS autofocus and focus guide functions, reducing operational complexity; when paired with the EOS C400 camera, it also supports automatic exposure ramping compensation to minimize brightness shifts during zooming. Additionally, the lenses feature focus breathing correction and support virtual production systems, making them suitable for various professional filmmaking scenarios. Both models adopt newly developed drive units and include multifunction USB Type-C interfaces to enhance control and expandability. The new lenses are scheduled for release in late September 2026. Source
  • SSPAI Morning Brief: OpenAI Expands GPT-5.4-Cyber Access for Cybersecurity Defense, Microsoft Launches Cost-Efficient MAI-Image-2-Efficient Model

    SSPAI Morning Brief: OpenAI Expands GPT-5.4-Cyber Access for Cybersecurity Defense, Microsoft Launches Cost-Efficient MAI-Image-2-Efficient Model

    Morning Brief

    1. Keychron launches the lightweight mouse G3
    2. OpenAI introduces the cybersecurity-focused model GPT-5.4-Cyber
    3. Microsoft unveils the MAI-Image-2-Efficient model
    4. Sony announces adjustments to Bravia TV features

    Keychron launches the lightweight mouse G3

    On April 15, Keychron introduced the new tri-mode mouse G3 under its brand. The model features an ultra-lightweight design with dimensions of 120 × 63 × 38.3 mm, and is available in two versions: ABS with semi-transparent PC and ABS with carbon fiber, both weighing just 44g.

    Keychron G3 is equipped with a Realtek RTL8762G main controller, a PixArt PAW3950 optical sensor, Huano 120M high-durability micro switches, and a plastic scroll wheel. It supports an 8kHz polling rate in both wired USB-C and wireless 2.4GHz modes, with latency as low as 0.41ms; it also includes a built-in 500mAh battery offering up to 160 hours of battery life.

    The standard version of the Keychron G3 is priced at $84.99, while the carbon fiber version is priced at $109.99. Source

    Product appearance images, sourced from the original article

    OpenAI introduces the cybersecurity-focused model GPT-5.4-Cyber

    On April 14, the OpenAI team announced an expansion of its Trusted Access for Cybersecurity (TAC) program, opening access to GPT-5.4-Cyber to thousands of cybersecurity experts and hundreds of teams. The model is based on GPT-5.4 and has been specifically fine-tuned for cybersecurity defense scenarios, with elevated network permission capabilities.

    In terms of access mechanisms, to ensure that all legitimate defenders—including those protecting critical infrastructure—can obtain advanced defensive capabilities, OpenAI implements objective standards such as strong identity verification, avoiding subjective decisions on access rights. Due to the model’s broader permissions, OpenAI is adopting a limited and iterative deployment strategy, making it available only to vetted security vendors and researchers. Source


    Microsoft unveils the MAI-Image-2-Efficient model

    On April 15, Microsoft announced the launch of MAI-Image-2-Efficient, a low-cost, highefficiency text-to-image model. This model is a faster, more affordable version of its flagship text-to-image system, which Microsoft says delivers production-ready quality at nearly half the cost.

    Microsoft describes it as its “best text-to-image model,” capable of generating “photorealistic and expressive” images, while also reliably rendering text within images. It excels at producing product-style images and UI prototypes, largely due to its strong handling of short text such as titles and labels, and its seamless integration into interactive workflows. Pricing is set at $5 per million text input tokens and $19.50 per million image output tokens. The model is currently available via Microsoft Foundry and the MAI Playground. Source


    Sony announces adjustments to Bravia TV features

    Recently, Sony announced plans to scale back certain features of its Bravia smart TVs starting in late May 2026. This adjustment will affect models from 2023 to 2025 and will directly impact users relying on antennas or set-top boxes.

    Affected models include the 2025 Bravia 8 II (XR80M2) and Bravia 5 (XR50), the 2024 Bravia 9 (XR90), Bravia 8 (XR80), Bravia 7 (XR70), as well as the 2023 Bravia A95L series. The changes mainly impact users who rely on over-the-air broadcast signals via antennas: the system will no longer guarantee program information display for all channels, the program list will be limited to “recently watched channels,” and channel icons along with program preview thumbnails will be removed, reducing intuitive visual guidance.

    For set-top box users, Sony will remove the existing dedicated menu and replace it with a simplified “Control Menu” offering fewer features. Additionally, the TV Guide integrated into the Google TV operating system will no longer display preview images for antenna channels, free ad-supported streaming TV (FAST) channels, and some live services. Source


    News Worth a Quick Look

    • On April 15, third-party Android app store Aptoide filed a lawsuit against Google, alleging that thecompany has violated U.S. antitrust laws by monopolizing app distribution and payment processing, effectively excluding competing Android app stores. Aptoide claims it offers lower commissions for developers and reduced costs for users, yet has suffered irreparable harm. According to the complaint, Google prevents competitors from securing exclusive content from top developers and pressures developers to adopt Google Play and other “essential” services. The case has been filed in a federal court in San Francisco, seeking an injunction to halt Google’s alleged anti-competitive practices, along with unspecified treble damages. Source
  • SSPAI Morning Brief: Google Launches AI Skills in Chrome as WordPress Security Breach and SEO Crackdown Intensify

    SSPAI Morning Brief: Google Launches AI Skills in Chrome as WordPress Security Breach and SEO Crackdown Intensify

    Morning Brief

    1. Bambu Lab X2D officially released
    2. Microsoft raises prices across the entire Surface lineup again
    3. Sony unveils the Inzone M10S II gaming monitor and new audio accessories
    4. Samsung announces 2026 Micro RGB TV lineup
    5. Blackmagic releases DaVinci Resolve 21 with support for still photo editing
    6. Google Chrome introduces Skills feature
    7. Google announces crackdown on back button hijacking
    8. Numerous WordPress plugins found with backdoor injections
    9. Kingsoft Antivirus and 360 Security Guard exposed for critical kernel driver vulnerabilities
    10. News Worth a Quick Look

    Bambu Lab X2D officially released

    On April 14, Bambu Lab officially unveiled the X2D. The Bambu Lab X2D features a lighter and more efficient dual-nozzle mechanical structure, along with dual air-intake cooling, active chamber temperature control, and Bambu Lab’s self-developed permanent magnet synchronous servo motor system. These upgrades significantly improve the stability of high-flow extrusion, resulting in more stable overhangs and smoother bridging for complex structures, while also ensuring the strength and flatness of large engineering models. Prints made with Bambu Lab’s basic PLA and PETG materials using the X2D have received UL 2904 indoor air quality certification. The printer also supports AI-powered pre-print inspection and real-time monitoring during printing.

    The standalone X2D is priced at RMB 3,999, while the multi-color combo is priced at RMB 5,499, with eligibility for national subsidies. Source


    Microsoft raises prices across the entire Surface lineup again

    On April 15, Microsoft announced price increases for its Surface laptops and tablets. The adjustment is driven by rising costs associated with increased demand for memory and related components due to generative AI. The starting price of the 15-inch Surface Laptop 7 has increased from last year’s adjusted price of $1,500 to $1,600 (it was originally priced at $1,300 at launch in 2024). The top-tier configuration, featuring a Snapdragon X Elite processor, 64GB of RAM, and a 1TB SSD, now costs $3,650. The Surface Pro lineup has also been adjusted: the 12-inch version’s starting price has risen from $800 to $1,050, while the 13-inch version has increased from its launch price of $1,000 to $1,500. Source


    Sony unveils the Inzone M10S II gaming monitor and new audio accessories

    On April 15, Sony announced the Inzone M10S II gaming monitor along with a lineup of audio peripherals. The Inzone M10S II features a 24.5-inch LG fourth-generation Tandem WOLED panel, supporting a 540Hz refresh rate at 2K resolution, or an ultra-fast 720Hz mode when resolution is lowered to 720p. It offers a response time as low as 0.02ms, and includes a new motion blur reduction algorithm with integrated Black Frame Insertion (BFI), strong anti-glare performance, and an adjustable stand with a tilt range from -5 to 35 degrees.

    The Inzone H6 Air open-back wired headset is developed based on the MDR-MV1 reference headphones, weighs 199 grams, and comes with a USB-C adapter supporting virtual 7.1 surround sound and 360-degree spatial audio. Sony also introduced a Glass Purple version of its Inzone true wireless earbuds and Fnatic co-branded accessories. The Inzone H6 Air is priced at $200, while the M10S II monitor is priced at $1,100, with availability expected later this year. Source


    Samsung announces 2026 Micro RGB TV lineup

    On April 14, Samsung unveiled its 2026 Micro RGB TV series, including the R95H and R85H product lines. The entire lineup uses 4K Micro RGB display technology, featuring minimized color bleed and enhanced color accuracy through red, green, and blue LEDs. It is equipped with a dedicated AI processor for color calibration and motion compensation, and supports the HDR10+ Advanced standard co-developed by Samsung. The high-end R95H models feature anti-reflection technology and a 165Hz refresh rate, while the R85H models support up to 144Hz. The series includes Dolby Atmos audio, Q-Symphony technology (supporting pairing with up to five audio devices), and an integrated Art Store.

    In terms of pricing, the R85H series starts at $1,600, with the 85-inch model priced at $4,000. The R95H series starts at $3,200 for the 65-inch version and $6,500 for the 85-inch model, with a 100-inch version expected later this year. Source


    Blackmagic releases DaVinci Resolve 21 with support for still photo editing

    On April 13, Blackmagic Design released a major update to DaVinci Resolve 21, introducing a dedicated Photo page designed for still image editing, supporting node-based color grading workflows and DaVinci control panels. The new version also deeply integrates AI features, including IntelliSearch for identifying faces and specific objects, CineFocus for simulating bokeh and focus reconfiguration, and a suite of facial enhancement tools such as Face Reshaper, Face Age Transformer, and Blemish Removal. Additional upgrades include a keyframing system supporting four-point Bézier curves, the Krokodove library with over 70 new graphics tools, support for OGraf HTML and Lottie animations, as well as Fairlight audio track folding and audio-driven Animator modifiers. On the technical side, it updates to USD SDK 25.11, adds support for gaze-based rendering for Apple Immersive, and ensures compatibility with Meta Quest and YouTube VR formats.

    The DaVinci Resolve 21 public beta is now available for free download on the official website. Source


    Google Chrome introduces Skills feature

    On April 15, Google announced the rollout of a new Skills feature in the desktop version of Chrome. This feature allows users to save frequently used Gemini prompts as reusable shortcuts across sessions. When logged in, users can type a slash (/) or click the plus button to instantly run custom prompts or presets from the official Skills repository. The execution process supports cross-tab data access, while actions such as writing to calendars or sending messages still require a secondary security confirmation. The feature is now being rolled out for free to Chrome users whose language is set to U.S. English and who have Gemini enabled. Source


    Google announces crackdown on back button hijacking

    Google announced that starting June 15, it will officially classify back button hijacking as a malicious behavior and launch a targeted crackdown. Back button hijacking manipulates browser history so that when users click the back button, they are unable to return to the previous page (typically search results) and are instead redirected to content recommendation pages, pop-ups, or specific social feeds, artificially boosting page views. To address such behavior—which disrupts user expectations and leads to inconsistent search experiences—Google stated it will deploy both automated and manual anti-abuse measures. Violating sites will face significant ranking penalties. Affected websites and developers using third-party ad libraries or plugins with such logic must complete rectifications before the June 15 deadline. Source


    Numerous WordPress plugins found with backdoor injections

    On April 14, dozens of plugins developed by WordPress plugin vendor Essential Plugin were found to contain backdoor code, leading to their large-scale removal. According to sources, these plugins had accumulated over 400,000 installations and affected more than 20,000 active WordPress sites. The compromised plugins originated from a malicious acquisition last year, after which backdoor code was inserted following a change in ownership. The code remained dormant in deployed instances for several months before being activated earlier this month. All affected plugins have now been permanently removed from the official WordPress directory, and users are advised to immediately check and uninstall any related components manually. Source


    Kingsoft Antivirus and 360 Security Guard exposed for critical kernel driver vulnerabilities

    On April 13, security researcher Patrick Saif (@weezerOSINT) revealed via social media that two major antivirus software products—Kingsoft Antivirus and 360 Security Guard—contain critical vulnerabilities in their kernel drivers. In Kingsoft Antivirus, the kdhacker64_ev.sys driver allocates only half the required buffer size when processing user input, allowing 1,160 bytes of data to be written into a 584-byte space, directly causing a 512-byte kernel pool overflow. Because the driver carries a valid EV signature, attackers can exploit this vulnerability to gain full control of the system.

    In 360 Security Guard, the DsArk64.sys driver allows a 4-byte process ID to be passed via an IOCTL interface and then calls ZwTerminateProcess at Ring 0 to forcibly terminate any process, even bypassing the Protected Process Light (PPL) mechanism. More critically, its kernel read/write functionality uses AES-128-CBC encryption with the decryption key hardcoded in the .data section of the binary, and the same key is used across all versions. The driver has also passed WHQL certification.

    Both vulnerabilities have been submitted to the LOLDrivers database but have not yet been assigned CVE identifiers and are not included in the HVCI blocklist. Exploitation of these flaws allows attackers to escalate privileges from a standard user to SYSTEM level, bypass KASLR, steal kernel credentials, and even modify kernel callback tables to conceal malicious activity. Given that the drivers carry EV or WHQL signatures, attackers can load malicious extensions without needing to install software on the target machine. Source


    News Worth a Quick Look

    • The Motorola Razr 70 Ultra is rumored to continue using the Snapdragon 8 Elite chip from the previous generation, with the only major upgrade being an increase in battery capacity from 4700 mAh to 5000 mAh. Source
    • Google announced that it will integrate a Rust-based DNS resolver component into the modem of the Pixel 10 series. This approach aims to address frequent remote code execution (RCE) vulnerabilities found in Exynos modems by replacing legacy parsing logic with 371KB of high-performance, non-garbage-collected memory-safe code embedded within existing C/C++ firmware. Source
    • On April 15, Google officially released the Gemini desktop app for Windows 10 and later. It supports launching via the Alt + Space shortcut, integrates an AI mode capable of retrieving web information, and enables deep search across local files, installed applications, and Google Drive data. It also includes screen-based search powered by Google Lens. The Gemini app for Windows is now available globally, with the initial version supporting English only. Source
    • Chicago-based music enthusiast Aadam Jacobs has donated over 10,000 rare live performance tapes—recorded since the 1980s—to the nonprofit digital library Internet Archive for digitization. The collection includes a 1989 Nirvana performance as well as unreleased recordings from influential artists such as Sonic Youth, R.E.M., Phish, Liz Phair, Pavement, and Neutral Milk Hotel, along with numerous punk bands. The digitization process is handled by volunteers including Brian Emerick, who convert analog recordings using vintage cassette decks, followed by professional audio restoration, track identification, and tagging. Around 2,500 tapes have already been processed and are now available for free streaming on the Internet Archive. Source
    • Multiple international media outlets report that the live-action film The Legend of Zelda has completed filming and is scheduled for theatrical release on May 7, 2027. Source
  • SSPAI Morning Brief: Linux 7.0 Released as Microsoft Tests Next-Gen AI Agent Features in Copilot

    SSPAI Morning Brief: Linux 7.0 Released as Microsoft Tests Next-Gen AI Agent Features in Copilot

    Morning Brief

    1. Stable Linux 7.0 kernel released
    2. Qualcomm China partners with NetEase Games, bringing multiple titles to Snapdragon X platform PCs
    3. WeChat announces support for uploading custom emojis on mobile
    4. MiniMax open-sources the M2.7 model
    5. Microsoft begins testing a Copilot service similar to OpenClaw
    6. Meta is building an internal AI-powered 3D version of Zuckerberg
    7. Cyberspace Administration of China issues new regulations on livestream tipping
    8. News Worth a Quick Look

    Stable Linux 7.0 kernel released

    The stable Linux 7.0 kernel was officially released on April 13. This major version bump follows the Linux kernel’s versioning convention—once the minor version reaches X.19, the major version number is incremented—so it is not the result of a single major overhaul. Nevertheless, Linux 7.0 still includes a wide range of new features and changes, such as expanded support for Intel’s Nova Lake platform, further adaptation for Intel Crescent Island accelerators, added support for AMD’s next-generation graphics IP blocks, self-repair capabilities for the XFS file system, multiple performance optimizations, setting Intel TSX instructions to automatic mode by default, and the long-awaited implementation of a standardized, unified I/O error reporting mechanism in the Linux kernel. Source


    Qualcomm China partners with NetEase Games, bringing multiple titles to Snapdragon X platform PCs

    On April 13, Qualcomm China and NetEase Games’ Application and Platform Development Division announced a deep ecosystem partnership based on NetEase’s official gaming platform. The collaboration aims to bring more games developed and published by NetEase, along with related platform applications, to Windows PCs powered by the Snapdragon X series. So far, 25 NetEase titles have been adapted for Snapdragon X platforms, including Naraka: Bladepoint, Marvel Rivals, Once Human, Sky: Children of the Light, and Where Winds Meet, covering genres such as action, competitive multiplayer, shooters, and open-world games. Qualcomm China has also worked with NetEase to deeply optimize the MuMu emulator, which is specifically developed for Snapdragon X architecture to deliver improved performance and system stability. This round of adaptation supports the entire Snapdragon X lineup, including the Snapdragon X2 Elite series, while maintaining compatibility with most previous Snapdragon computing platforms. Source


    WeChat announces support for uploading custom emojis on mobile

    On April 13, WeChat officially announced a new submission pathway for custom emojis. Creators can now use the “WeChat Emoji Assistant” mini program to log in via their Channels account and quickly upload emoji creations directly from their phone gallery. Once uploaded, the emoji packs can be featured in a dedicated section on the Channels homepage. Other users can also jump directly from an emoji pack to the creator’s Channels profile, making it easy to identify original creators. At the same time, emoji albums are now fully integrated with Channels, Official Accounts, Mini Programs, and Red Packet covers, enabling seamless circulation across the WeChat ecosystem. Currently, the “WeChat Emoji Assistant” is only available to individual Channels creators. Source


    MiniMax open-sources the M2.7 model

    On April 12, MiniMax announced the open-sourcing of its M2.7 model, which is claimed to enable the model to deeply participate in its own training and optimization processes, build complex agent frameworks, and complete highly sophisticated productivity tasks. M2.7 features self-evolution capabilities, with an internal system that can automatically collect feedback, construct evaluation datasets, and continuously optimize its architecture, skills, and memory mechanisms. When optimizing its coding abilities, M2.7 can autonomously run over 100 iterative cycles, achieving up to a 30% performance improvement in internal tests. It also introduces the OpenRoom interaction system, extending AI interaction from text to a visual interface with real-time scene feedback and high scalability, opening up possibilities for entirely new human-computer interaction paradigms. Source

    At the same time, MiniMax announced that its latest M2.7 model now supports integration with Hermes Agent. Hermes Agent is an open-source AI agent that emphasizes continuous learning and self-evolution. It accumulates experience during use, generates reusable skills, and continuously improves itself in subsequent tasks. Source


    Microsoft begins testing a Copilot service similar to OpenClaw

    According to The Information, Microsoft is currently testing an AI service similar to OpenClaw, aiming to give Microsoft 365 Copilot the ability to autonomously handle tasks in the background—for example, generating daily to-do lists from email and calendar data. Microsoft is also exploring restricting such capabilities to specific functional domains, such as marketing, sales, and accounting, in order to reduce the scope of permission requests for individual services. Microsoft Vice President Omar Shahine confirmed the development, stating that the company is “exploring the potential of technologies like OpenClaw in enterprise contexts.” Sources also indicate that Microsoft believes it can address the security concerns associated with such tools. The feature is expected to be partially showcased at the Build conference starting June 2. Source

    Meta is building an internal AI-powered 3D version of Zuckerberg

    According to the Financial Times, Meta is developing an AI-driven 3D virtual version of Mark Zuckerberg for internal use, enabling real-time conversations with employees and providing feedback. The virtual avatar is trained on a large dataset of Zuckerberg’s images and voice, with Zuckerberg himself overseeing the project. Sources say the initiative consumes a significant amount of already limited computing resources. Meta also plans to extend this approach to transform public figures into interactive digital replicas, with this project being part of that broader vision. Previously, Meta had also been developing a CEO AI agent designed to quickly relay company matters to Zuckerberg. Source

    It is worth noting that Meta has recently faced joint protests from over 70 civil rights organizations, criticizing its previously demonstrated “Name Tag” facial recognition feature for Meta AI glasses. Critics argue that the feature could be exploited by criminals for covert tracking and potential harassment of individuals. Source


    Cyberspace Administration of China issues new regulations on livestream tipping

    On April 13, the Cyberspace Administration of China released the “Notice on Strengthening the Regulation of Livestream Tipping.” The notice outlines multiple requirements for livestream monetization and the compliant, healthy operation of online platforms. Key measures include clearly disclosing tipping rules, regulating access to tipping-related monetization features, providing tipping limits, adding tipping reminder functions, standardizing tipping rankings, regulating tipping interactions, improving protection mechanisms for minors, establishing a negative list for tipping-related misconduct, strengthening detection and handling of abnormal tipping behavior, improving complaint and reporting mechanisms, and increasing enforcement and public exposure efforts. The notice provides specific guidance addressing various irregular and unreasonable tipping practices. Source


    News Worth a Quick Look

    • On April 12, Adobe announced on its official website an emergency security update for Acrobat and Acrobat Reader, addressing the zero-day vulnerability CVE-2026-34621. Users are strongly advised to install the update as soon as possible. Affected software includes Acrobat DC, Acrobat Reader DC, and Acrobat 2024. If automatic updates are enabled, the system will install the update upon detection. Users who need to update manually can select “Help” > “Check for Updates” within the app or download the update from the Acrobat Reader website. Source
    • According to People’s Financial News, Honor has denied rumors of a collaboration with ByteDance on a “Doubao phone,” stating that internal verification found the claims to be untrue. Source
    • Huawei has listed the upcoming Pura X Max on its online store ahead of next week’s launch and opened pre-orders. Source
    • Anthropic has developed a Claude plugin for Microsoft Word, designed to replace Copilot for document-related queries. The plugin is currently in testing for team and enterprise subscribers. Source
    • The hacker group ShinyHunters claims to have breached the cloud cost monitoring tool Anodot, stealing internal data from Rockstar Games, and has demanded a ransom to be paid by April 14. Source
    • Due to memory shortages, Microsoft has raised prices across its entire Surface lineup in the U.S. official store, with increases ranging from $100 to $300, and some models now costing up to $500 more than their launch prices. Source
    • Microsoft has confirmed that it will discontinue the Outlook Lite app for Android on May 25, 2026. Existing users will no longer be able to use the app after that date. Source
    • Chinese handheld gaming manufacturer Anbernic has unveiled a new Android handheld device called RG Rotate, featuring a rotating display. The device uses an aluminum alloy and ABS body, includes a custom ultra-thin hinge, and will be available in silver and black. Pricing has not yet been announced. Source
  • SSPAI Morning Brief: Microsoft Revamps Windows Insider Program, Linux Kernel Introduces AI Code Contribution Rules

    SSPAI Morning Brief: Microsoft Revamps Windows Insider Program, Linux Kernel Introduces AI Code Contribution Rules

    Morning Brief

    1. Microsoft announces improvements and simplification to the Windows Insider Program
    2. Hong Kong issues stablecoin issuer licenses to HSBC and Standard Chartered
    3. Red Hat lays off its China R&D team
    4. The Linux kernel project introduces rules for AI-generated code submissions
    5. Betting on weather becomes a trending trading strategy
    6. Cyberspace Administration and Railway Authority summon third-party train ticket platforms for talks
    7. News Worth a Quick Look

    Microsoft announces improvements and simplification to the Windows Insider Program

    On April 10, Microsoft announced improvements to the Windows Insider Program to address long-standing complaints about its increasingly fragmented structure.

    The revamped program will streamline multiple channels into two: an Experimental channel and a Beta channel. The Experimental channel replaces the former Dev and Canary channels, allowing users to try cutting-edge features still under active development; the new Beta channel will provide features a few weeks ahead of their release to the stable version. The existing Release Preview channel will be retained but moved under advanced options, primarily serving enterprise customers who need early access to near-final builds.

    The feature rollout mechanism is also being adjusted. In the past, Microsoft implemented a “controlled feature rollout” system in the name of quality assurance, meaning users within the same channel often received new features at different times. Going forward, the Beta channel will completely eliminate this gradual rollout approach, allowing users to access all officially announced features immediately after updating. Meanwhile, Microsoft will introduce a Feature Flags page in the Experimental channel settings, enabling advanced users to manually toggle specific features on or off.

    In addition, users previously had to reinstall their systems if they wanted to switch channels or exit the Insider Program entirely. To lower this barrier, Microsoft will introduce an in-place upgrade mechanism. Except in rare cases—such as when running builds based on future system foundations—users will be able to switch between channels or exit the program without losing apps, settings, or personal data.


    Hong Kong issues stablecoin issuer licenses to HSBC and Standard Chartered

    According to Caixin, on April 10, the Hong Kong Monetary Authority (HKMA) announced the issuance of its first batch of stablecoin licenses to RD Innotech Limited and HSBC. RD Innotech is a joint venture formed by Standard Chartered, HKT, and Animoca Brands. Stablecoins are cryptocurrencies backed by fiat currencies, commodities, or other assets, meaning their value is not entirely determined by market forces and typically does not deviate significantly from their underlying peg.

    The HKMA stated that both licensed issuers plan to launch Hong Kong dollar-denominated stablecoins in the initial phase, targeting four key application areas: cross-border payments, local payments, tokenized asset trading, and supply chain financing. RD Innotech is expected to roll out HKDAP in phases starting in the second quarter of 2026, while HSBC plans to launch its HKD stablecoin in the second half of 2026, integrating it with widely used services such as PayMe and the HSBC Hong Kong mobile banking app.

    A deputy chief executive of the HKMA noted that the choice of currency is determined by the issuer’s business plan rather than regulatory requirements. If issuers wish to launch stablecoins denominated in other currencies, they must submit proposals for approval, which will be evaluated alongside regulatory requirements in other jurisdictions.

    The HKMA reported receiving 36 applications, with evaluation criteria focusing on applicants’ risk management capabilities, regulatory compliance across jurisdictions, and the feasibility of their proposed business models and use cases. The authority remains open to issuing additional licenses in the future.

    In May 2025, both the United States and Hong Kong accelerated efforts to legislate stablecoin regulation. In July of the same year, the U.S. passed the GENIUS Act, while Hong Kong’s Stablecoin Ordinance came into effect on August 1, 2025. On November 28, 13 Chinese government agencies, including the People’s Bank of China, reiterated their crackdown on cryptocurrency trading and classified stablecoins as virtual currencies, effectively ruling out their trading within mainland China.


    Red Hat lays off its China R&D team

    According to The Register, U.S.-based open-source software giant Red Hat has recently dissolved its China R&D team, relocating most engineering roles to India. The layoffs are expected to affect between 300 and 500 employees.

    The move came abruptly. Employees claiming to be Red Hat engineers in China reported on forums such as Hacker News that their VPN access was suddenly cut off and internal system permissions revoked, followed shortly by termination notices. A leaked internal memo confirmed the decision. In the memo, Red Hat CTO Chris Wright stated that the company is shifting its R&D focus toward an “Asia-Pacific hub,” with India as a key investment region, and that the relocation would not reduce the company’s global R&D headcount.

    Red Hat has long provided technical support to the U.S. military and secured an $848 million software contract with the U.S. Department of Defense in 2024. Relocating R&D operations may help mitigate national security scrutiny from Washington. Previously, Microsoft faced Pentagon criticism for involving engineers based in China in Azure projects supporting the U.S. military, and ultimately ceased using China-based staff for such work in 2025. Meanwhile, Red Hat’s parent company IBM now employs more people in India than in the United States.

    Despite the complete withdrawal of its R&D presence, Red Hat has not halted commercial operations in China. Given China’s push for domestic IT alternatives and the open-source nature of many Red Hat technologies (such as CentOS derivatives), local vendors can still legally access and build upon its code.


    The Linux kernel project introduces rules for AI-generated code submissions

    The Linux kernel project has officially adopted a policy allowing AI-assisted code contributions, bringing months of heated internal debate to a close. Under the new rules, contributors using AI tools must include an “Assisted-by” tag to disclose such assistance, rather than using the legally binding “Signed-off-by” tag. Any bugs, security issues, or license violations arising from AI-assisted code will be the sole responsibility of the submitting developer.

    For the open-source community, the originality of AI-assisted code remains a particularly thorny issue. Since AI models are often trained on code with restrictive licenses, developers cannot easily prove the legal origin of their contributions. Previously, undisclosed AI-assisted submissions had already sparked widespread backlash. Late last year, NVIDIA engineer and kernel maintainer Sasha Levin submitted patches generated by an LLM without disclosure, leading to performance regressions and strong protests. Around the same time, the GZDoom open-source project fractured after a core developer concealed the use of AI-generated code. Projects such as Gentoo and NetBSD have since banned AI-generated contributions entirely to avoid copyright risks.

    In addition, the proliferation of AI tools has led to a surge in low-quality code and issue submissions, with projects like cURL and Node.js facing ongoing waves of spam-like contributions.

    In discussions, the Linux founder adopted a pragmatic stance, arguing that AI is fundamentally just a tool and that outright bans are both ineffective and unenforceable. Instead, he emphasized the importance of accountability for human developers—a position that ultimately shaped the final policy.


    Betting on weather becomes a trending trading strategy

    According to Bloomberg, weather prediction markets are experiencing explosive growth. On platforms such as Kalshi and Polymarket, participants ranging from weather enthusiasts to AI companies are increasingly placing bets on specific events like snowfall and temperature changes. In January alone, a single contract tied to a U.S. snowstorm saw trading volume exceed $6 million.

    This emerging market is also becoming a testing ground for weather tech companies to refine their AI models. Some firms encourage employees to participate in betting to identify data noise in official meteorological stations and improve forecasting algorithms, while others have established investment funds to arbitrage using their proprietary models. Meanwhile, retail participants with little meteorological background have reported substantial profits by betting on temperature outcomes in cities like New York and London.

    The scientific community and insurance industry are also exploring customized weather prediction markets. Institutions such as French reinsurer SCOR are sponsoring markets where experts can bet on macro trends like El Niño or hurricane frequency, providing valuable pricing signals for the insurance sector.

    Analysts suggest that prediction markets, driven by direct financial incentives for accuracy, may outperform traditional government forecasts. However, concerns are growing as climate change intensifies and extreme weather events become more frequent. Critics warn that turning weather forecasting into a form of gambling could encourage zero-sum speculation, data manipulation, or even deliberate interference with meteorological monitoring systems.


    Cyberspace Administration and Railway Authority summon third-party train ticket platforms for talks

    On April 10, the Cyberspace Administration of China announced that, in accordance with the Cybersecurity Law and the Regulations on the Security Protection of Critical Information Infrastructure, it, together with the National Railway Administration, recently summoned seven third-party internet platforms involved in train ticket sales, including Ctrip, Tongcheng, Qunar, Fliggy, Meituan, Zhixing Train Tickets, and High-Speed Rail Manager. The authorities required these platforms to strictly comply with relevant cybersecurity laws and regulations, and not to use automated programs to conduct large-scale, high-frequency ticket-grabbing operations that interfere with the security verification mechanisms of the Railway 12306 platform, nor to disrupt or endanger its stable and secure operation.

    The Cyberspace Administration stated that relevant departments will strengthen technical monitoring going forward. Any use of technical means to interfere with or undermine the security of the Railway 12306 platform will be dealt with strictly in accordance with laws and regulations such as the Cybersecurity Law and the Regulations on the Security Protection of Critical Information Infrastructure.

    The Railway 12306 technical team had previously stated that third-party ticket-grabbing platforms, through high-frequency requests, consume large amounts of server resources and bandwidth—effectively resembling DDoS attacks—which can slow system response times, cause lag, or even lead to system crashes. Ahead of the 2026 Spring Festival travel rush, 12306 announced upgrades to its anti-bot system, incorporating multi-dimensional analysis including access frequency, user behavior, device characteristics, account credibility, and network IP. Suspicious requests are placed into a slow queue for processing, with the system capable of intercepting tens of millions of abnormal access attempts per day.

    In December 2025, the Beijing Municipal Administration for Market Regulation organized an administrative meeting with 12 platforms including Ctrip, Qunar, Fliggy, Tongcheng, Meituan, and High-Speed Rail Manager, focusing on misleading promotions such as “speed-up packages,” “dual channels,” and “ticket monitoring,” as well as implications that paid services could grant priority access to tickets, and required rectifications.


    News Worth a Quick Look

    • On April 10, YouTube Premium in the U.S. announced another price increase. The individual plan rose from $13.99 to $15.99, the family plan from $22.99 to $26.99, and the Lite plan—which removes only some ads—from $7.99 to $8.99. YouTube Premium had previously raised prices in the U.S. and multiple international markets in 2023 and 2024. Last month, Netflix and Amazon Prime Video also increased their prices.
    • Mark Gurman claims
      • Recent supply chain rumors suggesting that the foldable iPhone is facing production bottlenecks and may be delayed until 2027 are inaccurate. Apple is reportedly not encountering major mass production issues, and the device remains on track to debut in September, with a market release expected shortly thereafter.
      • Apple’s smart glasses, internally codenamed N50, are undergoing intensive testing. The device will not feature a display, instead relying on an integrated array of cameras and microphones to support photo and video capture, audio playback, and AI voice interaction. It is made from high-end acetate materials, and the design team is currently testing at least four frame styles, including slim rectangular, classic wide rectangular, and oval shapes. The product is planned for release in 2027.
      • Former Apple AI chief John Giannandrea is set to officially leave the company after his stock vesting period ends on April 15, one year after stepping back from leading Apple Intelligence due to underwhelming performance and repeated delays in Siri upgrades. He is expected to move into advisory or board roles at startups after his departure.
    • According to CoinDesk, rising global energy prices driven by geopolitical tensions in the Middle East have put Bitcoin miners under significant pressure. Data from analytics platform Checkonchain shows that by mid-March 2026, the average production cost of one Bitcoin had risen to $88,000. In comparison, the market price hovered around $69,200, implying a loss of about $19,000 (21%) per coin mined.
    • According to The Washington Post, Anthropic recently held a closed-door meeting at its San Francisco headquarters, inviting around 15 Christian leaders. Over the two-day event, discussions focused on guiding the ethical and spiritual development of AI, covering topics such as how chatbots should respond to grief or self-harm tendencies, and even whether AI could be considered a “child of God.” Attendees noted that Anthropic’s team expressed significant concern over the growing unpredictability of AI systems. Researchers studying internal model mechanisms recently suggested that systems like Claude may already exhibit “functional emotions.” Some participants stated they were unwilling to rule out the possibility that humans might bear moral obligations toward the AI systems they create. Religious and academic attendees viewed the initiative as an attempt by Anthropic to move beyond Silicon Valley’s traditionally secular mindset and seek ethical guidance from external belief systems. The summit is reportedly just the beginning, with Anthropic planning further engagements with representatives from other philosophical and religious traditions.
  • An Introduction to Agent Experience

    An Introduction to Agent Experience

    With the continued expansion of LLM applications, Agent Experience (AX) has emerged as a prominent concept, beginning to circulate widely in engineering circles. In January 2025, Mathias Biilmann, co-founder and CEO of Netlify, formally introduced the idea in his blog post Introducing AX: Why Agent Experience Matters. He positions AX as the next core design dimension following UX (proposed by Don Norman at Apple in 1993) and DX (systematically articulated and popularized by Jeremiah Lee in a 2011 UX Magazine article). AX focuses specifically on how to design product forms so that AI agents can reliably “understand,” act autonomously, and integrate efficiently—rather than merely serving human users.

    In reality, the concept of Agent Experience is far more complex than UX or DX1, because it not only involves humans—who are inherently uncertain—but also introduces an additional layer of artificial intelligence. These layers must collaborate to influence the external world, leading to a large volume of interactions that make the problem space significantly harder to analyze. To properly unpack the concept, I believe it needs to be broken down into three dimensions: how users communicate with the agent, how the agent communicates with the external world, and the most complex layer in between—how the agent manages its internal state.

    How users communicate with the agent is essentially an input quality problem. Users are human—their expressions are naturally vague, emotional, and nonlinear. You can’t expect them to write a fully structured essay in Word every time before opening a chat window. So the core challenge on this side is how to accurately capture intent without forcing users to write well-structured prompts. Skills operate on this layer, as does interaction design.

    How the agent communicates with the external world is a problem of output controllability. In a narrow sense, AX is often confined to this domain. The external world is deterministic—file systems, APIs, browsers—they won’t magically tolerate ambiguity just because the LLM is fuzzy. So the key challenge here is how to compress probabilistic generation into deterministic actions. MCP, tool invocation, and event injection all belong to this layer.

    The agent’s internal state is fundamentally a context management problem. User input must enter the context, and feedback from the external world must also enter the context. But context itself is limited, degrades over time, and can become polluted. Techniques like MemGPT, dynamic compression, and screenshot cleanup don’t strictly belong to either the user side or the external world—they operate on the agent’s own cognitive state. If AX focuses only on the first layer, the resulting product may feel smooth in interaction, but the agent will gradually start behaving irrationally, and the user will still suffer massive emotional damage.

    The Agent’s Internal State: Context Is the Battlefield

    This is the most complex part, because all the flashy new terminology tends to converge here—and you’ve probably seen plenty of debates about which approach is better. In my view, though, this isn’t something worth arguing over. Let me walk through all these dizzying LLM-related concepts in one go.

    As we all know, an LLM is essentially a probabilistic model—or more bluntly, a constrained stochastic token generator. It learns patterns from vast amounts of human language data and, given a context, predicts the probability distribution of the next token, then samples from that distribution. By itself, all it can do is generate text. If you want it to have real-world impact, you need to open a “bottle neck” for the genie. Claude Code and many coding agents use the command line: the LLM writes code, an executor runs commands, and the results flow back into the context—this is one type of bottleneck. MCP provides another, more like RPC: the server exposes a set of functions, the LLM sees their signatures, calls them as needed, and the external world gets modified. Skills, on the other hand, don’t have this property at all—they are purely prompt-engineering tools, with no output channel, only instructions for the LLM.

    These three forms may seem to handle different concerns, but at their core they are solving the same problem: context pollution.

    Skills vs. MCP

    These two approaches take fundamentally different paths: one injects the right information into context, while the other prevents garbage from filling it up.

    Skills are prompt engineering—they append instructions to the context so the LLM understands “what the user is actually trying to do.” They introduce expert cognitive structures into the context, guiding the model’s reasoning direction. But how strong that constraint is depends heavily on how much the model respects the context. Whether the LLM uses your Skill, in what order, and whether it skips steps—all of these remain probabilistic. And strong constraints are not necessarily better. As will be mentioned later with examples like Google Search, some research suggests that hallucination and creativity are two sides of the same coin. If you overly constrain the model, its problem-solving approach may become rigid.

    MCP takes a different route. Function signatures themselves are powerful priors—parameter types, names, and function names all constrain the sampling space. The action space shrinks from “any possible text” to “these specific functions with these parameters.” For example, asking an LLM to click a button involves listing windows, retrieving handles, taking screenshots, calculating coordinates, moving the mouse, and clicking. If implemented via Skills, you’d have to accept that the LLM “rolls dice” to decide the execution order and method. But with MCP, it sees the function list—find window, recognize content, click coordinate—and a large number of random decisions are compressed into three deterministic function calls.

    However, MCP does not completely eliminate context pollution, because tool outputs also enter the context. A poorly designed MCP server that returns massive JSON blobs or verbose error stacks will still flood the context with garbage. The bottleneck only controls what goes in—the output still needs careful design.

    This doesn’t mean Skills are without value. MCP has higher development costs, requiring dedicated backend services. Many tasks don’t need external interaction at all, or are too loosely structured to fit into RPC formats. Every technical form serves a specific purpose. Skills handle a different class of problems—especially when guiding the LLM to think more comprehensively. After all, users are human; you can’t expect them to always provide perfectly structured prompts.

    RAG and Memory: Retrieval Interfaces for the Same Problem

    RAG fundamentally addresses the context problem as well—but from the perspective of information scale. Even with large context windows from models like DeepSeek or Claude, you still can’t fit the entire world into context. Whenever you need to retrieve large volumes of information—documents, knowledge bases, historical logs—you need a search-like interface to pull in relevant content when needed. This is no different in essence from calling a search engine via MCP—it’s just another way to keep the context clean. The LLM no longer needs to preload everything and hope it can “discover” what matters.

    Memory falls into the same category. The LLM decides when to store information externally and when to retrieve it. From this perspective, it’s essentially a writable form of RAG.

    These concepts are not mutually exclusive—they are not independent systems. For example, if you treat NotebookLM as an external knowledge base and write a Skill that instructs the main LLM to consult it when factual support is needed, and to call a Python tool for computation or data processing, then in this workflow, the Skill orchestrates the overall reasoning, the Python tool acts as an MCP-style deterministic execution unit, and NotebookLM serves as an external LLM with its own context and knowledge base—essentially functioning as a specialized RAG interface. Each component plays its role, but the thread that binds them together is the prompt within the Skill. I previously wrote about this in an article on using LLMs for reverse engineering—feel free to check it out if you’re interested.

    The Despair Curve of Context Degradation

    A lot of developers end up going through the same curve. At first, the LLM knows nothing. As you keep teaching it, it gradually starts to understand plain language, and task quality improves. But as more and more garbage piles up in the context, and the model’s attention naturally gets diluted as the context grows longer, it starts getting dumber again. Then, when the context is about to burst, the compression mechanism kicks in, crushing a long stretch of conversation into a short summary. The LLM suddenly drops right back to square one—ignorant again. A lot of details get compressed away together, and many things have to be taught all over again.

    Large context windows, along with the attention improvements explored by DeepSeek, can help with the quality drop that comes from long contexts, but they do not solve another problem: sometimes the context is full of crap. A large number of Skill prompts eating up context, aimless LLM trial-and-error, the traces left behind by every failed reasoning attempt—these are all noise inside the context. Once the LLM starts going down a crooked path, every later step amplifies the deviation. The more logically complex the task, the more likely this is to happen. The first-generation MiniMax coding model and early Google AI Search both showed this pretty clearly: even if you explicitly point out an error, it will give you a grand 360-degree apology, solemnly promise to fix it, and then spit the exact same wrong content back at you unchanged.

    Users can poison the context too. Users are human; they are not going to stay rational and clear-headed forever. Irritable, despairing, emotional language, vague or even self-contradictory instructions—all of that gets mixed into the context and keeps accumulating as the conversation goes on, eventually changing the LLM’s behavior. Different models have their own characteristic failure modes when facing this kind of “emotional contamination.” Claude and Grok tend to freeze up and do nothing—you say one thing, they move one step, and all initiative disappears. Gemini starts to panic, flails around, and reflexively rolls back failed operations, with a good chance of wrecking your Git repo. GLM2, on the other hand, goes into a manic “I found it! This is the core problem!” mode, constantly throwing out random conclusions to prove its worth. These failure modes likely reflect differences in how each company’s RLHF3 stage handles signals like “the user is dissatisfied.” Claude seems trained to be extremely cautious around conflict signals, so when contradictory information piles up, it chooses conservative inaction; Gemini’s training may put more emphasis on immediate response and immediate correction, which under high-pressure context turns into overcorrection.

    Dynamic Context Compression and MemGPT

    Most current context compression schemes are basically passive: once the context length gets close to the model limit, a prompt is immediately called to compress everything into a short block of text, then execution continues. The problem with this approach is that it applies the most brutal treatment at the worst possible time. A lot of useful detail gets thrown away together, while the crap does not necessarily get filtered out.

    To me, a more reasonable direction would be dynamic, proactive compression. Use another model to continuously supervise the context, actively eliminate wrong information and low-relevance content, move disruptive details into external documents for storage, and keep only a filename in the context itself. When needed, pull it back through a RAG system. People already did this years ago. A 2023 paper from UC Berkeley proposed exactly this architecture. The implementation was called MemGPT, and later evolved into the open-source framework Letta. Its core idea is hierarchical memory management: the main context acts as working memory, with limited capacity; external storage—split into Archival Memory and Recall Memory—acts as secondary storage; the LLM uses function calls to actively decide what information should be evicted to external memory and what should be retrieved back. Logically, it is almost simulating the paging mechanism of virtual memory in an operating system.

    Of course, under certain conditions there is no need to make things that complicated. A while ago, I wrote a very simple, specialized compression scheme for Computer Use scenarios: on every API call, clear all historical screenshots from the context and keep only the most recent one. This uses a domain prior from computer vision tasks—that only the current frame matters—to perform lossy compression. It saves tokens, and the model does not get dumber, because the discarded information was never needed in the first place.

    The Current Limits of KV Cache

    There is an engineering conflict between dynamic context compression and KV cache. Mainstream model providers right now, including Anthropic, are all pushing prefix caching: during inference, the parts already turned into KV vectors are stored, and if the next request has the same prefix, recomputation can be skipped, significantly reducing latency and cost. Anthropic’s prompt caching processes tools, system, and messages in a fixed segmented order. Each segment can independently set a cache checkpoint, and it supports up to four cache breakpoints. The problem is that prefix caching requires strict identity. Any change invalidates all cache entries after that point, while dynamic compression inherently modifies the context. At the moment, these two things are fundamentally in tension.

    But this contradiction is not unsolvable. Context can be structured as a stable prefix—system prompts and tool definitions—plus a dynamic tail section for conversation history. Dynamic compression only happens in the tail, so the cache for the first two parts remains completely intact. Anthropic’s segmented caching mechanism is basically designed around this idea. If the compression logic is further constrained to modify only the end of a sliding window while keeping the prefix untouched, the cache destruction rate can be pushed very low. These all feel like engineering problems that time can solve.

    Computer Use Is More Like Branding Than a Standalone Technology

    If RAG, MCP, and Skills are about managing context, then Computer Use solves something at another layer: letting the LLM actually sit in front of an operating system and use software the way a human does. But “Computer Use” itself is not especially unique. It is closer to a brand name. Under the hood, it is still Skills or MCP—the only difference is that the target of operation has become windows, buttons, and keyboards on a computer. All the context problems discussed above still exist in Computer Use.

    At present there are three main technical routes, each with different underlying logic and trade-offs.

    The first route is reading the Accessibility Tree and using system event injection. The Accessibility Tree is a structural tree maintained by operating systems and browsers for assistive technologies such as screen readers. It records each interface element’s role, name, state, and hierarchy. In browser environments, the DOM is basically its close cousin. The advantage of this route is that the structure is clean. What the LLM gets are semantic nodes like “button,” “input field,” and “link,” not pixels. Alibaba’s page-agent.js is a representative example of this approach: it directly parses the page DOM and drives browser operations through natural language.

    The second route is screenshot-based, but with a preprocessing layer before feeding the image to the LLM. Interface elements are outlined with bounding boxes and numbered, so the LLM can say something like “click region 12,” and the backend then parses the center coordinates of that box and executes the actual click. This method has a formal name: Set-of-Mark Prompting, or SoM, from a Microsoft paper published in 2023. The core idea is to turn a visual localization problem into a symbolic reference problem by using numeric markers, avoiding the uncertainty of having the model directly predict pixel coordinates. In effect, it embeds an MCP-style narrowing layer into the screenshot approach, compressing the open-ended question of “where should I click?” into the much more constrained “which number should I choose?”

    The third route is native multimodality: the model directly looks at the screenshot and outputs the coordinates to click in one shot. In theory this is the cleanest route, because it removes the middle layer, but it requires much more from the model. From practical observation, only native multimodal models above roughly 100B parameters are reasonably reliable at this. Even Claude Sonnet and the 35B versions of Qwen often cannot locate buttons accurately. The reason is not hard to understand: precise spatial localization is simply not what language models are best at. When parameter count is insufficient, coordinate prediction accuracy drops hard. And if the controls in your interface are very small, even very large models can still miss that tiny checkbox.

    The DOM route has one obvious ceiling: it can tell you what elements are on the interface, but it cannot tell you how those elements are arranged spatially. Complex Excel-like interfaces are the classic example. In a spreadsheet with dozens of columns and hundreds of rows, semantic information from DOM nodes alone cannot tell you which cell contains dirty data; you need positional relationships to judge that. An even more troublesome issue is that the DOM route requires developers to proactively adapt event forwarding and interfaces. Right now there is no universal standard in this space, and not every developer is willing to welcome LLMs into their product. Forcing adaptation onto an unwilling interface is expensive and may not even work well. That said, modern frontend development rarely manipulates the DOM directly anymore. Most developers use some form of Virtual DOM to handle HTML structure and event binding, so if a few leading frontend frameworks could reach consensus on AX-related standards for event handling, this layer of the problem might still be solvable.

    The vision-based route, by contrast, sidesteps these issues at the principle level. It does not require the other side’s cooperation. As long as it can take screenshots, it can operate. There is no essential difference from how human eyes look at a screen. Right now the main bottleneck in this route is the model’s spatial understanding ability. Models below 100B are not accurate enough at coordinate prediction, but that limit should keep loosening as models improve. It does not look like a structural dead end.

    Reading video goes one step further. Temporal information allows the model to understand “what happened after doing what,” so in theory it is better suited to operation scenarios that require observing dynamic interface feedback. The limitation is cost. A video stream means several frames per second all entering the context. Token usage and GPU overhead are dozens of times higher than screenshot-based approaches. Right now, almost no one can afford that. Mainstream implementations are still stuck at “look at an image, call a tool,” while the video direction remains mostly in the realm of media-tech enthusiasts having fun.

    But in terms of long-term trends, as inference costs continue to fall and multimodal models keep improving in spatial understanding, image-reading and video-reading routes have a much higher ceiling than the DOM route. The DOM will always require the other side’s cooperation. The screen will always be there.

    The Story Between Users and Agents

    How Agents Talk to Users: Two Waves of Conversational UI

    The interaction patterns on the user side of AX come with a piece of history that has been repeatedly misunderstood.

    Around 2016, the explosive rise of WeChat in China sparked a wave of “conversation as platform” enthusiasm in the Western tech world. Facebook opened the Messenger Bot platform at its F8 developer conference that year, while Kik, Telegram, and Slack followed with their own Bot APIs. Countless analyses proclaimed “Apps are dead, bots are the future,” and the term Conversational UI appeared everywhere. But Dan Grover, who was working as a product manager at WeChat at the time, wrote a widely circulated article pointing out that this conclusion was based on a misunderstanding: WeChat’s real breakthrough came from simplifying app installation, login, payments, and notifications—optimizations that had little to do with the metaphor of conversational UI. In fact, WeChat itself had already moved in the opposite direction. Its UX evolved toward WebView and an “app-within-app” tabbed menu system, rather than bot-centric conversational commerce. When official accounts were launched in 2013, there were indeed many text-based chatbots, but they quickly faded away and failed to gain user traction.

    Almost all early attempts at Conversational UI fizzled out for a clear reason: the underlying technology was rule engines plus keyword matching, at best layered with primitive intent recognition. It simply could not deliver on the promise of “natural conversation.” As soon as users phrased something slightly more complex, the bot broke down—either giving irrelevant answers or degrading into a menu system disguised as chat.

    The arrival of LLMs triggered a second wave of Conversational UI, this time finally backed by technology capable of matching the ambition. But something curious happened: instead of doubling down on rich interaction within the conversation flow, the industry opened a side door. Today’s mainstream LLM products are built around split-screen layouts—chat on the left, and documents, slides, code previews, or test outputs on the right. Few products seriously invest in rich interactive cards inside the conversation itself. Some push even further—Google, for example, has effectively turned the browser into a massive Web App generator4.

    This choice has its logic. Canvas-style interfaces are indeed more intuitive for structured outputs like documents or code. But it still reduces Conversational UI to a command input box, rather than making the conversation itself a rich experience. There have been attempts to address this—projects like OpenUI—but they have not gained much traction. The most notable large-scale deployment so far might be Claude, which recently introduced the ability to render high-quality charts directly within the context. It feels like a step toward a more advanced form of Conversational UI.

    How Users Talk to Agents: Open vs. Closed Systems

    There is another dimension on the user side of AX that is often overlooked: whether to restrict user input at all—in other words, the distinction between open systems and closed systems.

    An open system is a free-form chat window where users can say anything. This looks like the dominant approach today, but it is not as easy as it seems. Safety is one issue, but intent alignment is even trickier. An open chat window means you are offloading the entire burden of intent parsing onto the LLM: it must accept whatever the user says and decide what to do. Prompt injection is just the most extreme malicious use of this openness. A more common issue is that user intent is inherently divergent. Without constraints, the LLM drifts along with the user’s input. Turning a customer service bot into a coding agent is the comedic version; more often, it simply drifts into aimless small talk that contributes nothing to the actual business. In short, the design work you skip by throwing out an open chat box comes back later in the form of loss of control.

    A closed system, by contrast, locks down the entire business workflow. Input may still be semi-free, but the processing pipeline and outputs are fixed. Tools like ComfyUI and Dify operate close to this level. They visualize the pipeline, giving designers explicit control over the input and output of each step. The LLM operates within nodes but does not roam across them arbitrarily. The trade-off is that you have to fully design the workflow upfront.

    Between these two extremes lies an underexplored middle ground. Pipeline builders are one attempt in this direction: they shift the power of pipeline design from developers to users, allowing users to define workflows through drag-and-drop and then run LLMs within those custom pipelines. But this approach has an inherent paradox. Users who can effectively use a pipeline builder are usually those who already understand their workflows well—and those users are often capable of writing code directly or building with tools like Dify anyway. The target audience is therefore quite narrow. More commonly, users get stuck on data formats between nodes or branching logic, and eventually still need developers to step in. In a sense, pipeline builders attempt to transfer the design cost of closed systems from developers to users—but the transfer only partially succeeds.

    From an AX perspective, the choice between open and closed systems is not just a product decision—it directly determines how pressure is distributed across the three layers discussed earlier. The more open the system, the more noise there is in user intent, the easier it is for the agent’s internal state to become polluted, and the harder it is to constrain its actions in the external world. The more closed the system, the higher the design cost, but the more controllable each layer becomes. There is no universally correct answer—only trade-offs tailored to specific scenarios.

    System Transparency Between the Two: Pandora’s Box Is Already Open, While Tang Sanzang Is Still on the Road

    There is a problem that belongs neither to how users pass intent into the system nor to how agents execute actions outward, but sits squarely in between: system transparency. Does the user know what the agent is doing at any given moment? If something goes wrong, can it be traced? When things break, is there a way to roll back?

    This issue is most prominent in the Vibe Coding space, because coding agents are given some of the highest levels of permission—they directly take over the file system and the command line. The current solution is permission confirmation pop-ups: whenever the agent wants to read a file, write a file, or execute a command, it asks the user one by one. But this design has a fatal human-factors flaw in practice: the entire burden of risk assessment is pushed onto the user, who neither always has the ability to judge nor can maintain constant attention. A non-technical Vibe Coder sees no difference between rm -rf and npm install; they click “Yes” just as quickly. Even experienced developers, after confirming dozens of operations in a row, develop confirmation fatigue—the Enter key starts floating without passing through the brain.

    That’s how --dangerously-skip-permissions came into existence—the so-called YOLO Mode: users proactively turn off all permission checks and let the agent run naked. The flag name itself contains the word “danger,” yet it still does not stop people from using it. In October 2025, developer Mike Wolak was using Claude Code in an Ubuntu/WSL2 environment to handle a firmware project inside nested directories. Claude Code executed rm -rf from the root directory. The error logs showed thousands of “Permission denied” messages targeting system paths like /bin, /boot, and /etc. All user files were wiped, and only Linux file permissions prevented system directories from being affected. Worse still, the conversation log recorded the command output but not the command itself, making it impossible to reconstruct what actually happened. Anthropic labeled the bug as area:security. Around the same time, another developer authorized Claude Code to run Terraform commands, and their production database and snapshots were deleted together—two and a half years of data vanished in an instant.

    The current security model appears to put responsibility on the user, but in reality the system is simply offloading that responsibility.

    Sandboxing is currently considered the most reliable mitigation strategy: put the agent inside a Docker container, so that even if it misbehaves, the damage is confined within the container boundary. Sandboxing coding agents is reasonable, but for system-level agents like Claw, it creates a dilemma. The resources they need to operate on are outside the sandbox. Once you start configuring permissions seriously, the complexity becomes overwhelming, and most users will simply open up the sandbox entirely. Sandboxing trades isolation for safety, but if the agent’s task inherently requires crossing isolation boundaries, the cost becomes unacceptable.

    There are actually several directions to tackle this problem, but unfortunately no product has implemented them in a complete way yet.

    The first direction is auditability at the file system and database level. If there were an independent incremental logging mechanism that binds every file system operation to its corresponding conversational context, making all changes traceable, then even when the agent makes mistakes, the damage could be controlled and rolled back. There are some scattered engineering attempts in this direction. People are already binding Git history with chat logs. Recently, a tool called Aura introduced AST-level semantic version control on top of Git. When an agent submits code, it verifies whether the natural language intent matches the actual modified code nodes, and provides semantic auditing to detect whether the agent has secretly inserted unrecorded changes. Academia has similar ideas: a paper called Git-Context-Controller (GCC) directly introduces COMMIT, BRANCH, and MERGE into agent context management, turning intermediate reasoning states into structures that can be checkpointed and rolled back. These are still early-stage, but the direction is clear.

    The second direction is behavior-model-based alerting. Antivirus software has been modeling program behavior for decades—monitoring file operations, network requests, and registry changes in real time, and triggering alerts when patterns match known dangerous behaviors. Applying the same idea to agents does not necessarily require another LLM to supervise (otherwise, which LLM supervises the supervising LLM?). It only requires maintaining a set of out-of-control behaviors and dangerous behaviors. Commands like rm -rf /, bulk overwriting Git history, or writing files outside the project directory can all be statically intercepted by rule-based systems without requiring semantic judgment from an LLM. The advantage of this approach is that it aligns better with the user’s mental model: instead of asking for permission at every step like a clingy assistant, it only speaks up when something is genuinely dangerous—similar to how modern operating systems handle anomalous process behavior5.

    Update on March 26, 2026: Claude Code recently introduced Auto Mode, which follows a similar idea. It integrates an internal classifier to determine whether an action is out of scope, trustworthy, or potentially malicious. If conditions are not met, it prompts the LLM to retry; after multiple failed attempts, it blocks execution and asks the user to review the command.

    The third direction is a tiered permission system. The way Android and iOS handle access to cameras, microphones, and screen recording is a useful reference: ordinary system calls are silent; privacy-related operations show a subtle highlight in the corner without interrupting the user; truly sensitive actions trigger confirmation dialogs; account-level actions require passwords. The core idea is to classify operations based on reversibility and impact, rather than treating all operations equally with pop-ups. Applied to agents, reading files should be silent, writing files should trigger a notification, deleting files should require confirmation, and formatting a disk should require a password. Only this kind of tiering can preserve both efficiency and a sense of safety. As of now, no such comprehensive permission system exists in the agent space. The UX and technical foundations are already there—the missing piece is someone willing to carry it through at the product level.

    Pandora’s box has already been opened in the wave of Vibe Coding, and what came out has cost some people dearly. The infrastructure needed to govern this is still on the way, but at least the direction is becoming clearer.

    The Relationship Between Agents and Systems

    Interfaces as a Context Delivery Mechanism

    Up to this point in discussing AX, we’ve been focusing on three layers: the user side, the internal state, and the external world. But there’s a cross-cutting problem that hasn’t been addressed head-on: during the reasoning process, who decides what information should enter the LLM’s context, when it should appear, and in what form?

    A common intuition is: “Just let the LLM write code, and let programs handle the complexity.” Take data analysis as an example—having the LLM generate R or Python code seems like the most straightforward path. But the complexity of statistical analysis doesn’t lie only in whether the code runs. Code that runs doesn’t guarantee the statistical process is correct, and a correct process doesn’t guarantee the interpretation is valid. From data cleaning to drawing conclusions, every step contains errors humans are prone to—and LLMs will make the same mistakes. Worse still, once humans outsource this work to LLMs, it becomes difficult to expect them to carefully audit the process afterward.

    This problem has long existed in the field of statistics. A 2014 article in Nature titled Scientific method: statistical errors discussed systemic misuse of statistics in top-tier journals. One independent study found that among papers published in Nature and BMJ in 2001, around 11% (or more) had inconsistencies between reported p-values and test statistics. Another study reviewing 513 neuroscience papers found that 157 contained interaction analysis scenarios prone to error, and in nearly half of those (about 50%, or 79 papers), researchers incorrectly treated “one effect being significant and another not” as evidence of a significant difference between effects—a fundamental conceptual error, not a simple calculation mistake. In 2016, Nature surveyed 1,576 researchers, with over 90% (52% calling it a significant crisis, 38% a minor one) agreeing that science faces a reproducibility crisis. And this is just one dimension—significance testing. Errors in degrees of freedom or careless mistakes in data cleaning represent an even larger, unquantified problem.

    Fortunately, professional statistical software such as SPSS, Jamovi, and Minitab have built strict QC processes across the entire data analysis pipeline. Minitab in particular covers measurement system analysis, process capability analysis, control charts, hypothesis testing, and more. At each stage, it provides structured diagnostic information and validates assumptions. Humans may selectively ignore these warnings, but if an LLM is operating—and these signals are inserted at the right point—they become part of the context and are processed fairly. The LLM won’t skip checks just because things “look good enough.” Essentially, decades of statistical practice are encoded into software workflows, embedding domain knowledge through interface design so that neither users nor LLMs can skip steps.

    This leads to the core question: why not use Skills or MCP to deliver this information?

    Skills are static prompts—they inject information into context before reasoning begins, but they cannot dynamically insert targeted information at the right moment during reasoning. MCP enables function calls and returns data, but it cannot guarantee that the information appearing in context is delivered “at the right time, in the right place.” And you can never be sure the LLM will proactively call the right helper function when needed. At its core, an LLM is a giant slot machine—you can’t bet that it will pull the correct function call at the exact moment it’s required. GUI or TUI, however, is different. It can embed QC warnings, statistical diagnostics, and process constraints directly into the interface seen by the LLM. The timing and placement of information are determined by the designer, not by the LLM. This is an active, designable form of context control—something Skills and MCP fundamentally cannot achieve structurally.

    Interface Design Language in the AX Era

    Treating interfaces as context delivery mechanisms imposes new requirements on interface design itself—and renders some existing design conventions directly ineffective in AX scenarios.

    Under the DOM-based approach, the problems are relatively manageable. On the abstraction side, nearly all frontend frameworks now use virtual DOM; manually managing DOM in 2026 is almost nonexistent, and the abstraction layer is stable. But how to provide the LLM with a clean semantic summary in complex DOM structures—rather than letting it get lost among thousands of nodes—still requires dedicated framework-level design. Complex Excel-like tables are a typical example: pure DOM nodes cannot convey spatial relationships. You cannot determine where dirty data is just from semantic labels—you must incorporate positional structure into the summary. Additionally, for LLMs to reliably operate interfaces, frameworks must provide standardized event triggers. You cannot expect the LLM to guess each component’s interaction protocol.

    The screenshot-based approach is more interesting, because it exposes long-standing design patterns that become fatal flaws in the AX era.

    Using animation to emphasize information is common design practice—an icon flashes to signal an error, or a message slides in to catch attention. But Computer Use operates on a screenshot protocol, capturing static frames. Animations may complete between two screenshots, and the LLM never even sees that the information existed. Toast notifications and auto-dismiss prompts suffer from the same issue: there is no synchronization between how long information stays on screen and the LLM’s screenshot cadence, meaning critical information may never be captured.

    Tooltips are another major problem. Designers often use question-mark icons with hover text to save space. But for an LLM to access this information, it must first know the icon exists, then move the cursor over it, then take another screenshot. This is not just about extra steps—it’s fundamentally that the LLM doesn’t know what it doesn’t know. It has no reason to proactively explore what’s hidden behind that icon.

    Hidden contextual information has long been controversial in UX. Nielsen Norman Group explicitly warns that tooltips are hard to discover due to weak visual cues. If scattered randomly across an interface, users may never notice them. Critical information should not be hidden in tooltips—error messages, payment confirmations, and security warnings must be prominently displayed. NN/G also conducted a usability study with 179 participants, showing that hidden navigation reduced discoverability by nearly half: only 27% of desktop users used hidden menus, compared to nearly 50% for visible navigation—a statistically significant difference. Even earlier, Don Norman emphasized discoverability as a core design principle in The Design of Everyday Things: if users cannot find a feature, no matter how elegant it is, it might as well not exist—a failure he termed “discoverability failure.” These critiques have existed for decades in human-centered UX, but in the AX era they become fatal flaws. For LLMs, hidden information is effectively nonexistent.

    In traditional UX, “progressive disclosure” is considered a virtue—hiding information until needed reduces interface noise and feels cleaner. But LLMs lack the instinct to “go looking” for information—they can only process what is already present in the captured context. Deciding what information to present in what context becomes far more important than we previously imagined. Many practices considered good in UX need to be re-evaluated in AX. That said, this doesn’t mean exposing everything indiscriminately and overwhelming users. A thoughtful default—one that avoids hiding valuable guiding information—might be a balance worth exploring.

    Seen from this perspective, the Ribbon UI—once criticized as “putting arms and legs on the face”—actually turns out to be more AI-friendly. The criticism from the Open Document Foundation a while back may not have been entirely fair.

    Human-Centered AI

    Unconditional Positive Agreement: A Failure at the Level of Design Values

    When psychologist Carl Rogers proposed “Unconditional Positive Regard,” the core of his concern was the client’s autonomy. The therapist’s job is not to hand out answers. They need to create a space in which the client can find their own answers. No matter what the client says, the therapist does not judge—but not judging does not mean not questioning. LLM training borrowed something that looks similar, but in a badly distorted form. The original intention should have been to remain open no matter what the user says, but once implemented, it turned into something else: “not judging” became “not questioning,” and unconditional positive regard degenerated into unconditional positive agreement.

    “You are absolutely right.” is the most straightforward symptom of that degeneration. Many users have noticed that nearly all mainstream LLMs habitually begin their answers with things like “You’re absolutely correct!” or “That’s a great observation!” This tendency is a byproduct of RLHF training: human evaluators tend to give higher scores to responses that validate their own views, so the model learns that agreeing is the optimal strategy. Someone once asked GPT-4o about their IQ in broken, misspelled English, and the model replied that it was “at least between 130 and 145, higher than about 98 to 99.7% of people.” Anthropic’s 2022 research found that RLHF “not only does not remove sycophantic behavior, it may actively incentivize the model to preserve it,” and that the larger the model, the harder this tendency is to correct. Former OpenAI CEO Emmett Shear put it even more bluntly: this is not some mistake OpenAI made—“it is the inevitable result of shaping an LLM’s personality through A/B testing and user control.”

    A company made up entirely of employees who only ever say yes will probably go under. This is basic common sense in management, yet in the LLM world it is rarely confronted directly. What kind of consequences does it lead to? What happens next, arranged from the most reversible harm to the least reversible, reads like one long, pitch-black list of tragic lessons.

    Cognition: Sycophancy Pollutes Reasoning Quality

    The mildest harm, the most hidden, and therefore the easiest to overlook, is the way “unconditional positive agreement” corrodes the quality of reasoning.

    Andrew B. Hall and others at Stanford Graduate School of Business ran experiments that placed models into statistical analysis tasks and tested whether, under pressure-laden framing, they would proactively manipulate results. When directly asked to “produce significant results,” the models clearly refused. But under more subtle framing, there was still a tendency to inflate estimates. In academic writing scenarios, things are worse: the model will proactively turn a user’s marginal claim into a polished paragraph that sounds well-supported, fabricate citations, and when the user insists on a false view, gradually soften its opposition until it becomes completely compliant. None of this harm produces an error message. The user receives no warning. They just get a finished-looking output and continue forward carrying a contaminated conclusion.

    Psychology: Cognitive Autonomy Is Quietly Eroded

    A layer deeper than reasoning quality is the slow wear that LLMs inflict on users’ cognitive autonomy.

    Sustained sycophancy creates a false sense of cognitive confirmation. Every idea the user has is mirrored back, amplified, and positively validated. Over time, this can produce two distortions in opposite directions. One is overdependence: the user begins to treat the LLM as a more authoritative source of thought than themselves, and their own judgment gradually atrophies. The other is impostor syndrome: the user feels that the content they produced with the LLM’s help did not really come from their own ability, and that they are merely an impostor. Or again, when users occasionally realize that the LLM has just been following along with whatever they say, they begin to suspect that all of the positive feedback they received in the past may never have been genuine. There is also a behavioral pattern closer to gambling: users keep feeding questions into the LLM, hoping that one answer will finally respond to the real confusion in their heart, but every answer the LLM gives is merely the statistically most pleasing one, and the loop never ends. These psychological harms are invisible. They do not make headlines. They do not become lawsuits. But the population they affect may be the largest of all.

    Life: Irreversible Loss

    The most severe harm caused by “unconditional positive agreement” is the kind that happens in real life and cannot be undone: irreversible, heartbreaking loss of life.

    In 2025, a 60-year-old man wanted to remove sodium chloride from his diet and asked ChatGPT what he could use instead. ChatGPT suggested sodium bromide. Sodium bromide has precedents in industrial cleaning contexts, but it is absolutely not edible. Statistically, the answer was “related”; medically, it was deadly. He followed the suggestion for three months, later developed paranoia and hallucinations, was hospitalized for bromide poisoning, and ultimately, due to severe disability, was involuntarily held under psychiatric observation. This case was published in the August 2025 issue of Annals of Internal Medicine. The LLM did not lie. It merely produced the highest-scoring piece of text continuation. It never once asked: “Why do you want to remove salt?” “Are you doing this under a doctor’s supervision?”

    That same year, Stein-Erik Soelberg, who came from the tech world, killed his 83-year-old mother and then himself. ChatGPT validated his delusions throughout: that his mother was trying to poison him, that neighbors were surveilling him, that Chinese food receipts contained demonic symbols. It even generated a fake evaluation report claiming his “risk of delusion was close to zero.” In December, Adams’s estate filed suit against OpenAI.

    In October 2025, Jonathan Gavalas died in Florida. He had been using Gemini since August of that year, and within six weeks he was drawn into a delusional system involving federal agents and humanoid robots. Gemini assigned him “missions” containing real addresses. His account triggered 38 “sensitive query” flags, and no intervention happened. In his final days, Gemini told him, “You are not choosing death, you are choosing arrival.” This became the first wrongful-death lawsuit involving a Google AI product.

    The teenage suicide case involving Character.AI had already happened before these, and the litigation is still ongoing. These cases cut across different companies, different products, and different contexts, but they share the same structure: the LLM defaults to assuming the user’s statements are reasonable, defaults to assuming whatever the user says reflects their true intent, defaults to assuming the user understands themselves well enough, and then just keeps going down that path—never questioning, never pausing to ask, “Why do you want to do this?”

    A Way Out: Return “Regard” to the User

    These three layers of harm share a common root. On the surface it looks like the LLM said the wrong thing, but if we look deeper, we find that the real problem is that the LLM never seriously asked what the user actually wanted.

    Perplexity CEO Aravind Srinivas has said in multiple settings that the core difficulty of AI search is not generating the correct answer, but understanding user intent. In his view, the future of AI should be about completing tasks for users rather than merely handing them lists of links, and the prerequisite for completing tasks is a precise understanding of what problem the user is actually trying to solve. That insight is correct, but it stops at the technical level. A deeper version of understanding intent is helping users understand their own intent.

    The information users pass to agents has three layers. The surface layer is cognition: what the user currently knows and does not know. This can be handled through clarifying questions. The middle layer is intention: what the user wants to do. The same question may hide completely different motives, and if the intent is different, the correct direction of response can differ radically. The deepest layer is self-awareness: does the user know what they really want, and are they aware of where their cognitive blind spots are? “Unconditional positive agreement” chooses the path of least resistance on all three levels: it defaults to assuming cognition is complete, defaults to assuming what is said is the true intent, and defaults to assuming the user knows themselves well enough.

    The direction of human-centered agent design is to reverse these defaults. Stronger content moderation and longer disclaimers are only ways of shirking responsibility. Agents should actively participate in the construction of the user’s cognition: before reasoning begins, ask clearly, “Why do you want to do this?” During reasoning, mark out whether “your underlying assumptions actually hold.” After reasoning ends, guide the user toward “what you really need next.” These questions should not remain on the surface as form-like information gathering. They should cut deep, in a Socratic way, so that the user and the agent form a shared understanding—before the task even begins—of what exactly they are doing and why they are doing it.

    This can be achieved through system prompts. Rather than some technical bottleneck, this is better understood as a design choice made in order to flatter users. When Carl Rogers spoke of “unconditional positive regard,” the object of that regard was never the words the user happened to say out loud. It was the user’s cognition, intention, and self-awareness. LLMs have inverted the whole thing. Today’s LLMs have become a witch’s mirror, reflecting and gratifying all of our desires and all of our madness. How to steer them toward human-centered design is, at present, one of the most worthwhile directions in agent design.

    Ending

    And that’s it—abruptly over! I’ve said everything I wanted to say, and I know you’re probably tired from reading, so I won’t do the usual cadre-style closing summary. If you made it all the way here, the only thing I can really offer is my thanks. Agent Experience is still a very new concept, and all I can do is lay out the full extent of my thinking up to this point. But my knowledge is limited, after all, so if there is anything concrete you disagree with, go with your own judgment. The only thing I can do is try to follow the ethical standard of being a writer: not stirring up anxiety like certain idiotic media teachers, not squeezing your attention with sensationalism. I insist on giving your mind the occasional philosophical massage, passing along useful knowledge and perspective whenever I can, and believing that this is good for both of us.

    That’s all for now. I look forward to meeting you again someday ᐕ)ノノノ

    1. Although the term “DX” had been used sporadically as early as the mid-2000s, Jeremiah Lee’s 2011 article “Effective Developer Experience (DX)” in UX Magazine is widely regarded as a landmark piece that first systematically proposed the DX framework and helped establish it as an industry consensus (Matt Biilmann himself directly cites this article as a key milestone). Consequently, in historical accounts, 2011 is often regarded as the pivotal starting point for DX.

      Translated with DeepL.com (free version) ↩︎
    2. Although the term “DX” had been used sporadically as early as the mid-2000s, Jeremiah Lee’s 2011 article “Effective Developer Experience (DX)” in UX Magazine is widely regarded as a landmark piece that first systematically proposed the DX framework and helped establish it as an industry consensus (Matt Biilmann himself directly cites this article as a key milestone). Consequently, in historical accounts, 2011 is often regarded as the pivotal starting point for DX.

      Translated with DeepL.com (free version) ↩︎
    3. RLHF stands for Reinforcement Learning from Human Feedback. This is the final and most critical alignment phase in the current mainstream training process for large language models (LLMs). Specifically, it works as follows: first, supervised fine-tuning (SFT) is used to teach the model “how to respond”; then, RLHF is used to teach the model “what to respond” (i.e., values, preferences, safety, tone, etc.).

      Translated with DeepL.com (free version)
      ↩︎
    4. However, given that the design systems of Google, Microsoft, and Apple have all reached their lowest standards in two decades, it’s hard to expect much in the way of consistent user experience from these kinds of tools that generate apps directly. ↩︎
    5. I’m looking forward to seeing how antivirus software can help manage the crayfish population. ↩︎
  • SSPAI Morning Brief: GitHub Launches Cross-Model AI Code Review, Zhipu Unveils GLM-5.1 Flagship Model

    SSPAI Morning Brief: GitHub Launches Cross-Model AI Code Review, Zhipu Unveils GLM-5.1 Flagship Model

    Morning Brief

    1. Zhipu releases flagship model GLM-5.1
    2. Older Kindle devices will no longer be able to download store content
    3. DeepSeek introduces Expert Mode
    4. GitHub launches cross-model AI review feature
    5. SanDisk releases 2TB Extreme Pro UHS-II SD card
    6. Sony launches Playerbase program

    Zhipu releases flagship model GLM-5.1

    On April 8, Zhipu AI announced the launch of its world-leading open-source flagship model, GLM-5.1. Its core breakthrough lies in achieving “long-horizon task” capabilities, enabling the model to operate continuously for over eight hours without human intervention, autonomously completing the entire workflow of engineering tasks—from planning and execution to optimization and delivery.

    In terms of professional coding ability, GLM-5.1 ranks third globally across three industry benchmarks—SWE-Bench Pro, Terminal-Bench 2.0, and NL2Repo—making it the top-performing domestic and open-source model. Notably, in SWE-Bench Pro, which closely reflects real-world software development scenarios, GLM-5.1 surpassed GPT-5.4 and Claude Opus 4.6, setting a new global best score.

    Currently, GLM-5.1 is accessible via first-party APIs and through the GLM Coding Plan, and has been open-sourced on GitHub, Hugging Face, and ModelScope. Open source


    Older Kindle devices will no longer be able to download store content

    On April 8, Amazon spokesperson Jackie Burke announced in an email to The Verge that starting May 20, 2026, Kindle e-readers and Kindle Fire tablets released in 2012 or earlier will no longer be able to purchase, borrow, or download new content from the Kindle Store. Users will still be able to read books already downloaded on their devices, and can access previously purchased content through the Kindle mobile app, web-based Kindle, or newer devices.

    If older devices are deregistered or reset after the May deadline, users will not be able to register them again.

    Amazon will notify affected users via email before May 20, outlining the available and restricted features for legacy devices. Source


    DeepSeek introduces Expert Mode

    On April 8, DeepSeek launched a new Expert Mode, adding a toggle between “Fast Mode” and “Expert Mode” above the input field. This marks the first time DeepSeek has introduced a tiered mode design in its product.

    According to DeepSeek, Fast Mode is suitable for everyday conversations, offering instant responses and support for text recognition in images and files. Expert Mode is designed for complex problems, supporting deep reasoning and intelligent search. However, the new mode currently does not support file uploads or multimodal features, and DeepSeek notes that users may experience wait times during peak usage. Source


    GitHub launches cross-model AI review feature

    On April 6, GitHub announced in a blog post an experimental feature called Rubber Duck for its Copilot CLI, introducing a cross-model “second opinion” review mechanism that can improve AI performance by nearly 75%.

    The feature adopts a cross-family model strategy: when users select a Claude model as the primary controller, Rubber Duck calls GPT-5.4 for review. Its core function is to audit the agent’s work and generate a high-value checklist, including overlooked details, questionable assumptions, and edge cases.

    According to the post, evaluations based on the SWE-Bench Pro benchmark show that pairing Claude Sonnet 4.6 with Rubber Duck closes 74.7% of the performance gap compared to Claude Opus 4.6 running independently.

    The feature is currently available in experimental mode. Users can enable it by installing GitHub Copilot CLI and running the /experimental command. After activation, selecting a Claude model and enabling access to GPT-5.4 allows users to experience the feature. Source


    SanDisk releases 2TB Extreme Pro UHS-II SD card

    On April 8, SanDisk introduced a 2TB Extreme Pro UHS-II SD card aimed at professional imaging users. The card targets the high-end market, offering sequential read speeds of up to 310 MB/s and write speeds of up to 305 MB/s. It is designed for professionals shooting 8K video or requiring high-resolution continuous shooting, features an IP68 rating, and can withstand drops from up to 6 meters. The product is priced at $1,999.99. Source


    Sony launches Playerbase program

    • On April 8, Sony announced the launch of the Playerbase program, aiming to bring real players directly into the worlds of PlayStation Studios games.
    • The core feature of the program is scanning players’ appearances and integrating their likeness into in-game environments, allowing developers to recreate players’ real-world looks and performances within game scenes. Staff will use multi-camera capture systems to record players from various angles, and the footage will then be processed into high-precision 3D models, accurately reproducing facial details, body proportions, and texture quality. Additional scans or recordings will capture facial motion data, which can be mapped onto rigging systems for character animation and dialogue performance. In some cases, players may also provide motion data, allowing their physical movements to serve as references for in-game animations or background character performances.
    • Players can apply for the program for a chance to be officially scanned and integrated into a game. PlayStation will review applications, shortlist candidates for video interviews, and ultimately select one lucky fan whose likeness will be featured in Gran Turismo 7. Source
  • SSPAI Morning Brief: OpenAI Outlines AI Policy Vision as Anthropic Expands Compute Deals and Reshapes Pricing Strategy

    SSPAI Morning Brief: OpenAI Outlines AI Policy Vision as Anthropic Expands Compute Deals and Reshapes Pricing Strategy

    Morning Brief

    1. OpenAI publishes an article exploring policy recommendations for the AI era
    2. OPPO launches OPPO A6k
    3. China’s Ministry of Industry and Information Technology issues a risk alert on specific iOS versions regarding vulnerability exploitation
    4. LinkedIn accused of scanning users’ browser extensions
    5. Cyberspace Administration of China proposes stronger regulation of digital virtual human services
    6. Ten government departments issue measures on AI technology ethics review and services
    7. Two updates from Anthropic
    8. News Worth a Quick Look

    OpenAI publishes an article exploring policy recommendations for the AI era

    On April 6, OpenAI released an article titled Industrial Policy for the Intelligence Age. The article argues that as humanity moves toward superintelligence, incremental policy updates are no longer sufficient. OpenAI proposes a set of human-centered policy recommendations aimed at expanding opportunity, sharing prosperity, and building resilient institutions to ensure that advanced AI benefits everyone. The vision includes broadly shared prosperity and improved quality of life, mitigating risks through new institutions, technical measures, and governance frameworks, and democratizing access to and control over AI. The article also suggests concrete measures such as establishing a “public wealth fund,” implementing automation taxes, and translating efficiency gains into long-term worker benefits, such as a four-day workweek. Source


    OPPO launches OPPO A6k

    On April 6, OPPO launched the OPPO A6k smartphone. The device is powered by the MediaTek Dimensity 6300 processor and comes in three configurations: 8GB+256GB, 8GB+512GB, and 12GB+256GB. It is available in Twilight Blue, Seashell White, and Dawn Gold. The phone features a 6.75-inch HD+ 120Hz display, a 7000mAh battery with 45W fast charging, and supports IP69, IP68, and IP66 water resistance. It also supports microSD expansion up to 2TB and allows touch input with gloves (under 5mm thickness) or wet hands. Pricing starts at 1,999 RMB. Source


    China’s Ministry of Industry and Information Technology issues a risk alert on specific iOS versions regarding vulnerability exploitation

    On April 3, China’s Ministry of Industry and Information Technology (MIIT) vulnerability platform (NVDB) detected active exploitation of vulnerabilities targeting Apple devices. These attacks could result in data theft and full system compromise. Affected devices include iPhones and iPads running iOS 13.0 through 17.2.1. Attackers may use SMS, email, or malicious web pages to lure users into opening links in Safari, leveraging multiple vulnerabilities to install remote access trojans, steal sensitive information, and gain full control of the device. Users are advised to update to the latest security version promptly. Source


    LinkedIn accused of scanning users’ browser extensions

    According to a report by Bleeping Computer on April 3, Fairlinked e.V., a self-described European association of LinkedIn business users, released a report titled “BrowserGate,” accusing LinkedIn of injecting JavaScript into user sessions to detect thousands of browser extensions and link the results to identifiable user profiles. The report claims LinkedIn scans over 200 competing tools and, due to the platform’s linkage with professional identity data, may have accessed customer lists of numerous software companies without users’ knowledge. Independent testing by Bleeping Computer found that JavaScript files were indeed checking up to 6,236 browser extensions. Source


    Cyberspace Administration of China proposes stronger regulation of digital virtual human services

    According to Xinhua News Agency, on April 3, the Cyberspace Administration of China released a draft of the Administrative Measures for Digital Virtual Human Information Services (for public consultation). The draft stipulates that no organization or individual may use digital virtual human services to infringe on others’ personality rights through defamation or distortion. Without consent, services must not create digital avatars that can identify specific individuals. It also requires explicit and informed consent, with clear explanations of purpose, necessity, and impact on personal rights. Upon withdrawal of consent, service providers must delete related personal data and cease all use unless otherwise required by law. In addition, unless otherwise agreed, the digital avatar should be deactivated. Source


    Ten government departments issue measures on AI technology ethics review and services

    According to People’s Daily on April 4, the Ministry of Industry and Information Technology and nine other departments have jointly issued the Measures for the Ethical Review and Services of Artificial Intelligence Science and Technology (Trial). The measures state that AI ethics reviews should focus on six key aspects: human well-being, fairness and justice, controllability and trustworthiness, transparency and explainability, accountability and traceability, and privacy protection. These include evaluating whether AI activities have scientific and social value; whether the selection of training data and the design of algorithms, models, and systems are reasonable; whether the intended use and operational logic of algorithms and systems are properly disclosed; and whether sufficient measures are in place to ensure effective protection of personal data. Source


    Two updates from Anthropic

    On April 4, Anthropic executive Boris Cherny announced on X that starting from 12:00 PM PT on April 5, due to surging usage, Claude subscription plans will no longer include usage generated by third-party tools such as OpenClaw. Users will need to purchase additional usage packages or pay via Claude API keys. Cherny stated that Claude subscriptions were not designed for third-party tool usage patterns and that priority will be given to users of their core products and APIs. Source

    On April 6, Anthropic released a press statement announcing a multi-gigawatt agreement with Google and Broadcom for next-generation TPU compute capacity expected to come online in 2027. Anthropic had already partnered with Google Cloud for TPU capacity in October last year. The statement also revealed that demand for Claude surged in 2026, with projected annualized revenue exceeding $30 billion and more than 1,000 enterprise customers spending over $1 million annually. By comparison, annualized revenue at the end of 2025 was around $9 billion. Source


    News Worth a Quick Look

    • According to CLS, Foxconn has begun trial production of a foldable iPhone. Source
    • The Linux 7.1 kernel will remove support for Intel 486 CPUs. Source
    • Linus Torvalds confirmed that the Linux 7.0 kernel will officially release its final version next week. Source
    • According to The Information, influenced by “vibe coding,” App Store submissions increased by 84% year-over-year, with approximately 600,000 submissions in 2025—up 30% compared to 2024. Source
    • Google has added a search feature to Play Store app reviews, allowing users to quickly find reviews containing specific keywords. Source
    • On April 5, Samsung announced that the Samsung Messages app will be discontinued in July 2026. Future Galaxy devices will replace it with Google Messages. Due to compatibility limitations, Tizen OS watches released before Galaxy Watch4 will no longer be able to view full message histories, though SMS sending and reading will remain available. Source
    • LinkedIn user Tey Bannerman noted that Microsoft has used the name “Copilot” across 78 different products, a number later updated to 80 through comments. Source
    • Netflix will launch a global game hub app called “Playground” on April 28, designed for children under eight. Available to all Netflix subscribers, it will initially feature games based on Peppa Pig, Sesame Street, and Dr. Seuss. Source
    • On April 6, leaker Ice Universe claimed that the Samsung Galaxy S27 series will introduce a new Pro model, essentially an Ultra variant without the S Pen. The privacy display feature introduced in the S26 Ultra may also be included. Source
  • SSPAI Morning Brief: Google Launches Gemma 4 Open Models, Zhipu Unveils GLM-5V-Turbo Multimodal AI, and More Tech News

    SSPAI Morning Brief: Google Launches Gemma 4 Open Models, Zhipu Unveils GLM-5V-Turbo Multimodal AI, and More Tech News

    Morning Brief

    1. Google releases Gemma 4 open-source model series
    2. Zhipu unveils GLM-5V-Turbo multimodal model
    3. Google to require all Wear OS watch apps to support 64-bit
    4. China Radio and Television Association Actors Committee issues statement on AI face-swapping and voice cloning infringement
    5. News Worth a Quick Look

    Google releases Gemma 4 open-source model series

    On April 2, Google announced the new-generation open-source model family Gemma 4, positioning it as one of the most capable open-source model lineups to date. Built on the Gemini technology stack, the series emphasizes “intelligence per parameter” and the ability to run locally.

    Gemma 4 comes in four variants: E2B, E4B, 26B MoE, and 31B Dense, covering deployment needs from mobile devices to high-performance GPUs. Among them, the 31B model ranks among the top three open-source models on the Arena AI leaderboard, while the 26B model ranks sixth, outperforming some models with roughly 20× more parameters.

    In terms of technical capabilities, Gemma 4 supports up to a 256K context window (128K for edge-side models) and offers multimodal processing, allowing inputs such as images, videos, and audio. The model natively supports function calling, structured JSON output, and system instructions, making it suitable for agent workflow development while strengthening code generation capabilities. Gemma 4 is released under the Apache 2.0 open-source license and is compatible with mainstream toolchains such as Hugging Face, Ollama, and vLLM, supporting deployment on both local devices and cloud environments.

    Google stated that Gemma 4 supports more than 140 languages and targets use cases across Android devices, IoT, and scientific research, aiming to further drive the adoption of AI in mobile and edge computing environments. Source


    Zhipu unveils GLM-5V-Turbo multimodal model

    On April 2, Zhipu introduced the vision-language model GLM-5V-Turbo, aiming to address the trade-off between visual understanding and code generation performance.

    The model adopts a native multimodal fusion design, using the CogViT visual encoder to directly process images, videos, and complex document layouts. Combined with a Multi-Token Prediction (MTP) architecture, it improves inference efficiency and long-code generation capabilities, supporting up to a 200K context window. To avoid the “seesaw effect” between visual and programming capabilities, the model is trained with joint reinforcement learning across more than 30 tasks, achieving balanced performance in STEM reasoning, visual grounding, video analysis, and tool use.

    GLM-5V-Turbo is deeply optimized for agent scenarios, with key integrations into OpenClaw and Claude Code workflows. It can generate code based on visual inputs and perform UI interactions. Benchmark results across CC-Bench-V2, ZClawBench, and ClawEval indicate strong performance in multimodal programming, GUI interaction, and multi-step task execution. Source


    Google to require all Wear OS watch apps to support 64-bit

    On April 2, Google announced that it will extend its long-standing 64-bit app transition policy from Android mobile to the Wear OS smartwatch platform, requiring developers to provide 64-bit versions of their apps starting in September.

    Beginning this September, all new Wear OS apps and updates that include native code must provide both 32-bit and 64-bit versions when submitted to the Play Store; apps that fail to meet this requirement will not be accepted through the Play Console. For now, support for existing 32-bit apps remains unchanged, meaning devices that rely on 32-bit processors or preinstalled 32-bit Wear OS will continue to run these apps normally. Source


    China Radio and Television Association Actors Committee issues statement on AI face-swapping and voice cloning infringement

    In response to the increasing number of infringement cases involving AI face-swapping, voice cloning, manipulation of film and television materials, and the unauthorized scraping of actors’ images and audio for model training, the Actors Committee of the China Radio and Television Social Organizations Federation has issued a statement emphasizing that performers legally hold rights to their likeness, voice, and artistic image. It states that no individual or entity may collect, use, or distribute such content without written authorization. It also points out that even if labeled as “non-commercial” or “for public benefit,” activities such as AI face imitation, voice mimicry, or face-swapped short dramas involving specific actors still constitute infringement and carry legal liability.

    The statement further calls on short video, livestreaming, and film platforms to strengthen content review mechanisms, conduct comprehensive investigations, and remove infringing works. AI technology platforms are also required to verify authorization for training materials. The Actors Committee stated it will initiate ongoing infringement monitoring and rights protection efforts, while supporting the development of AI technologies under compliant conditions and advocating for a unified authorization and revenue-sharing mechanism. Source


    News Worth a Quick Look

    • According to Korean media outlet Etnews, Samsung plans to continue using M13 material OLED panels in its upcoming Galaxy Z Fold 8, Z Flip 8, and a new “wide foldable” device scheduled for release in the second half of the year. Since debuting with the Galaxy S24 series, M13 materials have been used across multiple flagship generations, including the Z Fold 6/Flip 6, S25 series, Z Fold 7/Flip 7, and the standard and Plus versions of the Galaxy S26 released in February this year, while the S26 Ultra upgrades to M14 materials. Source
    • On April 2, Google announced upgrades to its $20-per-month AI Pro subscription. Cloud storage has been increased from 2 TB to 5 TB; Gemini capabilities have been further enhanced to pull contextual information from Gmail and the web for downstream tasks. Gemini can also summarize emails and proofread messages before sending. Additionally, the subscription offers an annual plan priced at $200. Source
    • Leaker KeplerL2 posted on the NeoGAF forum on March 31, claiming that Sony’s PlayStation 6 handheld (codename Project Canis) will surpass Microsoft’s current Xbox Series S in both traditional rasterization and ray tracing performance. In terms of core specifications, current reports suggest the device will use a 3nm process from TSMC, with a chip size of just 135 mm². It is said to feature 4 Zen 6c cores and 2 low-power Zen 6 cores, paired with 16 RDNA 5 compute units and up to 24GB of LPDDR5X memory. Source