« All posts

» Summary

Aug 1, 2026

Aug 1, 2026
Today

221,000 Live Secrets Found in Hugging Face Datasets; Nvidia Vera CPU Details Emerge

A massive security scan across every public dataset on Hugging Face uncovered 221,303 live, unique credentials—ranging from cloud admin keys to CI/CD tokens—scattered through 7.6 PB of training data. Among the findings were 349 live GitHub tokens with write or admin scopes, 318 Docker Hub push tokens, and 8,557 active GCP service-account keys spanning thousands of projects. The scale of the leak highlights the risky practice of embedding secrets in publicly shared datasets.

In a separate cloud security disclosure, researchers detailed an AWS EKS privilege-escalation path. An attacker with a foothold in a single pod can query the instance metadata service for an IAM token, then exchange it for an EKS token carrying the system:node role. By using `kubectl create token` with bound-object parameters, they can obtain powerful service-account tokens and ultimately gain cluster-admin. The attack bypasses the NodeRestriction feature and demonstrates the need to lock down pod metadata access.

On the hardware front, Nvidia revealed architectural details of its Vera CPU—its first server chip sold independently of GPUs. The 88-core Armv9.2 design supports 176 threads and up to 1.5 TB of LPDDR5X memory, with a monolithic die approach rather than chiplets. Alibaba, ByteDance, Meta, Oracle, CoreWeave, and others have already committed to deployments, positioning Vera as a direct competitor to x86 chips from Intel and AMD.

AI inference efficiency received attention from two angles. Nvidia published guidance on co-designing attention kernels for long-context workloads, examining how group size, head dimension, and sequence length affect dense attention performance on its GPUs. Meanwhile, researchers observed that bursty request patterns can inadvertently speed up LLM inference by allowing token separation from large prefills—though the effect stems largely from a kernel optimization artifact rather than a fundamental workload change.

In other news, a Dutch court ruled against Meta under the Digital Services Act for failing to provide easily accessible non-profiled recommender options, a decision with implications for VLOPs across the EU. Spotify engineers shared Random Access Parquet, a key-to-file-location mapping technique to accelerate low-latency point queries on data lakes. The AT Protocol community proposed a permissioned data layer for social apps—offering access control but no end-to-end encryption. A UI library called Morphicons debuted with an API for seamless icon morphing without manual pairing. And an investigative report questioned the real-world efficiency of Bloom Energy’s fuel cells, noting that New York systems averaged only 20 months of acceptable performance over 15 years.

» Statistics

Posts
14
Reads
1
Avg. score
7.6

» Most read

  1. Bursty Arrivals Accelerate LLM Inference Times17.4
  2. 221,000 Live Secrets Found in 7.6PB of Hugging Face Training Data08.5
  3. Exploring Symmetries in a Go Network's Neural Architecture07.0
  4. Indexing the Data Lake for Online Point Queries07.6
  5. WordJS: A Node.js CMS with Secure Plugin Isolation07.3
  6. Morphicons: Seamless Icon Transitions Without Pairing07.8
  7. Ineffective by Design: The Bits of Freedom vs. Meta Case07.9
  8. Inside Nvidia's Vera CPU and its custom Olympus cores08.2
  9. Situated Software: The Future of Design in Social Contexts07.1
  10. Do Bloom's Fuel Cells Perform as Promised? Data Suggests No07.4

» Top scored

  1. 221,000 Live Secrets Found in 7.6PB of Hugging Face Training Data08.5
  2. Inside Nvidia's Vera CPU and its custom Olympus cores08.2
  3. AWS EKS Privilege Escalation: Pod Metadata to Cluster-Admin08.0
  4. Ineffective by Design: The Bits of Freedom vs. Meta Case07.9
  5. Morphicons: Seamless Icon Transitions Without Pairing07.8
  6. Indexing the Data Lake for Online Point Queries07.6
  7. Private Data in ATProto: Permissioned Data Proposal07.5
  8. Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference07.5
  9. Do Bloom's Fuel Cells Perform as Promised? Data Suggests No07.4
  10. Bursty Arrivals Accelerate LLM Inference Times17.4

» Sources

Hashnode #97Frontend Development Reddit1Hashnode #111Hashnode #151Hashnode #31Nvidia Developer Blog1Artificial Intelligence Reddit1The Register1

» Share