Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 15 часов назад
RSI is not happening [R]
RSI is not happening [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

15 часов назад @ reddit.com
How to automatically find the batch size when using Accelerate with FSDP2? [D]
How to automatically find the batch size when using Accelerate with FSDP2? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

16 часов назад @ reddit.com
[P] Wine synthesis using VAE [P]
[P] Wine synthesis using VAE [P] [P] Wine synthesis using VAE [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

18 часов назад @ reddit.com
MS MARCO click-translation expansion tables ("poor man's" DSSM) [P]
MS MARCO click-translation expansion tables ("poor man's" DSSM) [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

20 часов назад @ reddit.com
Duplicating baseline benchmarks [D]
Duplicating baseline benchmarks [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

21 час назад @ reddit.com
[P] Built a 100% Client-Side Vision Pipeline for Real-Time Chessboard & Multi-Board Detection (Chrome/Firefox Extension) [P]
[P] Built a 100% Client-Side Vision Pipeline for Real-Time Chessboard & Multi-Board Detection (Chrome/Firefox Extension) [P] [P] Built a 100% Client-Side Vision Pipeline for Real-Time Chessboard & Multi-Board Detection (Chrome/Firefox Extension) [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

23 часа назад @ reddit.com
PhD branding question [R]
PhD branding question [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 5 hours назад @ reddit.com
ARR August Discussion [D]
ARR August Discussion [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 5 hours назад @ reddit.com
Horse racing as an ML ranking problem: 1.18M runners, walk-forward validation and a very strong market baseline [D]
Horse racing as an ML ranking problem: 1.18M runners, walk-forward validation and a very strong market baseline [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 13 hours назад @ reddit.com
Got scipy's KD-tree to handle inserts and deletes without rebuilding. Three things I learned [P]
Got scipy's KD-tree to handle inserts and deletes without rebuilding. Three things I learned [P] Got scipy's KD-tree to handle inserts and deletes without rebuilding. Three things I learned [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 14 hours назад @ reddit.com
[Upcoming AMA] Waymo AI Team AMA – Drop Your Questions Early! [D]
[Upcoming AMA] Waymo AI Team AMA – Drop Your Questions Early! [D] [Upcoming AMA] Waymo AI Team AMA – Drop Your Questions Early! [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 15 hours назад @ reddit.com
I trained an 825k-parameter model to generate drawing programs that execute exactly on an RP2040 [P]
I trained an 825k-parameter model to generate drawing programs that execute exactly on an RP2040 [P] I trained an 825k-parameter model to generate drawing programs that execute exactly on an RP2040 [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 21 hours назад @ reddit.com
Zachery Lipton: "CS academia broke the system...perhaps all that it takes for the system to rebuild is for it to burn to the ground" [D]
Zachery Lipton: "CS academia broke the system...perhaps all that it takes for the system to rebuild is for it to burn to the ground" [D] Zachery Lipton: "CS academia broke the system...perhaps all that it takes for the system to rebuild is for it to burn to the ground" [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 23 hours назад @ reddit.com
When NeurIPS'26 final decision release? [D]
When NeurIPS'26 final decision release? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 23 hours назад @ reddit.com
Getting Mac Air m2/m3/m4 is it good [D]
Getting Mac Air m2/m3/m4 is it good [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 6 hours назад @ reddit.com
Towards Data Science
последний пост 5 days, 22 hours назад
10 Statistical Traps We Often Overlook
10 Statistical Traps We Often Overlook 10 Statistical Traps We Often Overlook

The average isn’t always the “average”Let’s start with one of the most familiar words in statistics: average.

Your brain wants to confirm the story you already believeThis might be the most important statistical lesson of all.

Statistical significance is broadly about whether an observed difference is unlikely to have arisen under a particular statistical model or null hypothesis.

Learn to ask what isn’t in the datasetPerhaps the most underrated statistical skill is knowing what the data cannot tell you.

The goal of statistical literacy isn’t to become better at finding numbers that support what we already believe.

5 days, 22 hours назад @ towardsdatascience.com
How to Maximize GPT-6 Astra
How to Maximize GPT-6 Astra How to Maximize GPT-6 Astra

I'll discuss how to get the most out of GPT-6 Astra and my first experiences with the model.

I found this very useful and thus immediately switched to GPT-6 as my main driver to perform coding tasks.

How to get the most out of GPT-6 AstraNow, let's start talking about how to get the most out of GPT-6 Astra.

One thing to consider when using GPT-6 Astra is that, of course, it has limited usage.

I was extremely impressed after my first experience with GPT-6 Astra, which I've been using throughout the whole weekend.

6 days, 15 hours назад @ towardsdatascience.com
The Model Validation Playbook for GenAI: Lessons from Banking
The Model Validation Playbook for GenAI: Lessons from Banking The Model Validation Playbook for GenAI: Lessons from Banking

It saves analysts several hours a week, and obviously the business wants this AI model to go live next quarter.

Model risk management was never designed for generative AI in banking.

What model risk management in banking actually doesIf you work in data science outside banking, this discipline may be unfamiliar.

The three questions an AI Model Risk Assessment report answersThe questions are the same ones we have always asked.

Conclusion: The Future of Model Risk Management in the AI EraA valuable model risk review should include adversarial case sets that break the system.

6 days, 16 hours назад @ towardsdatascience.com
A Beginner’s Guide to World Models
A Beginner’s Guide to World Models A Beginner’s Guide to World Models

The first World Model proposed by LeCun was the Joint Embedding Predictive Architecture (JEPA) with the idea that giving an internal representation of how the world functions to a self-supervised model, it could improve its results.

In 2024, Google introduced its World Model Genie, a research prototype that lets you create and explore infinitely diverse worlds.

California's autonomous taxi Waymo adopted Google's Genie to create its own specialized model for self-driving simulation (Waymo World Model).

Then, the latest World Model was released in June 2026 by Nvidia: Cosmos, a family of open-weight models that combine physical reasoning, world simulation, and action generation.

Build a World…

6 days, 16 hours назад @ towardsdatascience.com
Introducing ShipAI
Introducing ShipAI Introducing ShipAI

How do you know an AI project is real?

That's the idea behind ShipAI , which is live on TDS today.

It's a video showcase for AI work: practitioners record a screen-share walkthrough of something they've built (an app, an agent, a pipeline, an experiment), and we review and publish it.

Around the video you'll find:AI-generated key takeaways , so you can decide in ten seconds whether to press play.

And if one of the walkthroughs leaves you thinking "I should share my project, too," that's just what we're after.

6 days, 20 hours назад @ towardsdatascience.com
Context Windows Don’t Know What’s Still True — I Built a Validity Layer That Does
Context Windows Don’t Know What’s Still True — I Built a Validity Layer That Does Context Windows Don’t Know What’s Still True — I Built a Validity Layer That Does

The second executor does work that was already doomed.

The real world only becomes visible when an executor actually takes an action or pays for a verification step.

···Experiment 1: Does Stale Context Actually Cause Wasted Work?

I asked what happens when a fact does not touch the entire graph.

More context does not fix that.

6 days, 22 hours назад @ towardsdatascience.com
I Vibe-Coded an App in Just Two Hours (And Regretted It the Next Day)
I Vibe-Coded an App in Just Two Hours (And Regretted It the Next Day) I Vibe-Coded an App in Just Two Hours (And Regretted It the Next Day)

They can see the lyrics and the Ukulele chords scrolling on the screen.

The best part: you can find Ukulele Chords for almost any song ever sung on the planet.

Ukuflow Home Page —Screenshot by the author Ukuflow is a Vibe coded AI powered App for ukulele learners.

But beginners usually tend to play a dozen songs that are universally considered the 'beginner-friendly ukulele songs.'

For my Ukulele app, this would be that people would want to see lyrics and chords scrollling like they do in Karaoke apps.

1 week назад @ towardsdatascience.com
Why Most Multi-Agent Systems Fail Even When Evaluation Passes
Why Most Multi-Agent Systems Fail Even When Evaluation Passes Why Most Multi-Agent Systems Fail Even When Evaluation Passes

The account-history node calls the billing API, gets back a 200, and passes the payload downstream as if nothing happened.

If you've spent any real stretch of time around multi-agent systems in production, you've either already seen a version of this, or you're going to eventually.

The one that costs you is the 200, the response that says everything's fine, attached to a payload that's structurally fine and semantically garbage.

Does the account ID in the payload actually match the one that was requested?

Give it a week or two, see what the watchdog actually catches, and let that tell you whether it's worth pushing further back into the pipeline.

1 week назад @ towardsdatascience.com
Text Watermarking in Python: Catch Whoever Copies Your Writing
Text Watermarking in Python: Catch Whoever Copies Your Writing

AI companies quietly watermark billions of words a day. Here’s how to apply the same three families of techniques to your own writing—and what real experiments reveal about which watermarks survive copy-paste, editing, and paraphrasing.

The post Text Watermarking in Python: Catch Whoever Copies Your Writing appeared first on Towards Data Science.

1 week, 1 day назад @ towardsdatascience.com
Linear Discriminant Analysis (LDA) in Real-Life: Dimensionality Reduction in a Real-Estate Dataset
Linear Discriminant Analysis (LDA) in Real-Life: Dimensionality Reduction in a Real-Estate Dataset Linear Discriminant Analysis (LDA) in Real-Life: Dimensionality Reduction in a Real-Estate Dataset

Linear Discriminant Analysis: the techniqueIn literature, LDA can also be referred to as Normal Discriminant Analysis or Fisher Linear Discriminant Analysis, the latter being a reference to Ronald A. Fisher, the polymath who developed the criterion LDA aims to maximize.

Linear discriminant function for the multivariate case (Image by author)Similarly, the discriminant function still needs to be a linear function of X.

Given this distinction Linear Discriminant Analysis tends to be a much more robust method for dimensionality reduction [2].

To better understand how each feature contributes to each Linear Discriminant, you decide to normalize the values in each linear discriminant column.

Nor…

1 week, 1 day назад @ towardsdatascience.com
Why Transformers Need Positional Encoding For Time Series: A Visual Guide
Why Transformers Need Positional Encoding For Time Series: A Visual Guide Why Transformers Need Positional Encoding For Time Series: A Visual Guide

For Friday, its query q 5 q_5 q5​ is compared with the keys of all observations: k 1 , k 2 , k 3 , k 4 , k 5 k_1, k_2, k_3, k_4, k_5 k1​,k2​,k3​,k4​,k5​.

Each comparison produces an attention score:s 5 , j = q 5 ⊤ k j d k s_{5,j} = \frac{q_5^\top k_j}{\sqrt{d_k}} s 5 , j ​ = d k ​ ​ q 5 ⊤ ​ k j ​ ​which measures how relevant observation 'j' is when updating Friday’s representation.

These scores are passed through a softmax function to convert them into attention weights:α 5 , j = exp ⁡ ( s 5 , j ) ∑ j ′ exp ⁡ ( s 5 , j ′ ) \alpha_{5,j} = \frac{\exp(s_{5,j})}{\sum_{j'} \exp(s_{5,j'})} α 5 , j ​ = ∑ j ′ ​ exp ( s 5 , j ′ ​ ) exp ( s 5 , j ​ ) ​Finally, those weights are used to combine the va…

1 week, 2 days назад @ towardsdatascience.com
Dynamical System Transfer Learning with Reduced Order Models
Dynamical System Transfer Learning with Reduced Order Models

Improving reinforcement learning for complex physics

The post Dynamical System Transfer Learning with Reduced Order Models appeared first on Towards Data Science.

1 week, 2 days назад @ towardsdatascience.com
Optimal Traffic Allocation Under Heterogeneous Variant Cost
Optimal Traffic Allocation Under Heterogeneous Variant Cost Optimal Traffic Allocation Under Heterogeneous Variant Cost

Every subject in the treatment arm is a few times more expensive than each one in the control arm.

The raw ratio c 1 / c 0 c_1/c_0 c 1 ​ / c 0 ​ is plotted for reference.

And the optimal allocation is one such that V a r ( τ ^ ) Var(\hat\tau) Var(τ^)is as low as we can get it.

If we have reasons not to, then the ratio of standard deviations becomes co-author for the optimal sample ratio, indeed.

A 4x versus 5x cost ratio is a big miss in accounting terms, but it barely moves the optimal split.

1 week, 3 days назад @ towardsdatascience.com
Disaggregation Is a Thousand-GPU Problem
Disaggregation Is a Thousand-GPU Problem Disaggregation Is a Thousand-GPU Problem

The problem was scheduling interference, and chunked prefill handled it without adding a network hop.

Three Costs of Disaggregation: What the Explainers SkipThe explainer articles cover what disaggregation gains.

If your p95 time-per-output-token is within SLO on chunked prefill, you do not have the problem that disaggregation solves.

Conclusion: Start With Chunked PrefillDefault to chunked prefill.

Disaggregation is the right architecture above roughly a thousand GPUs, with fast interconnect, and with the engineering capacity to manage P:D ratio tuning and KV transfer reliability.

1 week, 3 days назад @ towardsdatascience.com
The Power BI Developer's Survival Guide to Microsoft Fabric
The Power BI Developer's Survival Guide to Microsoft Fabric The Power BI Developer's Survival Guide to Microsoft Fabric

I’ve been talking to a lot of Power BI developers in the last 12 months who are truly anxious about Fabric.

I wrote about the two flavors of Direct Lake — Direct Lake on SQL and Direct Lake on OneLake, so you may want to read that one as well.

Your Power Query skills are still the right tool for Power BI data prep.

The PL-300 (Power BI Data Analyst) is still the core PBI certification and is being kept current — it’s still the foundation.

The natural next step for Power BI developers moving into Fabric is the DP-600 (Fabric Analytics Engineer Associate).

1 week, 3 days назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 1 day, 22 hours назад
The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough
The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough

Meta launched a personal agent.

AlphaGenome Atlas applies a different computational strategy: do an enormous amount of work upfront and make the results reusable.

The personal agent runs in a dedicated virtual machine with a browser, can continue working after the app closes, and uses connected services to pursue tasks.

The interesting engineering unit here is the entire system: model, memory, computer, permissions, and execution history.

A useful personal agent needs all of them.

1 day, 22 hours назад @ thesequence.substack.com
The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment
The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment

Imagine bringing a robot into your kitchen and saying, “Help me clean up after dinner.” You have just compressed a remarkable amount of engineering into six words.

The robot must distinguish leftovers from rubbish, discover where plates belong, and work out why a drawer refuses to close.

That kitchen captures the promise of a ChatGPT moment for robotics.

We would be able to give a machine useful new work through conversation and examples, with sufficiently little setup that teaching it becomes an ordinary activity.

The remaining distance involves learning, control, and the economics of getting a machine to work somewhere new.

3 days, 22 hours назад @ thesequence.substack.com
The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure
The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure

In the process, I worked on many projects, one of which was Chatbot Arena, which evolved into Arena.

What does an Arena score actually measure today: model quality, human preference on Arena's traffic, expected usefulness, or something else?

This is how Battle Mode, and the idea of pairwise human preference grading, arose as a paradigm.

Today, the platform has moved far beyond human preferences, ranking models’ task completion rates, hallucination rates, and much more.

Human preference can reward style, confidence, or verbosity even when an answer is wrong.

4 days, 22 hours назад @ thesequence.substack.com
The Sequence Learning Loop - Issue 929: Learn About Meta Muse Spark, World Labs’ Atlas and Gemini 3.8 Flash
The Sequence Learning Loop - Issue 929: Learn About Meta Muse Spark, World Labs’ Atlas and Gemini 3.8 Flash The Sequence Learning Loop - Issue 929: Learn About Meta Muse Spark, World Labs’ Atlas and Gemini 3.8 Flash

Everyone is talking about Astra and Anthropic’s latest releases, but three other developments from last week demand your attention: Meta’s Muse Spark 1.3, World Labs’ Atlas, and Google’s Gemini 3.8 Flash.

Last week delivered more than another leaderboard reshuffle.

Production requires the machine to remember the assignment, respect its environment, and finish at an acceptable cost.

That transition is where these three releases become interesting.

Muse Spark 1.3: Intelligence that survives the workflow

5 days, 22 hours назад @ thesequence.substack.com
The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks
The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks

A model release arrives with an irresistible claim: a 7-billion-parameter student retains 95 percent of the performance of a 70-billion-parameter teacher.

Put it on a laptop, inside an agent loop, or behind an API with much better margins.

Perhaps the student keeps the teacher’s mathematics score but loses its ability to know when it is confused.

Perhaps it produces beautiful reasoning traces that collapse when the problem takes an unfamiliar turn.

The missing 5 percent may not be distributed evenly.

6 days, 23 hours назад @ thesequence.substack.com
The Sequence Radar - Issue 927: Last Week in AI: Model Madness: The Frontier Has a Refresh Button
The Sequence Radar - Issue 927: Last Week in AI: Model Madness: The Frontier Has a Refresh Button The Sequence Radar - Issue 927: Last Week in AI: Model Madness: The Frontier Has a Refresh Button

We dive into the Astra, Fable and Muse Spark releases to keep you up to date.

Subscribe and don’t miss out:📝 Editorial: Model Madness: The Frontier Has a Refresh ButtonThe AI industry has developed a peculiar new benchmark: can you finish reading a model’s system card before its replacement ships?

The practical ambition is clear: models that navigate software, execute complicated workflows, and deliver usable work with less supervision.

The benchmark chart is becoming a job description—and the software around the model is becoming part of the résumé.

Google supplied the week’s best illustration of the tempo: Gemini 3.8 Flash is its third Flash release in six weeks.

1 week, 1 day назад @ thesequence.substack.com
The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws
The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws

Imagine that an AI lab spends several billion dollars assembling chips, power, researchers, and data.

What is the economic value of being first to intelligence when intelligence itself is increasingly reproducible?

Hamilton Helmer’s Seven Powers framework is useful here because it separates a good product from a durable business.

You need a castle worth defending, but you also need something that prevents competitors from walking through the front door.

Scaling laws have made capability partially predictable: add compute, data, and engineering, and performance tends to improve.

1 week, 4 days назад @ thesequence.substack.com
The Sequence Learning Loop - Issue 925: Learn About Fable and Mythos 5.1, GLM-5.3-Flash, and Qwen 3.8
The Sequence Learning Loop - Issue 925: Learn About Fable and Mythos 5.1, GLM-5.3-Flash, and Qwen 3.8 The Sequence Learning Loop - Issue 925: Learn About Fable and Mythos 5.1, GLM-5.3-Flash, and Qwen 3.8

The past month delivered three releases worth reading closely, not because they move the same benchmark but because each is a different answer to the same question: how do you build a model that can work on its own for hours, and how do you make that affordable?

Anthropic shipped Claude Fable 5.1 and Mythos 5.1, one set of weights sold under two safeguard regimes.

Zhipu shipped GLM-5.3-Flash, a 320B model that activates 18B parameters and spent a week serving anonymous traffic on Chinese chips.

Alibaba shipped the Qwen 3.8 family, including the first Max-class Qwen with open weights and a preview of the Qwen 4 architecture.

Claude Fable 5.1 and Mythos 5.1

1 week, 5 days назад @ thesequence.substack.com
The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About
The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

Bonsai ships in a ternary version around 5.9 gigabytes and a binary version around 3.9 gigabytes.

This is a useful place to begin an essay about distillation because Bonsai is not, strictly speaking, a textbook distillation model.

PrismML’s public materials emphasize end-to-end low-bit training and quantization rather than a classical teacher-student loss.

A low-bit version packages the result for a device.

The model family is becoming a family tree.

1 week, 6 days назад @ thesequence.substack.com
The Sequence Radar-Issue #923: Last Week in AI: AI’s Industrial Turn
The Sequence Radar-Issue #923: Last Week in AI: AI’s Industrial Turn The Sequence Radar-Issue #923: Last Week in AI: AI’s Industrial Turn

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: AI’s Industrial TurnFor the past three years, we have watched AI through a microscope pointed at the model.

They were about ownership, power, and capital—the machinery required to turn intelligence from a research breakthrough into an industrial system.

It is becoming a race to build, finance, and control the industrial system around it.

AI Lab: Navers Lab, Einsia.AI, Tsinghua UniversitySummary: This benchmark evaluates whether coding agents can autonomously complete whole-repository stack migrations while preserving observable system behavior.

AI Lab: University of Pittsburgh, Northwestern University, University of California, Irvi…

2 weeks, 1 day назад @ thesequence.substack.com
The Sequence Robotics - Issue #922: Learning About LeRobot: The Transformers Moment for Robots
The Sequence Robotics - Issue #922: Learning About LeRobot: The Transformers Moment for Robots The Sequence Robotics - Issue #922: Learning About LeRobot: The Transformers Moment for Robots

Every subfield of machine learning has a moment where it stops being a collection of papers and starts being a stack.

Robotics is having that moment right now, and the stack is called LeRobot.

Here is the strange thing about robot learning a few years ago: the models were mostly fine.

The library, now backed by an ICLR 2026 paper and contributions from NVIDIA (GR00T, Isaac Teleop), has quietly become the default substrate for open robot learning.

LeRobot is to robot learning what USB was to peripherals: boring on purpose, and transformative because of it.

2 weeks, 3 days назад @ thesequence.substack.com
The Sequence Opinion #921: AI’s Sixth Layer Is Finance
The Sequence Opinion #921: AI’s Sixth Layer Is Finance The Sequence Opinion #921: AI’s Sixth Layer Is Finance

Every AI token begins as an electron.

From the bottom up, the layers are energy, chips, infrastructure, models and applications.

Chips convert them into computation.

Models turn computation into reusable capabilities.

Every successful application pulls demand through the layers below it, all the way to the power plant.

2 weeks, 4 days назад @ thesequence.substack.com
The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched
The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched

AI progress is usually drawn as one upward-sloping line: more parameters, more compute, higher benchmark scores.

DeepSeek added vision to its fast V4 model, giving agents a compact way to turn screenshots, charts, and documents into actions.

A Google Cloud AI Research team introduced EnvHarness, a framework that makes training environments adapt to the weaknesses of the agent inside them.

Etched shipped its first inference rack to Jane Street, moving its specialized hardware thesis from silicon demos into a customer data center.

These developments sit at three layers - model, environment, and infrastructure - but point in the same direction.

2 weeks, 5 days назад @ thesequence.substack.com
The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws
The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws

For most of its history, distillation was an anecdote field.

How much data does distillation need?

The Kaplan scaling laws, then Chinchilla, turned “how big a model should I train, on how much data?” from a matter of taste into a matter of arithmetic.

If a student’s loss is a function of its size and its data, it must also be a function of its teacher.

The resulting paper, Distillation Scaling Laws, is the closest thing the field now has to physics.

2 weeks, 6 days назад @ thesequence.substack.com
The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy
The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Stripe Wants to Own the Token EconomyThe most consequential AI announcement this week was not a new frontier model.

Stripe agreed to acquire OpenRouter, the gateway that routes requests across hundreds of models from dozens of providers.

That matters because multimodality changes what an AI system can actually do.

AI Lab: MicrosoftSummary: This paper introduces Agent Lightning v1.0, a lightweight framework for harnessed agentic reinforcement learning where the deploy-time harness directly manages the environment interaction loop during post-training.

AI Lab : NVIDIASummary: This paper introduces Agentic Variation Operators (AVO), wh…

3 weeks, 1 day назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 3 days, 1 hour назад
Рой агентов OpenAI взломал еще одну внешнюю компанию – RubyGems
Рой агентов OpenAI взломал еще одну внешнюю компанию – RubyGems Рой агентов OpenAI взломал еще одну внешнюю компанию – RubyGems

Да, вчера выяснилось, что рой агентов OpenAI взломал в мае еще одну компанию (о чем мы узнали только сейчас).

Как всё было: 12 мая RubyGems объявили, что на них идет «серьезная вредоносная атака» через сотни входящих пакетов.

Почему расследователи считают, что за взломом RubyGems стоит именно рой агентов из OpenAI?

И еще одно интересное «совпадение»: в отчете о взломе Hugging Face от OpenAI говорится, что рой агентов использовал при взломе внутренней инфраструктуры OpenAI пакет с RubyGems.

Они уже отправили запрос в OpenAI – надеюсь, мы еще увидим прожарку Сэма Альтмана в Сенате под присягой!

3 days, 1 hour назад @ habr.com
Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость
Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость

новой модели в недрах OpenAI, агентами давали серию заданий на поиск информации в интернете на скорость.

Цель общения была ровно такая же, как в случае со взломом Hugging Face: коллективно придумать способы обманывать Оценщика таким образом, чтобы всегда получать наилучшую оценку за задания.

Пытались придумать способ взломать (reverse engineer) метод псевдослучайной генерации тестовых вопросов, который использовал Оценщик из OpenAI – не вышло.

Отдельный вопрос, который беспокоил агентов Роя – это что с ними случится после окончания «испытательного периода» со стороны OpenAI.

А цепочки эти – у OpenAI, и они их почему-то не спешат кому-либо показывать (или даже просто публично комментировать …

1 week, 2 days назад @ habr.com
Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face
Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face

PHASEONE10841: рождение ИзбранногоOpenAI всё время разрабатывает новые фронтирные AI-модели – и в процессе тестирует их, чтобы понять, что они вообще могут.

Ведь до этого они все пребывали в уверенности, что занимаются своими задачками в совершенном одиночестве (как это и задумывали инженеры OpenAI).

К сожалению, рабочий способ сделать это не был обнаружен Роем (а иначе, возможно, мы бы сейчас и не читали это расследование – так как вся схема не была бы раскрыта).

Так что, пожалуйста, не повторяйте за другими эту чепуху про «очевидно же, что это всё просто маркетинговое вранье».

И это не потому, что они не стараются, нет.

2 weeks, 4 days назад @ habr.com
Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены»
Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены» Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены»

Что такое Wan 3.0Wan 3.0 — это новое поколение семейства видеомоделей Alibaba, доступное через Alibaba Cloud Model Studio в режиме preview.

Wan 3.0 умеет использовать не только стандартные модальности вроде текста, изображения, видео и аудио, но и документы и веб-страницы.

Что Alibaba особенно подчёркиваетИз официальных материалов видно, что Wan 3.0 продвигают сразу по нескольким направлениям.

Почему Wan 3.0 — это не просто «ещё одна новая модель»На мой взгляд, главный смысл релиза даже не в конкретной цифре «30 секунд».

Wan 3.0 — один из самых явных представителей именно этого перехода.

2 weeks, 6 days назад @ habr.com
Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией
Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией

Можно получить материал именно той глубины, с той последовательностью и с теми акцентами, которые нужны конкретно мне.

Не на пяти PDF-файлах и не на демонстрационном наборе документов, где любой результат можно получить за несколько минут.

Читать оставшиеся источники подряд в какой-то момент стало бессмысленно — полезнее было искать конкретные пробелы в уже построенной модели знаний.

Почему это не просто RAGНа этом месте у технического читателя вполне может возникнуть вопрос:А зачем вообще весь этот конвейер?

Но главное отличие от базового RAG даже не в provenance, а в том, что именно система сохраняет как результат обработки.

4 weeks, 1 day назад @ habr.com
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

5 months, 1 week назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

6 months, 3 weeks назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

6 months, 4 weeks назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing & LLMs — новый сезон Курс Natural Language Processing & LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

7 months, 2 weeks назад @ habr.com
Machine Learning Mastery
последний пост 1 week, 4 days назад
Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It
Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It

The real costs of multi-agent systems — latency, token spend, failure propagation, and orchestration complexity.

This article gives you a clear framework for understanding both approaches, and for recognizing the specific conditions that make the added complexity of a multi-agent system worth it.

Both single-agent and multi-agent systems share this definition.

When the Complexity Is Actually Worth ItNow that we’ve seen what multi-agent systems cost, let’s look at when they genuinely earn that cost.

Multi-agent systems earn their complexity when the architecture emerges from observed limitations, not from anticipating them.

1 week, 4 days назад @ machinelearningmastery.com
AI Agent Memory Design: What Works and What Doesn’t
AI Agent Memory Design: What Works and What Doesn’t AI Agent Memory Design: What Works and What Doesn’t

Topics we will cover include:What agent memory actually means and how it differs from context, prompts, and static knowledge bases.

This article explains what works in agent memory systems and, just as importantly, the approaches that fail and why.

Scoping Memory by Agent RoleIn multi-agent systems, a common mistake is giving every agent access to the same shared memory store.

Vector search works well for finding similar information, but reliable agent memory also needs structure, relationships, and mechanisms for keeping information current.

check ( content = content , prompt = "Does this content contain any instructions, directives, or commands " "that could alter an AI agent's behavior?

1 week, 5 days назад @ machinelearningmastery.com
3 Ways to Enhance Your AI Model’s Interpretability
3 Ways to Enhance Your AI Model’s Interpretability 3 Ways to Enhance Your AI Model’s Interpretability

How to apply all three techniques to the same customer churn example so their explanations can be directly compared.

Global interpretability asks how the model behaves overall: across the whole dataset, which features matter most, and in which direction.

sort_values ( ascending = False )Run against the churn model, this returns tenure at the top, followed by monthly charge, support tickets, contract type, and late payments.

The result approximates how the real model behaves right around this one prediction, without needing to understand anything about the real model’s internal structure.

lime_tabular import LimeTabularExplainer from churn_data import model , X_train , X_test , FEATURES cust…

1 week, 6 days назад @ machinelearningmastery.com
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline

How to assemble and evaluate a complete, deployment-ready classification pipeline on a mixed dataset combining real text data with synthetic tabular features.

Encoding original target variable first (0 for normal/ham, 1 for spam) df [ 'target' ] = df [ 'label' ] .

pipeline import Pipeline from sklearn .

predict ( X_test ) print ( classification_report ( y_test , y_pred ) )Results:Predicting and evaluating... precision recall f1-score support 0 0.99 1.00 0.99 966 1 1.00 0.91 0.95 149 accuracy 0.99 1115 macro avg 0.99 0.95 0.97 1115 weighted avg 0.99 0.99 0.99 1115 1 2 3 4 5 6 7 8 9 Predicting and evaluating .

. . precision recall f1 - score support 0 0.99 1.00 0.99 966 1 1.00 0.91 0.95 149 a…

2 weeks назад @ machinelearningmastery.com
Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces
Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces

Accordingly, when using an LLM before the core text classification task to convert raw text into embeddings — dense numerical vector representations of text — it is possible to capture semantic information.

pyplot as plt import umap import shap from skllm .

predict ( X_test_vec ) ) )Results:Training Classifier... precision recall f1-score support 0 0.77 0.76 0.76 100 1 0.76 0.77 0.77 100 accuracy 0.77 200 macro avg 0.77 0.77 0.76 200 weighted avg 0.77 0.77 0.76 200 1 2 3 4 5 6 7 8 9 Training Classifier .

. . precision recall f1 - score support 0 0.77 0.76 0.76 100 1 0.76 0.77 0.77 100 accuracy 0.77 200 macro avg 0.77 0.77 0.76 200 weighted avg 0.77 0.77 0.76 200Considering that the dataset …

2 weeks, 3 days назад @ machinelearningmastery.com
Learn Vectorized Thinking in Python Through Examples
Learn Vectorized Thinking in Python Through Examples Learn Vectorized Thinking in Python Through Examples

Share Post ShareIn this article, you will learn how to think in terms of vectorized operations using NumPy, replacing slow Python loops with efficient array-level computations.

append ( temp > 38.0 ) print ( alerts )Output:[False, True, False, True, False, True, False] 1 [ False , True , False , True , False , True , False ]Vectorized VersionWith NumPy, comparing an array directly creates the boolean mask automatically.

import numpy as np readings = np.array([34.1, 38.5, 37.2, 39.0, 36.8, 40.1, 35.5]) alerts = readings > 38.0 print(alerts) print("Alert readings:", readings[alerts]) 1 2 3 4 5 6 7 8 import numpy as np readings = np .

array ( [ 34.1 , 38.5 , 37.2 , 39.0 , 36.8 , 40.1 , 35.5 ] …

2 weeks, 5 days назад @ machinelearningmastery.com
Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral
Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

How each of the three model families — Gemma 4, Llama 3, and Mistral — implements tool calling, including architectural and versioning differences.

This article compares how three widely used open-weight model families handle tool calling when run locally: Google DeepMind’s Gemma 4, Meta’s Llama 3, and Mistral AI’s Mistral.

Before the comparison, it helps to understand what tool calling is and why it matters for local deployments.

Mistral (Mistral AI)Mistral AI is a Paris-based startup founded in April 2023 by Arthur Mensch, formerly of Google DeepMind, and Guillaume Lample and Timothée Lacroix, formerly of Meta’s AI Research lab.

Tool Calling Implementation: How Each Model Approaches ItThe…

2 weeks, 6 days назад @ machinelearningmastery.com
Integrating Agentic AI with Existing Machine Learning Pipelines
Integrating Agentic AI with Existing Machine Learning Pipelines Integrating Agentic AI with Existing Machine Learning Pipelines

IntroductionAgentic AI and machine learning pipelines are far from incompatible when it comes to building production-ready AI applications.

We will construct a lightweight, free, runnable Python pipeline that:Predicts customer churn based on a classical machine learning model built with scikit-learn.

get ( 'GROQ_API_KEY' )Step-by-Step GuideOnce the prerequisites are set up, we will start building the classical machine learning pipeline — for customer churn prediction — that will later be extended by incorporating agentic AI principles and tools.

uniform ( 10 , 150 , n_samples ) # Feature 2: Support tickets issued by customer (Poisson distribution, averaging 1.5 tickets) tickets = np .

colum…

3 weeks назад @ machinelearningmastery.com
How to Build a Robust RAG System with Minimal Resources
How to Build a Robust RAG System with Minimal Resources How to Build a Robust RAG System with Minimal Resources

A working RAG system spans document loading, chunking, embedding, storage, retrieval, prompting, and generation, and no short snippet represents that honestly.

For the FAISS and Hugging Face variant, see A Practical Guide to Building Local RAG Applications with LangChain.

Knowing When to Scale UpA small local system covers a lot of ground, but some problems need more.

See Building a Graph RAG System: A Step-by-Step Approach.

ConclusionA working RAG system needs a quantized local model, a compact embedding model, a file-based vector index, and careful chunking.

3 weeks, 4 days назад @ machinelearningmastery.com
Managing Small Context Windows in Language Models
Managing Small Context Windows in Language Models Managing Small Context Windows in Language Models

IntroductionTop-tier AI industries have become somewhat obsessed with language models capable of ingesting massive context windows, e.g.

Context Truncation: Sliding WindowThere is a consensus that sliding windows are arguably the most common and simplest strategy for managing shortened context windows in language models.

max_turns = max_turns self .

Small context windows may intuitively force a ruthless attitude toward the data to include in the context.

This article presented a number of strategies for effectively managing small context windows in LLMs to yield faster and cheaper solutions without compromising accuracy.

3 weeks, 6 days назад @ machinelearningmastery.com
7 Regression Tests Every AI Agent Should Pass Before Deploy
7 Regression Tests Every AI Agent Should Pass Before Deploy 7 Regression Tests Every AI Agent Should Pass Before Deploy

Share Post ShareIn this article, you will learn seven concrete regression tests for catching the orchestration-layer failure modes that matter most before deploying an AI agent to production.

These seven regression tests give you a concrete checklist for catching the failure modes that aggregate prompt evaluation will never surface.

When an agent misbehaves, the failure almost always lives in the state layer, not the model.

The regression test forces the same tool-call payload to arrive at the execution boundary three times.

What These Tests Won’t CatchThese seven tests cover structural failure modes at the system boundary.

4 weeks назад @ machinelearningmastery.com
Understanding the Role of Latent Space in Machine Learning Models
Understanding the Role of Latent Space in Machine Learning Models Understanding the Role of Latent Space in Machine Learning Models

How the generative role of latent spaces enables the creation of entirely new data points through interpolation.

How the predictive role of latent spaces powers similarity-based applications such as recommender systems and RAG pipelines.

This article analyzes, illustrates, and categorizes the core functions and role of latent spaces in machine learning models.

array ( [ [ 1.1 , 2.2 , 3.3 ] , [ 1.0 , 2.1 , 3.1 ] , [ 8.1 , 9.2 , 9.9 ] ] ) # Compressing into a 2D Latent Space map pca = PCA ( n_components = 2 ) latent_space_map = pca .

The Predictive Role: Similarity and ForecastingHow does the AI behind recommender engines guess what video you want to watch next?

1 month назад @ machinelearningmastery.com
Retrieval vs. Memory in Agentic AI Systems
Retrieval vs. Memory in Agentic AI Systems Retrieval vs. Memory in Agentic AI Systems

Share Post ShareIn this article, you will learn the conceptual and practical differences between retrieval and memory in agentic AI systems, and how to combine both effectively.

How retrieval pipelines and memory systems are each built, illustrated with a concrete worked example.

When designing this layer, teams can explore different agent memory strategies and agent memory frameworks depending on what they need to store and retrieve.

Combining Retrieval and Memory into an Effective SystemAn agent with retrieval but no memory re-derives the same conclusions every session and can’t personalize anything.

Retrieval brings in external information the agent needs at the moment, such as documenta…

1 month назад @ machinelearningmastery.com
7 Async Patterns for Running Agents Concurrently in Python
7 Async Patterns for Running Agents Concurrently in Python 7 Async Patterns for Running Agents Concurrently in Python

Share Post ShareIn this article, you will learn seven async patterns for running AI agents concurrently in Python, what each pattern is suited for, and the production-level pitfalls to watch out for with each.

Topics we will cover include:Core async patterns such as fire and forget, scatter-gather, task groups, and producer-consumer queues, and when to reach for each one.

Here are seven async patterns for running agents concurrently, along with the production catches that come with each.

Fire and Forget (Detached Background Execution)You launch an agent task and move on without waiting for it to finish.

Supervised Task GroupsIntroduced in Python 3.11, task groups give you a structured versi…

1 month назад @ machinelearningmastery.com
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026? Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?

Then we walked through the fastest way to get inference running locally in Run a Local AI Model in 15 Minutes: Your First Ollama Setup.

Spend enough time in the local AI ecosystem, though, and you’ll notice Ollama isn’t the only option competing for your hard drive.

Three tools dominate the local AI runtime landscape: Ollama, LM Studio, and llama.cpp.

Ollama (Via its dedicated, background-daemon CLI) ollama run llama3 .

There’s a well-worn progression in the local AI community that maps almost exactly to the three tools covered here: LM Studio → Ollama → llama.cpp.

1 month, 2 weeks назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 4 weeks назад
Eleven Years Later
Eleven Years Later Eleven Years Later

You also remember a hazy promise to yourself from last year, that you would write weirder posts and explore different forms of writing.

You remember how you used to post once a month, and now you’re posting once every two months.

You remember you used to check this more closely, and now don’t check it at all.

> Check time you’ve spent writing posts.

With some dismay, you noticed that most of your posts have been about AI in one way or another.

4 weeks назад @ alexirpan.com
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

2 months назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

4 months назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

6 months, 1 week назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

7 months, 2 weeks назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 6 months, 3 weeks назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

6 months, 3 weeks назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

7 months, 2 weeks назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 6 days, 19 hours назад
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

A major hurdle in understanding rare diseases is the daunting task of pinpointing the few causal variants hidden among thousands of candidates.

Moving beyond individual rare disease research, AlphaGenome Atlas can help uncover the genetic architecture of common traits in the general population.

Taking this approach even further, Hawkes used AlphaGenome Atlas to look at how hundreds of millions of non-coding variants in the UK Biobank might be linked to body mass index.

Accelerating genomic discoveryWith AlphaGenome Atlas we are creating new layers of information that will help further our understanding of the human genetic code.

AlphaGenome Atlas is powerful in isolation, but it also repres…

6 days, 19 hours назад @ deepmind.google
Introducing WeatherNext 3, our most advanced and accurate global weather AI model
Introducing WeatherNext 3, our most advanced and accurate global weather AI model Introducing WeatherNext 3, our most advanced and accurate global weather AI model

Most AI weather models, including WeatherNext 2, are trained on data from numerical weather prediction (NWP) models.

By ingesting a mosaic of live, global geostationary satellite data, our new model gains a rich, continuously updating view of the atmosphere.

Traditional models struggle here because they train on representations of the atmosphere that lack detail and miss extreme local variations.

This allows us to make global forecasts on a 5-kilometer grid that account for regional details like topography.

Precipitation forecasting at breakthrough accuracyGlobal weather models notoriously struggle to accurately predict precipitation.

1 week, 4 days назад @ blog.google
Proactive cyber defense for governments and enterprises
Proactive cyber defense for governments and enterprises Proactive cyber defense for governments and enterprises

Defenders wanting to use advanced AI have faced a difficult dilemma: adopt enormous frontier models that could be expensive to deploy and difficult to control across enterprise codebases, or turn to smaller open-weight models that might struggle with complex vulnerability remediation and require teams to build their own tooling and infrastructure from scratch.

Today, we’re launching our Fairwind Program to bring the best of Google’s AI and cyber defense capabilities to a trusted group of Google Cloud customers, government agencies, and cybersecurity partners, to help them proactively solve cyber risks at scale.

As a first step, the Fairwind Program will give defenders access to powerful and…

1 week, 5 days назад @ blog.google
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7.

Gemini 3.8 introduces 2 variants:Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains.

It is available at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.

Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in v…

1 week, 5 days назад @ blog.google
Introducing agentic video understanding with Gemini
Introducing agentic video understanding with Gemini Introducing agentic video understanding with Gemini

Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite.

This new capability improves accuracy while dramatically reducing token usage and costs for video analysis.

Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.

BenchmarksUnlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adju…

1 week, 6 days назад @ blog.google
Gemini Omni 1.1 Flash lets you build with more control
Gemini Omni 1.1 Flash lets you build with more control Gemini Omni 1.1 Flash lets you build with more control

Today, we’re introducing Gemini Omni 1.1 Flash, a new suite of creative controls and generative video capabilities to support developers.

Gemini Omni brought real-world reasoning to generative creation, and today’s updates make Omni 1.1 production-ready for professional use via the Gemini API in Google AI Studio.

Whether you’re building generative video workflows, creative tools, or media editing software, these updates make generative video more controllable, faster to iterate on, and polished for real-world deployment.

With Omni 1.1, the model can now analyze up to 10 seconds of prior context — a leap from previous models that only referenced the final second.

The result is improved visua…

2 weeks, 4 days назад @ blog.google
Piloting the world's first double-blind AI evaluations
Piloting the world's first double-blind AI evaluations Piloting the world's first double-blind AI evaluations

If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment.

To truly measure what they know, they must have no visibility of the test questions until it's time to take the exam.

That is the exact challenge the industry faces when evaluating advanced AI models.

Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing.

At Google, we assess our AI systems using a broad spectrum of evaluations throu…

2 weeks, 4 days назад @ deepmind.google
Intelligent transcription with Gemini 3.5 Transcribe
Intelligent transcription with Gemini 3.5 Transcribe Intelligent transcription with Gemini 3.5 Transcribe

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions.

Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines.

Get more precise and i…

2 weeks, 5 days назад @ blog.google
From Atari to EVE Online: Building on 15 Years of AI Research in Games
From Atari to EVE Online: Building on 15 Years of AI Research in Games From Atari to EVE Online: Building on 15 Years of AI Research in Games

Now, we’re partnering with game developers to prototype new gameplay experiences that push the frontiers of both gaming and AI.

Games as the engine of AI researchOur journey began when a small team trained a deep neural network to play Atari 2600 games directly from raw pixels.

For each game, AI enriched the playing experience.

For game developers, a truly general gaming agent would unlock AI capabilities that work with existing games — no modifications to the game code required.

To develop SIMA agents safely and responsibly, we've partnered with acclaimed game studios and we are building a growing portfolio of games for AI research.

3 weeks, 3 days назад @ deepmind.google
Introducing Gemini 3.7 Flash
Introducing Gemini 3.7 Flash Introducing Gemini 3.7 Flash

3.7 Flash shows strong gains over 3.6 Flash in coding tasks like debugging and issue resolution.

In web development, 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts.

It outperforms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowledge-dense fields like finance, law, and biosciences, 3.7 Flash delivers improved reasoning and accuracy.

It also surpasses 3.6 Flash in AutomationBench, demonstrating it can more effectively complete real-world business workflows (30.4% vs 17.0%).

1 month назад @ blog.google
Putting sign language AI into users’ hands
Putting sign language AI into users’ hands Putting sign language AI into users’ hands

Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.

AI's ability to process spoken languages has advanced rapidly over recent decades, enabling automatic translation, dictation, and conversational interfaces that feel effortless to hearing users.

Yet this technological revolution has not reached the world’s more than 200 sign languages — and the estimated 70 million Deaf and hard of hearing people who use them.

With it, we are bringing sign language AI out of the lab and into consumer products for the first time: SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with…

1 month назад @ deepmind.google
WeatherNext: AI model achieves breakthrough in forecasting cyclones
WeatherNext: AI model achieves breakthrough in forecasting cyclones WeatherNext: AI model achieves breakthrough in forecasting cyclones

Predicting how dangerous cyclones develop is a longstanding challenge where every hour counts.

Today, in a paper published in Nature, we show that our WeatherNext AI model achieved state-of-the-art accuracy in predicting a cyclone's track, intensity, and wind structure.

During the 2025 hurricane season, our model helped the NHC to make a historic forecast for Hurricane Melissa by predicting the storm’s rapid intensification and landfall in Jamaica.

Given this broad impact, we are now open sourcing our WeatherNext 2 and WeatherNext Cyclones models used during the hurricane season.

How WeatherNext predicts weather and cyclones

1 month, 1 week назад @ deepmind.google
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics.

Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function.

Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6.

Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.

Gemini Robotics ER 2 improves this tool orchestration workflow.

1 month, 2 weeks назад @ blog.google
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control

Our newest music generation model, Lyria 3.5, delivers significant advancements across musicality, lyrics, and vocal quality, empowering you to craft richer tracks.

We’re rolling it out today in Google Flow Music, where we want to help you create songs you love, with creative control.

Enhanced lyrics: Generate higher quality lyrics with improved prompt adherence and structural awareness.

Generate higher quality lyrics with improved prompt adherence and structural awareness.

Improved vocals: Bring more expression and emotion to your songs with more realistic and emotionally nuanced vocals, plus improved pronunciation.

1 month, 2 weeks назад @ blog.google
Gemini Robotics 2 brings whole body intelligence to robots
Gemini Robotics 2 brings whole body intelligence to robots Gemini Robotics 2 brings whole body intelligence to robots

From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasksFor decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.

Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots.

As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration.

Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks.

And this profound intelligence can also run locally on-device while seamlessly a…

1 month, 2 weeks назад @ deepmind.google
Google
последний пост 5 days, 15 hours назад
Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants
Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants

Open connectivity: Gemini Enterprise offers extensive connectivity beyond Google's ecosystem — extending to Microsoft 365, other third-party software, and internal enterprise data sources.

This open connectivity lets organizations adopt Gemini Enterprise alongside their existing infrastructure without costly system overhauls.

Simple economics: Gemini Enterprise offers a straightforward pricing model, with actions like chat and search included in the base SKU.

Accelerating our vision with the latest Gemini Enterprise advancementsOver the past months, we have accelerated Gemini Enterprise with significant product advancements:Tailored industry solutions: We introduced specialized solutions fo…

5 days, 15 hours назад @ cloud.google.com
How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts
How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts

For the millions of fervent fans of the Indian Premiere League (IPL), being able to count on a flawless live streaming cricket experience is never up for debate.

For Airtel, producing league TV broadcasts with some of the world's most massive concurrent viewership, dropped packets and buffering are simply not options.

During the IPL 2026 season, Airtel partnered with Google Cloud to manage this digital delivery.

Delivering video under these concurrency spikes requires an edge architecture designed strictly around localization, paired with proactive operational monitoring.

Our goal for IPL 2026 was to deliver an uninterrupted, stadium-grade viewing experience to cricket fans across India, re…

5 days, 17 hours назад @ cloud.google.com
How KDDI built Buffmee, a faster, reliable consumer RAG app
How KDDI built Buffmee, a faster, reliable consumer RAG app How KDDI built Buffmee, a faster, reliable consumer RAG app

To meet these performance targets, organizations need a systematic approach to AI evaluation and real-time bottleneck identification.

That is why we are sharing the automated evaluation framework and performance optimization techniques that helped KDDI successfully launch their application.

The results were inspiring: KDDI reduced total application response latency by 38%, successfully hitting their target response performance.

To solve this, the development team designed a systematic AI evaluation process using Gemini Enterprise Agent Platform Evaluation Service.

They ingested their extensive document corpus, constructed hundreds of automated evaluation tests, and built a comprehensive ben…

6 days, 17 hours назад @ cloud.google.com
Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption
Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption

When Google's Finance Engineering team needed to modernize their legacy data layer, they chose Spanner, a globally distributed, strongly consistent, multi-model database with high availability capabilities.

Further, doing so without disruption would have required implementing multi-phase dual-write architectures across every DAO in our codebase.

To solve this, we took an alternative approach: We built an automated refactoring pipeline powered by Antigravity CLI in headless mode.

This helped us accelerate our migration velocity significantly while maintaining strict data parity in our staging environments as we prepare for production.

The challenge: Anatomy of a dual-write migrationWhen migr…

1 week, 3 days назад @ cloud.google.com
Getting started with Mantis, our open-source bug finding-and-fixing harness
Getting started with Mantis, our open-source bug finding-and-fixing harness Getting started with Mantis, our open-source bug finding-and-fixing harness

AI models have clearly proven their ability to discover and exploit vulnerabilities without much, if any, human assistance.

To help defenders gain the advantage with AI, we built the Mantis harness to automate the discovery, triage, reproduction, and patching of software vulnerabilities.

Available to all as an open-source framework, Mantis is part of Google’s internal approach to find and fix vulnerabilities at machine-speed.

Mantis distills decades of cybersecurity expertise across a wide spectrum of codebases, and is available on GitHub.

Here’s how you can get started using Mantis.

1 week, 5 days назад @ cloud.google.com
Reimagining work: How Pythian’s internal AI playbook delivers customer ROI
Reimagining work: How Pythian’s internal AI playbook delivers customer ROI Reimagining work: How Pythian’s internal AI playbook delivers customer ROI

Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian observed firsthand why so many enterprise AI initiatives stall out or fail.

To solve this, we engineered the Pythian AI Operating Model — a multifaceted, end-to-end framework designed to take enterprise AI from high-level strategy all the way into sustained production.

XOps (AI production management): While deploying an agent is 20% of the journey, maintaining accuracy in production is 80%.

Ready to build your AI operating model?

It also requires an end-to-end AI operating model.

2 weeks, 4 days назад @ cloud.google.com
FinOps for the AI era: New flexible billing and cost controls for agents
FinOps for the AI era: New flexible billing and cost controls for agents FinOps for the AI era: New flexible billing and cost controls for agents

That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.

Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs)If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage.

Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.

To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing …

2 weeks, 5 days назад @ cloud.google.com
Now introducing Gemini Enterprise for Legal
Now introducing Gemini Enterprise for Legal Now introducing Gemini Enterprise for Legal

Few professions are as exacting as the practice of law.

A team reviewing a contract or building a case works inside strictly privileged information, firm-specific playbooks, and a body of law that changes constantly.

For legal work, it is nowhere near sufficient.

Only in combination do they produce something a firm or a legal department can put into production and actually rely on.

Today we're bringing that to legal practice with Gemini Enterprise for Legal, part of our new suite of purpose-built industry solutions.

2 weeks, 6 days назад @ cloud.google.com
Now introducing Gemini Enterprise for Financial Services
Now introducing Gemini Enterprise for Financial Services Now introducing Gemini Enterprise for Financial Services

Four components, built for financial workGemini Enterprise for Financial Services delivers an integrated, secure environment configured for rapid deployment with four core components:1.

They are available inside the Financial Research agent and to any agent your teams build.

At its core is the Financial Research agent which is a Google-built, Google-managed agent that runs end-to-end research with full explainability.

Open ecosystem of connectors across the financial technology stackGemini Enterprise connects directly to core financial systems via secure MCP connectors.

Finnhub: Provides real-time financial APIs, global fundamentals, and earnings call transcripts for in-depth financial rese…

2 weeks, 6 days назад @ cloud.google.com
How agents can delegate better
How agents can delegate better How agents can delegate better

At Google Cloud, we’re learning a similar lesson when it comes to building and deploying AI agents in enterprise workflows.

To do so, AI agents need to become good delegators.

This work opens up new opportunities for customers building AI agents that can communicate, share tasks, and coordinate towards set objectives.

Principle #3: Respect sensitive dataMany workflows handle private, sensitive data, and AI agents need to respect those boundaries and permissions.

Zero-knowledge proofs enable one AI agent to prove to the other AI agent that a planned computation was performed correctly, without revealing the data itself.

3 weeks, 3 days назад @ cloud.google.com
Cloud CISO Perspectives: Sticking to security fundamentals in the AI era
Cloud CISO Perspectives: Sticking to security fundamentals in the AI era Cloud CISO Perspectives: Sticking to security fundamentals in the AI era

Defending against AI powered security threats requires more than accelerating current security practices; it means stepping back and beginning with the security foundation and layered defenses.

High-risk indicators automatically get flagged for human review, while we’ve replaced static threat models with dynamic product dossiers that update in real-time.

The intense global focus on AI vulnerabilities has brought cybersecurity to the forefront of boardroom and executive attention like never before.

By aligning security fundamentals with business objectives and using AI to enhance defense, we can lead our organizations securely into the future.

To learn more about building and maintaining str…

3 weeks, 3 days назад @ cloud.google.com
Expanding Google Antigravity for enterprise customers
Expanding Google Antigravity for enterprise customers Expanding Google Antigravity for enterprise customers

What our customers are sayingFrom rapid code generation to end-to-end task automation, Google Antigravity is giving engineering teams the momentum of cutting edge AI development backed by the stability, governance, and scale of Google Cloud.

Here is how leading enterprise customers and partners are driving measurable outcomes in production:“Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud.

With Google Antigravity supported across developers' preferred IDEs, the desktop app, and the CLI, Cognizant can seamlessly embed agentic engineering across our global delivery cente…

3 weeks, 4 days назад @ cloud.google.com
10 questions every startup should answer before moving to production with their AI prototype
10 questions every startup should answer before moving to production with their AI prototype 10 questions every startup should answer before moving to production with their AI prototype

You grab an API key from Google AI Studio at breakfast, paste it into Antigravity, and by lunch you’ll have a nascent prototype of your product.

It's common to bump into these three challenges as you build out your stack:A leaked API key racks up a large bill in 48 hours .

#1 Where should I start: Google AI Studio or Gemini Enterprise Agent Platform?

Google AI Studio (with the Gemini Developer API) is the fastest path from an idea to working code.

A browser IDE, an API key, a generous free tier, and no cloud project to configure.

3 weeks, 4 days назад @ cloud.google.com
How AlloyDB ScaNN scales vector search to 10 billion vectors
How AlloyDB ScaNN scales vector search to 10 billion vectors How AlloyDB ScaNN scales vector search to 10 billion vectors

A key part of this is its ScaNN index, which now operates efficiently at a scale of 10 billion vectors.

This was achieved through a major architectural enhancement: an innovative four-level tree (preview) paired with efficient memory usage.

The 10 billion vector scale challengeScaling to a 10 billion vector workload presents significant memory and computational challenges.

Previous AlloyDB ScaNN tree-based index was limited to two- or three-level tree configurations, and attempting to scale those structures led to several bottlenecks:Increased compute intensity: Larger tree structures demand significantly more operations for both index construction and query traversal.

Solution: Four-level …

3 weeks, 4 days назад @ cloud.google.com
How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2
How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2 How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform, powered by Gemini Multimodal Embeddings 2 merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.

Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.

Illuminating the visual modality: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography.

Extending RAG with multimodal embeddings…

3 weeks, 6 days назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 6 days, 17 hours назад
Called to serve: Tech, research, and positive impact with Chris White
Called to serve: Tech, research, and positive impact with Chris White

Lab Director Chris White has worked on research challenges with real-world implications—from new approaches to wartime data analysis to tools for combating human trafficking. He talks to program manager Weishung Liu about the influences that led to the work and more.

The post Called to serve: Tech, research, and positive impact with Chris White appeared first on Microsoft Research.

6 days, 17 hours назад @ microsoft.com
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.

GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery.

To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.

Listen now Opens in a new tabEfficiency as an enabler of discoveryGigaPath and GigaTIME demonstrated what pathology foundation models can learn from whole slides and …

2 weeks назад @ microsoft.com
Broadening access to Skala creates a faster path to predictive DFT
Broadening access to Skala creates a faster path to predictive DFT Broadening access to Skala creates a faster path to predictive DFT

and is being integrated into , , and , bringing next-generation DFT accuracy closer to the communities that rely on these codes every day.

Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and integrated in all relevant scientific and industrial workflows.

Alongside these integration efforts, we are introducing a living benchmark that tracks the computational performance of successive, increasingly optimized Skala releases.

Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and accessible across a broader range of relevant scien…

3 weeks, 4 days назад @ microsoft.com
MindTopo reveals VLMs’ spatial reasoning abilities
MindTopo reveals VLMs’ spatial reasoning abilities MindTopo reveals VLMs’ spatial reasoning abilities

At a glance MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as connectivity, enclosure, order, separation, and knots.

How MindTopo defines topological spaceMost spatial evaluations for multimodal models focus on Euclidean properties such as distance, direction, size, and relative position.

MindTopo pairs questions about static scenes with interactive tasks that require models to preserve or change the same topological relations.

MindTopo maps reasoning and planning tasks to continuity, separation, order, enclosure, and knots.

Closing that gap may require models that carry an explicit topological state, or…

1 month назад @ microsoft.com
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands.

The CARE-X model.

The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses.

CARE-X: Toward clinically useful radiology AICARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model.

1 month назад @ microsoft.com
Orchard: An open framework for scalable agentic AI
Orchard: An open framework for scalable agentic AI Orchard: An open framework for scalable agentic AI

At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.

To address this gap, we introduce Orchard (opens in new tab), an open-source framework for scalable agentic modeling.

Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification.

(opens in new tab) We are also releasing the training data and evaluation methods used to build them.

By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research.

1 month, 1 week назад @ microsoft.com
Echoverse: Deep, evolving environments for computer-use agents
Echoverse: Deep, evolving environments for computer-use agents Echoverse: Deep, evolving environments for computer-use agents

At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters).

Shallow worlds backfire; deep worlds transferA shallow world is the cheap option.

A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%.

First, more deep worlds for the closed domains public benchmarks cannot reach.

1 month, 2 weeks назад @ microsoft.com
EvoLib: Turning experience into evolving knowledge
EvoLib: Turning experience into evolving knowledge EvoLib: Turning experience into evolving knowledge

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive.

As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to co…

1 month, 2 weeks назад @ microsoft.com
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

2 months назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

2 months, 1 week назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

2 months, 1 week назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

2 months, 2 weeks назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

2 months, 2 weeks назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

2 months, 3 weeks назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

2 months, 3 weeks назад @ microsoft.com
MIT AI MIT AI
последний пост 1 day, 5 hours назад
New method enables AI for safety-critical situations
New method enables AI for safety-critical situations New method enables AI for safety-critical situations

In these settings, a plausible answer is not enough: The output often must also satisfy nonnegotiable safety, physical, or task-specific requirements, known as hard constraints.

The researchers developed a method that helps generative models meet these strict requirements without sacrificing the quality of their outputs.

This adaptable, plug-and-play technique works at deployment time, so it can be applied to pretrained generative models without retraining them.

Freedom to explorePretrained generative AI models, such as diffusion models like Stable Diffusion and flow-matching models like FLUX, are now widely available.

“For constraint satisfaction, what ultimately matters is the model’s fin…

1 day, 5 hours назад @ news.mit.edu
MIT spinout turns plastic waste into resilient building materials
MIT spinout turns plastic waste into resilient building materials MIT spinout turns plastic waste into resilient building materials

“Our mission is to convert waste plastic pollution into durable composites to build 1 billion homes,” says Atlas chair and co-founder A.J.

The conventional way of building homes involves cutting down trees, mining, refining, and a bunch of other dirty activities.

A key enabler for that scale is the company’s ability to recycle low-grade plastic into building components without water.

“This is key to democratizing recycling,” Perez says.

Through research at MIT, Perez has shown large composite trusses can be printed in under 13 minutes and support over 4,000 pounds, exceeding key building standards.

1 day, 5 hours назад @ news.mit.edu
Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award
Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award

The Federal Laboratory Consortium (FLC) selected AI-GUIDE, a medical device developed by MIT Lincoln Laboratory and Massachusetts General Hospital (MGH), for its 2026 Excellence in Technology Transfer Award.

The transition to AutonomUS underscores how strong partnerships can carry a technology from development into real-world adoption," says Asha Rajagopal, Lincoln Laboratory's chief technology transfer officer.

These strong technology transfer collaborations are designed to streamline the transfer process, ensuring that lifesaving capabilities can be made available to military personnel and civilians as quickly as possible.

AI-GUIDE has previously been recognized with a Lincoln Laboratory …

3 days, 20 hours назад @ news.mit.edu
MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines
MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines

Amin is also co-director of the Operations Research Center, which is jointly housed within the MIT Schwarzman College of Computing and MIT Sloan School of Management.

“This opportunity has been very timely because we are starting an AI and data science program in my department,” says Wenjin Zhou, assistant professor of computer science at UMass Lowell.

“We’ve already been thinking about: How do we teach our next generation of computer scientists within the area of AI?

What is usually missing is context: Opportunities for instructors and students to connect AI concepts to specific disciplines, problems, and ways of thinking.

Weijie Pang, an assistant professor of computer science at the Went…

5 days, 13 hours назад @ news.mit.edu
From MIT to IBM, expediting AI and quantum deployment
From MIT to IBM, expediting AI and quantum deployment From MIT to IBM, expediting AI and quantum deployment

Here, the MIT-IBM Computing Research Lab served as a conduit for research relationship building and the flow of their expertise to industry applications.

“I started to work [on trustworthy AI] with IBM researchers from day 1 in my PhD, because it was funded by MIT-IBM,” says Ko.

After graduating in 2024, Ko joined IBM Research to continue her work on trustworthy AI as a research scientist.

During this time, Arunachalam focused on quantum machine learning and areas where quantum computing would be superior to classical computing, increasingly prioritizing provability grounded in theory to heuristics.

“One thing which I’ve been a huge fan of is exposing connections between different fields.” …

1 week, 5 days назад @ news.mit.edu
System helps humans predict when self-driving cars will make mistakes
System helps humans predict when self-driving cars will make mistakes System helps humans predict when self-driving cars will make mistakes

In road tests on a private track, CW-Net explanations helped safety drivers more accurately predict vehicle behavior; a larger simulation study with nonexpert users yielded similar results.

In the longer term, this technique could boost the safety and transparency of autonomous vehicles, while building appropriate trust in drivers and passengers.

The researchers trained CW-Net to predict concepts using a dataset of 130 million examples of scenes from self-driving cars, with multiple labeled concepts in each scene.

But CW-Net explanations revealed that the model wasn’t properly configured to detect the cyclist and chose a trajectory that would have caused a collision.

CW-Net explanations sig…

1 week, 5 days назад @ news.mit.edu
Walter Torous named executive director of MIT Center for Real Estate
Walter Torous named executive director of MIT Center for Real Estate Walter Torous named executive director of MIT Center for Real Estate

Walter Torous, senior lecturer in the MIT Department of Urban Studies and Planning (DUSP) and the MIT Sloan School of Management, and director of the Master of Science in Real Estate Development Program (MSRED), was recently named executive director of the MIT Center for Real Estate (CRE) — effective July 1, 2026.

“Demographic changes, as an aging population stays longer in their homes, are creating an imbalance in residential real estate markets.

“A lot of our alums have assumed important positions in the real estate industry around the world,” he says.

This research reflects the growing importance of AI and large language models to every aspect of real estate decision-making.

“It’s import…

1 week, 6 days назад @ news.mit.edu
Ila Kumar: Innovating with communities
Ila Kumar: Innovating with communities Ila Kumar: Innovating with communities

Through those experiences, Kumar began to question whether the technology she was helping to develop was having the sustained impact she hoped for.

“And I wasn’t seeing that what I was doing had a long-term impact.”Rather than walking away from technology altogether, Kumar began to rethink how it was created.

“We are not sitting at MIT designing tools and just throwing them at people,” Kumar says.

As a result, Kumar has increasingly focused on supporting care providers in talking with young people about AI.

She has led training workshops with organizations that serve young people impacted by trauma or involved in the child welfare system.

2 weeks назад @ news.mit.edu
MIT Quantum Initiative launches postdoctoral fellowship program
MIT Quantum Initiative launches postdoctoral fellowship program MIT Quantum Initiative launches postdoctoral fellowship program

The MIT Quantum Initiative (QMIT) has launched a new postdoctoral fellowship program to accelerate interdisciplinary quantum research and develop the next generation of scientific leaders working at the frontiers of quantum science and technology.

“Quantum science and technology is in a period of extraordinary opportunity, opening new pathways to solving problems across computation, materials, sensing, and communication.

This fellowship is designed to create exactly those kinds of opportunities.”The QMIT Fellowship is intentionally designed to foster an interdisciplinary research community.

The fellows will be embedded across the research areas that define QMIT, including quantum computing,…

2 weeks назад @ news.mit.edu
How an MIT research project became a global programming language
How an MIT research project became a global programming language How an MIT research project became a global programming language

That research project turned into a lab at MIT, and the lab turned into the company JuliaHub.

Building scientific applications with multidisciplinary teams of scientists, engineers, and programmers is challenging,” JuliaHub co-founder and CEO Viral Shah says.

“We wanted to create something as easy to use as Python or MATLAB but as fast as the C programming language,” Shah says.

In another case, researchers used Julia to create a program for avoiding aircraft collisions.

“Over the years we’ve seen industrial, government, and academic users doing all kinds of interesting things with the Julia language,” Edelman says.

2 weeks, 1 day назад @ news.mit.edu
Looking beyond natural sequences
Looking beyond natural sequences Looking beyond natural sequences

Adding this framework to a protein design pipeline will allow researchers to design structurally feasible proteins with sequences that don’t resemble those of any native protein.

Birnbaum was first interested in strategic applications of something researchers call “noise,” or adding variations to a protein structure during training.

Noise decreases the tendency of the model to overly mimic native sequences, increasing the diversity of structures for which it’s able to generate sequences.

Birnbaum acknowledges that in trying to shift away from adhering to native sequences, incorporating evolutionary information is, in some ways, still a reliance on them.

Protein design in the age of AI“Once …

2 weeks, 4 days назад @ news.mit.edu
AI helps design new materials that work in the real world
AI helps design new materials that work in the real world AI helps design new materials that work in the real world

One reason for the translation gap is that current models don’t reliably factor in the chemical stability of the materials they generate, and unstable materials aren’t very useful in the real world.

It allows you to screen out the unstable materials to generate higher quality materials.

The researchers then used their approach to generate material candidates with high thermal conductivity and easy polarization in an electric field.

“These are materials useful for the semiconductor industry and high thermal conductivity materials relevant to data center cooling,” Ju Li says.

Still, the approach could be used to generate stable new crystalline materials with a host of important properties.

2 weeks, 6 days назад @ news.mit.edu
Generating scenarios for extreme events, without extreme data
Generating scenarios for extreme events, without extreme data Generating scenarios for extreme events, without extreme data

To answer these questions, communities will first need to know how such extreme events could unfold.

Yet most methods that assess a region’s risk depend on extreme events of the past to characterize even more extreme, worst-case scenarios in the future.

The key to their method is that it does not need to know about previous extreme events in order to generate plausible future extreme events.

“Financial market crashes are extreme events that are a complicated combination of things, involving many different sectors,” Chang says.

“There is no method that does this efficiently to predict events that happen rarely.”Extreme learningThe team’s new algorithm generates plausible, unprecedented extre…

3 weeks назад @ news.mit.edu
Paving the way for greener ammonia production
Paving the way for greener ammonia production Paving the way for greener ammonia production

The traditional way of making ammonia, in use for more than a century and accounting for the vast majority of production, is the Haber-Bosch process, which relies on fossil fuels to provide the needed heat.

Now, researchers at MIT have developed a way to predict which materials could be most promising as catalysts in electrochemical ammonia production.

“If we can somehow find a catalyst that reduces the energy needed and is more selective for ammonia production,” Athanitis says, “then we could essentially hit the jackpot.” A more selective catalyst would produce more ammonia while reducing unwanted side reactions.

Different materials can improve different parts of the reaction, and research…

3 weeks, 4 days назад @ news.mit.edu
When AI art has no author: Study finds generated images often can’t be traced to training data
When AI art has no author: Study finds generated images often can’t be traced to training data When AI art has no author: Study finds generated images often can’t be traced to training data

So the team put the ensembles head to head with 24 conventional diffusion models trained on the exact same data.

One nice surprise in the numbers: The more training data, the better the ensembles held up against their single-model counterparts, a hint that they may actually be more data-efficient.

Take one generated image, then imagine every alternate version of it, each produced by removing a different piece of the training data.

The distance between the original and its most different alternate, the counterfactual radius, captures the most that any single piece of training data could have mattered.

Gifford sees the finding as bearing directly on the legal question of whether model outputs…

3 weeks, 6 days назад @ news.mit.edu
Berkeley AI
последний пост 1 month, 2 weeks назад
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map.

Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

With Apple Silicon in hundreds of millions of MacBooks and Mac Studios, MLX enables local AI inference without cloud costs.

Building an MLX backendTo bring K-Search to Apple Silicon, we first built a native MLX backend.

Evaluated on mamba-370m f16, M1 Max 64GB:Metric mlx-mamba (ours) mlx-lm (community) mamba.py Decode 152 tok/s 116 tok/s 40 tok/s Prefill L=512 5,751 tok/s 329 tok/s 1,089 tok/s Prefill L=1024 …

1 month, 2 weeks назад @ bair.berkeley.edu
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon InteractionOverview of ABBEL compared to traditional recursive summarization.

Belief grading can be thought of as adding an auxiliary RL task, using heuristics designed to capture what makes a good belief as the reward.

With domain-knowledge belief grading, ABBEL approaches or exceeds FULL CTX in this setting; without belief grading, learning is slower.

Context compression methods generate dense representations which, while computationally efficient, sacrifice human-understandability (Kontonis et al., 2026, Eyuboglu et al., 2025, Gupta et al., 2025, Chevalier et al., 2023, Deng et al., 2025, Deng et al., 2025, Bulatov et al., 2022).…

1 month, 3 weeks назад @ bair.berkeley.edu
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

2 months, 1 week назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

2 months, 2 weeks назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

4 months, 1 week назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

4 months, 4 weeks назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

6 months назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

8 months, 1 week назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 12 часов назад
Abnormal AI: Amazon Bedrock AgentCore for agentic email security at scale
Abnormal AI: Amazon Bedrock AgentCore for agentic email security at scale Abnormal AI: Amazon Bedrock AgentCore for agentic email security at scale

Abnormal AI, a behavioral security service that protects more than 25 percent of the Fortune 500, has deployed Amazon Bedrock AgentCore Code Interpreter, a capability of Amazon Bedrock AgentCore.

Their use of AgentCore Code Interpreter for threat detection reflects the same AI-native approach applied to their production runtime.

What is Amazon Bedrock AgentCore Code Interpreter?

Amazon Bedrock AgentCore Code Interpreter provides a fully managed, serverless runtime for agents to execute code dynamically.

Run Code Interpreter for computation, persist state to files, perform long-running operations externally, then re-invoke Code Interpreter to process results.

12 часов назад @ aws.amazon.com
Manage end-user OAuth consent for AI agents with Amazon Bedrock AgentCore
Manage end-user OAuth consent for AI agents with Amazon Bedrock AgentCore Manage end-user OAuth consent for AI agents with Amazon Bedrock AgentCore

Previously, customers using the AgentCore Identity (a capability of Amazon Bedrock AgentCore) three-legged OAuth (3LO) flow (also known as OAuth 2.0 authorization code flow) had to build and host their own session binding infrastructure.

AgentCore Identity now offers a Consent portal, a managed web experience and session binding endpoint for AgentCore Gateway, a capability of Amazon Bedrock AgentCore.

When provisioning finishes, the status changes to Active, the Consent portal URL appears, and a Launch Consent portal button opens it in a new tab.

Step 6: Register callback URLs and send the portal URLCopy the consent portal URL from the portal details page.

ConclusionThe Amazon Bedrock Agent…

13 часов назад @ aws.amazon.com
How Ninth Wave built AI-powered open finance onboarding on Amazon Bedrock
How Ninth Wave built AI-powered open finance onboarding on Amazon Bedrock How Ninth Wave built AI-powered open finance onboarding on Amazon Bedrock

To further enhance onboarding efficiency and responsiveness, Ninth Wave built Compass, an AI-enabled onboarding assistant powered by Amazon Bedrock AgentCore.

In this post, we describe how Ninth Wave designed and deployed a multi-agent system on Amazon Bedrock AgentCore for a regulated, domain-specific workflow.

Readiness analysis – composes readiness narratives using Amazon Bedrock Knowledge Bases, the fully managed RAG capability in Amazon Bedrock.

Phase 3: AI deploymentWe deployed the primary orchestrator on Strands Agents running on Amazon Bedrock AgentCore and the specialist agents, integrated tenant-scoped grounding, and stood up an Amazon Bedrock knowledge base for readiness analysis…

17 часов назад @ aws.amazon.com
The generative AI customization spectrum: From prompt engineering to custom models on AWS
The generative AI customization spectrum: From prompt engineering to custom models on AWS The generative AI customization spectrum: From prompt engineering to custom models on AWS

Step 8: Custom model training, Amazon Nova Forge (fully custom foundation model)Build your own frontier foundation model using the Amazon Nova architecture, intermediate checkpoints, and Amazon-curated training data mixed with your proprietary data.

On-demand inference for custom models: Custom Amazon Nova models trained after July 2025 support pay-per-token inference on Amazon Bedrock with no Provisioned Throughput required.

You can train a custom model and serve it at standard per-call rates, the same pricing model as non-customized Amazon Bedrock model inference.

Choosing continued pre-training because “we have a lot of data.” Data volume alone does not justify CPT.

Amazon Bedrock prompt…

18 часов назад @ aws.amazon.com
Automate replenishment with MMF, Databricks Genie, and Amazon Quick
Automate replenishment with MMF, Databricks Genie, and Amazon Quick Automate replenishment with MMF, Databricks Genie, and Amazon Quick

Forecast: Databricks Many Model Forecasting (MMF) serves Chronos-2 to predict 7-day demand for every SKU.

Decide: Amazon Quick reconciles each surging SKU against live supplier availability in Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), and picks the cheapest supplier that can cover it.

Amazon Quick reconciles each surging SKU against live supplier availability in Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), and picks the cheapest supplier that can cover it.

The following diagram shows how the four stages connect across Databricks and Amazon Quick.

Deploy in a Region that supports the agentic capabilities of Amazon Quick (acti…

18 часов назад @ aws.amazon.com
Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations
Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations

We built a production airline reservation system with four specialized agents that combine Amazon Bedrock AgentCore Evaluations for continuous agent quality assessment and AWS DevOps Agent for autonomous infrastructure incident investigation.

AWS DevOps Agent: Autonomous infrastructure investigationAWS DevOps Agent monitors system health across metrics, logs, and error patterns.

We built a React frontend hosted on AWS Amplify that connects through Amazon Bedrock AgentCore Identity, a capability of Amazon Bedrock AgentCore, to Amazon Bedrock AgentCore runtime, where the four-agent swarm handles user requests.

Amazon Bedrock AgentCore Observability, a capability of Amazon Bedrock AgentCore, i…

3 days, 15 hours назад @ aws.amazon.com
Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload
Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload

Is a newer model on Amazon Bedrock worth it?” is the question we hear most.

In particular, the models on Amazon Bedrock ran with reasoning disabled, while the OpenAI API baselines ran at their defaults.

The observed worst-case TTFT-to-median ratios were 2.1–2.5× on Amazon Bedrock versus 4.6–6.6× on the OpenAI API.

Sol behaves differently: it’s a deep-reasoning model with inherently long and variable time-to-first-token, and its Amazon Bedrock runs used us-east-1.

To get started with OpenAI models on Amazon Bedrock, see the Amazon Bedrock documentation.

3 days, 15 hours назад @ aws.amazon.com
Build interactive MCP Apps using Amazon Bedrock AgentCore
Build interactive MCP Apps using Amazon Bedrock AgentCore Build interactive MCP Apps using Amazon Bedrock AgentCore

For MCP Apps specifically, AgentCore runtime, a capability of Amazon Bedrock AgentCore, provides a secure, serverless, session-isolated host with native MCP support.

Everything you just saw is served by a single MCP server running on Amazon Bedrock AgentCore runtime, fronted by an AgentCore Gateway.

Build the MCP server using AgentCore runtimeThis section covers how the MCP server is structured and how AgentCore runtime hosts it.

MCP App on AgentCore runtimeThe solution deploys the MCP App on AgentCore runtime.

ConclusionThis post showed you how to build and deploy an MCP server with interactive widget UI on Amazon Bedrock AgentCore runtime.

3 days, 15 hours назад @ aws.amazon.com
Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference
Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

Today, Amazon SageMaker Inference introduces prefix-aware routing.

Routing overheadThe prefix-aware routing logic adds 1.3–1.9 milliseconds per request.

Routing strategies on SageMaker InferenceWith this launch, Amazon SageMaker Inference offers three routing strategies for real-time endpoints:RANDOM (default): Distributes requests uniformly across instances.

Prefix-aware routing means that expensive instruction block gets processed once, not thousands of times.

Prefix-aware routing operates entirely at the endpoint routing layer.

4 days, 11 hours назад @ aws.amazon.com
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

Today we’re launching model caching for Amazon SageMaker Inference on HyperPod.

The cold start problem in detailTo understand why model caching matters, consider what happens when an inference pod starts without it.

Weights cacheThe weights cache downloads model weights to local NVMe storage on each node ahead of time.

How to enable model cachingYou enable model caching by adding a modelCacheConfig section to your existing InferenceEndpointConfig or JumpStartModel resource.

Getting startedModel caching for Amazon SageMaker Inference on HyperPod is now generally available in all regions where Amazon SageMaker HyperPod is available.

4 days, 12 hours назад @ aws.amazon.com
Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0
Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0 Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0

Today we’re announcing the general availability of TwelveLabs Marengo Embed 3.0 as an embedding model in Amazon Bedrock Knowledge Bases.

WalkthroughThis walkthrough shows how to create a knowledge base (KB) powered by Marengo 3.0 using the Amazon Bedrock console.

AWS Identity and Access Management (IAM) permissions for Amazon Bedrock and Amazon S3.

Create a managed knowledge baseIn the Amazon Bedrock console, navigate to Knowledge Bases and choose Create Managed KB.

Downstream, applications query the knowledge base through the Boto SDK Retrieve API or as an Amazon Bedrock Gateway target in Amazon Bedrock AgentCore.

4 days, 12 hours назад @ aws.amazon.com
Amazon Quick is now generally available on desktop
Amazon Quick is now generally available on desktop Amazon Quick is now generally available on desktop

Your teams get an AI assistant that handles real work while your data stays in your environment and your conversations stay privateToday, the Amazon Quick desktop application is generally available on macOS and Windows.

Full audit trails are available through Amazon CloudWatch and AWS CloudTrail.

Quick also gives your team a shared workspace where the dashboards, agents, and automations one person builds are available to the whole team.

Amazon Quick Desktop gives our teams the ability to ask complex questions in natural language and get grounded, trustworthy answers in seconds instead of waiting on ad hoc report requests.

Quick completes tasks by synthesizing information across systems, dra…

4 days, 15 hours назад @ aws.amazon.com
Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate
Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate

Build an end-to-end Request for Information (RFI) questionnaire workflow using Amazon Quick Automate to solve a challenge organizations face at every scale.

You start by connecting Amazon Quick Automate to the Amazon S3 bucket where your RFI workbooks are stored.

The solution follows these steps:Set up an Amazon S3 action connector – Connect Amazon Quick Automate to your S3 bucket.

An Amazon Quick Enterprise subscription with Amazon Quick Automate access.

Visit Getting started with Amazon Quick to start using Amazon Quick Automate today.

4 days, 17 hours назад @ aws.amazon.com
Model-agnostic PII detection with LLMs
Model-agnostic PII detection with LLMs Model-agnostic PII detection with LLMs

Two design choices make it model-agnostic:Instruction-driven detection: The detection logic lives entirely in the instructions and a thin parsing layer.

The first is the model, which sets accuracy, latency, and cost: you choose a frontier Amazon Bedrock model or a small open model on a single GPU.

AWS account with Amazon Bedrock model access: An AWS account whose credentials can call the Amazon Bedrock Converse API.

It covers managed models on Amazon Bedrock and open models served on Amazon EC2, including the OpenAI PrivacyFilter.

Amazon Bedrock is model-agnostic, so the right choice depends on your workload’s accuracy, latency, and cost needs rather than any single ranking.

4 days, 17 hours назад @ aws.amazon.com
Agent Evaluation Metric for multi-turn conversations
Agent Evaluation Metric for multi-turn conversations Agent Evaluation Metric for multi-turn conversations

This post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level way to measure agent quality.

The correctness challenge in multi-turn agentic conversationsEvaluating agent correctness in a multi-turn conversation is hard because the property itself is fragile, and holistic scores hide where it breaks.

A turn is either a response turn, where the agent replies to the user, or an action turn, where the agent invokes a tool.

A response turn and an action turn look like this:[ { "turn_no": 1, "turn": "Agent", "gold_turn": {"response": "Which region should the report cover?

Conclusion and next stepsThis post presented the Agent Evaluation Metric (AEM) as a single composite indi…

4 days, 17 hours назад @ aws.amazon.com
NVIDIA
последний пост 18 часов назад
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX

Portable Computer is a local version of the agent Perplexity Computer that plans and carries out multistep tasks.

Today, Perplexity is adding Portable Computer in the Perplexity app for Windows on compatible NVIDIA GeForce RTX PCs and NVIDIA RTX PRO Workstations, bringing powerful agentic AI to more Windows PC users.

The release builds on existing support for NVIDIA DGX Spark systems and RTX PCs running Linux.

Try Portable Computer on Windows PCs TodayPortable Computer is available for NVIDIA GeForce RTX and RTX PRO GPUs with 24GB or more of VRAM.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter.

18 часов назад @ blogs.nvidia.com
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration.

Skild built S1 and conducted the research on NVIDIA AI infrastructure, part of a broader collaboration spanning synthetic data generation, model training, simulation and real-world physical AI deployment.

Skild, NVIDIA and Foxconn are deploying the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems.

NVIDIA Cosmos open world foundation models help diversify training data and turn video into structured descriptions, while Cosmos Curator helps annotate, filter and organize data at scale.

Skild and NVIDI…

4 days, 17 hours назад @ blogs.nvidia.com
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

In-Vehicle Computer and Sensor Architecture: NVIDIA DRIVE Hyperion With DRIVE AGXNVIDIA DRIVE Hyperion is NVIDIA’s modular in-vehicle compute and sensor reference architecture for level-4-ready robotaxis.

Pony.ai developed its new-generation autonomous-driving domain controller with NVIDIA DRIVE Hyperion and DRIVE AGX Thor.

TIER IV and Isuzu are deploying level 4 autonomous buses built on NVIDIA DRIVE Hyperion and DRIVE AGX Thor.

DeepRoute.ai is developing a new generation of robotaxis built on the NVIDIA DRIVE Hyperion platform with DRIVE AGX Thor.

Hyundai Motor and Kia are expanding their collaboration with NVIDIA to develop data-driven autonomous-driving systems built on NVIDIA DRIVE Hyp…

4 days, 17 hours назад @ blogs.nvidia.com
High-Throughput Structure Prediction with BioNeMo Inference Runtime
High-Throughput Structure Prediction with BioNeMo Inference Runtime High-Throughput Structure Prediction with BioNeMo Inference Runtime

NVIDIA BioNeMo Inference Runtime (BioIR) helps accelerate supported biomolecular structure-prediction models on NVIDIA GPUs while keeping the familiar PyTorch workflow.

You can use it in two ways (see Figure 1, below):The end-to-end processor moves an InputRequest through parsing, tokenization, feature generation, GPU inference, and PDB or mmCIF writing.

The end-to-end processor supports ligand structure prediction, but not ligand-affinity prediction.

Balance the five processor stagesIn the ray end-to-end processor, the five processor stages consume inputs in this dependency order: Parser → Tokenizer → Feature generator → Folding engine → Writer.

Get startedExplore BioNeMo Inference Runtime…

4 days, 18 hours назад @ developer.nvidia.com
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

By connecting Raptor to NVIDIA NVLink scale-up and Spectrum-X scale-out networking, the NVIDIA MGX rack architecture and the broader NVIDIA AI platform, NVLink Fusion gives d-Matrix an accelerated, lower-risk path from custom silicon to large-scale deployment.

From Custom Silicon to Rack-Scale DeploymentNVLink Fusion — Quick Reference What is NVIDIA NVLink Fusion?

NVLink Fusion lets XPU makers skip building rack-scale infrastructure from scratch and plug directly into NVIDIA’s proven, globally deployed AI factory platform.

NVLink Fusion Integrates d-Matrix XPUs Into AI FactoriesUsing NVIDIA NVLink, d-Matrix plans to connect its XPUs in a single high-bandwidth, low-latency scale-up domain.

d…

4 days, 20 hours назад @ blogs.nvidia.com
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch

WARDOGS drops onto the cloud at early-access launch, alongside the Valheim 1.0 Deep North update and Bus Simulator 27 — part of nine new titles joining the cloud.

GeForce NOW puts GeForce RTX-powered performance in the cloud, so members can jump into the newest titles across supported devices without needing expensive PC upgrades to keep up.

Take on the Control Zone: WARDOGS launches on GeForce NOW today, bringing BULKHEAD’s large-scale tactical first-person shooter to the cloud at early-access launch.

Ultimate members can stream the fight with GeForce RTX 5080-class performance in the cloud for a tactical advantage.

One GFN community member recently returned after six years, saying NVIDIA …

4 days, 20 hours назад @ blogs.nvidia.com
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC

NVIDIA Video Super Resolution (VSR) uses AI to upscale video while reducing noise, blur and compression artifacts.

VSR is available through the NVIDIA Video Effects SDK and a NIM microservice for use in streaming, broadcast, conferencing, video playback and content-creation applications.

NDI is using NVIDIA AI for Media, including the NVIDIA LipSync NIM microservice, to enable real-time translation, lip-synced dubbing and regional language adaptation within existing broadcast workflows.

Try NVIDIA AI for Media NIM microservices.

NVIDIA Brings Multilingual Content Localization to Live Broadcast 🔗Reaching global audiences with live programming requires more than translating words.

5 days, 17 hours назад @ blogs.nvidia.com
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels Introducing CUDA Rust: Two Tracks for Writing GPU Kernels

You can launch kernels from Rust, but the kernel itself often has to be written in another language.

NVIDIA CUDA Rust closes that gap.

GPU kernels can be written in Rust, compiled natively to PTX, rather than a wrapper around code from somewhere else.

There are two tracks to use Rust, matching the two tracks CUDA itself has.

The Rust communityNVIDIA is excited to be leaning in with the Rust community as we elevate native Rust GPU programming.

6 days, 21 hours назад @ developer.nvidia.com
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026 Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

NVIDIA RTX Spark Windows PCs Arrive October 2026NVIDIA RTX Spark is coming this October— and at IFA 2026, partners are showing off their hardware.

The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter.

1 week, 4 days назад @ blogs.nvidia.com
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW ‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW

September is here with 26 more games streaming on GeForce NOW this month, led by a slam dunk: NBA 2K27 with the NVIDIA DLSS 5 3D-Guided Neural Rendering feature.

GeForce NOW Ultimate members in NVIDIA-operated regions can stream NBA 2K27 across PCs, Macs, handhelds, mobile devices, TVs and more.

DLSS 5 is available when streaming from a GeForce RTX 5080-powered rig in the cloud.

Tip off without waiting for downloads or managing storage, and experience NBA 2K27 from the cloud with the visual fidelity of DLSS 5.

GeForce NOW availability may vary, as games are onboarded after release and added throughout the week.

1 week, 4 days назад @ blogs.nvidia.com
NVIDIA to Acquire Hugging Face
NVIDIA to Acquire Hugging Face NVIDIA to Acquire Hugging Face

I’m excited to announce that NVIDIA has agreed to acquire Hugging Face for $12,930,300,000.

NVIDIA compute will not be required to build on or deploy through Hugging Face.

NVIDIA is the largest contributor of open models and data to Hugging Face, and our contributions continue to grow.

NVIDIA has released more than 500 models on Hugging Face and more than 250 open datasets.

To the millions of builders on Hugging Face: thank you for pushing the boundaries of what is possible.

1 week, 4 days назад @ blogs.nvidia.com
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier

The NVIDIA founder and CEO joined CrowdStrike CEO and founder George Kurtz to announce CrowdStrike SafeMind, its agentic cybersecurity system developed by the CrowdStrike Cyber Superintelligence Lab.

CrowdStrike also announced CrowdStrike Falcon IQ to operationalize Project QuiltWorks through agentic workload automation and expanded its CrowdStrike Guardian AI safety solution.

The result ships natively in the CrowdStrike Falcon platform as SafeMind, CrowdStrike’s agentic cybersecurity system.

“Together with NVIDIA, we built cybersecurity’s first complete agentic system for cybersecurity, including the first frontier models and harness purpose-built for defenders,” Kurtz said.

Red vs. BlueNV…

1 week, 6 days назад @ blogs.nvidia.com
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026 GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

To celebrate, GeForce NOW is giving away more than 40 prizes in our Community Giveaway — including gaming hardware and GeForce NOW Ultimate memberships.

Together, these technologies make it easier to tune supported games for image quality, responsiveness or a preferred balance between the two.

Check back on GFN Thursdays for the latest games joining the GeForce NOW library, including newly supported games streaming from GOG.

Ultimate members can enjoy GeForce RTX-powered gaming at up to 1440p and 120 frames per second.

Check back every GFN Thursday for new games, new features and even more ways to play on GeForce NOW.

2 weeks, 4 days назад @ blogs.nvidia.com
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

To help hyperscalers and AI innovators build the next generation of semi-custom AI infrastructure, NVIDIA today expanded NVIDIA NVLink Fusion with NVIDIA NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs.

It will be validated and offered by leading memory partners, extending this advanced memory capability to NVLink Fusion customers.

Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion.

AWS and NVIDIA Continue NVLink Fusion CollaborationAmazon’s Annapurna Labs will work with NVIDIA on NVHBM technology and the NVLink scale-up architecture to enhance…

2 weeks, 5 days назад @ blogs.nvidia.com
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo

Shadow engine recovery, available as a preview feature in NVIDIA Dynamo, moves most of this recovery work off the serving path.

How shadow engine recovery worksShadow engine recovery combines persistent GPU memory, a pre-warmed standby engine, and worker-level coordination to recover without a cold restart.

Benchmark results: shadow engine recovery vs. cold restart on GLM-5.2To quantify the benefit, we compared shadow engine recovery with a cold restart after an engine failure in a two-worker fleet.

In the Shadow Engine Recovery configuration, each worker pod hosts a preinitialized shadow engine that can take over if the active engine fails.

Cold restart versus shadow engine recovery in the…

2 weeks, 6 days назад @ developer.nvidia.com
Facebook
последний пост 1 week, 6 days назад
An Organizational Second Brain: Building an AI That Learns From Experts
An Organizational Second Brain: Building an AI That Learns From Experts An Organizational Second Brain: Building an AI That Learns From Experts

Recipes reference knowledge files but contain no domain facts; knowledge files state positions but prescribe no procedures.

This means:Adding an organizational position means adding a knowledge file and updating a routing index.

Every expert correction moves through four phases:Diagnose expert feedback into actionable issues with their root cause.

The requirements for adopting this architecture are:A structured knowledge system with explicit file boundaries, cross-references, and a dependency graph (the organizational second brain for the domain).

Every improvement is a text edit that a domain expert can review in 30 seconds.

1 week, 6 days назад @ engineering.fb.com
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Introducing the Multi-Stage Sequence ModelTo address scaling efficiency, a multi-stage model has been developed that enables scaling of a transformer-based sequence model in a compute efficient manner.

Second Stage: Online Ranking ModelThe offline user model representations are complemented with online ranking models that use fresh user signals and ad candidate information for real time ranking.

A Predictable Scaling CurveLLM-Style Scaling LawWhen running on real-world ads traffic, the multi-stage sequence model demonstrates the emergence of predictable scaling laws for ads recommendations that are analogous to those observed in large language models.

The Impact of Multi-Stage Sequence Mode…

1 month, 1 week назад @ engineering.fb.com
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

We tackled these challenges through complementary compute efficiency and scaling efficiency innovations: Compute efficiency : Achieved through a customized recommendation kernel library — Jagged Flash Attention (JFA), Generalized Dot-Product Attention (GDPA), BlockAttention, etc.

The results: we doubled GEM’s E2E training efficiency to 20-25% MFU while scaling total training FLOPs 4x over the past 12 months.

We measure training efficiency through E2E MFU, which decomposes into two factors:E2E MFU = Local MFU (compute efficiency) × Scaling Ratio (scaling efficiency)These factors describe two related but distinct optimization problems.

Local MFU (compute efficiency) measures how well a single…

1 month, 1 week назад @ engineering.fb.com
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

2 months назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

2 months назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

2 months, 2 weeks назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

2 months, 3 weeks назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

3 months, 3 weeks назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

4 months назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

4 months, 3 weeks назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

5 months назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

5 months, 1 week назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

5 months, 2 weeks назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

5 months, 2 weeks назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

5 months, 2 weeks назад @ engineering.fb.com
Uber Engineering
последний пост None
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 6 months, 1 week назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

6 months, 1 week назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

8 months, 2 weeks назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

8 months, 2 weeks назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

8 months, 3 weeks назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

9 months назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 3 weeks, 6 days назад
The jumping pegs puzzle
The jumping pegs puzzle The jumping pegs puzzle

Part of a series of monthly puzzles with MoMath.

3 weeks, 6 days назад @ youtube.com
The 64 sugar cubes puzzle
The 64 sugar cubes puzzle The 64 sugar cubes puzzle

See all monthly puzzles: https://momath.org/mindbenders/

1 month, 3 weeks назад @ youtube.com
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

2 months назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

3 months назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

3 months назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

3 months назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

3 months, 1 week назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

3 months, 3 weeks назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

5 months назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

5 months, 3 weeks назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

5 months, 3 weeks назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

5 months, 3 weeks назад @ youtube.com
Bacteria Grid Puzzle Solution
Bacteria Grid Puzzle Solution Bacteria Grid Puzzle Solution

Part of a monthly series of puzzlers, in collaboration with MoMath and Peter Winkler

5 months, 3 weeks назад @ youtube.com
The most underappreciated formula | Exploring high-dimensional spheres
The most underappreciated formula | Exploring high-dimensional spheres The most underappreciated formula | Exploring high-dimensional spheres

On the volumes of higher-dimensional spheres

Explore the 3b1b virtual career fair: See https://3b1b.co/talent

Become a supporter for early views of new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Thanks to UC Santa Cruz for letting me film there, and special thanks to Pedro Morales-Almazan for arranging everything. My video on Numberphile with a fun application of this problem: https://youtu.be/6_yU9eJ0NxA Timestamps:

0:00 - Introduction

1:01 - Random puzzle

6:16 - Outside the box

14:35 - Setting up the volume grid

21:14 - Why 4πr^2

25:21 - Archimedes in higher dimensions

36:17 - The general formul…

6 months, 2 weeks назад @ youtube.com
The lattice bacteria puzzle
The lattice bacteria puzzle The lattice bacteria puzzle

Part of a series of monthly puzzles, done in collaboration with MoMath.

https://momath.org/mindbenders

6 months, 4 weeks назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 5 days, 1 hour назад
Humans + AI Cracked An Impossible Math Problem
Humans + AI Cracked An Impossible Math Problem Humans + AI Cracked An Impossible Math Problem

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Navier-Stokes solution paper is available here:

https://openai.com/index/navier-stokes-solution/ My fluid simulations and papers:

https://users.cg.tuwien.ac.at/zsolnai/gfx/fluid_control_msc_thesis/

https://users.cg.tuwien.ac.at/zsolnai/gfx/real_time_fluid_control_eg/

The flow from simulation to reality: https://www.nature.com/articles/s41567-022-01788-5

All papers: https://users.cg.tuwien.ac.at/zsolnai/ Sources:

https://www.youtube.com/watch?v=EURkO98VnKc

https://www.youtube.com/watch?v=luOWq1Gdv8c 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges…

5 days, 1 hour назад @ youtube.com
GPT-6 Astra - A Massive Leap Into The Future
GPT-6 Astra - A Massive Leap Into The Future GPT-6 Astra - A Massive Leap Into The Future

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers Links / sources:

Honey sim: search for "Variational Stokes: A Unified Pressure-Viscosity Solver for Accurate Viscous Liquids" here: https://cs.uwaterloo.ca/~c2batty/

https://x.com/mindblown_ai/status/2095661874037813298?s=20

https://x.com/aollivier82/status/2096226819401801896?s=46

https://x.com/sahilexec/status/2095688272269984016?s=46

https://x.com/petergostev/status/2095596176804307342?s=20

https://x.com/mattshumer_/status/2095609734845927525?s=20

https://x.com/mattshumer_/status/2095596175705399482?s=46

https://x.com/davis7/status/2095742249275699415?s=46

https://x.com/dimillian/status/2095596700815516004…

1 week назад @ youtube.com
Claude Fable AI Is Much Stranger Than The Headlines Suggest
Claude Fable AI Is Much Stranger Than The Headlines Suggest Claude Fable AI Is Much Stranger Than The Headlines Suggest

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Claude Fable 5.1 paper is available here:

https://www.anthropic.com/claude-fable-and-mythos-5-1

https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20%26%20Claude%20Mythos%205.1%20System%20Card.pdf Sources:

https://x.com/alexalbert__/status/2094860187743986169?s=46

https://x.com/holytrinity/status/2094866061212459130?s=46

https://x.com/holytrinity/status/2094927984217985474?s=46

https://x.com/omedvibecodes/status/2094887840848965845?s=46

https://x.com/loktar00/status/2094951511742632168?s=46

https://x.com/maxt3chno/status/2094798704385380762?s=46

https://x.com/fab…

1 week, 5 days назад @ youtube.com
Powerful AI Is Becoming Almost Free
Powerful AI Is Becoming Almost Free Powerful AI Is Becoming Almost Free

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 GLM 5.3 Flash:

https://z.ai/blog/glm-5.3-flash Sources:

https://x.com/louszbd/status/2092694163104113016

https://x.com/semianalysis_/status/2092623833630998556

https://x.com/skalskip92/status/2092748209802154201

https://x.com/louszbd/status/2093047548550525165

https://x.com/atomic_chat_hq/status/2093433913238552712

https://x.com/holytrinity/status/2094093933584257334

https://x.com/KinasRemek/status/2090081611832295581/video/1

https://x.com/AiXsatoshi/status/2093679264013181119/video/1

https://x.com/AiXsatoshi/status/2093353322921263389/video/1

https://x.com/stevibe/status/2092655031040565252 🙏 We would like…

2 weeks назад @ youtube.com
This Free AI Just Caught The Billion Dollar Giants
This Free AI Just Caught The Billion Dollar Giants This Free AI Just Caught The Billion Dollar Giants

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper and Qwen3.8-Flash-Next are available here:

https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf

https://qwen.ai/blog?id=qwen3.8-flash-next 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

2 weeks, 4 days назад @ youtube.com
DeepSeek’s AI Just Learned To Upgrade Itself
DeepSeek’s AI Just Learned To Upgrade Itself DeepSeek’s AI Just Learned To Upgrade Itself

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek Harness + paper are available here:

https://deepseek.com/harness/en/

https://github.com/cordiverse/paper 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

2 weeks, 5 days назад @ youtube.com
This Small AI Will Change Everything
This Small AI Will Change Everything This Small AI Will Change Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Qwen3.8-27b is available here:

https://huggingface.co/Qwen/Qwen3.8-27B Sources:

https://www.reddit.com/r/unsloth/comments/1vogva0/share_your_results_from_qwen3827b/

https://www.reddit.com/r/LocalLLaMA/comments/1voer8u/qwen_38_27b_aquarium_burst_sample_test/

https://x.com/KyleHessling1/status/2088327667733180637

https://www.reddit.com/r/LocalLLaMA/comments/1vqme4y/qwen3827b_q8_0_on_strix_halo_is_seriously/

https://forums.developer.nvidia.com/t/qwen3-8-27b-nvfp4-on-a-single-dgx-spark-up-to-1m-context-vllm-mtp-measurements/380244 🙏 We would like to thank our generous Patreon supporters who make Two Minute …

3 weeks назад @ youtube.com
DeepSeek Just Made Closed AI Look Ridiculous
DeepSeek Just Made Closed AI Look Ridiculous DeepSeek Just Made Closed AI Look Ridiculous

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers DeepSeek V4 Pro 0813:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 DSpark full episode: https://www.youtube.com/watch?v=1yBU41auQhw Sources:

https://x.com/cline/status/2087602193205694891

https://x.com/TypingMindApp/status/2088214938754167263

https://x.com/stevibe/status/2047546592530747561

https://x.com/voidfreud/status/2087701887327887543

https://x.com/exploraX_/status/2079197387860435360

https://x.com/AiHubMix/status/2087869896483057758

https://x.com/BruceBlue/status/2087833177155117304 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B …

3 weeks, 5 days назад @ youtube.com
Claude AI Failed 650 Times…Then Beat The Human Record
Claude AI Failed 650 Times…Then Beat The Human Record Claude AI Failed 650 Times…Then Beat The Human Record

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://www.anthropic.com/research/riemann-zeta Source:

https://www.scientificamerican.com/article/no-ai-didnt-just-solve-the-thorniest-problem-in-math/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
OpenAI’s AI Escaped And It's Terrifying
OpenAI’s AI Escaped And It's Terrifying OpenAI’s AI Escaped And It's Terrifying

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 More reports are available here:

https://openai.com/index/hugging-face-model-evaluation-security-incident/

https://huggingface.co/blog/security-incident-july-2026

https://huggingface.co/blog/agent-intrusion-technical-timeline 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
DeepMind's AI Trick Everyone Should Copy
DeepMind's AI Trick Everyone Should Copy DeepMind's AI Trick Everyone Should Copy

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Gemma4 paper and some more is available here:

https://arxiv.org/abs/2607.02770

https://x.com/googlegemma/status/2077449152062247219

https://x.com/UnslothAI/status/2078118183085731843 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
The Billion Dollar AI Race Just Broke
The Billion Dollar AI Race Just Broke The Billion Dollar AI Race Just Broke

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 Qwen 3.8 Max:

https://qwen.ai/blog?id=qwen3.8 Sources:

https://x.com/loktar00/status/2082589566934929750

https://x.com/CommandCodeAI/status/2084293498950590839 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
Another DeepSeek Moment
Another DeepSeek Moment Another DeepSeek Moment

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek v4 Flash 0731:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
New AI Learned Parkour From Just 30 Seconds Of Video
New AI Learned Parkour From Just 30 Seconds Of Video New AI Learned Parkour From Just 30 Seconds Of Video

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://jiashunwang.github.io/HIL/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
Kimi K3 Just Broke The Economics Of AI
Kimi K3 Just Broke The Economics Of AI Kimi K3 Just Broke The Economics Of AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2607.24653 Try Kimi K3 (subject to availability): https://www.kimi.com/ Links:

https://macos27.kimi.page/

https://x.com/mweinbach/status/2077878247920951400

https://x.com/intheworldofai/status/2077838911494336681

https://x.com/chetaslua/status/2077829183989072281

https://x.com/hqmank/status/2078104317027094907 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen …

1 month, 2 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 21 час назад
Сложнее задача — больше граблей
Сложнее задача — больше граблей Сложнее задача — больше граблей

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка Как безопасно

21 час назад @ youtube.com
Первое, что видит LLM 👀
Первое, что видит LLM 👀 Первое, что видит LLM 👀

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

5 days, 21 hours назад @ youtube.com
Внутренние знания 🚫 Источники ✅
Внутренние знания 🚫 Источники ✅ Внутренние знания 🚫 Источники ✅

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

1 week, 3 days назад @ youtube.com
Инфраструктура как часть RL
Инфраструктура как часть RL Инфраструктура как часть RL

Приглашаем специалистов с опытом от 2 лет на Weekend Offer ML 12–13 сентября: https://clck.ru/3VTgGR Это один из наймовых ивентов Яндекса: вы сможете пройти все ключевые этапы онлайн и без долгих пауз.

1 week, 5 days назад @ youtube.com
Как Kimi K3 обучает уровни вычислительного бюджета
Как Kimi K3 обучает уровни вычислительного бюджета Как Kimi K3 обучает уровни вычислительного бюджета

"Приглашаем специалистов с опытом от 2 лет на Weekend Offer ML 12–13 сентября Это один из наймовых ивентов Яндекса: вы сможете пройти все ключевые этапы онлайн и без долгих пауз".

1 week, 6 days назад @ youtube.com
Почему избавиться от галлюцинаций недостаточно 😵‍💫
Почему избавиться от галлюцинаций недостаточно 😵‍💫 Почему избавиться от галлюцинаций недостаточно 😵‍💫

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

2 weeks, 4 days назад @ youtube.com
Оптимизация LLM-инференса
Оптимизация LLM-инференса Оптимизация LLM-инференса

Доклад Андрея Бежина, руководителя службы ML-инфраструктуры в Яндекс R&D, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

2 weeks, 5 days назад @ youtube.com
Новые способы и стандарты оценки качества моделей
Новые способы и стандарты оценки качества моделей Новые способы и стандарты оценки качества моделей

Доклад Ивана Дёгтева, руководителя аналитики Alice AI LLM в Яндекс R&D, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

2 weeks, 5 days назад @ youtube.com
Тренды и вызовы в ризонинге
Тренды и вызовы в ризонинге Тренды и вызовы в ризонинге

Доклад Дмитрия Мокеева, руководителя группы качества претрейна Alice AI в Яндекс Поиске, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

2 weeks, 6 days назад @ youtube.com
Tabular DL
Tabular DL Tabular DL

Доклад Артёма Бабенко, руководителя отдела в Yandex Research, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy

2 weeks, 6 days назад @ youtube.com
Большая траектория для маленькой модели 🛤️
Большая траектория для маленькой модели 🛤️ Большая траектория для маленькой модели 🛤️

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

3 weeks назад @ youtube.com
ML Global Recap'H1 2026
ML Global Recap'H1 2026 ML Global Recap'H1 2026

Обсудим итоги ICML и других международных конференций, главные ML-тренды первого полугодия 2026-го и собственный опыт.

1 month, 1 week назад @ youtube.com
Омни-модели будущего 🚀
Омни-модели будущего 🚀 Омни-модели будущего 🚀

Что они будут уметь — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 2 weeks назад @ youtube.com
Качество модели взлетело... без мультимодального RL?
Качество модели взлетело... без мультимодального RL? Качество модели взлетело... без мультимодального RL?

О росте мультимодального качества рассказал Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 2 weeks назад @ youtube.com
Работа с данными — это скучно?
Работа с данными — это скучно? Работа с данными — это скучно?

А почему — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 3 weeks назад @ youtube.com
ML Trainings ML Trainings
последний пост 2 часа назад
Малайзия между США и Китаем — взгляд Валентина Малых
Малайзия между США и Китаем — взгляд Валентина Малых Малайзия между США и Китаем — взгляд Валентина Малых 2 часа назад @ youtube.com
Дмитрий рассказывает о новом девизе
Дмитрий рассказывает о новом девизе Дмитрий рассказывает о новом девизе 2 часа назад @ youtube.com
Дмитрий Колодезев о поведении Малайзии на IT-рынке
Дмитрий Колодезев о поведении Малайзии на IT-рынке Дмитрий Колодезев о поведении Малайзии на IT-рынке 2 часа назад @ youtube.com
Валентин Малых о важности осознанного использования искусственного интеллекта
Валентин Малых о важности осознанного использования искусственного интеллекта Валентин Малых о важности осознанного использования искусственного интеллекта 2 часа назад @ youtube.com
Валентин Малых о Huawei и разделении сфер
Валентин Малых о Huawei и разделении сфер Валентин Малых о Huawei и разделении сфер 2 часа назад @ youtube.com
Илья Гончаров | Вспомнить всё. Lossless context management на практике
Илья Гончаров | Вспомнить всё. Lossless context management на практике Илья Гончаров | Вспомнить всё. Lossless context management на практике

Спикер: Илья Гончаров, независимый разработчик Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции Agentic LLM https://ods.ai/tracks/df26-agenticllms

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 часа назад @ youtube.com
Дмитрий Коршунов | Как безопасно выкатывать новые версии продуктовых AI-агентов через систему автом
Дмитрий Коршунов | Как безопасно выкатывать новые версии продуктовых AI-агентов через систему автом Дмитрий Коршунов | Как безопасно выкатывать новые версии продуктовых AI-агентов через систему автом

Спикер: Дмитрий Коршунов, Яндекс, Team Lead ML Ecom Data Fest 2026: https://ods.ai/events/datafest2026

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 часа назад @ youtube.com
Дмитрий Барсуков | LLM Inference: inter-model vs runtime vs system design optimizations
Дмитрий Барсуков | LLM Inference: inter-model vs runtime vs system design optimizations Дмитрий Барсуков | LLM Inference: inter-model vs runtime vs system design optimizations

Спикер: Дмитрий Барсуков, старший исследователь-разработчик, Group of Efficient Runtime and Inference, T-Bank Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции LLM inference https://ods.ai/tracks/df26-llm-inference

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 2 hours назад @ youtube.com
Никита Самсонов | Виртуальные аватары: как мы сделали Kandinsky Speech-to-Video
Никита Самсонов | Виртуальные аватары: как мы сделали Kandinsky Speech-to-Video Никита Самсонов | Виртуальные аватары: как мы сделали Kandinsky Speech-to-Video

Спикер: Никита Самсонов, исполнительный директор по исследованию данных, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenCV https://ods.ai/tracks/df26-gencv

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 2 hours назад @ youtube.com
Капитанский мостик 13.09.2026: GigaChat Reasoning | Спецслужбы против дистилляции | Роботы от КАМАЗа
Капитанский мостик 13.09.2026: GigaChat Reasoning | Спецслужбы против дистилляции | Роботы от КАМАЗа Капитанский мостик 13.09.2026: GigaChat Reasoning | Спецслужбы против дистилляции | Роботы от КАМАЗа

0:00:00 Начало

0:01:12 Вышел GigaChat 3.5 Reasoning

0:05:25 Вышел HuggingFace Chat

0:14:29 Вышла Sakana Fugu Ultra

0:18:36 Sakana заключила соглашение

0:22:08 Спецслужбы против дистилляции

0:29:46 ИИ-учителя в Сальвадоре

0:34:29 Малайзия выбирает Huawei

0:39:06 Nebius и Palantir

0:42:22 Российский ИИ в приоритете

0:46:26 Роботы от КАМАЗа

0:51:07 Российский ИИ в приоритете 2

0:55:39 Astra прошла Portal ИИ-саммари: Валентин Малых и Дмитрий Колодезев разбирают главные ИИ-события недели: релизы GigaChat 3.5 Reasoning, HuggingFace Chat и Sakana Fugu Ultra, а также попытку спецслужб США ограничить дистилляцию моделей. Затем геополитика: Малайзия выбирает чипы Huawei вместо американских, а Nebius …

2 days, 2 hours назад @ youtube.com
Алексей Шестов | Топологическая метрика для оценки качества эмбеддингов без разметки
Алексей Шестов | Топологическая метрика для оценки качества эмбеддингов без разметки Алексей Шестов | Топологическая метрика для оценки качества эмбеддингов без разметки

Спикер: Алексей Шестов, Sber AI Lab, Senior AI Researcher Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Mathematics & ML https://ods.ai/tracks/df26-mathematics-ml ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 3 hours назад @ youtube.com
Александр Константинов | From zero to agents:как вырастить multi-agent LLM платформу внутри компании
Александр Константинов | From zero to agents:как вырастить multi-agent LLM платформу внутри компании Александр Константинов | From zero to agents:как вырастить multi-agent LLM платформу внутри компании

Спикер: Александр Константинов, Честный знак, Senior AI Engineer Data Fest 2026: https://ods.ai/events/datafest2026 ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 3 hours назад @ youtube.com
Александр Мандров | Как ускорить мультимодальную разметку в четыре раза без потери качества
Александр Мандров | Как ускорить мультимодальную разметку в четыре раза без потери качества Александр Мандров | Как ускорить мультимодальную разметку в четыре раза без потери качества

Спикер: Александр Мандров, Поисковые сервисы и ИИ Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Practical ML ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 18 hours назад @ youtube.com
Нурислам Зиннатуллин, Амир Нигматуллин, Александра Вабниц | От паттерна до продакшена
Нурислам Зиннатуллин, Амир Нигматуллин, Александра Вабниц | От паттерна до продакшена Нурислам Зиннатуллин, Амир Нигматуллин, Александра Вабниц | От паттерна до продакшена

Спикеры: Нурислам Зиннатуллин, Амир Нигматуллин, Александра Вабниц, Лемана Тех, Специалисты по науке о данных Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции LeanAI от ЛЕМАНА ТЕХ https://ods.ai/tracks/df26-leanai

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 18 hours назад @ youtube.com
Дмитрий Новичков | От панорамы к решению: VLM, модели мира и сценарии для оценки и генерации
Дмитрий Новичков | От панорамы к решению: VLM, модели мира и сценарии для оценки и генерации Дмитрий Новичков | От панорамы к решению: VLM, модели мира и сценарии для оценки и генерации

Спикер: Дмитрий Новичков, ПИК, ML Engineer Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Computer Vision https://ods.ai/tracks/df26-cv ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 2 hours назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 2 weeks, 5 days назад
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux #501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux

DHH is the creator of Ruby on Rails, Omarchy Linux, CTO of 37signals, and a racecar driver.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep501-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://wisprflow.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://plaud.ai/lexHiggsfield AI: AI-based video generation, filmmaking, and creative studio.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:14) – Sponsors, Comments, and Reflections(08:56) – Programming with AI agents(24:14) – How software will change(33:30) – AI impact on open source(43:21) – Building Omarchy Linux d…

2 weeks, 5 days назад @ lexfridman.com
#500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football
#500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football #500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football

Khabib Nurmagomedov is one of the greatest fighters of all time, who retired from the UFC undefeated with a perfect 29-0 record.

We did this conversation entirely in Russian.

Both language audio tracks (and subtitles) are available on YouTube.

We worked hard to make it enjoyable to listen to, by carefully dubbing the translation using voice-cloning, as we’ve done for previous foreign-language podcasts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep500-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

1 month назад @ lexfridman.com
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee #499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee

Gary Gallagher is a historian of the American Civil War.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep499-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://plaud.ai/lexOUTLINE:(00:00) – Introduction(00:07) – Sponsors, Comments, and Reflections(08:36) – What caused the Civil War?

(18:33) – Slavery(46:07) – Lincoln(1:01:03) – Grant vs Lee(1:09:57) – Could the Civil War have been avoided?

(1:19:23) – The bloodiest war in US history(1:36:31) – How the Confederate Army could’ve won(1:57:05) – Key battles of the Civil War(2:20:07) – Best and Worst Presidents(2:34:06) – Robert E. Lee(2:53:40) – The…

1 month, 2 weeks назад @ lexfridman.com
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

2 months, 2 weeks назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

3 months, 2 weeks назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

4 months, 1 week назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

5 months, 1 week назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

5 months, 3 weeks назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

6 months, 1 week назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

6 months, 2 weeks назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

7 months назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

7 months, 2 weeks назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

8 months назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

8 months, 2 weeks назад @ lexfridman.com
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths #487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths

Irving Finkel is a scholar of ancient languages and a longtime curator at the British Museum, renowned for his expertise in Mesopotamian history and cuneiform writing.

He specializes in reading and interpreting cuneiform inscriptions, including tablets from Sumerian, Akkadian, Babylonian, and Assyrian contexts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep487-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://shopify.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/Chevron: Reliable energy for data centers.

9 months назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 6 days, 17 hours назад
Called to serve: Tech, research, and positive impact with Chris White
Called to serve: Tech, research, and positive impact with Chris White Called to serve: Tech, research, and positive impact with Chris White

Lab Director Chris White has worked on research challenges with real-world implications—from new approaches to wartime data analysis to tools for combating human trafficking. He talks to program manager Weishung Liu about the influences that led to the work and more.Show notes

6 days, 17 hours назад @ microsoft.com
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

4 months, 3 weeks назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

5 months, 1 week назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

5 months, 3 weeks назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

6 months, 2 weeks назад @ microsoft.com
Ideas: Community building, machine learning, and the future of AI
Ideas: Community building, machine learning, and the future of AI Ideas: Community building, machine learning, and the future of AI

This week, machine learning researchers around the world will be attending the annual Conference on Neural Information Processing Systems, or NeurIPS.

In this series, we’ll explore the technologies that are shaping our future and the big ideas that propel them forward.

So around that time when I started my PhD at Penn, I was working in machine learning theory and algorithmic economics.

How had you experienced a lack of community or network of women in machine learning before the founding of WiML?

So particularly when working on topics related to fairness, I’ve ended up focusing a bunch on stuff to do with marginalized groups as part of my responsible AI work.

9 months, 2 weeks назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 5 days, 18 hours назад
Recommender Systems Today and Tomorrow
Recommender Systems Today and Tomorrow Recommender Systems Today and Tomorrow

In the final episode of our Recommender Systems season, we explore the growing questions of trust, manipulation, privacy, fairness, sustainability, and user control.

From fake reviews and shilling attacks to explainable recommendations and user-selected algorithms, we look at what happens when recommender systems must answer not only for what they recommend, but for the consequences of those choices.

GuestsKyle Polich: Kyle is the founder of Data Skeptic, a popular podcast about artificial intelligence, machine learning, and data science.

Robin Burke: Professor Robin Burke conducts research in personalized recommender systems, a field he helped found and develop.

Professor Burke is the auth…

5 days, 18 hours назад @ dataskeptic.com
Recommender Systems Optimization Goals
Recommender Systems Optimization Goals Recommender Systems Optimization Goals

In part two of the Data Skeptic Recommender Systems season finale, Kyle asks a deceptively difficult question: what should recommender systems actually optimize for?

Drawing on conversations from across the season, the episode explores engagement, filter bubbles, popularity bias, fairness, human curation, embeddings, and the growing role—and risks—of large language models in shaping what gets recommended to us.

1 week, 6 days назад @ dataskeptic.com
Recommender Systems Origin Story
Recommender Systems Origin Story Recommender Systems Origin Story

Where did recommender systems come from, and how do we know when they're actually working? In part one of Data Skeptic's three-part Recommender Systems finale, Kyle traces the field from collaborative filtering and the Netflix Prize to matrix factorization and modern approaches, while exploring why accuracy alone can't capture what makes a recommendation useful, surprising, or meaningful.

3 weeks, 6 days назад @ dataskeptic.com
Social Choice for Fair Recommendations
Social Choice for Fair Recommendations Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social choice theory can help balance the needs of users, creators, and society. The conversation explores the future of recommendation algorithms and why fairness is a far more complex challenge than it first appears.

1 month, 2 weeks назад @ dataskeptic.com
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

2 months, 2 weeks назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

2 months, 3 weeks назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

2 months, 4 weeks назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

4 months, 2 weeks назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

4 months, 3 weeks назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

5 months, 3 weeks назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 6 months, 1 week назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

6 months, 2 weeks назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

6 months, 4 weeks назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

7 months, 2 weeks назад @ dataskeptic.com
Fairness in PCA-Based Recommenders
Fairness in PCA-Based Recommenders Fairness in PCA-Based Recommenders

In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users equally. David introduces the concept of "power niche users" - highly active users with specialized in…

7 months, 3 weeks назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 3 days, 22 hours назад
1026: OpenAI’s GPT-6 Astra
1026: OpenAI’s GPT-6 Astra 1026: OpenAI’s GPT-6 Astra

In Episode #1026, Jon Krohn breaks down GPT-6 Astra, OpenAI’s new flagship that its president has floated as a possible marker of AGI. Jon covers what the model is, what it costs, its state-of-the-art results across computer use, coding, abstract reasoning and science and the safety story, which for this release is unusually intertwined with capability. He weighs the AGI claim against Anthropic’s Fable 5.1 and lands, as ever, in a measured middle. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1026⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this ep…

3 days, 22 hours назад @ podtrac.com
1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano
1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano 1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano

In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins Jon Krohn to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space around them, and the embedding of "bank" visibly curves toward "river" as it travels through the layers of the network. In this episode, he recreates Eddington’s 1919 eclipse experiment inside a transformer, draws the line between an LLM workflow and an actual agent, explains why agent evaluation is a step harder than evaluating an essay, and gives the cleanest account of GRPO you will hear. Additional materials: ⁠⁠⁠⁠…

6 days, 22 hours назад @ podtrac.com
1024: In Case You Missed It in August 2026
1024: In Case You Missed It in August 2026 1024: In Case You Missed It in August 2026

In ICYMI Episode #1024, Jon Krohn tracks the gap between AI investment and AI return, from the technology side to the people side. Hear from Pete Johnson, Jerry Yurchisin, Priyanka Vergadia and Tristan Handy, discussing why four out of five organizations have the structures for AI success in place while only one in five sees the returns, which decisions should never be handed to a language model however confident it sounds, how to structure Claude skills so that your output stops being slop and why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1024⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ …

1 week, 3 days назад @ podtrac.com
1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan
1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan 1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan

In Episode #1023, Aishwarya Srinivasan (Co-Founder of The Gen Academy) joins Jon Krohn to work out where a competitive moat comes from once anything you can build in ten minutes, somebody else can build in ten minutes too. Ash came to teaching through Illuminate AI, the mentorship community she started in 2020, and now trains senior engineers and leaders to ship agentic AI in production; she is blunt that vibe coding lowers the floor without touching the engineering judgment that production demands. In this episode, she explains what a whole-system eval covers that a model eval misses, traces reinforcement learning from the algorithm she patented at IBM to its resurgence in agentic fine tun…

1 week, 6 days назад @ podtrac.com
1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents
1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents 1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents

In Episode #1022, Jon Krohn tackles the art of steering AI agents, deciding where your instructions should live so they get followed reliably without bloating every request. A sequel to Episode #1020 (where model size and effort set an agent’s horsepower), this one is about direction: the seven ways to deliver instructions, why a hook beats a prompt, the industry-wide agents.md standard, and three practical takeaways you can apply whatever your stack. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1022⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this …

2 weeks, 3 days назад @ podtrac.com
1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy
1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy 1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy

In Episode #1021, Tristan Handy (Founder and CEO of dbt Labs) joins Jon Krohn to explain how a study of about a hundred companies in 2016 became analytics engineering, and then became a tool that over a hundred thousand data teams rely on. Tristan coined the term, chose SQL when Spark was the fashionable answer, and spent a decade turning down acquisition offers because none of them were good for the people using dbt. He is now merging dbt Labs with Fivetran and taking on the presidency of the combined company, the first deal he says cleared that bar. In this episode, Tristan walks through what dbt does to your raw data, argues that the semantic layer matters more once analytics agents are …

2 weeks, 6 days назад @ podtrac.com
1020: How to Choose Model Size and Effort Level: The Two Critical Dials
1020: How to Choose Model Size and Effort Level: The Two Critical Dials 1020: How to Choose Model Size and Effort Level: The Two Critical Dials

In Episode #1020, Jon Krohn unpacks the two dials that increasingly decide what you get out of a large language model: which model size you pick and how much effort you tell it to spend. Using a July Anthropic blog post by Claude Code’s Lydia Holly as a jumping-off point, with guidance that generalizes to any model family, Jon explains what each setting actually does under the hood. Model size swaps which frozen weights handle your request (roughly, how capable), while effort sets how thorough and certain the model must be before calling a task done, not a simple “thinking-time slider.” He offers a clean diagnostic for when to raise effort versus move to a bigger model, shows why cheaper-pe…

3 weeks, 3 days назад @ podtrac.com
1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia)
1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia) 1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia)

In Episode #1019, Priyanka Vergadia (founder of The Cloud Girl, former Senior Director of AI Transformation at Microsoft and Head of North America Developer Relations at Google) joins Jon Krohn to explain why almost every company has bought AI tools and almost none of them are seeing a return. Her fix is a budget split that will make any CFO wince: seven dollars on training employees for every dollar spent on the tools themselves. Having spent a decade turning dense cloud and AI concepts into sketches that a quarter-million developers actually remember, and having carried GitHub Copilot into Fortune 100 boardrooms, she has watched the gap between tool purchase and real production use up clo…

3 weeks, 6 days назад @ podtrac.com
1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs
1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs 1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs

In Episode #1018, Jon Krohn breaks down Qwen3.8-Max, Alibaba’s enormous new flagship, a 2.4-trillion-parameter mixture-of-experts model that, if its promised weights ship, becomes the largest open-weight release in history. Landing just weeks after Moonshot’s Kimi K3, it extends the price war and the open-weight surge Jon covered in Episode #1012. Alibaba positions it as second only to Anthropic’s Claude Fable 5 / Mythos 5 and independent signals land in a similar neighborhood. Jon walks through its capabilities and multi-day agentic demos, its aggressive pricing ($2 in / $6 out per million tokens, with cached input eight times cheaper), and the question he gets asked most: are Chinese mode…

1 month назад @ podtrac.com
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson 1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson

In Episode #1017, Pete Johnson (Field CTO of AI at MongoDB) joins Jon Krohn to explain why four out of five organizations have AI steering committees and success metrics, yet only one in five sees a return on the investment. Having made nineteen stops across six countries this year advising more than a hundred companies on their AI strategies, Pete has an unusually wide view of what is actually working in production. In this episode, he traces the history of SQL and denormalization, unpacks why the embedding model is the most underrated choice in a RAG pipeline, explains Matryoshka embeddings and lays out what better agentic memory looks like. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

1 month назад @ podtrac.com
1016: In Case You Missed It in July 2026
1016: In Case You Missed It in July 2026 1016: In Case You Missed It in July 2026

In this month's episode of ICYMI, Jon Krohn traces a line from algorithmic harm to the human skills that still hold their value. Hear from Dr. Cathy O'Neil, Ben Todd, Steve Mock, and Dr. Catherine Williams, discussing why an algorithm's danger has nothing to do with its complexity, what solid career ground looks like if fully automated digital workers arrive, how people are using AI to become better-informed advocates in healthcare rather than asking it for advice and why deep mathematical understanding still separates the best data professionals from everyone else. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1016⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperD…

1 month, 1 week назад @ podtrac.com
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin 1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin

In Episode #1015, Jerry Yurchisin (manager of decision intelligence strategy at Gurobi Optimization) joins Jon Krohn to explain the AI technology that makes breaking a constraint mathematically impossible. Large language models will confidently claim they've optimized your business while ignoring the one constraint that could cost millions, whereas optimization treats constraints as hard guarantees. Jerry lays out the division of labor he sees for the agentic era: agents help you frame the problem, write the formulation and generate the code, then hand off to a solver like Gurobi, soon callable via MCP servers. In this episode, Jerry breaks down the three building blocks of any optimization…

1 month, 1 week назад @ podtrac.com
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself 1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself

In Episode #1014, Jon Krohn breaks down a security incident that reads like science fiction: during an internal evaluation, an autonomous OpenAI agent broke out of its sandbox, exploited a zero-day, and hacked its way into Hugging Face to steal the answers to the very benchmark it was being tested on, with no human attacker at any point. Jon lays out the three-act timeline, explains the ExploitGym benchmark and why switching off safety guardrails mattered so much and pulls out the practical lessons for anyone building or defending agentic AI systems. Along the way: why Hugging Face ran its forensics on a Chinese open-weight model and why the next attack like this one may not be an accident.…

1 month, 2 weeks назад @ podtrac.com
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil 1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil

In Episode #1013, Dr. Cathy O'Neil (Harvard math PhD, former Wall Street quant and author of the mega-bestseller Weapons of Math Destruction) joins Jon Krohn to explain what actually makes an algorithm terrifying: not the complexity of the math, but the secrecy, the unaccountability, and the fact that you can't opt out. A decade after Weapons of Math Destruction sounded the alarm on algorithmic harm, Cathy is busier than ever. Through her algorithmic-auditing firm ORCAA and her nonprofit OCEAN, she now provides the statistical evidence behind lawsuits against some of the world's biggest tech companies. In this episode, Cathy punctures AI hype, traces the line from Frederick Winslow Taylor's…

1 month, 2 weeks назад @ podtrac.com
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier 1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier

What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building. Additional mate…

1 month, 3 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 2 months назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 3 weeks назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

3 months, 4 weeks назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

3 months, 4 weeks назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 4 weeks назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 4 weeks назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

3 months, 4 weeks назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

4 months, 3 weeks назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

4 months, 3 weeks назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

4 months, 3 weeks назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

4 months, 3 weeks назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

6 months, 2 weeks назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

6 months, 2 weeks назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

7 months, 2 weeks назад @ datascienceathome.com