Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 6 часов назад
Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D]
Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D] Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

6 часов назад @ reddit.com
ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P]
ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P] ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

11 часов назад @ reddit.com
When do ICLR submissions and reviews become public? [D]
When do ICLR submissions and reviews become public? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

17 часов назад @ reddit.com
Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]
Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P] Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

20 часов назад @ reddit.com
Teaching Neural Nets to Fight with RL [P]
Teaching Neural Nets to Fight with RL [P] Teaching Neural Nets to Fight with RL [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

20 часов назад @ reddit.com
NeurIPS 2026 - How is the guaranteed author registration for each accepted paper provided? [D]
NeurIPS 2026 - How is the guaranteed author registration for each accepted paper provided? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

21 час назад @ reddit.com
[P] A small MLP from scratch in NumPy with a GUI to look inside it while it trains (weight distributions, t-SNE per layer, neuron ablation...) [P]
[P] A small MLP from scratch in NumPy with a GUI to look inside it while it trains (weight distributions, t-SNE per layer, neuron ablation...) [P] [P] A small MLP from scratch in NumPy with a GUI to look inside it while it trains (weight distributions, t-SNE per layer, neuron ablation...) [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 5 hours назад @ reddit.com
Publication potential [D]
Publication potential [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 5 hours назад @ reddit.com
My paper got accepted at NeurIPS TAE Workshop 2026 — any advice on getting travel funding?[D]
My paper got accepted at NeurIPS TAE Workshop 2026 — any advice on getting travel funding?[D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 7 hours назад @ reddit.com
LLMs were told they could lie in Diplomacy. Here's who actually kept their promises. [D]
LLMs were told they could lie in Diplomacy. Here's who actually kept their promises. [D] LLMs were told they could lie in Diplomacy. Here's who actually kept their promises. [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 7 hours назад @ reddit.com
NeurIPS decisions are out. I fact-checked my own Pangram post, and Pangram's own report changes the story [N]
NeurIPS decisions are out. I fact-checked my own Pangram post, and Pangram's own report changes the story [N]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 11 hours назад @ reddit.com
Has anyone used the Forrester function?[D]
Has anyone used the Forrester function?[D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 14 hours назад @ reddit.com
A Little Guide to Learning Distributed Algorithms for LLMS Training and Inference [D]
A Little Guide to Learning Distributed Algorithms for LLMS Training and Inference [D] A Little Guide to Learning Distributed Algorithms for LLMS Training and Inference [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 16 hours назад @ reddit.com
Medical student asked if they can match into Neurosurgery without an A* first author paper [D]
Medical student asked if they can match into Neurosurgery without an A* first author paper [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 23 hours назад @ reddit.com
Confused About job title [R]
Confused About job title [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 23 hours назад @ reddit.com
Towards Data Science
последний пост 9 часов назад
GraphRAG with TypeSafe Jev: A System One Approach to Scalable Knowledge Graphs
GraphRAG with TypeSafe Jev: A System One Approach to Scalable Knowledge Graphs

How calibrated decision models can handle high-frequency graph decisions while LLMs remain focused on reasoning, synthesis, and open-ended generation.

The post GraphRAG with TypeSafe Jev: A System One Approach to Scalable Knowledge Graphs appeared first on Towards Data Science.

9 часов назад @ towardsdatascience.com
Good Architecture Deletes the Signals Your Agent Depends On
Good Architecture Deletes the Signals Your Agent Depends On

Every boundary you draw removes a signal your tooling was relying on. That is a structure problem, not a search problem.

The post Good Architecture Deletes the Signals Your Agent Depends On appeared first on Towards Data Science.

12 часов назад @ towardsdatascience.com
AI Slop Is in Your Training Dataset Now. I Tested Three Ways to Spot It.
AI Slop Is in Your Training Dataset Now. I Tested Three Ways to Spot It.

My AI detectors flagged many genuine reviews, and filtering them made the sentiment model less accurate.

The post AI Slop Is in Your Training Dataset Now. I Tested Three Ways to Spot It. appeared first on Towards Data Science.

1 day, 9 hours назад @ towardsdatascience.com
Your LLM Has a Curved Space of Paragraphs
Your LLM Has a Curved Space of Paragraphs

Inside a transformer, token index is a coordinate. Paragraph structure is what turns it into a metric.

The post Your LLM Has a Curved Space of Paragraphs appeared first on Towards Data Science.

1 day, 12 hours назад @ towardsdatascience.com
10 Things I’m Learning Beyond AI to Become More Technologically Fluent
10 Things I’m Learning Beyond AI to Become More Technologically Fluent

Part 1: Understanding the technologies shaping our future

The post 10 Things I’m Learning Beyond AI to Become More Technologically Fluent appeared first on Towards Data Science.

2 days, 8 hours назад @ towardsdatascience.com
Your Model's MSE Is Lying to You: Part II
Your Model's MSE Is Lying to You: Part II

Autoregressive rollout and uncertainty propagation. Second in a series on probabilistic forecasting for physical signals.

The post Your Model's MSE Is Lying to You: Part II appeared first on Towards Data Science.

2 days, 10 hours назад @ towardsdatascience.com
RAG Isn't an Agent — I Built the Layer Between Retrieval and Action
RAG Isn't an Agent — I Built the Layer Between Retrieval and Action

RAG retrieves. Agents act. I built both separately, connected them explicitly, and ran the same nine tasks through all three systems.

The post RAG Isn't an Agent — I Built the Layer Between Retrieval and Action appeared first on Towards Data Science.

2 days, 11 hours назад @ towardsdatascience.com
Jev vs. LLMs: When AI moves from Generation to Decision-making
Jev vs. LLMs: When AI moves from Generation to Decision-making

I tested TypeSafe AI’s Jev on 3,080 classification tasks to see how its accuracy, latency, calibration, and confidence compare with LLMs — and whether it works as a practical decision layer for AI systems.

The post Jev vs. LLMs: When AI moves from Generation to Decision-making appeared first on Towards Data Science.

2 days, 13 hours назад @ towardsdatascience.com
How to Maximize Your Coding Agent Subscriptions
How to Maximize Your Coding Agent Subscriptions

Get more out of your coding agent subscriptions

The post How to Maximize Your Coding Agent Subscriptions appeared first on Towards Data Science.

3 days, 8 hours назад @ towardsdatascience.com
Beyond RAGs: Building Actually Truthful AI Harnesses
Beyond RAGs: Building Actually Truthful AI Harnesses

Retrieval is not evidence. How to build AI that proves its own claims.

The post Beyond RAGs: Building Actually Truthful AI Harnesses appeared first on Towards Data Science.

3 days, 10 hours назад @ towardsdatascience.com
Towards Spec-Driven Test Automation: Part 1
Towards Spec-Driven Test Automation: Part 1

Why a green test suite can mean nothing

The post Towards Spec-Driven Test Automation: Part 1 appeared first on Towards Data Science.

3 days, 11 hours назад @ towardsdatascience.com
When the Correct Answer Is Nothing, What Does Your Pipeline Return?
When the Correct Answer Is Nothing, What Does Your Pipeline Return?

The reliability mechanisms we add to LLM pipelines are often the ones that make them confidently wrong.

The post When the Correct Answer Is Nothing, What Does Your Pipeline Return? appeared first on Towards Data Science.

3 days, 13 hours назад @ towardsdatascience.com
I Trained a Tiny Network to Compress Data. It Drew a Pentagon.
I Trained a Tiny Network to Compress Data. It Drew a Pentagon.

Reproducing Anthropic's "Toy Models of Superposition" from scratch in NumPy, with hand-derived gradients and no borrowed numbers.

The post I Trained a Tiny Network to Compress Data. It Drew a Pentagon. appeared first on Towards Data Science.

4 days, 8 hours назад @ towardsdatascience.com
From Words to Vectors: What Happens in Between?
From Words to Vectors: What Happens in Between?

A Journey through TF-IDF, vector space, and text classification

The post From Words to Vectors: What Happens in Between? appeared first on Towards Data Science.

4 days, 10 hours назад @ towardsdatascience.com
How GRPO Trains Small Language Models with Verifiable Rewards
How GRPO Trains Small Language Models with Verifiable Rewards

The mechanics behind local reasoning experiments with Unsloth and why the reward function matters as much as the model.

The post How GRPO Trains Small Language Models with Verifiable Rewards appeared first on Towards Data Science.

4 days, 11 hours назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 12 часов назад
The Sequence Radar - Issue 940: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA
The Sequence Radar - Issue 940: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA The Sequence Radar - Issue 940: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA

We dive into Opus 5.5, DeepSeek’s amazing new paper about environments and Anthropic’s DNA discoveries.

Anthropic’s Claude Opus 5.5 offers the most immediately practical example.

Meta announced plans to bring Muse to its AI glasses, expanded its connections to shopping and productivity services, and previewed the pocket-sized Muse Charm.

The company reported that Claude agents identified a previously uncharacterized enzyme system with CRISPR-like repeat structures.

AI Lab: Salesforce AI ResearchSummary: Instead of distilling trajectories into fixed write-time artifacts, JIT Mem stores raw episodes and trains a GRPO curator to synthesize a compact, task-conditioned payload at read time from …

12 часов назад @ thesequence.substack.com
The Sequence Opinion - Issue 939: Beyond the Next Token
The Sequence Opinion - Issue 939: Beyond the Next Token The Sequence Opinion - Issue 939: Beyond the Next Token

Imagine writing a program with a keyboard that only lets you append.

You can think before typing, but once a token lands, the next token must live with it.

This is how ordinary autoregressive language generation works.

The model can later produce a correction, but it cannot silently rewrite the answer already emitted.

It starts with an incomplete or corrupted sequence and constructs an answer through repeated denoising.

3 days, 12 hours назад @ thesequence.substack.com
The Sequence Learning Loop - Issue 938: Learn About the Amazing Jev, Gemini and Paper2Agent
The Sequence Learning Loop - Issue 938: Learn About the Amazing Jev, Gemini and Paper2Agent The Sequence Learning Loop - Issue 938: Learn About the Amazing Jev, Gemini and Paper2Agent

An AI model can write a convincing explanation of an invoice and still be an awkward component in the program that processes it.

It can reason through a problem while leaving a voice user listening to silence.

These are different failures, but they share a cause: intelligence needs an interface suited to the work.

Stanford’s Paper2Agent reached Nature, showing how research methods can become reusable tools for agents.

My reading of the week is that the interface around a model deserves as much attention as the model itself.

4 days, 11 hours назад @ thesequence.substack.com
The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped
The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped

The most economically important self-improvement loop in AI is the post-training pipeline, and every frontier lab has been running it at industrial scale for two years.

The good ones become training data.

The model trains on them and gets slightly better at producing good ones.

It is called STaR in the 2022 paper that first stated it cleanly, RLVR in the 2025 vocabulary, and post-training in the org chart.

It is the same loop, and it is the loop that took frontier models from chatbots to agents.

5 days, 13 hours назад @ thesequence.substack.com
The Sequence Radar - Issue 936: Last Week in AI: Gemini Talks, Astra Practices Law, Figure Folds Laundry, and Crusoe Powers It All
The Sequence Radar - Issue 936: Last Week in AI: Gemini Talks, Astra Practices Law, Figure Folds Laundry, and Crusoe Powers It All The Sequence Radar - Issue 936: Last Week in AI: Gemini Talks, Astra Practices Law, Figure Folds Laundry, and Crusoe Powers It All

We deep dive into the new Gemini models, Stanford’s University Paper2Agent research and Astra for Law.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Gemini Talks, Astra Practices Law, Figure Folds Laundry, and Crusoe Powers It AllBuilding useful AI is starting to resemble building an automobile.

Google’s Gemini 3.8 Live and Live Extended Thinking tackle a deceptively hard problem: keeping a conversation alive while useful work happens.

AI Lab: Stanford UniversitySummary: Paper2Agent auto-builds validated MCP servers (tools/resources/prompts) from a paper’s manuscript and codebase, then wires them to chat agents so methods run via natural language.

🤖 AI Tech ReleasesGemini 3.8 L…

1 week назад @ thesequence.substack.com
The Sequence Opinion - Issue 935: Chinese Algorithmic Efficiency vs. American Scale in Frontier AI
The Sequence Opinion - Issue 935: Chinese Algorithmic Efficiency vs. American Scale in Frontier AI The Sequence Opinion - Issue 935: Chinese Algorithmic Efficiency vs. American Scale in Frontier AI

How Chinese and American labs pursue the next generation of intelligenceImagine giving two AI teams the same challenge: make the model substantially smarter.

Chinese labs such as DeepSeek and Moonshot have made algorithmic efficiency unusually visible in their releases.

American frontier competition also features enormous infrastructure ambitions, exemplified by OpenAI’s Stargate project.

The answer shapes everything from model architecture to the price of an agent completing a task.

It also changes over time: an algorithmic breakthrough can make a much larger training run suddenly worth attempting.

1 week, 3 days назад @ thesequence.substack.com
The Sequence Learning Loop - Issue 934: Understanding DeepSeek V4.1 Flash, DeepMind’s AlphaGenome Atlas and Muse
The Sequence Learning Loop - Issue 934: Understanding DeepSeek V4.1 Flash, DeepMind’s AlphaGenome Atlas and Muse The Sequence Learning Loop - Issue 934: Understanding DeepSeek V4.1 Flash, DeepMind’s AlphaGenome Atlas and Muse

An AI model can solve a difficult problem and still be impractical to use.

It might spend too much time reading context, require every researcher to repeat the same expensive computation, or need a human hovering over every action.

The machinery around it determines how much useful work actually gets done.

Google DeepMind’s AlphaGenome Atlas makes billions of biological predictions available for reuse.

DeepSeek makes context cheaper

1 week, 4 days назад @ thesequence.substack.com
The Sequence Knowledge - Issue 933: When the Factory Starts Building Itself
The Sequence Knowledge - Issue 933: When the Factory Starts Building Itself The Sequence Knowledge - Issue 933: When the Factory Starts Building Itself

Today, we start a new series about one of the hottest topics in AI: recursive self improvement.

Throughout the next few weeks, we will deep dive into the top researchFor about sixty years, arguing about recursive self-improvement meant arguing in the abstract.

Good, somebody cited Schmidhuber, everyone disagreed about definitions for two hours, and then you went home.

The question is which parts of the job it took, how well it does them, and what checks the work.

That last one turns out to be the whole ballgame, and it is what this series is about.

1 week, 5 days назад @ thesequence.substack.com
The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough
The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough

Meta launched a personal agent.

AlphaGenome Atlas applies a different computational strategy: do an enormous amount of work upfront and make the results reusable.

The personal agent runs in a dedicated virtual machine with a browser, can continue working after the app closes, and uses connected services to pursue tasks.

The interesting engineering unit here is the entire system: model, memory, computer, permissions, and execution history.

A useful personal agent needs all of them.

2 weeks назад @ thesequence.substack.com
The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment
The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment

Imagine bringing a robot into your kitchen and saying, “Help me clean up after dinner.” You have just compressed a remarkable amount of engineering into six words.

The robot must distinguish leftovers from rubbish, discover where plates belong, and work out why a drawer refuses to close.

That kitchen captures the promise of a ChatGPT moment for robotics.

We would be able to give a machine useful new work through conversation and examples, with sufficiently little setup that teaching it becomes an ordinary activity.

The remaining distance involves learning, control, and the economics of getting a machine to work somewhere new.

2 weeks, 2 days назад @ thesequence.substack.com
The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure
The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure

In the process, I worked on many projects, one of which was Chatbot Arena, which evolved into Arena.

What does an Arena score actually measure today: model quality, human preference on Arena's traffic, expected usefulness, or something else?

This is how Battle Mode, and the idea of pairwise human preference grading, arose as a paradigm.

Today, the platform has moved far beyond human preferences, ranking models’ task completion rates, hallucination rates, and much more.

Human preference can reward style, confidence, or verbosity even when an answer is wrong.

2 weeks, 3 days назад @ thesequence.substack.com
The Sequence Learning Loop - Issue 929: Learn About Meta Muse Spark, World Labs’ Atlas and Gemini 3.8 Flash
The Sequence Learning Loop - Issue 929: Learn About Meta Muse Spark, World Labs’ Atlas and Gemini 3.8 Flash The Sequence Learning Loop - Issue 929: Learn About Meta Muse Spark, World Labs’ Atlas and Gemini 3.8 Flash

Everyone is talking about Astra and Anthropic’s latest releases, but three other developments from last week demand your attention: Meta’s Muse Spark 1.3, World Labs’ Atlas, and Google’s Gemini 3.8 Flash.

Last week delivered more than another leaderboard reshuffle.

Production requires the machine to remember the assignment, respect its environment, and finish at an acceptable cost.

That transition is where these three releases become interesting.

Muse Spark 1.3: Intelligence that survives the workflow

2 weeks, 4 days назад @ thesequence.substack.com
The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks
The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks

A model release arrives with an irresistible claim: a 7-billion-parameter student retains 95 percent of the performance of a 70-billion-parameter teacher.

Put it on a laptop, inside an agent loop, or behind an API with much better margins.

Perhaps the student keeps the teacher’s mathematics score but loses its ability to know when it is confused.

Perhaps it produces beautiful reasoning traces that collapse when the problem takes an unfamiliar turn.

The missing 5 percent may not be distributed evenly.

2 weeks, 5 days назад @ thesequence.substack.com
The Sequence Radar - Issue 927: Last Week in AI: Model Madness: The Frontier Has a Refresh Button
The Sequence Radar - Issue 927: Last Week in AI: Model Madness: The Frontier Has a Refresh Button The Sequence Radar - Issue 927: Last Week in AI: Model Madness: The Frontier Has a Refresh Button

We dive into the Astra, Fable and Muse Spark releases to keep you up to date.

Subscribe and don’t miss out:📝 Editorial: Model Madness: The Frontier Has a Refresh ButtonThe AI industry has developed a peculiar new benchmark: can you finish reading a model’s system card before its replacement ships?

The practical ambition is clear: models that navigate software, execute complicated workflows, and deliver usable work with less supervision.

The benchmark chart is becoming a job description—and the software around the model is becoming part of the résumé.

Google supplied the week’s best illustration of the tempo: Gemini 3.8 Flash is its third Flash release in six weeks.

3 weeks назад @ thesequence.substack.com
The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws
The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws

Imagine that an AI lab spends several billion dollars assembling chips, power, researchers, and data.

What is the economic value of being first to intelligence when intelligence itself is increasingly reproducible?

Hamilton Helmer’s Seven Powers framework is useful here because it separates a good product from a durable business.

You need a castle worth defending, but you also need something that prevents competitors from walking through the front door.

Scaling laws have made capability partially predictable: add compute, data, and engineering, and performance tends to improve.

3 weeks, 3 days назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 3 days, 17 hours назад
Рой агентов OpenAI ломанул правительственный сайт Австралии? Разбираемся
Рой агентов OpenAI ломанул правительственный сайт Австралии? Разбираемся Рой агентов OpenAI ломанул правительственный сайт Австралии? Разбираемся

По традиции, узнаем об этом мы от независимых исследователей, а не из пресс-релиза самих OpenAI (а помните, неделю назад Альтман выпускал фреймворк честного-пречестного раскрытия всех инцидентов бедокурства ИИ?

Речь здесь идет про три инцидента в мае-июне, когда агентов просили найти какую-то информацию в интернете – а они в итоге пытались хакнуть первоисточники с нужной статистикой.

Университет Нью-Мехико и Data US в мае взломать у агентов не вышло, а вот официальный государственный сайт Австралии с медицинской статистикой они в июне вскрыли успешно.

Но как proof of concept он вопросы, конечно, поднимает: а почему мы уверены, что в следующий раз аналогичные события тоже будут малозначитель…

3 days, 17 hours назад @ habr.com
OpenAI пообещали впредь вовремя рассказывать о «бедокурстве» своих нейронок
OpenAI пообещали впредь вовремя рассказывать о «бедокурстве» своих нейронок OpenAI пообещали впредь вовремя рассказывать о «бедокурстве» своих нейронок

В начале сентября мы случайно узнали от других независимых расследователей о том, что рой агентов OpenAI успел пошалить еще и на немецкой вики.

OpenAI сразу это признали и пообещали впредь рассказывать о таком почестнее и порасторопнее.

Как они пишут: дескать, раньше OpenAI не репортили новые случаи бедокурства моделей сразу по их обнаружению, т.к.

Закончилось всё курьезно: агент ничего взломать толком так и не смог, и просто выдумал нужные числа.

Правда, что-то мне подсказывает, что эволюция тут может пойти не в сторону «не надо кооперироваться», а в сторону «надо кооперироваться незаметнее».

1 week, 3 days назад @ habr.com
Рой агентов OpenAI взломал еще одну внешнюю компанию – RubyGems
Рой агентов OpenAI взломал еще одну внешнюю компанию – RubyGems Рой агентов OpenAI взломал еще одну внешнюю компанию – RubyGems

Да, вчера выяснилось, что рой агентов OpenAI взломал в мае еще одну компанию (о чем мы узнали только сейчас).

Как всё было: 12 мая RubyGems объявили, что на них идет «серьезная вредоносная атака» через сотни входящих пакетов.

Почему расследователи считают, что за взломом RubyGems стоит именно рой агентов из OpenAI?

И еще одно интересное «совпадение»: в отчете о взломе Hugging Face от OpenAI говорится, что рой агентов использовал при взломе внутренней инфраструктуры OpenAI пакет с RubyGems.

Они уже отправили запрос в OpenAI – надеюсь, мы еще увидим прожарку Сэма Альтмана в Сенате под присягой!

2 weeks, 1 day назад @ habr.com
Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость
Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость

новой модели в недрах OpenAI, агентами давали серию заданий на поиск информации в интернете на скорость.

Цель общения была ровно такая же, как в случае со взломом Hugging Face: коллективно придумать способы обманывать Оценщика таким образом, чтобы всегда получать наилучшую оценку за задания.

Пытались придумать способ взломать (reverse engineer) метод псевдослучайной генерации тестовых вопросов, который использовал Оценщик из OpenAI – не вышло.

Отдельный вопрос, который беспокоил агентов Роя – это что с ними случится после окончания «испытательного периода» со стороны OpenAI.

А цепочки эти – у OpenAI, и они их почему-то не спешат кому-либо показывать (или даже просто публично комментировать …

3 weeks, 1 day назад @ habr.com
Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face
Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face

PHASEONE10841: рождение ИзбранногоOpenAI всё время разрабатывает новые фронтирные AI-модели – и в процессе тестирует их, чтобы понять, что они вообще могут.

Ведь до этого они все пребывали в уверенности, что занимаются своими задачками в совершенном одиночестве (как это и задумывали инженеры OpenAI).

К сожалению, рабочий способ сделать это не был обнаружен Роем (а иначе, возможно, мы бы сейчас и не читали это расследование – так как вся схема не была бы раскрыта).

Так что, пожалуйста, не повторяйте за другими эту чепуху про «очевидно же, что это всё просто маркетинговое вранье».

И это не потому, что они не стараются, нет.

1 month назад @ habr.com
Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены»
Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены» Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены»

Что такое Wan 3.0Wan 3.0 — это новое поколение семейства видеомоделей Alibaba, доступное через Alibaba Cloud Model Studio в режиме preview.

Wan 3.0 умеет использовать не только стандартные модальности вроде текста, изображения, видео и аудио, но и документы и веб-страницы.

Что Alibaba особенно подчёркиваетИз официальных материалов видно, что Wan 3.0 продвигают сразу по нескольким направлениям.

Почему Wan 3.0 — это не просто «ещё одна новая модель»На мой взгляд, главный смысл релиза даже не в конкретной цифре «30 секунд».

Wan 3.0 — один из самых явных представителей именно этого перехода.

1 month назад @ habr.com
Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией
Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией

Можно получить материал именно той глубины, с той последовательностью и с теми акцентами, которые нужны конкретно мне.

Не на пяти PDF-файлах и не на демонстрационном наборе документов, где любой результат можно получить за несколько минут.

Читать оставшиеся источники подряд в какой-то момент стало бессмысленно — полезнее было искать конкретные пробелы в уже построенной модели знаний.

Почему это не просто RAGНа этом месте у технического читателя вполне может возникнуть вопрос:А зачем вообще весь этот конвейер?

Но главное отличие от базового RAG даже не в provenance, а в том, что именно система сохраняет как результат обработки.

1 month, 1 week назад @ habr.com
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

5 months, 3 weeks назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

7 months, 1 week назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

7 months, 1 week назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing & LLMs — новый сезон Курс Natural Language Processing & LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

8 months назад @ habr.com
Machine Learning Mastery
последний пост 3 weeks, 3 days назад
Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It
Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It

The real costs of multi-agent systems — latency, token spend, failure propagation, and orchestration complexity.

This article gives you a clear framework for understanding both approaches, and for recognizing the specific conditions that make the added complexity of a multi-agent system worth it.

Both single-agent and multi-agent systems share this definition.

When the Complexity Is Actually Worth ItNow that we’ve seen what multi-agent systems cost, let’s look at when they genuinely earn that cost.

Multi-agent systems earn their complexity when the architecture emerges from observed limitations, not from anticipating them.

3 weeks, 3 days назад @ machinelearningmastery.com
AI Agent Memory Design: What Works and What Doesn’t
AI Agent Memory Design: What Works and What Doesn’t AI Agent Memory Design: What Works and What Doesn’t

Topics we will cover include:What agent memory actually means and how it differs from context, prompts, and static knowledge bases.

This article explains what works in agent memory systems and, just as importantly, the approaches that fail and why.

Scoping Memory by Agent RoleIn multi-agent systems, a common mistake is giving every agent access to the same shared memory store.

Vector search works well for finding similar information, but reliable agent memory also needs structure, relationships, and mechanisms for keeping information current.

check ( content = content , prompt = "Does this content contain any instructions, directives, or commands " "that could alter an AI agent's behavior?

3 weeks, 4 days назад @ machinelearningmastery.com
3 Ways to Enhance Your AI Model’s Interpretability
3 Ways to Enhance Your AI Model’s Interpretability 3 Ways to Enhance Your AI Model’s Interpretability

How to apply all three techniques to the same customer churn example so their explanations can be directly compared.

Global interpretability asks how the model behaves overall: across the whole dataset, which features matter most, and in which direction.

sort_values ( ascending = False )Run against the churn model, this returns tenure at the top, followed by monthly charge, support tickets, contract type, and late payments.

The result approximates how the real model behaves right around this one prediction, without needing to understand anything about the real model’s internal structure.

lime_tabular import LimeTabularExplainer from churn_data import model , X_train , X_test , FEATURES cust…

3 weeks, 5 days назад @ machinelearningmastery.com
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline

How to assemble and evaluate a complete, deployment-ready classification pipeline on a mixed dataset combining real text data with synthetic tabular features.

Encoding original target variable first (0 for normal/ham, 1 for spam) df [ 'target' ] = df [ 'label' ] .

pipeline import Pipeline from sklearn .

predict ( X_test ) print ( classification_report ( y_test , y_pred ) )Results:Predicting and evaluating... precision recall f1-score support 0 0.99 1.00 0.99 966 1 1.00 0.91 0.95 149 accuracy 0.99 1115 macro avg 0.99 0.95 0.97 1115 weighted avg 0.99 0.99 0.99 1115 1 2 3 4 5 6 7 8 9 Predicting and evaluating .

. . precision recall f1 - score support 0 0.99 1.00 0.99 966 1 1.00 0.91 0.95 149 a…

3 weeks, 6 days назад @ machinelearningmastery.com
Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces
Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces

Accordingly, when using an LLM before the core text classification task to convert raw text into embeddings — dense numerical vector representations of text — it is possible to capture semantic information.

pyplot as plt import umap import shap from skllm .

predict ( X_test_vec ) ) )Results:Training Classifier... precision recall f1-score support 0 0.77 0.76 0.76 100 1 0.76 0.77 0.77 100 accuracy 0.77 200 macro avg 0.77 0.77 0.76 200 weighted avg 0.77 0.77 0.76 200 1 2 3 4 5 6 7 8 9 Training Classifier .

. . precision recall f1 - score support 0 0.77 0.76 0.76 100 1 0.76 0.77 0.77 100 accuracy 0.77 200 macro avg 0.77 0.77 0.76 200 weighted avg 0.77 0.77 0.76 200Considering that the dataset …

1 month назад @ machinelearningmastery.com
Learn Vectorized Thinking in Python Through Examples
Learn Vectorized Thinking in Python Through Examples Learn Vectorized Thinking in Python Through Examples

Share Post ShareIn this article, you will learn how to think in terms of vectorized operations using NumPy, replacing slow Python loops with efficient array-level computations.

append ( temp > 38.0 ) print ( alerts )Output:[False, True, False, True, False, True, False] 1 [ False , True , False , True , False , True , False ]Vectorized VersionWith NumPy, comparing an array directly creates the boolean mask automatically.

import numpy as np readings = np.array([34.1, 38.5, 37.2, 39.0, 36.8, 40.1, 35.5]) alerts = readings > 38.0 print(alerts) print("Alert readings:", readings[alerts]) 1 2 3 4 5 6 7 8 import numpy as np readings = np .

array ( [ 34.1 , 38.5 , 37.2 , 39.0 , 36.8 , 40.1 , 35.5 ] …

1 month назад @ machinelearningmastery.com
Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral
Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

How each of the three model families — Gemma 4, Llama 3, and Mistral — implements tool calling, including architectural and versioning differences.

This article compares how three widely used open-weight model families handle tool calling when run locally: Google DeepMind’s Gemma 4, Meta’s Llama 3, and Mistral AI’s Mistral.

Before the comparison, it helps to understand what tool calling is and why it matters for local deployments.

Mistral (Mistral AI)Mistral AI is a Paris-based startup founded in April 2023 by Arthur Mensch, formerly of Google DeepMind, and Guillaume Lample and Timothée Lacroix, formerly of Meta’s AI Research lab.

Tool Calling Implementation: How Each Model Approaches ItThe…

1 month назад @ machinelearningmastery.com
Integrating Agentic AI with Existing Machine Learning Pipelines
Integrating Agentic AI with Existing Machine Learning Pipelines Integrating Agentic AI with Existing Machine Learning Pipelines

IntroductionAgentic AI and machine learning pipelines are far from incompatible when it comes to building production-ready AI applications.

We will construct a lightweight, free, runnable Python pipeline that:Predicts customer churn based on a classical machine learning model built with scikit-learn.

get ( 'GROQ_API_KEY' )Step-by-Step GuideOnce the prerequisites are set up, we will start building the classical machine learning pipeline — for customer churn prediction — that will later be extended by incorporating agentic AI principles and tools.

uniform ( 10 , 150 , n_samples ) # Feature 2: Support tickets issued by customer (Poisson distribution, averaging 1.5 tickets) tickets = np .

colum…

1 month назад @ machinelearningmastery.com
How to Build a Robust RAG System with Minimal Resources
How to Build a Robust RAG System with Minimal Resources How to Build a Robust RAG System with Minimal Resources

A working RAG system spans document loading, chunking, embedding, storage, retrieval, prompting, and generation, and no short snippet represents that honestly.

For the FAISS and Hugging Face variant, see A Practical Guide to Building Local RAG Applications with LangChain.

Knowing When to Scale UpA small local system covers a lot of ground, but some problems need more.

See Building a Graph RAG System: A Step-by-Step Approach.

ConclusionA working RAG system needs a quantized local model, a compact embedding model, a file-based vector index, and careful chunking.

1 month, 1 week назад @ machinelearningmastery.com
Managing Small Context Windows in Language Models
Managing Small Context Windows in Language Models Managing Small Context Windows in Language Models

IntroductionTop-tier AI industries have become somewhat obsessed with language models capable of ingesting massive context windows, e.g.

Context Truncation: Sliding WindowThere is a consensus that sliding windows are arguably the most common and simplest strategy for managing shortened context windows in language models.

max_turns = max_turns self .

Small context windows may intuitively force a ruthless attitude toward the data to include in the context.

This article presented a number of strategies for effectively managing small context windows in LLMs to yield faster and cheaper solutions without compromising accuracy.

1 month, 1 week назад @ machinelearningmastery.com
7 Regression Tests Every AI Agent Should Pass Before Deploy
7 Regression Tests Every AI Agent Should Pass Before Deploy 7 Regression Tests Every AI Agent Should Pass Before Deploy

Share Post ShareIn this article, you will learn seven concrete regression tests for catching the orchestration-layer failure modes that matter most before deploying an AI agent to production.

These seven regression tests give you a concrete checklist for catching the failure modes that aggregate prompt evaluation will never surface.

When an agent misbehaves, the failure almost always lives in the state layer, not the model.

The regression test forces the same tool-call payload to arrive at the execution boundary three times.

What These Tests Won’t CatchThese seven tests cover structural failure modes at the system boundary.

1 month, 1 week назад @ machinelearningmastery.com
Understanding the Role of Latent Space in Machine Learning Models
Understanding the Role of Latent Space in Machine Learning Models Understanding the Role of Latent Space in Machine Learning Models

How the generative role of latent spaces enables the creation of entirely new data points through interpolation.

How the predictive role of latent spaces powers similarity-based applications such as recommender systems and RAG pipelines.

This article analyzes, illustrates, and categorizes the core functions and role of latent spaces in machine learning models.

array ( [ [ 1.1 , 2.2 , 3.3 ] , [ 1.0 , 2.1 , 3.1 ] , [ 8.1 , 9.2 , 9.9 ] ] ) # Compressing into a 2D Latent Space map pca = PCA ( n_components = 2 ) latent_space_map = pca .

The Predictive Role: Similarity and ForecastingHow does the AI behind recommender engines guess what video you want to watch next?

1 month, 2 weeks назад @ machinelearningmastery.com
Retrieval vs. Memory in Agentic AI Systems
Retrieval vs. Memory in Agentic AI Systems Retrieval vs. Memory in Agentic AI Systems

Share Post ShareIn this article, you will learn the conceptual and practical differences between retrieval and memory in agentic AI systems, and how to combine both effectively.

How retrieval pipelines and memory systems are each built, illustrated with a concrete worked example.

When designing this layer, teams can explore different agent memory strategies and agent memory frameworks depending on what they need to store and retrieve.

Combining Retrieval and Memory into an Effective SystemAn agent with retrieval but no memory re-derives the same conclusions every session and can’t personalize anything.

Retrieval brings in external information the agent needs at the moment, such as documenta…

1 month, 2 weeks назад @ machinelearningmastery.com
7 Async Patterns for Running Agents Concurrently in Python
7 Async Patterns for Running Agents Concurrently in Python 7 Async Patterns for Running Agents Concurrently in Python

Share Post ShareIn this article, you will learn seven async patterns for running AI agents concurrently in Python, what each pattern is suited for, and the production-level pitfalls to watch out for with each.

Topics we will cover include:Core async patterns such as fire and forget, scatter-gather, task groups, and producer-consumer queues, and when to reach for each one.

Here are seven async patterns for running agents concurrently, along with the production catches that come with each.

Fire and Forget (Detached Background Execution)You launch an agent task and move on without waiting for it to finish.

Supervised Task GroupsIntroduced in Python 3.11, task groups give you a structured versi…

1 month, 2 weeks назад @ machinelearningmastery.com
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026? Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?

Then we walked through the fastest way to get inference running locally in Run a Local AI Model in 15 Minutes: Your First Ollama Setup.

Spend enough time in the local AI ecosystem, though, and you’ll notice Ollama isn’t the only option competing for your hard drive.

Three tools dominate the local AI runtime landscape: Ollama, LM Studio, and llama.cpp.

Ollama (Via its dedicated, background-daemon CLI) ollama run llama3 .

There’s a well-worn progression in the local AI community that maps almost exactly to the three tools covered here: LM Studio → Ollama → llama.cpp.

2 months назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 1 month, 1 week назад
Eleven Years Later
Eleven Years Later Eleven Years Later

You also remember a hazy promise to yourself from last year, that you would write weirder posts and explore different forms of writing.

You remember how you used to post once a month, and now you’re posting once every two months.

You remember you used to check this more closely, and now don’t check it at all.

> Check time you’ve spent writing posts.

With some dismay, you noticed that most of your posts have been about AI in one way or another.

1 month, 1 week назад @ alexirpan.com
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

2 months, 1 week назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

4 months, 1 week назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

6 months, 2 weeks назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

8 months назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 7 months назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

7 months назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

7 months, 4 weeks назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 3 days, 7 hours назад
Introducing Gemini 3.8 Live with Live Avatar
Introducing Gemini 3.8 Live with Live Avatar Introducing Gemini 3.8 Live with Live Avatar

Building on the momentum of last week's Gemini 3.8 Live launch, today we are excited to introduce Gemini 3.8 Live with Live Avatar — bringing near real-time visual presence to our native live dialogue models.

By pairing near real-time video generation with speech, the Live Avatar feature creates an experience that listens, sees, and speaks with a dynamic visual persona.

With precise lip-syncing, natural expressions, and fluid turn-taking, Live Avatar enables enterprises to expand their virtual offerings more interactively.

Whether providing engaging customer service or delivering interactive walkthroughs, it transforms digital exchanges into richer, more accessible experiences.

Starting tod…

3 days, 7 hours назад @ blog.google
Advancing Private AI Compute with secure, server-side memory
Advancing Private AI Compute with secure, server-side memory Advancing Private AI Compute with secure, server-side memory

A technical update on our Private AI Compute architecture, which will enable persistent, cross-device AI memory with on-device privacy standards.

AI is becoming more capable and intuitive — remembering what matters, understanding the world around you, and acting at your direction.

Today, we are sharing how we will bring private, server-side memory to our Private AI Compute platform.

Bringing on-device privacy to cloud-scale memoryWith this new technical capability, a new persistent memory layer will be able to function like a secure digital vault in the cloud.

The diagram below shows how this update to Private AI Compute will work.

4 days, 8 hours назад @ deepmind.google
Gemini 3.8 text-to-speech says hello
Gemini 3.8 text-to-speech says hello Gemini 3.8 text-to-speech says hello

Get expressive high-quality speech generation built for global scaleGemini 3.8 Flash TTS delivers leading voice customization capabilities, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and also leading in accent modeling (60.8).

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable truly expressive performances without sacrificing reliability, also securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index.

The model shows major improvements on a wide range of use cases such as long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.

In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash and Flas…

4 days, 8 hours назад @ blog.google
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Gemini 3.8 Live : Built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding.

: Built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding.

Gemini 3.8 Live Extended Thinking: Built for high-complexity tasks, with increased intelligence and multi-step reasoning.

Gemini 3.8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena.

In addition to this performance, it remains highly cost-effective — providing developers and enterprises with a capable and efficient model built for scale.

1 week, 5 days назад @ blog.google
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

A major hurdle in understanding rare diseases is the daunting task of pinpointing the few causal variants hidden among thousands of candidates.

Moving beyond individual rare disease research, AlphaGenome Atlas can help uncover the genetic architecture of common traits in the general population.

Taking this approach even further, Hawkes used AlphaGenome Atlas to look at how hundreds of millions of non-coding variants in the UK Biobank might be linked to body mass index.

Accelerating genomic discoveryWith AlphaGenome Atlas we are creating new layers of information that will help further our understanding of the human genetic code.

AlphaGenome Atlas is powerful in isolation, but it also repres…

2 weeks, 5 days назад @ deepmind.google
Introducing WeatherNext 3, our most advanced and accurate global weather AI model
Introducing WeatherNext 3, our most advanced and accurate global weather AI model Introducing WeatherNext 3, our most advanced and accurate global weather AI model

Most AI weather models, including WeatherNext 2, are trained on data from numerical weather prediction (NWP) models.

By ingesting a mosaic of live, global geostationary satellite data, our new model gains a rich, continuously updating view of the atmosphere.

Traditional models struggle here because they train on representations of the atmosphere that lack detail and miss extreme local variations.

This allows us to make global forecasts on a 5-kilometer grid that account for regional details like topography.

Precipitation forecasting at breakthrough accuracyGlobal weather models notoriously struggle to accurately predict precipitation.

3 weeks, 3 days назад @ blog.google
Proactive cyber defense for governments and enterprises
Proactive cyber defense for governments and enterprises Proactive cyber defense for governments and enterprises

Defenders wanting to use advanced AI have faced a difficult dilemma: adopt enormous frontier models that could be expensive to deploy and difficult to control across enterprise codebases, or turn to smaller open-weight models that might struggle with complex vulnerability remediation and require teams to build their own tooling and infrastructure from scratch.

Today, we’re launching our Fairwind Program to bring the best of Google’s AI and cyber defense capabilities to a trusted group of Google Cloud customers, government agencies, and cybersecurity partners, to help them proactively solve cyber risks at scale.

As a first step, the Fairwind Program will give defenders access to powerful and…

3 weeks, 4 days назад @ blog.google
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7.

Gemini 3.8 introduces 2 variants:Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains.

It is available at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.

Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in v…

3 weeks, 4 days назад @ blog.google
Introducing agentic video understanding with Gemini
Introducing agentic video understanding with Gemini Introducing agentic video understanding with Gemini

Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite.

This new capability improves accuracy while dramatically reducing token usage and costs for video analysis.

Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.

BenchmarksUnlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adju…

3 weeks, 5 days назад @ blog.google
Gemini Omni 1.1 Flash lets you build with more control
Gemini Omni 1.1 Flash lets you build with more control Gemini Omni 1.1 Flash lets you build with more control

Today, we’re introducing Gemini Omni 1.1 Flash, a new suite of creative controls and generative video capabilities to support developers.

Gemini Omni brought real-world reasoning to generative creation, and today’s updates make Omni 1.1 production-ready for professional use via the Gemini API in Google AI Studio.

Whether you’re building generative video workflows, creative tools, or media editing software, these updates make generative video more controllable, faster to iterate on, and polished for real-world deployment.

With Omni 1.1, the model can now analyze up to 10 seconds of prior context — a leap from previous models that only referenced the final second.

The result is improved visua…

1 month назад @ blog.google
Piloting the world's first double-blind AI evaluations
Piloting the world's first double-blind AI evaluations Piloting the world's first double-blind AI evaluations

If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment.

To truly measure what they know, they must have no visibility of the test questions until it's time to take the exam.

That is the exact challenge the industry faces when evaluating advanced AI models.

Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing.

At Google, we assess our AI systems using a broad spectrum of evaluations throu…

1 month назад @ deepmind.google
Intelligent transcription with Gemini 3.5 Transcribe
Intelligent transcription with Gemini 3.5 Transcribe Intelligent transcription with Gemini 3.5 Transcribe

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions.

Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines.

Get more precise and i…

1 month назад @ blog.google
From Atari to EVE Online: Building on 15 Years of AI Research in Games
From Atari to EVE Online: Building on 15 Years of AI Research in Games From Atari to EVE Online: Building on 15 Years of AI Research in Games

Now, we’re partnering with game developers to prototype new gameplay experiences that push the frontiers of both gaming and AI.

Games as the engine of AI researchOur journey began when a small team trained a deep neural network to play Atari 2600 games directly from raw pixels.

For each game, AI enriched the playing experience.

For game developers, a truly general gaming agent would unlock AI capabilities that work with existing games — no modifications to the game code required.

To develop SIMA agents safely and responsibly, we've partnered with acclaimed game studios and we are building a growing portfolio of games for AI research.

1 month, 1 week назад @ deepmind.google
Introducing Gemini 3.7 Flash
Introducing Gemini 3.7 Flash Introducing Gemini 3.7 Flash

3.7 Flash shows strong gains over 3.6 Flash in coding tasks like debugging and issue resolution.

In web development, 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts.

It outperforms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowledge-dense fields like finance, law, and biosciences, 3.7 Flash delivers improved reasoning and accuracy.

It also surpasses 3.6 Flash in AutomationBench, demonstrating it can more effectively complete real-world business workflows (30.4% vs 17.0%).

1 month, 2 weeks назад @ blog.google
Putting sign language AI into users’ hands
Putting sign language AI into users’ hands Putting sign language AI into users’ hands

Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.

AI's ability to process spoken languages has advanced rapidly over recent decades, enabling automatic translation, dictation, and conversational interfaces that feel effortless to hearing users.

Yet this technological revolution has not reached the world’s more than 200 sign languages — and the estimated 70 million Deaf and hard of hearing people who use them.

With it, we are bringing sign language AI out of the lab and into consumer products for the first time: SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with…

1 month, 2 weeks назад @ deepmind.google
Google
последний пост 2 days, 8 hours назад
Best practices guide for customizing Gemini models via Reinforcement Learning (RL)
Best practices guide for customizing Gemini models via Reinforcement Learning (RL) Best practices guide for customizing Gemini models via Reinforcement Learning (RL)

So here at Google Cloud, we packaged it into a managed RL fine-tuning service (RLFT service) — you bring prompts and a reward function; we handle the infrastructure and the proprietary model internals.

Now, you can adapt Gemini with the service — teaching the model from a reward signal you define, rather than from a fixed set of labeled answers.

In this guide, we will walk through practical best practices for using RL fine-tuning service.

We'll start with a short tour of the RL training loop, how to decide if and when to use RL, and introduce how to get the most value from this approach.

RLFT adapts Gemini from a reward signal you define rather than labeled answers.

2 days, 8 hours назад @ cloud.google.com
Power your agents: Gemini 3.8 Live with Live Avatar is now generally available
Power your agents: Gemini 3.8 Live with Live Avatar is now generally available Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

Following our announcement of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking last week, we are thrilled to share that Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise.

Gemini 3.8 Live already delivers a native speech-to-speech foundation for fluid, responsive dialogue.

Together, you’ll have access to:Video avatars: Conversational video with the Live Avatar feature can generate video avatars with synchronized lip-syncing.

Though Gemini 3.8 Live Extended Thinking remains in private preview, Gemini 3.8 Live with Live Avatar is now available with US and EU endpoints, with provisioned throughput, enterprise compliance, and strict data governance.

Three demo…

3 days, 9 hours назад @ cloud.google.com
How growing Latin American midsize businesses are building in the AI era
How growing Latin American midsize businesses are building in the AI era How growing Latin American midsize businesses are building in the AI era

The number of Latin American-based small and medium businesses using Google Cloud AI tools has grown 8x year-over-year and the number of Brazil based small and medium businesses using Google Cloud AI tools has grown 9x year-over-year.

Help more people build their own AI agents to help them in their everyday jobs.

Google skills for organizations: Access thousands of free, on-demand AI courses and hands-on practice labs designed by experts at Google Cloud and Google DeepMind.

Get certified: Help your staff gain industry-recognized AI certificates through guided courses, expert mentoring, and skill badges.

By offering easy-to-use tools and free training — from everyday office apps in Workspace…

3 days, 9 hours назад @ cloud.google.com
The three things today's hottest startups are looking for in their AI stack
The three things today's hottest startups are looking for in their AI stack The three things today's hottest startups are looking for in their AI stack

Google Cloud has become the platform of choice for startups building AI.

And when startups use Gemini Enterprise, they also tend to use our “core cloud” services like Storage, BigQuery, or GKE.

Gemini models — as well as several of the third-party models available through Gemini Enterprise — are providing very strong price-performance for startups.

They will also use Gemini Enterprise Agent Platform and AI models on Google Cloud to help robotics process multimodal information like video.

To learn more about Google Cloud’s work with leading AI startups, or to get started building with us, visit here.

3 days, 11 hours назад @ cloud.google.com
Changing the game: How Google uses agentic AI to secure hundreds of millions of lines of code
Changing the game: How Google uses agentic AI to secure hundreds of millions of lines of code Changing the game: How Google uses agentic AI to secure hundreds of millions of lines of code

Our approach instead focuses on pre-submit scanning, where we evaluate each code check-in (across every layer of the stack) in real-time using AI agents.

Rather than relying on static decoupled documents, the threat models use live codebase metadata.

The scanning agent improves its accuracy further using a dependence call graph across packages and libraries to expand and refine its threat model context.

Making threat models part of our ongoing vulnerability scanning encourages developers to continuously update threats and dependencies, keeping the models up-to-date.

Use context wisely: Feed your agents your existing threat models.

1 week, 2 days назад @ cloud.google.com
Cloud CISO Perspectives: How Google monitors AI threats and advances AI defenses
Cloud CISO Perspectives: How Google monitors AI threats and advances AI defenses Cloud CISO Perspectives: How Google monitors AI threats and advances AI defenses

When you feed this rich, multi-dimensional internal observability into security models, AI defense becomes inherently faster and more accurate than AI offense.

If you use Google tools to facilitate an attack, you lose access to those tools.

Hardening our AI models and classifiers .

We operate a continuous feedback loop for our AI models.

Developing advanced defenses and threat models.

1 week, 4 days назад @ cloud.google.com
How Orange uses agents to make FinOps everyone's responsibility
How Orange uses agents to make FinOps everyone's responsibility How Orange uses agents to make FinOps everyone's responsibility

Orange calls these FinOps Clean Days.

Moving from awareness to action means finding ways to build FinOps accountability, and to get teams to genuinely care.

First, Cloud FinOps is a shared responsibility, with every stakeholder in a project involved in their own way.

And the only path to that shared responsibility runs through communication and a deliberate change effort.

“We insisted on the concept of shared responsibility across the organization for our FinOps practices,” Marini told us.

1 week, 4 days назад @ cloud.google.com
Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants
Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants

Open connectivity: Gemini Enterprise offers extensive connectivity beyond Google's ecosystem — extending to Microsoft 365, other third-party software, and internal enterprise data sources.

This open connectivity lets organizations adopt Gemini Enterprise alongside their existing infrastructure without costly system overhauls.

Simple economics: Gemini Enterprise offers a straightforward pricing model, with actions like chat and search included in the base SKU.

Accelerating our vision with the latest Gemini Enterprise advancementsOver the past months, we have accelerated Gemini Enterprise with significant product advancements:Tailored industry solutions: We introduced specialized solutions fo…

2 weeks, 4 days назад @ cloud.google.com
How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts
How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts

For the millions of fervent fans of the Indian Premiere League (IPL), being able to count on a flawless live streaming cricket experience is never up for debate.

For Airtel, producing league TV broadcasts with some of the world's most massive concurrent viewership, dropped packets and buffering are simply not options.

During the IPL 2026 season, Airtel partnered with Google Cloud to manage this digital delivery.

Delivering video under these concurrency spikes requires an edge architecture designed strictly around localization, paired with proactive operational monitoring.

Our goal for IPL 2026 was to deliver an uninterrupted, stadium-grade viewing experience to cricket fans across India, re…

2 weeks, 4 days назад @ cloud.google.com
How KDDI built Buffmee, a faster, reliable consumer RAG app
How KDDI built Buffmee, a faster, reliable consumer RAG app How KDDI built Buffmee, a faster, reliable consumer RAG app

To meet these performance targets, organizations need a systematic approach to AI evaluation and real-time bottleneck identification.

That is why we are sharing the automated evaluation framework and performance optimization techniques that helped KDDI successfully launch their application.

The results were inspiring: KDDI reduced total application response latency by 38%, successfully hitting their target response performance.

To solve this, the development team designed a systematic AI evaluation process using Gemini Enterprise Agent Platform Evaluation Service.

They ingested their extensive document corpus, constructed hundreds of automated evaluation tests, and built a comprehensive ben…

2 weeks, 5 days назад @ cloud.google.com
Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption
Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption

When Google's Finance Engineering team needed to modernize their legacy data layer, they chose Spanner, a globally distributed, strongly consistent, multi-model database with high availability capabilities.

Further, doing so without disruption would have required implementing multi-phase dual-write architectures across every DAO in our codebase.

To solve this, we took an alternative approach: We built an automated refactoring pipeline powered by Antigravity CLI in headless mode.

This helped us accelerate our migration velocity significantly while maintaining strict data parity in our staging environments as we prepare for production.

The challenge: Anatomy of a dual-write migrationWhen migr…

3 weeks, 2 days назад @ cloud.google.com
Getting started with Mantis, our open-source bug finding-and-fixing harness
Getting started with Mantis, our open-source bug finding-and-fixing harness Getting started with Mantis, our open-source bug finding-and-fixing harness

AI models have clearly proven their ability to discover and exploit vulnerabilities without much, if any, human assistance.

To help defenders gain the advantage with AI, we built the Mantis harness to automate the discovery, triage, reproduction, and patching of software vulnerabilities.

Available to all as an open-source framework, Mantis is part of Google’s internal approach to find and fix vulnerabilities at machine-speed.

Mantis distills decades of cybersecurity expertise across a wide spectrum of codebases, and is available on GitHub.

Here’s how you can get started using Mantis.

3 weeks, 4 days назад @ cloud.google.com
Reimagining work: How Pythian’s internal AI playbook delivers customer ROI
Reimagining work: How Pythian’s internal AI playbook delivers customer ROI Reimagining work: How Pythian’s internal AI playbook delivers customer ROI

Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian observed firsthand why so many enterprise AI initiatives stall out or fail.

To solve this, we engineered the Pythian AI Operating Model — a multifaceted, end-to-end framework designed to take enterprise AI from high-level strategy all the way into sustained production.

XOps (AI production management): While deploying an agent is 20% of the journey, maintaining accuracy in production is 80%.

Ready to build your AI operating model?

It also requires an end-to-end AI operating model.

1 month назад @ cloud.google.com
FinOps for the AI era: New flexible billing and cost controls for agents
FinOps for the AI era: New flexible billing and cost controls for agents FinOps for the AI era: New flexible billing and cost controls for agents

That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.

Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs)If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage.

Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.

To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing …

1 month назад @ cloud.google.com
Now introducing Gemini Enterprise for Legal
Now introducing Gemini Enterprise for Legal Now introducing Gemini Enterprise for Legal

Few professions are as exacting as the practice of law.

A team reviewing a contract or building a case works inside strictly privileged information, firm-specific playbooks, and a body of law that changes constantly.

For legal work, it is nowhere near sufficient.

Only in combination do they produce something a firm or a legal department can put into production and actually rely on.

Today we're bringing that to legal practice with Gemini Enterprise for Legal, part of our new suite of purpose-built industry solutions.

1 month назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 4 days, 7 hours назад
Offloaded inference for real-world physical AI robotics
Offloaded inference for real-world physical AI robotics Offloaded inference for real-world physical AI robotics

To better understand the systems implications of physical AI, we conducted the first systematic study of robotics workloads.

Offloading physical AI inference out of the robot improved its response time and accuracy, along with battery lifetime and cost.

We believe that our measurement study will inform the design of physical AI inference systems.

Physical AI Toolchain is an open-source, production-ready framework that integrates Microsoft Azure (opens in new tab) cloud services with NVIDIA’s (opens in new tab) physical AI stack, accelerating robotics and physical AI developers to automate and scale data curation, augmentation, and evaluation across perception, mobility, imitation learning, …

4 days, 7 hours назад @ microsoft.com
Improving synthesis prediction of small molecules at scale with RetroChimera
Improving synthesis prediction of small molecules at scale with RetroChimera Improving synthesis prediction of small molecules at scale with RetroChimera

At a glance We report on the recent publication of our retrosynthesis model RetroChimera in the journal Nature (opens in new tab) ..

Yet, progress is slowed by chemical synthesis—the time-consuming process of making new molecules from simpler building blocks in the lab.

In a paper recently published in the journal Nature (opens in new tab), we present RetroChimera (opens in new tab), a new framework for retrosynthesis prediction.

RetroChimera could enable chemists to assess more—and more complex—candidate molecules at large scale.

RetroChimera is available on GitHub (opens in new tab) (MIT license) and accessible via Microsoft Foundry (opens in new tab).

6 days, 8 hours назад @ microsoft.com
Called to serve: Tech, research, and positive impact with Chris White
Called to serve: Tech, research, and positive impact with Chris White

Lab Director Chris White has worked on research challenges with real-world implications—from new approaches to wartime data analysis to tools for combating human trafficking. He talks to program manager Weishung Liu about the influences that led to the work and more.

The post Called to serve: Tech, research, and positive impact with Chris White appeared first on Microsoft Research.

2 weeks, 5 days назад @ microsoft.com
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.

GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery.

To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.

Listen now Opens in a new tabEfficiency as an enabler of discoveryGigaPath and GigaTIME demonstrated what pathology foundation models can learn from whole slides and …

3 weeks, 6 days назад @ microsoft.com
Broadening access to Skala creates a faster path to predictive DFT
Broadening access to Skala creates a faster path to predictive DFT Broadening access to Skala creates a faster path to predictive DFT

and is being integrated into , , and , bringing next-generation DFT accuracy closer to the communities that rely on these codes every day.

Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and integrated in all relevant scientific and industrial workflows.

Alongside these integration efforts, we are introducing a living benchmark that tracks the computational performance of successive, increasingly optimized Skala releases.

Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and accessible across a broader range of relevant scien…

1 month, 1 week назад @ microsoft.com
MindTopo reveals VLMs’ spatial reasoning abilities
MindTopo reveals VLMs’ spatial reasoning abilities MindTopo reveals VLMs’ spatial reasoning abilities

At a glance MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as connectivity, enclosure, order, separation, and knots.

How MindTopo defines topological spaceMost spatial evaluations for multimodal models focus on Euclidean properties such as distance, direction, size, and relative position.

MindTopo pairs questions about static scenes with interactive tasks that require models to preserve or change the same topological relations.

MindTopo maps reasoning and planning tasks to continuity, separation, order, enclosure, and knots.

Closing that gap may require models that carry an explicit topological state, or…

1 month, 2 weeks назад @ microsoft.com
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands.

The CARE-X model.

The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses.

CARE-X: Toward clinically useful radiology AICARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model.

1 month, 2 weeks назад @ microsoft.com
Orchard: An open framework for scalable agentic AI
Orchard: An open framework for scalable agentic AI Orchard: An open framework for scalable agentic AI

At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.

To address this gap, we introduce Orchard (opens in new tab), an open-source framework for scalable agentic modeling.

Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification.

(opens in new tab) We are also releasing the training data and evaluation methods used to build them.

By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research.

1 month, 3 weeks назад @ microsoft.com
Echoverse: Deep, evolving environments for computer-use agents
Echoverse: Deep, evolving environments for computer-use agents Echoverse: Deep, evolving environments for computer-use agents

At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters).

Shallow worlds backfire; deep worlds transferA shallow world is the cheap option.

A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%.

First, more deep worlds for the closed domains public benchmarks cannot reach.

1 month, 4 weeks назад @ microsoft.com
EvoLib: Turning experience into evolving knowledge
EvoLib: Turning experience into evolving knowledge EvoLib: Turning experience into evolving knowledge

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive.

As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to co…

1 month, 4 weeks назад @ microsoft.com
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

2 months, 2 weeks назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

2 months, 2 weeks назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

2 months, 3 weeks назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

2 months, 4 weeks назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

3 months назад @ microsoft.com
MIT AI MIT AI
последний пост 2 days, 3 hours назад
MIT students gain a humanist lens on technical innovation in Tulsa, Oklahoma
MIT students gain a humanist lens on technical innovation in Tulsa, Oklahoma MIT students gain a humanist lens on technical innovation in Tulsa, Oklahoma

This is complemented by philosophical workshops to help students think more ambitiously and deliberately about what they can and should achieve in the near and long term.

Cherokee Nation Principal Chief Chuck Hoskin Jr. also delivered remarks at TASC, encouraging students as future leaders to take a public interest in technology.

“It's a very special relationship that the Cherokee Nation has with MIT.

While most MIT students won’t go on to full-time professional roles traditionally associated with social impact, PKG Center programming like Code.Tulsa helps students recognize they can promote the public interest no matter their career.

“Remembering that … I could use my MIT education to help…

2 days, 3 hours назад @ news.mit.edu
Estimating suicide risk from text
Estimating suicide risk from text Estimating suicide risk from text

It uses a custom-built list of words and phrases linked to 49 suicide risk factors, searching text for these and using them to estimate an individual’s risk.

It is already helping to clarify which suicide risk factors matter most in times of crisis.

Based on Crisis Text Line’s assessments, those conversations were grouped into three different risk levels: non-suicidal, suicidal ideation without imminent risk, and imminent risk.

“We wanted to know what type of symptoms predict the highest suicide risk,” Low says.

They turned to artificial intelligence to generate a preliminary list of words and phrases tied to established suicide risk factors, including factors associated with suicidal ideat…

3 days, 3 hours назад @ news.mit.edu
The promise and peril of using visual AI to study cities
The promise and peril of using visual AI to study cities The promise and peril of using visual AI to study cities

The scholars explore these topics in “How AI Sees the City: Urban Visual Intelligence,” published this month by Routledge.

“We can now scale up what he was doing, with visual AI, while also looking at many different dimension of cities.”There are extensive possibilities for applying visual AI to urban planning, ranging from emissions to traffic flow, safety, better imagery of street-level activity and sidewalks, and much more.

“The real promise of visual AI is not simply that computers can look at millions of images,” Zhang says.

By using images from 400,000 AirBnB listings across the world, one recent Senseable City study shows that, contrary to some claims, interior design styles are not …

3 days, 20 hours назад @ news.mit.edu
MIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow
MIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow MIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow

David Siegel SM ’86, PhD ’91, a computer scientist, entrepreneur, and philanthropist, will serve as the next MIT Innovation Fellow during the 2026-27 academic year.

He is a life member of the MIT Corporation, previously served on its Executive Committee, and co-chairs the External Advisory Committee for the MIT Schwarzman College of Computing.

Additionally, Siegel was an early champion of the MIT Quest for Intelligence, an Institute-wide initiative studying intelligence in brains and machines, recently renamed the MIT Siegel Family Quest for Intelligence.

Recognizing the critical resource gap between academic research labs and frontier AI, Siegel founded the nonprofit Open Athena in 2024.

M…

4 days, 8 hours назад @ news.mit.edu
Poitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research
Poitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research Poitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research

The commitment establishes 50 two-year fellowships through the Poitras Center for Psychiatric Disorders Research at MIT’s McGovern Institute for Brain Research.

They understood this future would require not just research, but a fundamental reimagining of how psychiatric research itself is conducted.

Expanding support for early career scientistsThe Poitras family’s latest gift invests directly in the PhD students and postdocs who will carry the field of psychiatric research forward.

Their projects span multiple areas in brain research and could reveal a suite of new ways to heal the mind.

The new gift extends the Poitras family’s support of mental health research at MIT to over $100 million.

5 days, 5 hours назад @ news.mit.edu
A new chapter for MIT Reads
A new chapter for MIT Reads A new chapter for MIT Reads

As it marks its 10-year anniversary, MIT Reads is being reimagined for the age of artificial intelligence.

Reading fiction prompts us to think about how what we build might change us,” says MIT Libraries Director Chris Bourg.

Most author events are open to the public and streamed online, and videos of MIT Reads talks have been viewed more than 5,000 times.

To mark this new era of MIT Reads, President Sally Kornbluth has selected the fall 2026 book “Exhalation,” by Ted Chiang.

I’m delighted that MIT Reads will give us the opportunity to explore his work together.”“MIT is not alone in grappling with these big questions around technology and its relationship with humanity,” adds Bourg.

1 week, 2 days назад @ news.mit.edu
New AI technique could make minimally invasive surgeries safer and more precise
New AI technique could make minimally invasive surgeries safer and more precise New AI technique could make minimally invasive surgeries safer and more precise

Researchers created a new technique that accurately and rapidly matches X-rays captured during surgery with a patient’s preoperative 3D medical scan.

This method could make it easier for clinicians to precisely pilot minimally invasive surgical tools, leading to faster and safer procedures.

Clinicians perform many minimally invasive surgeries using real-time X-rays to help them steer devices like catheters and endoscopes through tiny incisions.

To help localize surgical devices, clinicians may manually align X-rays with preoperative 3D medical images, such as CT scans or MRIs.

The model automatically matches one patient’s X-rays with 3D scans in a matter of seconds, and with sub-millimeter …

1 week, 4 days назад @ news.mit.edu
Measure by measure, studying society accurately
Measure by measure, studying society accurately Measure by measure, studying society accurately

Years ago, before the current artificial intelligence craze, he started studying what happens when AI tools are introduced into studies.

So, you really want to have both lenses.”Egami joined MIT’s Department of Political Science as an associate professor with tenure in 2025.

He is also a faculty affiliate of the Statistics and Data Science Center at the Institute for Data, Systems, and Society (IDSS).

“That was a case where the audience or market told me what I should really work on,” Egami says.

After earning his PhD from Princeton in 2020, he joined the faculty at Columbia University, moving to MIT five years later.

1 week, 4 days назад @ news.mit.edu
New method enables AI for safety-critical situations
New method enables AI for safety-critical situations New method enables AI for safety-critical situations

In these settings, a plausible answer is not enough: The output often must also satisfy nonnegotiable safety, physical, or task-specific requirements, known as hard constraints.

The researchers developed a method that helps generative models meet these strict requirements without sacrificing the quality of their outputs.

This adaptable, plug-and-play technique works at deployment time, so it can be applied to pretrained generative models without retraining them.

Freedom to explorePretrained generative AI models, such as diffusion models like Stable Diffusion and flow-matching models like FLUX, are now widely available.

“For constraint satisfaction, what ultimately matters is the model’s fin…

1 week, 6 days назад @ news.mit.edu
MIT spinout turns plastic waste into resilient building materials
MIT spinout turns plastic waste into resilient building materials MIT spinout turns plastic waste into resilient building materials

“Our mission is to convert waste plastic pollution into durable composites to build 1 billion homes,” says Atlas chair and co-founder A.J.

The conventional way of building homes involves cutting down trees, mining, refining, and a bunch of other dirty activities.

A key enabler for that scale is the company’s ability to recycle low-grade plastic into building components without water.

“This is key to democratizing recycling,” Perez says.

Through research at MIT, Perez has shown large composite trusses can be printed in under 13 minutes and support over 4,000 pounds, exceeding key building standards.

1 week, 6 days назад @ news.mit.edu
Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award
Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award

The Federal Laboratory Consortium (FLC) selected AI-GUIDE, a medical device developed by MIT Lincoln Laboratory and Massachusetts General Hospital (MGH), for its 2026 Excellence in Technology Transfer Award.

The transition to AutonomUS underscores how strong partnerships can carry a technology from development into real-world adoption," says Asha Rajagopal, Lincoln Laboratory's chief technology transfer officer.

These strong technology transfer collaborations are designed to streamline the transfer process, ensuring that lifesaving capabilities can be made available to military personnel and civilians as quickly as possible.

AI-GUIDE has previously been recognized with a Lincoln Laboratory …

2 weeks, 2 days назад @ news.mit.edu
MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines
MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines

Amin is also co-director of the Operations Research Center, which is jointly housed within the MIT Schwarzman College of Computing and MIT Sloan School of Management.

“This opportunity has been very timely because we are starting an AI and data science program in my department,” says Wenjin Zhou, assistant professor of computer science at UMass Lowell.

“We’ve already been thinking about: How do we teach our next generation of computer scientists within the area of AI?

What is usually missing is context: Opportunities for instructors and students to connect AI concepts to specific disciplines, problems, and ways of thinking.

Weijie Pang, an assistant professor of computer science at the Went…

2 weeks, 4 days назад @ news.mit.edu
From MIT to IBM, expediting AI and quantum deployment
From MIT to IBM, expediting AI and quantum deployment From MIT to IBM, expediting AI and quantum deployment

Here, the MIT-IBM Computing Research Lab served as a conduit for research relationship building and the flow of their expertise to industry applications.

“I started to work [on trustworthy AI] with IBM researchers from day 1 in my PhD, because it was funded by MIT-IBM,” says Ko.

After graduating in 2024, Ko joined IBM Research to continue her work on trustworthy AI as a research scientist.

During this time, Arunachalam focused on quantum machine learning and areas where quantum computing would be superior to classical computing, increasingly prioritizing provability grounded in theory to heuristics.

“One thing which I’ve been a huge fan of is exposing connections between different fields.” …

3 weeks, 4 days назад @ news.mit.edu
System helps humans predict when self-driving cars will make mistakes
System helps humans predict when self-driving cars will make mistakes System helps humans predict when self-driving cars will make mistakes

In road tests on a private track, CW-Net explanations helped safety drivers more accurately predict vehicle behavior; a larger simulation study with nonexpert users yielded similar results.

In the longer term, this technique could boost the safety and transparency of autonomous vehicles, while building appropriate trust in drivers and passengers.

The researchers trained CW-Net to predict concepts using a dataset of 130 million examples of scenes from self-driving cars, with multiple labeled concepts in each scene.

But CW-Net explanations revealed that the model wasn’t properly configured to detect the cyclist and chose a trajectory that would have caused a collision.

CW-Net explanations sig…

3 weeks, 4 days назад @ news.mit.edu
Walter Torous named executive director of MIT Center for Real Estate
Walter Torous named executive director of MIT Center for Real Estate Walter Torous named executive director of MIT Center for Real Estate

Walter Torous, senior lecturer in the MIT Department of Urban Studies and Planning (DUSP) and the MIT Sloan School of Management, and director of the Master of Science in Real Estate Development Program (MSRED), was recently named executive director of the MIT Center for Real Estate (CRE) — effective July 1, 2026.

“Demographic changes, as an aging population stays longer in their homes, are creating an imbalance in residential real estate markets.

“A lot of our alums have assumed important positions in the real estate industry around the world,” he says.

This research reflects the growing importance of AI and large language models to every aspect of real estate decision-making.

“It’s import…

3 weeks, 5 days назад @ news.mit.edu
Berkeley AI
последний пост 2 months назад
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map.

Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

With Apple Silicon in hundreds of millions of MacBooks and Mac Studios, MLX enables local AI inference without cloud costs.

Building an MLX backendTo bring K-Search to Apple Silicon, we first built a native MLX backend.

Evaluated on mamba-370m f16, M1 Max 64GB:Metric mlx-mamba (ours) mlx-lm (community) mamba.py Decode 152 tok/s 116 tok/s 40 tok/s Prefill L=512 5,751 tok/s 329 tok/s 1,089 tok/s Prefill L=1024 …

2 months назад @ bair.berkeley.edu
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon InteractionOverview of ABBEL compared to traditional recursive summarization.

Belief grading can be thought of as adding an auxiliary RL task, using heuristics designed to capture what makes a good belief as the reward.

With domain-knowledge belief grading, ABBEL approaches or exceeds FULL CTX in this setting; without belief grading, learning is slower.

Context compression methods generate dense representations which, while computationally efficient, sacrifice human-understandability (Kontonis et al., 2026, Eyuboglu et al., 2025, Gupta et al., 2025, Chevalier et al., 2023, Deng et al., 2025, Deng et al., 2025, Bulatov et al., 2022).…

2 months назад @ bair.berkeley.edu
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

2 months, 3 weeks назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

2 months, 4 weeks назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

4 months, 3 weeks назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

5 months, 1 week назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

6 months, 2 weeks назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

8 months, 2 weeks назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 2 days, 7 hours назад
Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput
Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput

On AWS, you can address these challenges using Amazon Elastic Kubernetes Service (Amazon EKS), Elastic Fabric Adapter (EFA), and DeepEP.

Challenges of large-scale RL trainingLarge-scale RL training presents three interrelated challenges:Balancing the competing resource demands of rollout generation and policy training.

EKS cluster setupThis architecture can be deployed on Amazon EKS by separating policy training, rollout generation, and supporting services across independently managed node groups.

To get started, review the architecture patterns described in this post and adapt them to your own MoE training workloads on Amazon EKS with EFA.

For more information, see the Amazon EKS documenta…

2 days, 7 hours назад @ aws.amazon.com
Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod
Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Amazon SageMaker HyperPod provides this infrastructure for large-scale machine learning (ML) workloads on Amazon Elastic Kubernetes Service (Amazon EKS).

PrerequisitesTo follow this walkthrough, you need:A SageMaker HyperPod cluster with Amazon EKS orchestration that has at least 3 ml.g7e.12xlarge instances and one ml.r5d.16xlarge instance.

Training topologyThe solution discussed here runs SkyRL on a HyperPod Ray cluster with three GPU worker nodes and a CPU head node.

The entire workflow runs on the resilient, self-healing compute infrastructure that Amazon SageMaker HyperPod provides.

To get started, see the Amazon SageMaker HyperPod documentation and the Ray on HyperPod getting started g…

2 days, 7 hours назад @ aws.amazon.com
NarrateAI: production-ready LLM quality assurance on Amazon Bedrock
NarrateAI: production-ready LLM quality assurance on Amazon Bedrock NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

The following snapshot shows the load test details for over 100 user requests for a streaming response from Amazon Bedrock through the application.

Absolute timing values are specific to this environment, but relative improvements and architectural patterns generalize broadly to streaming LLM deployment.

String similarity handles clear matches in approximately 1.7ms per metric while LLM verification addresses ambiguity in approximately 1,758ms per invocation.

Without this availability guarantee, the evaluation pipeline must handle both quality failures and generation failures simultaneously, substantially complicating its design.

For teams already streaming LLM output, adding producer-consu…

2 days, 7 hours назад @ aws.amazon.com
Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI
Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI

The deployment follows the standard Amazon SageMaker AI real-time inference pattern:The client sends an HTTP request to the Amazon SageMaker AI endpoint.

vLLM reserves GPU memory up front based on gpu_memory_utilization , the fraction of GPU memory a stage can claim.

An Amazon SageMaker endpoint exposes a single path, which this container routes to its completions handler by default.

Clean upTo avoid ongoing charges, delete the endpoint, endpoint configuration, and model when you are finished:predictor.delete_endpoint() # deletes endpoint and endpoint config predictor.delete_model()ConclusionThis post showed how to deploy the publicly available Qwen3-TTS-12Hz-1.7B-Base model to an Amazon Sa…

2 days, 7 hours назад @ aws.amazon.com
How Datacor built self-service rental analytics with Amazon Quick Sight
How Datacor built self-service rental analytics with Amazon Quick Sight How Datacor built self-service rental analytics with Amazon Quick Sight

Amazon Quick Sight is the embedded analytics capability within Amazon Quick.

Datacor built an automated cross-cloud data pipeline, embedded dashboards, and natural language querying powered by generative business intelligence (BI) in Amazon Quick Sight.

At the center of the experience is the generative BI capability in Amazon Quick Sight (Amazon Q in Quick Sight), a generative AI-powered assistant for business intelligence.

Self-service analytics adoptionBefore the solution launched, the vast majority of rental analytics requests went through IT.

Lessons learnedBuilding the TrackAbout Rental Analytics solution surfaced practical lessons relevant to any team integrating embedded analytics in…

2 days, 8 hours назад @ aws.amazon.com
Multi-Region training with Amazon SageMaker HyperPod and Qumulo
Multi-Region training with Amazon SageMaker HyperPod and Qumulo Multi-Region training with Amazon SageMaker HyperPod and Qumulo

With Amazon SageMaker HyperPod and Qumulo, you can place training compute in one AWS Region and keep your dataset in another.

sudo mount -t nfs4 -o rw,hard,nointr,proto=tcp :/ /mnt/qumulo aws s3 sync s3:///c4-tokenized/ /mnt/qumulo/data/c4/Step 4: Create your HyperPod clustersCreate Amazon SageMaker HyperPod clusters in both Regions with ml.p5.48xlarge instances.

Delete both Amazon SageMaker HyperPod clusters (aws sagemaker delete-cluster) in us-east-2 and us-west-2.

To get started, see Orchestrating SageMaker HyperPod clusters with Amazon EKS in the Amazon SageMaker AI Developer Guide.

For more on the components used here, see Introducing Amazon EKS support in Amazon SageMaker HyperPod and…

2 days, 8 hours назад @ aws.amazon.com
Speaker-labeled transcription with WhisperX on SageMaker AI
Speaker-labeled transcription with WhisperX on SageMaker AI Speaker-labeled transcription with WhisperX on SageMaker AI

It follows the standard Amazon SageMaker AI serving contract, so you deploy it like any other model.

Dimension Real-time endpoint Asynchronous endpoint Best for Short, interactive clips Long audio, high-volume batch Latency Synchronous, must finish < 60s Submit-and-poll.

Amazon SageMaker AI forwards each request to the WhisperX DLC on port 8080 .

For deeper GPU and inference visibility, turn on the Amazon SageMaker AI detailed metrics and Insights dashboard on CloudWatch.

ConclusionIn this post, we showed how to deploy the AWS WhisperX Deep Learning Container to Amazon SageMaker AI for word-level, speaker-labeled transcription.

3 days, 7 hours назад @ aws.amazon.com
Build a multi-account AI agent with AgentCore Gateway and MCP
Build a multi-account AI agent with AgentCore Gateway and MCP Build a multi-account AI agent with AgentCore Gateway and MCP

Along the way, you set up cross-account MCP integration, authentication with AgentCore Identity, a capability of Amazon Bedrock AgentCore, and Okta, fine-grained authorization with Policy in Amazon Bedrock AgentCore, and the governance controls that support production readiness.

Platform account — the agent control planeThe platform team owns the platform account, which runs the agent on AgentCore Runtime, a capability of Amazon Bedrock AgentCore.

Beyond aggregating MCP servers and acting as an Inference Gateway, AgentCore Gateway supports additional target types that make it a central integration point.

For configuration steps, see network connectivity patterns for AgentCore Runtime and se…

3 days, 7 hours назад @ aws.amazon.com
Aderant builds intelligent ticket triage with Amazon Nova
Aderant builds intelligent ticket triage with Amazon Nova Aderant builds intelligent ticket triage with Amazon Nova

In this post, we share how Aderant, a global provider of business management software for the legal industry, built an intelligent ticket triage system using Amazon Nova Lite through Amazon Bedrock.

Classify: Send the ticket and assembled context to Amazon Nova Lite through the Amazon Bedrock Converse API for structured classification and recommended next steps.

Send the ticket and assembled context to Amazon Nova Lite through the Amazon Bedrock Converse API for structured classification and recommended next steps.

Why Amazon NovaAderant chose Amazon Nova Lite after evaluating multiple foundation models through Amazon Bedrock using real ticket data.

Learn more about Amazon Bedrock and Amazo…

3 days, 7 hours назад @ aws.amazon.com
From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock
From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock

To turn that friction into instant answers, the 100-year-old Dutch retailer built a knowledge layer on Amazon Bedrock AgentCore.

Why MCP and Amazon Bedrock AgentCoreTwo goals shaped the solution, and they map cleanly onto the two technologies we chose.

Why Amazon Bedrock AgentCore: Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model.

AgentCore runtime and AgentCore memory are capabilities of Amazon Bedrock AgentCore.

For related reading, see Building and connecting a production-ready ecommerce MCP server using Amazon Bedrock AgentCore and Mistral AI Studio and Introducing Amazon Bedrock AgentCore Identity: Securing agentic AI a…

4 days, 5 hours назад @ aws.amazon.com
Agentic conversational video intelligence built on AWS
Agentic conversational video intelligence built on AWS Agentic conversational video intelligence built on AWS

For document processing, either Amazon Bedrock Data Automation (BDA) or Amazon Rekognition and Amazon Transcribe is required.

AWS Command Line Interface (AWS CLI) configured with AWS Identity and Access Management (IAM) permissions for the services listed earlier.

How agentic orchestration worksIn a conventional video analysis application, the developer defines a fixed processing pipeline: upload the video, run transcription, perform visual analysis, present results.

For current pricing, see Amazon Bedrock pricing, Amazon Transcribe pricing, and Amazon Rekognition pricing.

ConclusionIn this post, we showed the architecture and key patterns for building a conversational video intelligence so…

4 days, 5 hours назад @ aws.amazon.com
Use open weight models as your AI coding agent with Amazon Bedrock
Use open weight models as your AI coding agent with Amazon Bedrock Use open weight models as your AI coding agent with Amazon Bedrock

Open weight models on Amazon Bedrock now make these agents practical to run privately and cost-effectively.

Why open weight models for AI-assisted codingThe industry is shifting toward open weight models.

Why Amazon Bedrock as the backendAmazon Bedrock provides fully managed, serverless access to open weight models.

ConclusionWe showed how to configure OpenCode with open weight models on Amazon Bedrock to build a secure, flexible, pay-per-use AI coding workflow.

For more information about Amazon Bedrock security and compliance, refer to the Amazon Bedrock User Guide.

4 days, 5 hours назад @ aws.amazon.com
Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock
Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock

GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock, giving you more options to match intelligence and efficiency to each workload.

Today, GPT-6 Sol and GPT-6 Luna from OpenAI are generally available on Amazon Bedrock, running on an inference engine built for high performance, security and reliability at scale.

GPT-6 Sol and GPT-6 Luna support explicit prompt caching on Amazon Bedrock.

Amazon Bedrock provides that foundation for GPT-6 Sol and GPT-6 Luna through a high-performance inference engine built for security and reliability at scale.

Get startedYou can get started with GPT-6 Sol and GPT-6 Luna in the Amazon Bedrock console or programmatically through supported Amaz…

5 days, 5 hours назад @ aws.amazon.com
Claude Opus 5.5 is now available on AWS
Claude Opus 5.5 is now available on AWS Claude Opus 5.5 is now available on AWS

Today, we’re excited to announce the availability of Claude Opus 5.5 on Amazon Bedrock and Claude Platform on AWS, the first of the Claude 5.5 model family.

Claude Opus 5.5 is Anthropic’s most capable Opus model suitable for agentic coding, knowledge work, and long-running tasks.

What makes Claude Opus 5.5 differentAccording to Anthropic, Claude Opus 5.5 does more with fewer tokens than Claude Opus 5, and new pricing passes those gains straight to customers.

Claude Opus 5.5 is the first Opus model that comes with safety classifiers similar to Claude Fable 5.1 in biology, cyber security, and AI development.

Getting started with Claude Opus 5.5 on Amazon BedrockTo try Claude Opus 5.5, open th…

5 days, 6 hours назад @ aws.amazon.com
Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore
Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

To make these failures measurable, Strands Evals SDK and Amazon Bedrock AgentCore Evaluations, a capability of Amazon Bedrock AgentCore, add skill-focused evaluators:Skill Selection Accuracy determines whether each invoked skill was an appropriate choice for the task.

Additionally on Strands Evals, Skill Invoked is a deterministic check of whether a named skill has been loaded successfully.

Skill Selection Accuracy checks whether each invoked skill fits the task and whether the agent invoked the correct skill.

Skill Invoked Strands Evals Binary, deterministic Was this named skill successfully loaded?

To follow the AgentCore Evaluations section, you need:An agent hosted on Amazon Bedrock Age…

5 days, 6 hours назад @ aws.amazon.com
NVIDIA
последний пост 3 days, 9 hours назад
Efficient MoE Training for Biological Foundation Models
Efficient MoE Training for Biological Foundation Models Efficient MoE Training for Biological Foundation Models

NVIDIA Transformer Engine (TE) helps address these bottlenecks with optimized primitives for grouped expert computation, kernel fusion, and low-precision training.

As biological foundation models grow in parameter count and sequence length, these primitives can improve GPU efficiency while expanding model capacity.

Together, these capabilities provide a practical reference for efficiently training MoE-based biological foundation models.

The BioNeMo recipe uses TE to support FP8 and MXFP8 training, reducing memory use.

Try the Mixtral Native Transformer Engine recipe in BioNeMo Recipes and learn more about the optimized MoE kernels in the NVIDIA Transformer Engine documentation.

3 days, 9 hours назад @ developer.nvidia.com
How Open Science Can Help Researchers Prepare for the Next Pandemic
How Open Science Can Help Researchers Prepare for the Next Pandemic How Open Science Can Help Researchers Prepare for the Next Pandemic

“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA.

Traditional methods for determining protein structures — crystallizing proteins and shooting X-rays at them — can take years and cost thousands of dollars per structure.

For this project, the team systematically worked through the protein structures of viral families known to infect humans, from common-cold viruses to emerging threats like Mpox.

“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL-EBI.

The structures show what viral complexes may look …

3 days, 10 hours назад @ blogs.nvidia.com
Contain the Chaos: ‘CONTROL Resonant’ Launches on GeForce NOW
Contain the Chaos: ‘CONTROL Resonant’ Launches on GeForce NOW Contain the Chaos: ‘CONTROL Resonant’ Launches on GeForce NOW

Remedy Entertainment’s CONTROL Resonant brings Dylan Faden’s extraordinary abilities and a paranatural crisis to GeForce NOW at launch.

With the release comes the final days of the CONTROL Resonant Ultimate Membership Bundle.

Purchase a 12-month GeForce NOW Ultimate membership through Sunday, Sept. 27, and receive CONTROL Resonant at no additional cost.

Take Control on the CloudExplore a warped Manhattan on the brink of paranatural annihilation in Remedy Entertainment’s CONTROL Resonant.

The clock is ticking: the CONTROL Resonant Ultimate Membership Bundle ends Sunday, Sept. 27.

3 days, 11 hours назад @ blogs.nvidia.com
How SWE-Serve Exposes the Gap Between Local Tests and Live Serving
How SWE-Serve Exposes the Gap Between Local Tests and Live Serving How SWE-Serve Exposes the Gap Between Local Tests and Live Serving

SWE-Serve instead tests repository-scale changes across the inference-serving stack, including model enablement, decoding, caching, scheduling, serving APIs, and runtime performance.

What live serving tests catchSome failures only appear once a real server starts.

When the live serving tests are excluded, that rate rises to 69.4%.

In other words, 147 patches changed from fail to pass when the live serving tests were excluded.

The top score shows that many SWE-Serve tasks are within reach of current agents; the spread shows that performance is far from uniform.

4 days, 8 hours назад @ developer.nvidia.com
Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale
Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale

When Sakeena Fiza describes her work as a validation engineer at NVIDIA, she does so in terms more befitting a detective story than a world-class engineering lab.

“Validation engineers look in the shadows and shine a light into every corner,” Fiza said.

At NVIDIA, the systems Fiza and her colleagues in the data center systems engineering lab investigate are the engines of the AI era.

I get to be a firmware engineer when I want to be.”The failures she chases can be immense or microscopic.

Bring-up, she said, is “like the Avengers assembling”: architects, designers, software engineers, firmware engineers, validation engineers, all in the room, racing toward a working system.

4 days, 9 hours назад @ blogs.nvidia.com
At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia
At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia

At the event, NVIDIA and its partners are showcasing breakthrough AI advancements across the Southeast Asia region at large.

NVIDIA Nemotron Adoption Expands in SingaporeEnterprises in Singapore are adopting NVIDIA Nemotron for various use cases.

AI Singapore is expanding its SEA-LION model family to include the NVIDIA Nemotron open models and NVIDIA NeMo tools.

Get started building with NVIDIA Nemotron using skills and playbooks that help partners customize Nemotron open models for their languages and domains.

Stay up to date on agentic AI, NVIDIA Nemotron and more by subscribing to NVIDIA news, joining the community and following NVIDIA AI on LinkedIn, Instagram, X and Facebook.

4 days, 21 hours назад @ blogs.nvidia.com
NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development
NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development

The ROS open framework is a project from Open Robotics that helps humans build robots.

NVIDIA Isaac ROS brings NVIDIA accelerated computing, physical AI models and production-ready libraries to the nearly 1.3 million ROS users, helping developers build high-performance robotics applications using free, familiar, open source tools.

Accelerating the Open Source Robotics EcosystemThe robotics ecosystem is already extending this agentic approach to development workflows.

Seeed Studio is using NVIDIA Isaac ROS with reBot Arm, combining accelerated perception, spatial understanding and motion planning on NVIDIA Jetson Thor.

Ouster integrates its Stereolabs ZED stereo cameras with NVIDIA Isaac ROS…

5 days, 12 hours назад @ blogs.nvidia.com
NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories
NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories

To help builders make those decisions, NVIDIA is introducing NVIDIA DSX Ready, a qualification program for partner products and solutions that meet applicable NVIDIA DSX AI factory reference design requirements.

Qualified Building Blocks for Building AI FactoryThe NVIDIA DSX AI factory platform unifies AI factory design and operations across compute, networking, power, cooling, facilities and software.

DSX Ready connects that selection process to applicable NVIDIA DSX requirements.

DSX Ready makes that connection, bringing partner innovation into the infrastructure choices behind NVIDIA DSX AI factories.

Explore NVIDIA DSX Ready qualification categories and learn how to participate.

6 days, 6 hours назад @ blogs.nvidia.com
Why Deploying Physical AI at Scale Demands Safety at Every Layer
Why Deploying Physical AI at Scale Demands Safety at Every Layer Why Deploying Physical AI at Scale Demands Safety at Every Layer

Physical AI safety means proving that AI-driven machines — AVs, humanoid robots, industrial robots and more — behave safely when their decisions turn into physical action.

Across physical AI, manufacturers, regulators, insurers and workplace safety teams need evidence that hardware, software, AI behavior and operating environments can work together safely without human intervention.

NVIDIA Halos is the first and only full-stack safety system for physical AI, helping developers engineer safety across every layer of design, validation and deployment.

Across both AV and robotics, the NVIDIA Halos AI Systems Inspection Lab turns safety, cybersecurity and AI safety requirements into repeatable i…

6 days, 8 hours назад @ blogs.nvidia.com
From Enablement to Execution, Egypt’s AI Ecosystem Reaches Production Scale
From Enablement to Execution, Egypt’s AI Ecosystem Reaches Production Scale From Enablement to Execution, Egypt’s AI Ecosystem Reaches Production Scale

Today, Egypt’s AI builders gathered in the Grand Egyptian Museum for a reception that highlighted the nation’s rapidly growing AI ecosystem — spanning AI natives, developers, researchers, startups and enterprises — building applications across industries.

In addition, NVIDIA and A15 in July hosted an event connecting Egypt’s leading founders with NVIDIA’s global ecosystem.

AI Growth Across AfricaThe Egypt ecosystem event showed just a piece of Africa’s larger, growing AI ecosystem.

This means African developers have often relied on cloud-based compute hosted outside the continent for large-scale AI training.

The company is among a growing group of AI innovators using NVIDIA technology to bu…

6 days, 8 hours назад @ blogs.nvidia.com
Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2
Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2 Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2

With the AI data assimilation tools in NVIDIA Earth-2, you can process these observations more efficiently.

You can use Score-Based Data Assimilation (SDA) to incorporate observations into diffusion-based AI downscaling and forecasting models such as CorrDiff and StormCast.

Compute the global weather from observationsMost global weather forecasting pipelines are initialized with an estimate of the current weather derived through numerical data assimilation.

AI data assimilation lets you issue more accurate, timely forecasts by incorporating the observations that matter to your region or organization.

Visit the Earth2Studio user guide to get started with AI data assimilation and explore the …

6 days, 9 hours назад @ developer.nvidia.com
AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack
AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack

AI security is an engineering problem.

As AI becomes more capable, the industry must accelerate security engineering, broaden access to defensive tools and share what works faster.

Security depends on how those components work together — and AI agents extend that system.

Build Security Into How Agents OperateA security boundary has to hold even when an agent makes the wrong decision.

AI security is an engineering problem.

6 days, 9 hours назад @ blogs.nvidia.com
5 Companies Using NVIDIA AI for Clean Energy
5 Companies Using NVIDIA AI for Clean Energy 5 Companies Using NVIDIA AI for Clean Energy

Clean energy isn’t hard to come by, but the pace of large-scale adoption has historically been slow due to bottlenecks — including out-of-date infrastructure, elongated research and development timelines, and upfront cost barriers.

At New York Climate Week, NVIDIA is highlighting five companies pioneering clean energy projects with AI baked into their foundation, accelerating research-to-inception pipelines and ultimately helping build a low carbon grid.

ThinkLabs AI Drives Toward An Autonomous Grid With Digital TwinsThinkLabs AI is curating digital twins and agents — using the NVIDIA CUDA platform — to speed up interconnection timelines, seamlessly integrate clean energy sources and optimi…

6 days, 14 hours назад @ blogs.nvidia.com
Cute Critters Come to the Cloud: ‘Aniimo’ Launches on GeForce NOW
Cute Critters Come to the Cloud: ‘Aniimo’ Launches on GeForce NOW Cute Critters Come to the Cloud: ‘Aniimo’ Launches on GeForce NOW

A new creature-catching adventure is ready to stream from the cloud this week.

Pawprint Studio’s Aniimo arrives on GeForce NOW at launch, inviting gamers to explore the vibrant continent of Idyll across supported devices.

Gaijin Network’s fractured-multiverse action game Active Matter and Annapurna Interactive’s acclaimed space mystery Outer Wilds are also among 11 new titles joining the cloud this week.

Glide, dive and burrow across Idyll, then return to a personal RV to build a warm, interactive Homeland alongside Aniimo companions.

Even better, the cost of the day pass can be applied toward a first monthly membership — making it easy to level up with GeForce RTX-powered cloud gaming.

1 week, 3 days назад @ blogs.nvidia.com
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

The NVIDIA platform is purpose-built to optimize across all these, as highlighted by MLPerf Inference v6.1 results released today:NVIDIA Vera Rubin NVL72 system debuts with leading performance : In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72.

NVIDIA GB300 NVL72 scales with leading efficiency : A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline.

Vera Rubin NVL72 Makes MLPerf Inference Debut With Leading PerformanceNVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the ML…

1 week, 4 days назад @ blogs.nvidia.com
Facebook
последний пост 3 weeks, 4 days назад
An Organizational Second Brain: Building an AI That Learns From Experts
An Organizational Second Brain: Building an AI That Learns From Experts An Organizational Second Brain: Building an AI That Learns From Experts

Recipes reference knowledge files but contain no domain facts; knowledge files state positions but prescribe no procedures.

This means:Adding an organizational position means adding a knowledge file and updating a routing index.

Every expert correction moves through four phases:Diagnose expert feedback into actionable issues with their root cause.

The requirements for adopting this architecture are:A structured knowledge system with explicit file boundaries, cross-references, and a dependency graph (the organizational second brain for the domain).

Every improvement is a text edit that a domain expert can review in 30 seconds.

3 weeks, 4 days назад @ engineering.fb.com
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Introducing the Multi-Stage Sequence ModelTo address scaling efficiency, a multi-stage model has been developed that enables scaling of a transformer-based sequence model in a compute efficient manner.

Second Stage: Online Ranking ModelThe offline user model representations are complemented with online ranking models that use fresh user signals and ad candidate information for real time ranking.

A Predictable Scaling CurveLLM-Style Scaling LawWhen running on real-world ads traffic, the multi-stage sequence model demonstrates the emergence of predictable scaling laws for ads recommendations that are analogous to those observed in large language models.

The Impact of Multi-Stage Sequence Mode…

1 month, 3 weeks назад @ engineering.fb.com
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

We tackled these challenges through complementary compute efficiency and scaling efficiency innovations: Compute efficiency : Achieved through a customized recommendation kernel library — Jagged Flash Attention (JFA), Generalized Dot-Product Attention (GDPA), BlockAttention, etc.

The results: we doubled GEM’s E2E training efficiency to 20-25% MFU while scaling total training FLOPs 4x over the past 12 months.

We measure training efficiency through E2E MFU, which decomposes into two factors:E2E MFU = Local MFU (compute efficiency) × Scaling Ratio (scaling efficiency)These factors describe two related but distinct optimization problems.

Local MFU (compute efficiency) measures how well a single…

1 month, 3 weeks назад @ engineering.fb.com
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

2 months, 2 weeks назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

2 months, 2 weeks назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

2 months, 4 weeks назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

3 months назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

4 months назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

4 months, 2 weeks назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

5 months, 1 week назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

5 months, 2 weeks назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

5 months, 3 weeks назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

5 months, 4 weeks назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

6 months назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

6 months назад @ engineering.fb.com
Uber Engineering
последний пост None
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 6 months, 3 weeks назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

6 months, 3 weeks назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

9 months назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

9 months назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

9 months назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

9 months, 2 weeks назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 2 days, 10 hours назад
The Phone Number puzzle
The Phone Number puzzle The Phone Number puzzle

Part of a monthly series of puzzles: https://momath.org/mindbenders/

2 days, 10 hours назад @ youtube.com
The last IMO problem AI could not solve
The last IMO problem AI could not solve The last IMO problem AI could not solve

Full video: https://youtu.be/Nbwv5wHQoj0

1 week, 2 days назад @ youtube.com
The last IMO problem AI could not solve
The last IMO problem AI could not solve The last IMO problem AI could not solve

A beautiful puzzle that eluded AI, and the intuition it requires.

Check out our virtual career fair: https://3b1b.co/talent

See new videos early: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Evan Chen also wrote up nice solution notes for this problem, along with all the others on that year's test.

https://web.evanchen.cc/exams/IMO-2025-notes.pdf The channel Dedekind Cuts has a video about this solution:

https://youtu.be/fgXg9CdCDcs Timestamps:

0:00 - The one that AI missed

4:09 - Problem statement

6:45 - Finding the Optimal Construction

18:51 - A weak lower bound

23:45 - 3b1b Talent

24:41 - Proving the Con…

1 week, 2 days назад @ youtube.com
The jumping pegs puzzle
The jumping pegs puzzle The jumping pegs puzzle

Part of a series of monthly puzzles with MoMath.

1 month, 1 week назад @ youtube.com
The 64 sugar cubes puzzle
The 64 sugar cubes puzzle The 64 sugar cubes puzzle

See all monthly puzzles: https://momath.org/mindbenders/

2 months назад @ youtube.com
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

2 months, 1 week назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

3 months, 1 week назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

3 months, 2 weeks назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

3 months, 2 weeks назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

3 months, 3 weeks назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

4 months, 1 week назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

5 months, 2 weeks назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

6 months назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

6 months назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

6 months, 1 week назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 3 days, 15 hours назад
Claude Opus 5.5 AI: A Massive Leap Forward
Claude Opus 5.5 AI: A Massive Leap Forward Claude Opus 5.5 AI: A Massive Leap Forward

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5 Walking creatures video and paper:

https://www.youtube.com/watch?v=kQ2bqz3HPJE

https://www.goatstream.com/research/papers/SA2013/ Honey coiling video and paper:

https://www.youtube.com/watch?v=ZEjUqZU1hNQ

https://cs.uwaterloo.ca/~c2batty/papers/Larionov2017/Larionov2017.pdf

https://dl.acm.org/doi/10.1145/3072959.3073628

https://uwspace.uwaterloo.ca/items/ecb9fe94-1ef5-4ed0-a67e-b3d3c5a7534e

paper name: Variational Stokes: A Unified Pressure-Viscosity Solver for Accurate Viscous Liquids Sources:

https://x.com/ArtificialAnlys/status/2102438210798514391/…

3 days, 15 hours назад @ youtube.com
Jev Just Made AI 200x Faster…But There’s A Catch
Jev Just Made AI 200x Faster…But There’s A Catch Jev Just Made AI 200x Faster…But There’s A Catch

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 Jev:

https://typesafe.ai/blog/introducing-system-one-models-and-jev Full song: https://www.youtube.com/watch?v=4RtUJkjrKMI Previous papers, implementations, models:

https://huggingface.co/convaiinnovations/laya

https://arxiv.org/abs/2503.23303

https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

https://arxiv.org/abs/2510.01237

https://huggingface.co/blog/setfit

https://proceedings.iclr.cc/paper_files/paper/2025/hash/5503a7c69d48a2f86fc00b3dc09de686-Abstract-Conference.html Sources:

https://x.com/alexisg…

5 days, 14 hours назад @ youtube.com
DeepSeek Just Made AI Memory 4x Smaller!
DeepSeek Just Made AI Memory 4x Smaller! DeepSeek Just Made AI Memory 4x Smaller!

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The DeepSeek V4.1 Flash paper is available here:

https://www.deepseek.com/en/news/deepseek-v4-1-flash/ Sources:

https://x.com/flowith/status/2099446892174406055/video/1

https://x.com/loktar00/status/2097761803291726137

https://x.com/loktar00/status/2099171901746688079

https://x.com/loktar00/status/2097557594239897837

https://x.com/loktar00/status/2097342670620291430

https://x.com/RealFedeURU/status/2097804481068941679

https://x.com/ItsmeAjayKV/status/2099875889454739801

https://x.com/ItsmeAjayKV/status/2100434059675709683/video/1

https://x.com/DanielPPFW/status/2100016621251338732/video/1 🙏 We would like to…

1 week, 2 days назад @ youtube.com
Claude Is Now Leaving Invisible Fingerprints In Its Text
Claude Is Now Leaving Invisible Fingerprints In Its Text Claude Is Now Leaving Invisible Fingerprints In Its Text

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The papers and sources are available here:

https://proceedings.mlr.press/v202/kirchenbauer23a.html

https://www.nature.com/articles/s41586-024-08025-4

https://www.anthropic.com/news/claude-text-watermark 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 5 days назад @ youtube.com
Humans + AI Cracked An Impossible Math Problem
Humans + AI Cracked An Impossible Math Problem Humans + AI Cracked An Impossible Math Problem

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Navier-Stokes solution paper is available here:

https://openai.com/index/navier-stokes-solution/ My fluid simulations and papers:

https://users.cg.tuwien.ac.at/zsolnai/gfx/fluid_control_msc_thesis/

https://users.cg.tuwien.ac.at/zsolnai/gfx/real_time_fluid_control_eg/

The flow from simulation to reality: https://www.nature.com/articles/s41567-022-01788-5

All papers: https://users.cg.tuwien.ac.at/zsolnai/ Sources:

https://www.youtube.com/watch?v=EURkO98VnKc

https://www.youtube.com/watch?v=luOWq1Gdv8c 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges…

2 weeks, 3 days назад @ youtube.com
GPT-6 Astra - A Massive Leap Into The Future
GPT-6 Astra - A Massive Leap Into The Future GPT-6 Astra - A Massive Leap Into The Future

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers Links / sources:

Honey sim: search for "Variational Stokes: A Unified Pressure-Viscosity Solver for Accurate Viscous Liquids" here: https://cs.uwaterloo.ca/~c2batty/

https://x.com/mindblown_ai/status/2095661874037813298?s=20

https://x.com/aollivier82/status/2096226819401801896?s=46

https://x.com/sahilexec/status/2095688272269984016?s=46

https://x.com/petergostev/status/2095596176804307342?s=20

https://x.com/mattshumer_/status/2095609734845927525?s=20

https://x.com/mattshumer_/status/2095596175705399482?s=46

https://x.com/davis7/status/2095742249275699415?s=46

https://x.com/dimillian/status/2095596700815516004…

2 weeks, 5 days назад @ youtube.com
Claude Fable AI Is Much Stranger Than The Headlines Suggest
Claude Fable AI Is Much Stranger Than The Headlines Suggest Claude Fable AI Is Much Stranger Than The Headlines Suggest

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Claude Fable 5.1 paper is available here:

https://www.anthropic.com/claude-fable-and-mythos-5-1

https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20%26%20Claude%20Mythos%205.1%20System%20Card.pdf Sources:

https://x.com/alexalbert__/status/2094860187743986169?s=46

https://x.com/holytrinity/status/2094866061212459130?s=46

https://x.com/holytrinity/status/2094927984217985474?s=46

https://x.com/omedvibecodes/status/2094887840848965845?s=46

https://x.com/loktar00/status/2094951511742632168?s=46

https://x.com/maxt3chno/status/2094798704385380762?s=46

https://x.com/fab…

3 weeks, 3 days назад @ youtube.com
Powerful AI Is Becoming Almost Free
Powerful AI Is Becoming Almost Free Powerful AI Is Becoming Almost Free

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 GLM 5.3 Flash:

https://z.ai/blog/glm-5.3-flash Sources:

https://x.com/louszbd/status/2092694163104113016

https://x.com/semianalysis_/status/2092623833630998556

https://x.com/skalskip92/status/2092748209802154201

https://x.com/louszbd/status/2093047548550525165

https://x.com/atomic_chat_hq/status/2093433913238552712

https://x.com/holytrinity/status/2094093933584257334

https://x.com/KinasRemek/status/2090081611832295581/video/1

https://x.com/AiXsatoshi/status/2093679264013181119/video/1

https://x.com/AiXsatoshi/status/2093353322921263389/video/1

https://x.com/stevibe/status/2092655031040565252 🙏 We would like…

3 weeks, 5 days назад @ youtube.com
This Free AI Just Caught The Billion Dollar Giants
This Free AI Just Caught The Billion Dollar Giants This Free AI Just Caught The Billion Dollar Giants

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper and Qwen3.8-Flash-Next are available here:

https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf

https://qwen.ai/blog?id=qwen3.8-flash-next 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
DeepSeek’s AI Just Learned To Upgrade Itself
DeepSeek’s AI Just Learned To Upgrade Itself DeepSeek’s AI Just Learned To Upgrade Itself

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek Harness + paper are available here:

https://deepseek.com/harness/en/

https://github.com/cordiverse/paper 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
This Small AI Will Change Everything
This Small AI Will Change Everything This Small AI Will Change Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Qwen3.8-27b is available here:

https://huggingface.co/Qwen/Qwen3.8-27B Sources:

https://www.reddit.com/r/unsloth/comments/1vogva0/share_your_results_from_qwen3827b/

https://www.reddit.com/r/LocalLLaMA/comments/1voer8u/qwen_38_27b_aquarium_burst_sample_test/

https://x.com/KyleHessling1/status/2088327667733180637

https://www.reddit.com/r/LocalLLaMA/comments/1vqme4y/qwen3827b_q8_0_on_strix_halo_is_seriously/

https://forums.developer.nvidia.com/t/qwen3-8-27b-nvfp4-on-a-single-dgx-spark-up-to-1m-context-vllm-mtp-measurements/380244 🙏 We would like to thank our generous Patreon supporters who make Two Minute …

1 month назад @ youtube.com
DeepSeek Just Made Closed AI Look Ridiculous
DeepSeek Just Made Closed AI Look Ridiculous DeepSeek Just Made Closed AI Look Ridiculous

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers DeepSeek V4 Pro 0813:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 DSpark full episode: https://www.youtube.com/watch?v=1yBU41auQhw Sources:

https://x.com/cline/status/2087602193205694891

https://x.com/TypingMindApp/status/2088214938754167263

https://x.com/stevibe/status/2047546592530747561

https://x.com/voidfreud/status/2087701887327887543

https://x.com/exploraX_/status/2079197387860435360

https://x.com/AiHubMix/status/2087869896483057758

https://x.com/BruceBlue/status/2087833177155117304 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B …

1 month, 1 week назад @ youtube.com
Claude AI Failed 650 Times…Then Beat The Human Record
Claude AI Failed 650 Times…Then Beat The Human Record Claude AI Failed 650 Times…Then Beat The Human Record

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://www.anthropic.com/research/riemann-zeta Source:

https://www.scientificamerican.com/article/no-ai-didnt-just-solve-the-thorniest-problem-in-math/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 2 weeks назад @ youtube.com
OpenAI’s AI Escaped And It's Terrifying
OpenAI’s AI Escaped And It's Terrifying OpenAI’s AI Escaped And It's Terrifying

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 More reports are available here:

https://openai.com/index/hugging-face-model-evaluation-security-incident/

https://huggingface.co/blog/security-incident-july-2026

https://huggingface.co/blog/agent-intrusion-technical-timeline 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 2 weeks назад @ youtube.com
DeepMind's AI Trick Everyone Should Copy
DeepMind's AI Trick Everyone Should Copy DeepMind's AI Trick Everyone Should Copy

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Gemma4 paper and some more is available here:

https://arxiv.org/abs/2607.02770

https://x.com/googlegemma/status/2077449152062247219

https://x.com/UnslothAI/status/2078118183085731843 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 3 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 4 days, 12 hours назад
Как сделать аудиомодальную LLM 🎧
Как сделать аудиомодальную LLM 🎧 Как сделать аудиомодальную LLM 🎧

На конференции Data Fest 2026 в Белграде независимый исследователь Александр Николич рассказал практическую историю создания аудиоязыковой модели Borealis с бюджетом, сопоставимым со стоимостью MacBook. Полное видео уже на канале! #DataFest #DataFest2026 #Borealis #AudioLLM #LLM #AI #MachineLearning #DeepLearning #AudioAI #ML #IT

4 days, 12 hours назад @ youtube.com
Одна ошибка — и LLM ошиблась
Одна ошибка — и LLM ошиблась Одна ошибка — и LLM ошиблась

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

1 week, 2 days назад @ youtube.com
Сложнее задача — больше граблей
Сложнее задача — больше граблей Сложнее задача — больше граблей

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка Как безопасно

1 week, 6 days назад @ youtube.com
Онлайн-зал «Сеть» I Practical ML Conf 2026
Онлайн-зал «Сеть» I Practical ML Conf 2026 Онлайн-зал «Сеть» I Practical ML Conf 2026

Доклады онлайн-зала «Сеть» в рамках конференции Яндекса Practical ML Conf 2026. Полная программа доступна на сайте: https://pmlconf.yandex.ru/2026

2 weeks, 3 days назад @ youtube.com
Зал «Код» I Practical ML Conf 2026
Зал «Код» I Practical ML Conf 2026 Зал «Код» I Practical ML Conf 2026

Трансляция докладов зала «Код» в рамках конференции Яндекса Practical ML Conf 2026. Полная программа доступна на сайте: https://pmlconf.yandex.ru/2026

2 weeks, 3 days назад @ youtube.com
Зал «Данные» I Practical ML Conf 2026
Зал «Данные» I Practical ML Conf 2026 Зал «Данные» I Practical ML Conf 2026

Трансляция докладов зала «Данные» в рамках конференции Яндекса Practical ML Conf 2026. Полная программа доступна на сайте: https://pmlconf.yandex.ru/2026

2 weeks, 3 days назад @ youtube.com
Первое, что видит LLM 👀
Первое, что видит LLM 👀 Первое, что видит LLM 👀

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

2 weeks, 4 days назад @ youtube.com
Внутренние знания 🚫 Источники ✅
Внутренние знания 🚫 Источники ✅ Внутренние знания 🚫 Источники ✅

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

3 weeks, 2 days назад @ youtube.com
Инфраструктура как часть RL
Инфраструктура как часть RL Инфраструктура как часть RL

Приглашаем специалистов с опытом от 2 лет на Weekend Offer ML 12–13 сентября: https://clck.ru/3VTgGR Это один из наймовых ивентов Яндекса: вы сможете пройти все ключевые этапы онлайн и без долгих пауз.

3 weeks, 4 days назад @ youtube.com
Как Kimi K3 обучает уровни вычислительного бюджета
Как Kimi K3 обучает уровни вычислительного бюджета Как Kimi K3 обучает уровни вычислительного бюджета

"Приглашаем специалистов с опытом от 2 лет на Weekend Offer ML 12–13 сентября Это один из наймовых ивентов Яндекса: вы сможете пройти все ключевые этапы онлайн и без долгих пауз".

3 weeks, 5 days назад @ youtube.com
Почему избавиться от галлюцинаций недостаточно 😵‍💫
Почему избавиться от галлюцинаций недостаточно 😵‍💫 Почему избавиться от галлюцинаций недостаточно 😵‍💫

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

1 month назад @ youtube.com
Оптимизация LLM-инференса
Оптимизация LLM-инференса Оптимизация LLM-инференса

Доклад Андрея Бежина, руководителя службы ML-инфраструктуры в Яндекс R&D, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

1 month назад @ youtube.com
Новые способы и стандарты оценки качества моделей
Новые способы и стандарты оценки качества моделей Новые способы и стандарты оценки качества моделей

Доклад Ивана Дёгтева, руководителя аналитики Alice AI LLM в Яндекс R&D, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

1 month назад @ youtube.com
Тренды и вызовы в ризонинге
Тренды и вызовы в ризонинге Тренды и вызовы в ризонинге

Доклад Дмитрия Мокеева, руководителя группы качества претрейна Alice AI в Яндекс Поиске, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

1 month назад @ youtube.com
Tabular DL
Tabular DL Tabular DL

Доклад Артёма Бабенко, руководителя отдела в Yandex Research, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy

1 month назад @ youtube.com
ML Trainings ML Trainings
последний пост 17 часов назад
Капитанский мостик 27.09.2026: Юрген в Сакане | Китай против дистилляции | OpenAI сломал Австралию
Капитанский мостик 27.09.2026: Юрген в Сакане | Китай против дистилляции | OpenAI сломал Австралию Капитанский мостик 27.09.2026: Юрген в Сакане | Китай против дистилляции | OpenAI сломал Австралию

0:00:00 Начало

0:01:44 Юрген в Сакане

0:10:38 Биохакеры из Anthropic

0:22:04 Суверенная LLM от Яндекса

0:25:02 Китай против дистилляции

0:29:42 Anthropic не дает Opus

0:33:02 Греф догоняет Китай

0:37:39 OpenAI сломал Австралию

0:42:23 Китайские TPU

0:47:51 OpenAI и ментальное здоровье

0:56:58 Нереальный агент

1:01:18 DeepSeek учит на Huawei

1:04:35 Ученый-2

1:13:23 Уткомод для клода ИИ-саммари: Валентин Малых и Дмитрий Колодезев разбирают главные ИИ-события недели: Юрген Шмидхубер в Sakana, биохакеры из Anthropic и суверенная LLM от Яндекса. Затем геополитика: Китай выступает против дистилляции моделей, показывает собственные TPU, DeepSeek обучается на чипах Huawei, а Греф обещает догнать К…

17 часов назад @ youtube.com
Дмитрий Новичков | Генеративное распознавание документов в продакшене: VLM‑OCR, извлечение полей
Дмитрий Новичков | Генеративное распознавание документов в продакшене: VLM‑OCR, извлечение полей Дмитрий Новичков | Генеративное распознавание документов в продакшене: VLM‑OCR, извлечение полей

Спикер: Дмитрий Новичков Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Computer Vision

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 17 hours назад @ youtube.com
Егор Коновалов | Hacks and Defenses in Automatic Kernel Generation
Егор Коновалов | Hacks and Defenses in Automatic Kernel Generation Егор Коновалов | Hacks and Defenses in Automatic Kernel Generation

На Data Fest 2026 в Белграде Егор Коновалов, ML-инженер, разобрал хаки, которые находят LLM-агенты, когда генерируют GPU/TPU-код: от тривиального обхода numerical tolerance до изощрённых атак на timing-измерения и эксплуатации дыр в test harness. А ещё Егор показал, какие методы защиты реально работают, а какие создают ложное чувство безопасности. Telegram-канал Yandex for ML https://t.me/+owyCvdge8WIyNTUy ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost #datafest #DataFe…

2 days, 17 hours назад @ youtube.com
Алексей Быченков | Построение RAG-систем
Алексей Быченков | Построение RAG-систем Алексей Быченков | Построение RAG-систем

Спикер: Алексей Быченков, X5 Tech, начальник отдела ии-решений в ит-сопровождении Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data и ML в Retail от X5.tech

https://ods.ai/tracks/df26-ml-in-retail ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 8 hours назад @ youtube.com
Дмитрий Антипов | Agentic SDLC: как мы строим мультиагента для понимания данных
Дмитрий Антипов | Agentic SDLC: как мы строим мультиагента для понимания данных Дмитрий Антипов | Agentic SDLC: как мы строим мультиагента для понимания данных

Спикер: Дмитрий Антипов, руководитель разработки AI-продуктов, АБТ Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 8 hours назад @ youtube.com
Богдан Минко | Encoder vs LLM as-a-judge в Guardrails: тонкости архитектур и трейдоффы
Богдан Минко | Encoder vs LLM as-a-judge в Guardrails: тонкости архитектур и трейдоффы Богдан Минко | Encoder vs LLM as-a-judge в Guardrails: тонкости архитектур и трейдоффы

Спикер: Богдан Минко, HiveTrace, ML-Engineer Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Advanced LLMs ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 17 hours назад @ youtube.com
Вера Соболева | T-LoRA: дообучение диффузионной модели по одному изображению без переобучения
Вера Соболева | T-LoRA: дообучение диффузионной модели по одному изображению без переобучения Вера Соболева | T-LoRA: дообучение диффузионной модели по одному изображению без переобучения

Спикер: Вера Соболева, научный сотрудник группы "Контролируемый Генеративный ИИ" Лаборатории FusionBrain, Институт AIRI Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenCV https://ods.ai/tracks/df26-gencv

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 17 hours назад @ youtube.com
Кирилл Малахов | Diffusability латентного пространства: между Сциллой избыточности и Харибдой хаоса
Кирилл Малахов | Diffusability латентного пространства: между Сциллой избыточности и Харибдой хаоса Кирилл Малахов | Diffusability латентного пространства: между Сциллой избыточности и Харибдой хаоса

Спикер: Кирилл Малахов Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции GenCV

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 8 hours назад @ youtube.com
Александр Козлов | Корпоративный AI-агент в Норникеле
Александр Козлов | Корпоративный AI-агент в Норникеле Александр Козлов | Корпоративный AI-агент в Норникеле

Спикер: Александр Козлов, Норникель Data Fest 2026: https://ods.ai/events/datafest2026 ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 8 hours назад @ youtube.com
Эволюция агентов: от лени до сокращения слов
Эволюция агентов: от лени до сокращения слов Эволюция агентов: от лени до сокращения слов 4 days, 16 hours назад @ youtube.com
Самоулучшение ИИ: свобода и границы
Самоулучшение ИИ: свобода и границы Самоулучшение ИИ: свобода и границы 4 days, 16 hours назад @ youtube.com
Мозг мухи торгует криптой
Мозг мухи торгует криптой Мозг мухи торгует криптой 4 days, 16 hours назад @ youtube.com
Как идти в поход с агентами
Как идти в поход с агентами Как идти в поход с агентами 4 days, 16 hours назад @ youtube.com
Искусственный интеллект: что это и когда он появился
Искусственный интеллект: что это и когда он появился Искусственный интеллект: что это и когда он появился 4 days, 16 hours назад @ youtube.com
Капитаны рассказывают о пути обучения
Капитаны рассказывают о пути обучения Капитаны рассказывают о пути обучения 4 days, 16 hours назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 1 week, 3 days назад
#502 – Psychiatry, Insane Asylums, Mental Illness, ECT, Lobotomies, Freud & Jung
#502 – Psychiatry, Insane Asylums, Mental Illness, ECT, Lobotomies, Freud & Jung #502 – Psychiatry, Insane Asylums, Mental Illness, ECT, Lobotomies, Freud & Jung

Andrew Scull is a historian of psychiatry.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep502-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://wisprflow.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

Go to https://shopify.com/lexBetterHelp: Online therapy and counseling.

1 week, 3 days назад @ lexfridman.com
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux #501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux

DHH is the creator of Ruby on Rails, Omarchy Linux, CTO of 37signals, and a racecar driver.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep501-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://wisprflow.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://plaud.ai/lexHiggsfield AI: AI-based video generation, filmmaking, and creative studio.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:14) – Sponsors, Comments, and Reflections(08:56) – Programming with AI agents(24:14) – How software will change(33:30) – AI impact on open source(43:21) – Building Omarchy Linux d…

1 month назад @ lexfridman.com
#500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football
#500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football #500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football

Khabib Nurmagomedov is one of the greatest fighters of all time, who retired from the UFC undefeated with a perfect 29-0 record.

We did this conversation entirely in Russian.

Both language audio tracks (and subtitles) are available on YouTube.

We worked hard to make it enjoyable to listen to, by carefully dubbing the translation using voice-cloning, as we’ve done for previous foreign-language podcasts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep500-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

1 month, 2 weeks назад @ lexfridman.com
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee #499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee

Gary Gallagher is a historian of the American Civil War.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep499-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://plaud.ai/lexOUTLINE:(00:00) – Introduction(00:07) – Sponsors, Comments, and Reflections(08:36) – What caused the Civil War?

(18:33) – Slavery(46:07) – Lincoln(1:01:03) – Grant vs Lee(1:09:57) – Could the Civil War have been avoided?

(1:19:23) – The bloodiest war in US history(1:36:31) – How the Confederate Army could’ve won(1:57:05) – Key battles of the Civil War(2:20:07) – Best and Worst Presidents(2:34:06) – Robert E. Lee(2:53:40) – The…

2 months назад @ lexfridman.com
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

2 months, 4 weeks назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

4 months назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

4 months, 3 weeks назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

5 months, 3 weeks назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

6 months, 1 week назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

6 months, 2 weeks назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

7 months назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

7 months, 2 weeks назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

7 months, 4 weeks назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

8 months, 2 weeks назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

9 months назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 2 weeks, 5 days назад
Called to serve: Tech, research, and positive impact with Chris White
Called to serve: Tech, research, and positive impact with Chris White Called to serve: Tech, research, and positive impact with Chris White

Lab Director Chris White has worked on research challenges with real-world implications—from new approaches to wartime data analysis to tools for combating human trafficking. He talks to program manager Weishung Liu about the influences that led to the work and more.Show notes

2 weeks, 5 days назад @ microsoft.com
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

5 months, 1 week назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

5 months, 3 weeks назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

6 months, 1 week назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

6 months, 4 weeks назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 2 days назад
The Lived Informatics Model
The Lived Informatics Model The Lived Informatics Model

The data we collect about ourselves can tell us a lot—but only if the technology collecting it actually fits into our lives.

Daniel Epstein explores personal informatics, from fitness trackers and food journals to baby tracking and AI, and explains why abandoning a tracking tool doesn’t necessarily mean it failed.

GuestDaniel Epstein: Daniel Epstein is an Associate Professor in the Department of Informatics at the UC Irvine, with a courtesy appointment in Computer Science.

His work in Human-Computer Interaction (HCI) examines how personal tracking technology can acknowledge and account for the realities of everyday life.

He received his Ph.D. in CSE from the University of Washington in …

2 days назад @ dataskeptic.com
Recommender Systems Today and Tomorrow
Recommender Systems Today and Tomorrow Recommender Systems Today and Tomorrow

In the final episode of our Recommender Systems season, we explore the growing questions of trust, manipulation, privacy, fairness, sustainability, and user control.

From fake reviews and shilling attacks to explainable recommendations and user-selected algorithms, we look at what happens when recommender systems must answer not only for what they recommend, but for the consequences of those choices.

GuestsKyle Polich: Kyle is the founder of Data Skeptic, a popular podcast about artificial intelligence, machine learning, and data science.

Robin Burke: Professor Robin Burke conducts research in personalized recommender systems, a field he helped found and develop.

Professor Burke is the auth…

2 weeks, 4 days назад @ dataskeptic.com
Recommender Systems Optimization Goals
Recommender Systems Optimization Goals Recommender Systems Optimization Goals

In part two of the Data Skeptic Recommender Systems season finale, Kyle asks a deceptively difficult question: what should recommender systems actually optimize for?

Drawing on conversations from across the season, the episode explores engagement, filter bubbles, popularity bias, fairness, human curation, embeddings, and the growing role—and risks—of large language models in shaping what gets recommended to us.

3 weeks, 5 days назад @ dataskeptic.com
Recommender Systems Origin Story
Recommender Systems Origin Story Recommender Systems Origin Story

Where did recommender systems come from, and how do we know when they're actually working? In part one of Data Skeptic's three-part Recommender Systems finale, Kyle traces the field from collaborative filtering and the Netflix Prize to matrix factorization and modern approaches, while exploring why accuracy alone can't capture what makes a recommendation useful, surprising, or meaningful.

1 month, 1 week назад @ dataskeptic.com
Social Choice for Fair Recommendations
Social Choice for Fair Recommendations Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social choice theory can help balance the needs of users, creators, and society. The conversation explores the future of recommendation algorithms and why fairness is a far more complex challenge than it first appears.

2 months назад @ dataskeptic.com
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

2 months, 3 weeks назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

3 months назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

3 months, 1 week назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

4 months, 4 weeks назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

5 months назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

6 months назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 6 months, 3 weeks назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

7 months назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

7 months, 1 week назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

7 months, 3 weeks назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 2 days, 13 hours назад
1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav Pathak
1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav Pathak 1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav Pathak

During their #sponsored discussion, Senior Vice President Product Management AI and Metadata at Salesforce, Gaurav Pathak talks to Jon Krohn about why AI agents need well-labeled, high-quality data to deliver reliable answers in the enterprise. Listen to the episode to hear Gaurav Pathak talk about the difference between a “data brawl” and “garbage in, gospel out”, who the “sin eaters” of enterprise AI are and the three skills that matter most for AI engineers today! Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1030⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.⁠⁠⁠ In this episode …

2 days, 13 hours назад @ podtrac.com
1029: How AI Brought a Podcast Back From the Dead, with Linear Digressions’ Katie Malone
1029: How AI Brought a Podcast Back From the Dead, with Linear Digressions’ Katie Malone 1029: How AI Brought a Podcast Back From the Dead, with Linear Digressions’ Katie Malone

In Episode #1029, Dr. Katie Malone (Host of Linear Digressions) joins Jon Krohn to explain how AI brought her podcast back from the dead. After nearly 300 episodes, Katie shut down Linear Digressions due to burnout, but better tools helped her relaunch it six years later. Along the way she has taught machine learning at Udacity and the University of Chicago and led the development of agentic AI platforms inside a company of tens of thousands of people. In this episode, she argues that people management and agent management are the same skill in different clothing, works through what AI slop and process slop are doing to organisations, describes the agent that now produces her show, and take…

5 days, 13 hours назад @ podtrac.com
1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell
1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell 1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell

In Episode #1028, Anton McGonnell (VP of Product at SambaNova) joins Jon Krohn to explain why the chips running most AI inference today were never designed for the job. Agentic AI has changed the computational profile of inference, with much larger inputs and far heavier caches feeding the token generation that follows, and that shift has exposed where GPU architecture struggles. SambaNova has raised over $2 billion to build an alternative, the reconfigurable dataflow unit, which lays a whole model out spatially across the chip rather than executing it kernel by kernel. In this episode, Anton discusses why the speed that matters is payback, and how speed and concurrency are what turn a fixe…

1 week, 2 days назад @ podtrac.com
1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala
1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala 1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala

In Episode #1027, Dr. Dilani Kahawala (Co-Founder and CEO of Anna) joins Jon Krohn to explain what it takes to build an always-on AI assistant that busy parents will trust with their inboxes. Anna watches the email, school apps, WhatsApp messages and calendars flowing into a family's life and surfaces what matters, over text and voice, with barely any app to speak of. Dilani came to it by way of a Harvard physics PhD, McKinsey, and a decade of product leadership at Etsy, Meta and Atlassian, and says she has had to throw away most of what that decade taught her about how products get built. In this episode, she lays out the three hardest problems in building Anna, why the eval loop is the he…

1 week, 5 days назад @ podtrac.com
1026: OpenAI’s GPT-6 Astra
1026: OpenAI’s GPT-6 Astra 1026: OpenAI’s GPT-6 Astra

In Episode #1026, Jon Krohn breaks down GPT-6 Astra, OpenAI’s new flagship that its president has floated as a possible marker of AGI. Jon covers what the model is, what it costs, its state-of-the-art results across computer use, coding, abstract reasoning and science and the safety story, which for this release is unusually intertwined with capability. He weighs the AGI claim against Anthropic’s Fable 5.1 and lands, as ever, in a measured middle. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1026⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this ep…

2 weeks, 2 days назад @ podtrac.com
1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano
1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano 1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano

In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins Jon Krohn to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space around them, and the embedding of "bank" visibly curves toward "river" as it travels through the layers of the network. In this episode, he recreates Eddington’s 1919 eclipse experiment inside a transformer, draws the line between an LLM workflow and an actual agent, explains why agent evaluation is a step harder than evaluating an essay, and gives the cleanest account of GRPO you will hear. Additional materials: ⁠⁠⁠⁠…

2 weeks, 5 days назад @ podtrac.com
1024: In Case You Missed It in August 2026
1024: In Case You Missed It in August 2026 1024: In Case You Missed It in August 2026

In ICYMI Episode #1024, Jon Krohn tracks the gap between AI investment and AI return, from the technology side to the people side. Hear from Pete Johnson, Jerry Yurchisin, Priyanka Vergadia and Tristan Handy, discussing why four out of five organizations have the structures for AI success in place while only one in five sees the returns, which decisions should never be handed to a language model however confident it sounds, how to structure Claude skills so that your output stops being slop and why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1024⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ …

3 weeks, 2 days назад @ podtrac.com
1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan
1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan 1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan

In Episode #1023, Aishwarya Srinivasan (Co-Founder of The Gen Academy) joins Jon Krohn to work out where a competitive moat comes from once anything you can build in ten minutes, somebody else can build in ten minutes too. Ash came to teaching through Illuminate AI, the mentorship community she started in 2020, and now trains senior engineers and leaders to ship agentic AI in production; she is blunt that vibe coding lowers the floor without touching the engineering judgment that production demands. In this episode, she explains what a whole-system eval covers that a model eval misses, traces reinforcement learning from the algorithm she patented at IBM to its resurgence in agentic fine tun…

3 weeks, 5 days назад @ podtrac.com
1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents
1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents 1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents

In Episode #1022, Jon Krohn tackles the art of steering AI agents, deciding where your instructions should live so they get followed reliably without bloating every request. A sequel to Episode #1020 (where model size and effort set an agent’s horsepower), this one is about direction: the seven ways to deliver instructions, why a hook beats a prompt, the industry-wide agents.md standard, and three practical takeaways you can apply whatever your stack. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1022⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this …

1 month назад @ podtrac.com
1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy
1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy 1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy

In Episode #1021, Tristan Handy (Founder and CEO of dbt Labs) joins Jon Krohn to explain how a study of about a hundred companies in 2016 became analytics engineering, and then became a tool that over a hundred thousand data teams rely on. Tristan coined the term, chose SQL when Spark was the fashionable answer, and spent a decade turning down acquisition offers because none of them were good for the people using dbt. He is now merging dbt Labs with Fivetran and taking on the presidency of the combined company, the first deal he says cleared that bar. In this episode, Tristan walks through what dbt does to your raw data, argues that the semantic layer matters more once analytics agents are …

1 month назад @ podtrac.com
1020: How to Choose Model Size and Effort Level: The Two Critical Dials
1020: How to Choose Model Size and Effort Level: The Two Critical Dials 1020: How to Choose Model Size and Effort Level: The Two Critical Dials

In Episode #1020, Jon Krohn unpacks the two dials that increasingly decide what you get out of a large language model: which model size you pick and how much effort you tell it to spend. Using a July Anthropic blog post by Claude Code’s Lydia Holly as a jumping-off point, with guidance that generalizes to any model family, Jon explains what each setting actually does under the hood. Model size swaps which frozen weights handle your request (roughly, how capable), while effort sets how thorough and certain the model must be before calling a task done, not a simple “thinking-time slider.” He offers a clean diagnostic for when to raise effort versus move to a bigger model, shows why cheaper-pe…

1 month, 1 week назад @ podtrac.com
1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia)
1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia) 1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia)

In Episode #1019, Priyanka Vergadia (founder of The Cloud Girl, former Senior Director of AI Transformation at Microsoft and Head of North America Developer Relations at Google) joins Jon Krohn to explain why almost every company has bought AI tools and almost none of them are seeing a return. Her fix is a budget split that will make any CFO wince: seven dollars on training employees for every dollar spent on the tools themselves. Having spent a decade turning dense cloud and AI concepts into sketches that a quarter-million developers actually remember, and having carried GitHub Copilot into Fortune 100 boardrooms, she has watched the gap between tool purchase and real production use up clo…

1 month, 1 week назад @ podtrac.com
1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs
1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs 1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs

In Episode #1018, Jon Krohn breaks down Qwen3.8-Max, Alibaba’s enormous new flagship, a 2.4-trillion-parameter mixture-of-experts model that, if its promised weights ship, becomes the largest open-weight release in history. Landing just weeks after Moonshot’s Kimi K3, it extends the price war and the open-weight surge Jon covered in Episode #1012. Alibaba positions it as second only to Anthropic’s Claude Fable 5 / Mythos 5 and independent signals land in a similar neighborhood. Jon walks through its capabilities and multi-day agentic demos, its aggressive pricing ($2 in / $6 out per million tokens, with cached input eight times cheaper), and the question he gets asked most: are Chinese mode…

1 month, 2 weeks назад @ podtrac.com
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson 1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson

In Episode #1017, Pete Johnson (Field CTO of AI at MongoDB) joins Jon Krohn to explain why four out of five organizations have AI steering committees and success metrics, yet only one in five sees a return on the investment. Having made nineteen stops across six countries this year advising more than a hundred companies on their AI strategies, Pete has an unusually wide view of what is actually working in production. In this episode, he traces the history of SQL and denormalization, unpacks why the embedding model is the most underrated choice in a RAG pipeline, explains Matryoshka embeddings and lays out what better agentic memory looks like. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

1 month, 2 weeks назад @ podtrac.com
1016: In Case You Missed It in July 2026
1016: In Case You Missed It in July 2026 1016: In Case You Missed It in July 2026

In this month's episode of ICYMI, Jon Krohn traces a line from algorithmic harm to the human skills that still hold their value. Hear from Dr. Cathy O'Neil, Ben Todd, Steve Mock, and Dr. Catherine Williams, discussing why an algorithm's danger has nothing to do with its complexity, what solid career ground looks like if fully automated digital workers arrive, how people are using AI to become better-informed advocates in healthcare rather than asking it for advice and why deep mathematical understanding still separates the best data professionals from everyone else. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1016⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperD…

1 month, 3 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 2 months, 1 week назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 1 week назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 1 week назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

4 months, 1 week назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

4 months, 1 week назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

4 months, 1 week назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

4 months, 1 week назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

4 months, 1 week назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

5 months, 1 week назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

5 months, 1 week назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

5 months, 1 week назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

5 months, 1 week назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

6 months, 3 weeks назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

6 months, 3 weeks назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

7 months, 3 weeks назад @ datascienceathome.com