Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 8 часов назад
AAAI 2027 Review: No code submission? [D]
AAAI 2027 Review: No code submission? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

8 часов назад @ reddit.com
HyperSAE: Decoupled Poincaré Geometry for Sparse Autoencoders -- 9.8% MSE reduction, 0.2% dead latents on Gemma-2-2B [P]
HyperSAE: Decoupled Poincaré Geometry for Sparse Autoencoders -- 9.8% MSE reduction, 0.2% dead latents on Gemma-2-2B [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

8 часов назад @ reddit.com
We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P]
We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P] We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

10 часов назад @ reddit.com
Prospects of Finding a ML Engineering Job [D]
Prospects of Finding a ML Engineering Job [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

14 часов назад @ reddit.com
Planning/RL for a stochastic single-player merge puzzle: afterstates, previewed chance events, and long-horizon throughput [D]
Planning/RL for a stochastic single-player merge puzzle: afterstates, previewed chance events, and long-horizon throughput [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

15 часов назад @ reddit.com
fru - Fast Random Forest Implementation [P]
fru - Fast Random Forest Implementation [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 9 hours назад @ reddit.com
Transformers are famously bad at arithmetic, so I set one's weights by hand (no training) and it multiplies with 100% accuracy [P]
Transformers are famously bad at arithmetic, so I set one's weights by hand (no training) and it multiplies with 100% accuracy [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 9 hours назад @ reddit.com
How to file a complaint about a published CVPR paper? [R]
How to file a complaint about a published CVPR paper? [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 12 hours назад @ reddit.com
Semi Edge Inference Idea [D]
Semi Edge Inference Idea [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 16 hours назад @ reddit.com
Comparing embedding models with synthetic query probing [R]
Comparing embedding models with synthetic query probing [R] Comparing embedding models with synthetic query probing [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 16 hours назад @ reddit.com
3 Collapsing models [R]
3 Collapsing models [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 17 hours назад @ reddit.com
A Mechanistic Explanation of Prompt Injection (and why you should study roles) [R]
A Mechanistic Explanation of Prompt Injection (and why you should study roles) [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 9 hours назад @ reddit.com
I never understood positional encoding until I read this article. [D]
I never understood positional encoding until I read this article. [D] I never understood positional encoding until I read this article. [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 10 hours назад @ reddit.com
Non-Physical Intelligence Has A Ceiling [D]
Non-Physical Intelligence Has A Ceiling [D] Non-Physical Intelligence Has A Ceiling [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 11 hours назад @ reddit.com
ECCV workshop, camera ready instructions? [D]
ECCV workshop, camera ready instructions? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 12 hours назад @ reddit.com
Towards Data Science
последний пост 10 часов назад
Stop Calling the First Significant Day a Win
Stop Calling the First Significant Day a Win Stop Calling the First Significant Day a Win

But if the test was checked every day and stopped as soon as p dropped below 0.05, the false-positive rate went to 27.7 percent.

If the result looks good on day four, or day eight, or day twelve, the pressure to stop becomes very real.

If the stopping rule is not valid, the p-value at the stopping day does not mean what the team thinks it means.

In the simulation:How often you look False-positive rate 1 look, end only 5.0% 2 looks 8.3% 5 looks 14.0% 10 looks 19.1% Daily, 30 looks 27.7%The intuition is simple.

Calibrated here, it used a z cutoff of 2.73 rather than the usual 1.96, and held the false-positive rate at 4.9 percent under daily monitoring.

10 часов назад @ towardsdatascience.com
Should AI Developers Make the Switch from Polars to Pandas?
Should AI Developers Make the Switch from Polars to Pandas? Should AI Developers Make the Switch from Polars to Pandas?

Many developers have adopted Polars as a faster option than Pandas.

See, Pandas and Polars were based on different philosophies, and it is much more valuable to understand those philosophies than to make a decision based on benchmark figures.

But traditional Pandas operations normally run on a single core.

The data used by Polars is stored in the Apache Arrow columnar format.

Analysts develop their ideas using Pandas, while production pipelines are increasingly turning to Polars in order to process larger datasets more efficiently.

12 часов назад @ towardsdatascience.com
The Budget Split That Explains Itself
The Budget Split That Explains Itself The Budget Split That Explains Itself

A plain LP will dump the whole budget on the single best channel and starve the rest.

BRACKETS = [(0.25, 1.00), (0.35, 0.65), (0.40, 0.35)] # (budget fraction, marginal yield)The first slice of a strong channel is worth a lot.

No binary variables, so the familiar LP shadow prices remain available.

Adding decreasing marginal yields lets the same LP model diversify without introducing binary decisions.

That holds anywhere a fixed pool gets split under rules that bite, whether the pool is a budget, a team, compute, or shelf space.

13 часов назад @ towardsdatascience.com
Can a Local LLM Run My AI Assistant?
Can a Local LLM Run My AI Assistant? Can a Local LLM Run My AI Assistant?

TL;DR — I replayed 27 real tasks from my own AI agent against two local models, one hardware upgrade apart, scoring both against the same frozen Claude baseline.

Only the local model is re-run.

The local model reasons over the same real inbox and calendar content Claude saw, not invented placeholder text.

per task vs Claude Claude $0.763106 — qwen3-coder:30b $0.00014792 5,159× cheaper qwen3.5 122B $0.000969 787× cheaperChart by the author.

The local model is still roughly three orders of magnitude cheaper than the API.

15 часов назад @ towardsdatascience.com
How to Effectively Deploy Code With Claude Code
How to Effectively Deploy Code With Claude Code How to Effectively Deploy Code With Claude Code

It basically refers to the pipeline you have from after you’ve written the code and you want to deploy the code.

However, now this has completely shifted because code can be written super quickly now that we have coding agents to write all the code for us.

Snapshot the dev branch to merge to mainAnother technique I recently had to implement is snapshotting the dev branch before taking it to prod.

Deploying code has become quite different now that we produce so much more code and because of coding agents.

While a previous bottleneck was to write the code itself, bottlenecks have now shifted from writing the code to other software engineering tasks, such as deploying the code and the CI/CD pi…

1 day, 10 hours назад @ towardsdatascience.com
Building an Agent-Ready Data Warehouse: What Traditional Architectures Do Wrong
Building an Agent-Ready Data Warehouse: What Traditional Architectures Do Wrong Building an Agent-Ready Data Warehouse: What Traditional Architectures Do Wrong

A data agent can check metadata, select data sources, write SQL, and use the results to recommend next steps.

This is precisely where a problem arises that many older data warehouse architectures are not designed to address.

A Queryable Warehouse Is Not Automatically Agent-ReadyAt first glance, a sophisticated cloud data warehouse might appear AI-ready.

Access to lower levels may still be necessary, but raw data tables require stricter controls because they expose implementation details and only partially processed data records.

Start with One Decision, Not the Entire Data WarehouseDo not expose all warehouse data and add controls later.

1 day, 12 hours назад @ towardsdatascience.com
Variational Autoencoders (VAEs) Explained: From Theory to ELBO and the Reparameterization Trick
Variational Autoencoders (VAEs) Explained: From Theory to ELBO and the Reparameterization Trick Variational Autoencoders (VAEs) Explained: From Theory to ELBO and the Reparameterization Trick

Diagram illustrating a good scenario in which the similarity between objects is preserved in the latent space.

An analogous scenario would occur if you took a point midway between two embedded objects in the latent space.

Given all of that, the main motivations for a new autoencoder version are related to how the latent space is constructed.

The training objective will remain coherent: to maximize log p(x), we need to maximize the ELBO.

Based on that, the KL divergence loss is calculated to estimate how close the latent image distribution is to a normal distribution.

1 day, 13 hours назад @ towardsdatascience.com
SPP-Net Paper Walkthrough: Breaking the Fixed-Size Constraint
SPP-Net Paper Walkthrough: Breaking the Fixed-Size Constraint SPP-Net Paper Walkthrough: Breaking the Fixed-Size Constraint

The flow of the conventional CNN architecture (top) vs the flow of a CNN with SPP layer (bottom) [1].

How SPP Layer WorksFigure 3 below displays the detailed steps of how SPP-Net processes an image.

According to the paper, this part will then be connected to the SPP layer followed by two FC layers and the classification layer.

Here I am trying to pass the dummy tensor x through the network, which simulates an RGB image of size 224×224.

Here I challenge you to try implementing SPP layer on modern architectures like ResNet, ConvNeXt, etc.

1 day, 15 hours назад @ towardsdatascience.com
I Thought Loading Data Was the Finish Line. It Was the Starting Point.
I Thought Loading Data Was the Finish Line. It Was the Starting Point. I Thought Loading Data Was the Finish Line. It Was the Starting Point.

, I gave myself a 12-month roadmap to go from data analyst to data engineer.

Here’s what my articles table actually looked like once I stopped and paid attention to it.

With that, I had a dbt project sitting on top of the same Postgres instance my RSS pipeline had been writing to for weeks.

My raw articles table isn’t something dbt built, it’s external data that already exists, so dbt calls it a source.

The Postgres database, the dbt project, all of it lives locally in Docker, which means none of this exists anywhere the moment my laptop is off.

2 days, 12 hours назад @ towardsdatascience.com
How to Implement Structured Output with Local LLMs
How to Implement Structured Output with Local LLMs How to Implement Structured Output with Local LLMs

We’ll use Gemma 4 as our local LLM, Ollama as the serving runtime, and Pydantic to define and validate the output schema.

How Do We Implement Structured Output with a Local LLM?

Since these notes contain private information, a local LLM is a natural fit as the first step.

That’s the shape we want the local LLM to output.

With structured output, we can integrate local LLMs into a larger workflow, where the downstream components can easily consume LLMs’ results.

2 days, 14 hours назад @ towardsdatascience.com
Before Q, K, and V: Reconstructing the Transformer
Before Q, K, and V: Reconstructing the Transformer Before Q, K, and V: Reconstructing the Transformer

We want to break this symmetry so let’s keep only the middle interaction function v, which I’ll call the “attention” function from now on.

(One caveat is that the Transformer architecture adds scaling for computational stability, hence the term “scaled dot product attention”.

Okay, all of this is great—but where are the matrices Q, K, and V that the article title promised us?

The product between q and K^T creates a vector containing every dot product between q and a key in K, and the softmax on top normalizes the final dot product scores.

We needed to reduce our cache size for the reusable matrix-vector multiplies in the dot product, which required computing the dot product in lower dimensi…

3 days, 12 hours назад @ towardsdatascience.com
Building a Streamlit UI for My LangGraph AI Agent
Building a Streamlit UI for My LangGraph AI Agent Building a Streamlit UI for My LangGraph AI Agent

In this article, we will build a clean, interactive Streamlit UI on top of the existing LangGraph agent.

In LangGraph, the agent architecture is defined and executed as a compiled state graph object so they essentially mean the same thing in this article.

Both serve as a wrapper for the LangGraph agent.

This function sends the customer action to the LangGraph agent and saves the resulting state for the Streamlit interface.

Finally, we have some render functions ( _render...() ) defined in streamlit_app.py to convert the current LangGraph state into visible Streamlit components.

3 days, 14 hours назад @ towardsdatascience.com
Matplotlib vs Plotly: Which Python Chart Tool Should You Choose?
Matplotlib vs Plotly: Which Python Chart Tool Should You Choose? Matplotlib vs Plotly: Which Python Chart Tool Should You Choose?

In this article, I’ll provide several examples of using Plotly and Matplotlib on the same datasets to illustrate their key differences.

Now, let’s create the same plot using Plotly Express, which provides a high-level interface similar to Seaborn.

Although the code complexity is comparable, the user experience when using Plotly is greatly enhanced.

Example 3: Saving and SharingHow you save and share plots using the two tools differs significantly in one key aspect.

With the high-level plotly.express module, creating these interactive plots often requires minimal code changes compared to Matplotlib/Seaborn, while providing a much richer user experience.

4 days, 10 hours назад @ towardsdatascience.com
Loop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top One
Loop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top One Loop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top One

Listing questions break the one assumption retrieval is built on, that the answer is the top passage.

For a factual question, top-k retrieval works because the answer is in one passage.

For a listing question, the answer is in N passages, where N is unknown ahead of time.

1.3 Detecting the listing intentThe first thing the pipeline needs is to recognize that the question is a listing question, not a factual one.

Sources and further readingThe benchmark showing top-k retrieval ceiling far below 100% on list questions is Amouyal et al.

4 days, 12 hours назад @ towardsdatascience.com
The Problem with pandas Isn’t Performance. It’s Cognitive Overhead.
The Problem with pandas Isn’t Performance. It’s Cognitive Overhead. The Problem with pandas Isn’t Performance. It’s Cognitive Overhead.

AI agents can generate the pandas code for us now, so what does it matter?

Let’s be honest: reading other people’s pandas code, especially code that was built through an interactive session can be rather painful.

The enduring popularity of visual data toolsIf you’re still not convinced that any of this matters, think for a minute about the popularity of visual data tools.

But if “code-as-documentation” is kind of the point, why are we using boilerplate-laden pandas code as our language of choice?

Since then, the Python data ecosystem has become entrenched across academia, industry and government IT environments.

4 days, 13 hours назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 16 часов назад
The Sequence Knowledge - Issue 911: Distilling Diffusion and Multimodal Models
The Sequence Knowledge - Issue 911: Distilling Diffusion and Multimodal Models The Sequence Knowledge - Issue 911: Distilling Diffusion and Multimodal Models

Text distillation teaches a smaller model to imitate an answer.

Diffusion and multimodal distillation must compress trajectories, distributions, motion, and the semantic geometry between different worlds.

The teacher says “Paris”; the student learns to say “Paris.” The teacher writes a good explanation; the student learns the shape of the explanation.

Diffusion distillation is stranger.

A diffusion model does not emit an image in one clean forward pass.

16 часов назад @ thesequence.substack.com
The Sequence Radar - Issue 910: Last Week in AI: Google Rewires Its Brain and Meta Hires a Coding Swarm
The Sequence Radar - Issue 910: Last Week in AI: Google Rewires Its Brain and Meta Hires a Coding Swarm The Sequence Radar - Issue 910: Last Week in AI: Google Rewires Its Brain and Meta Hires a Coding Swarm

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Google Rewires Its Brain and Meta Hires a Coding SwarmAI weeks are usually measured in parameter counts.

Meanwhile, Meta released Muse Code, a terminal-based coding agent powered by Muse Spark 1.2.

Let’s review this week’s developments:🔎 AI ResearchAI Lab: FAIR, Meta, Reality Labs, Meta, University of Oxford.

AI Lab: Google Cloud AI Research, University of California, Los Angeles.

🤖 AI Tech ReleasesMuse CodeMeta AI released the beta version of Muse Code, a terminal coding agent optimized for tasks across large repositories.

2 days, 16 hours назад @ thesequence.substack.com
The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering
The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering

For most of software history, engineering capacity was easy to sketch on a whiteboard.

The modern engineering organization now has a second, elastic workforce.

The human workforce is measured in headcount.

The machine workforce is measured, imperfectly, in tokens.

And because companies love measurable things—especially things that produce dashboards—we have entered the era of token maxing.

5 days, 15 hours назад @ thesequence.substack.com
The Sequence AI of the Week #908: You Need to Learn About Gemini Robotics
The Sequence AI of the Week #908: You Need to Learn About Gemini Robotics The Sequence AI of the Week #908: You Need to Learn About Gemini Robotics

You ask Apptronik’s Apollo 2 to put the watering can into the green bin on the bottom shelf.

It walks to the table, picks up the can, takes a few steps to the shelves, and places it where you asked.

Nothing in that sentence sounds hard until you remember what the previous Gemini Robotics models actually were.

They were a torso bolted to a fixed base doing tabletop work.

That is the real news in Gemini Robotics 2, and it is a bigger deal than the b-roll makes it look.

6 days, 15 hours назад @ thesequence.substack.com
The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures
The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures

Every form of distillation in this series so far has quietly preserved one thing: teacher and student spoke the same dialect.

The student was a compressed copy, then a more capable apprentice, then a reasoner trained on traces — but underneath, it was always the same kind of machine, attention layers stacked on attention layers, differing only in size.

You take a fully trained transformer and pour its capability into a fundamentally different computational substrate, and somehow the capability survives the transplant.

The first time you see it work, it feels a little illicit, like recovering a person’s memories after swapping out their brain for different hardware.

So it’s worth understandi…

1 week назад @ thesequence.substack.com
The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction
The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction

The AI of the week dives into Gemini Robotics 2.

Google DeepMind pushed the frontier in a different direction with Gemini Robotics 2.

AI Lab: MistralAISummary: This paper presents Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that simplifies content moderation into a unified binary question-answering task.

🤖 AI Tech ReleasesGemini Robotics 2Google DeepMind released Gemini Robotics 2 , a three-model suite of intelligence for robotics.

LFM2.5Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, two encoder models that can be easily adaped to downstream tasks.

1 week, 2 days назад @ thesequence.substack.com
The Sequence Robotics #905: Who Builds the Robot Brain?
The Sequence Robotics #905: Who Builds the Robot Brain? The Sequence Robotics #905: Who Builds the Robot Brain?

This is the first post of a new section of TheSequence focused on advancements in robotics.

Our goal is to keep you up to date with the most important developments in AI robotics which is an area that is not well covered by other newsletters.

For this first post, I wanted to discuss the current landscape of AI models for robotics.

Robotics is what happens when an AI model leaves the library and discovers physics.

This is why the race for the robot foundation model will not simply replay the LLM market.

1 week, 4 days назад @ thesequence.substack.com
TheSequence Opinion #904: The Age of Research Is Overrated. AI Engineering Is Winning
TheSequence Opinion #904: The Age of Research Is Overrated. AI Engineering Is Winning TheSequence Opinion #904: The Age of Research Is Overrated. AI Engineering Is Winning

From roughly 2012 to 2020, he argued, the field lived in an age of research.

From 2020 to 2025, it entered an age of scaling.

Open the release notes for almost any frontier model in 2026 and the architecture diagram looks strangely familiar.

The headline improvements are usually elsewhere: better data, longer context, stronger reinforcement learning, synthetic tasks, tool use, memory, verification, adaptive reasoning budgets, and agent orchestration.

The age of research has returned, but much of that research is now expressed as industrial-scale engineering.

1 week, 5 days назад @ thesequence.substack.com
The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar
The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar

Take every open-weight model that discloses its parameter count, put total parameters on a log x-axis, put Terminal-Bench 2.1 score on the y-axis, and you get a reasonably tidy cloud sloping up and to the right.

Disclosed-size open-weight models on Terminal-Bench 2.1.

Laguna S 2.1 scores 70.2%.

On DeepSWE, a harder and less saturated benchmark, the gap stops being subtle at all: Laguna S 2.1 scores 40.4 against DeepSeek-V4-Pro-Max’s 9.0.

A 13x parameter deficit paired with a 4x score advantage is the kind of result that usually means somebody broke the eval.

1 week, 6 days назад @ thesequence.substack.com
The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher
The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher

For most of machine learning history, data was treated as geology.

It already existed somewhere in the world—in books, websites, code repositories, conversations, photographs, and databases.

The researcher’s job was to excavate it, clean it, tokenize it, and feed it into a model.

The dataset trains the student.

The teacher disappears at inference time, but some of its behavior remains embedded in the student.

2 weeks назад @ thesequence.substack.com
The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack
The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack

Next Week in The Sequence:Our series about AI model distillation continues with another exciting technique.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI StackWhen I started The Sequence years ago, AI was still a relatively niche field, followed closely by researchers, a small group of builders, and a few overly enthusiastic people like me.

AI Lab: Meta AISummary: This paper introduces GAMUT, a multimodal benchmark designed to evaluate the factual completeness of long-form generations rather than just their factual precision.

AI Lab: Microsoft ResearchSummary: This paper presents Experiential Learning (EL)…

2 weeks, 2 days назад @ thesequence.substack.com
The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA?
The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA? The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA?

THESIS Google is the closest strategic mirror of NVIDIA’s full-stack method, but not a universal drop-in replacement; AMD and AWS make a literal “only” claim too strong.

This is useful, but incomplete in the same way that comparing airlines by engine thrust is incomplete.

It is an industrial system that turns models into running software with unusually little friction.

That is the strongest form of the case for Google as NVIDIA’s only viable competitor.

Google is the closest full-stack strategic rival, not a universal drop-in replacement, and AWS and AMD make the word “only” uncomfortable.

2 weeks, 5 days назад @ thesequence.substack.com
The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time
The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time

Inkling is best understood not as a single 975-billion-parameter brain that fires all at once, but as a giant warehouse of specialist capacity.

The more interesting number is the one beside it: 41 billion parameters active per token.

A good mental picture is a university with 256 specialist departments on each relevant floor.

It selects six departments that appear useful for this token, adds two general-purpose departments that always attend, combines their work, and moves on.

A line of Python might summon one set of specialists; a phrase in Greek, a diagram label, or a piece of audio may summon another.

2 weeks, 6 days назад @ thesequence.substack.com
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models

The distilled 32B model started solving competition math it had no business solving.

The 7B model began verifying its own work and branching its reasoning mid-stream — emergent behaviors nobody trained into it directly.

A grab-bag of small dense models suddenly reasoned like something ten times their size.

We spent an entire installment establishing why naive sequence-level imitation is the wrong tool, and then the single most important reasoning-distillation result of the decade is naive sequence-level imitation.

The answer is the whole story of this installment, and it turns out to be more interesting than either “imitation works” or “imitation doesn’t.”

3 weeks назад @ thesequence.substack.com
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race

In the AI of the Week , we discuss Thinking Machine first open weights model.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: China, Compression and the Open-Model RaceFor years, AI progress has been narrated as a horse race: larger models, higher benchmark scores, more expensive clusters.

Moonshot describes it as the first open model in the three-trillion-parameter class, although the full weights are not due until later this month.

This week, AI stopped looking like a single race.

AI Lab: Shanghai AI LaboratorySummary: ADVANCED MATHBENCH introduces a rigorous evaluation suite focusing on the generation and process-level verification of advanced, natural-language mathematical pr…

3 weeks, 2 days назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 4 months, 1 week назад
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

4 months, 1 week назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

5 months, 3 weeks назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

5 months, 3 weeks назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing & LLMs — новый сезон Курс Natural Language Processing & LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

6 months, 1 week назад @ habr.com
Machine Learning Mastery
последний пост 1 week, 6 days назад
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026? Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?

Then we walked through the fastest way to get inference running locally in Run a Local AI Model in 15 Minutes: Your First Ollama Setup.

Spend enough time in the local AI ecosystem, though, and you’ll notice Ollama isn’t the only option competing for your hard drive.

Three tools dominate the local AI runtime landscape: Ollama, LM Studio, and llama.cpp.

Ollama (Via its dedicated, background-daemon CLI) ollama run llama3 .

There’s a well-worn progression in the local AI community that maps almost exactly to the three tools covered here: LM Studio → Ollama → llama.cpp.

1 week, 6 days назад @ machinelearningmastery.com
5 Architectural Patterns for Persistent Memory and State in AI Agents
5 Architectural Patterns for Persistent Memory and State in AI Agents 5 Architectural Patterns for Persistent Memory and State in AI Agents

The fix isn’t a bigger context window; it’s treating memory and state as deliberate architectural decisions, not afterthoughts.

Memory feeds into state; state feeds back into memory.

Also worth calling out explicitly: credentials and secrets are not semantic memory.

Episodic Event Logs (Historical Reflection)Semantic memory stores what the agent knows; episodic memory stores what the agent did.

The moment your system serves more than one user, memory has to be siloed.

2 weeks, 1 day назад @ machinelearningmastery.com
Stateful vs. Stateless Agent Design: Tradeoffs for Scalable Agentic Systems
Stateful vs. Stateless Agent Design: Tradeoffs for Scalable Agentic Systems Stateful vs. Stateless Agent Design: Tradeoffs for Scalable Agentic Systems

environ [ "GROQ_API_KEY" ] = "PASTE_YOUR_GROQ_API_KEY_HERE" # Initializing the client client = Groq ( ) # Using an efficient model from Groq: Llama 3.1 8B Instant MODEL_ID = "llama-3.1-8b-instant"An important setup decision here is the choice of a specific model.

strip ( )To understand the limitations of a stateless agent, we simulate a simple user-model conversation through it:# --- Testing the Stateless Agent --- print("--- Turn 1 ---") prompt_1 = "Hi, my name is Alice and I am learning about API infrastructure."

--- Turn 2 (Without Client Context) --- Agent: Unfortunately, I don't have any information about you, including your name.

-- - Turn 2 ( Without Client Context ) -- - Agent : Unf…

2 weeks, 4 days назад @ machinelearningmastery.com
An Introduction to Loop Engineering
An Introduction to Loop Engineering An Introduction to Loop Engineering

Topics we will cover include:The origin and definition of loop engineering, and how it fits into the broader progression from prompt engineering to context engineering to harness engineering.

The three hardest problems in loop engineering — context management, termination, and verification — and the failure modes that result from getting any one of them wrong.

Prompt engineering effectively became one ingredient within context engineering rather than a separate discipline.

Loop engineering is simply the part where all of that gets put into motion and given a rhythm.

It’s that “loop engineering” is a product name and a rallying phrase for a research direction that’s been quietly accumulating…

2 weeks, 5 days назад @ machinelearningmastery.com
The Current State of Agentic AI
The Current State of Agentic AI The Current State of Agentic AI

How the Model Context Protocol, persistent memory graphs, and emerging security patterns define the current production landscape.

IntroductionLook back at how we built AI agents just a year ago, and the dominant paradigm was brute-force orchestration.

This tutorial breaks down the current state of agentic AI architecture, covers the three major shifts defining production systems today, and walks through how to design a modern agent swarm.

The current state of tool calling is increasingly defined by the Model Context Protocol (MCP).

This open standard acts as a universal adapter between AI models and local or remote data sources.

3 weeks назад @ machinelearningmastery.com
Building Agentic Workflows in Python with LangGraph
Building Agentic Workflows in Python with LangGraph Building Agentic Workflows in Python with LangGraph

Managing Conversation History with MessagesStateEvery node in a LangGraph graph reads the current state and writes updates back to it.

add_node ( "run_model" , run_model ) builder .

Define a tool with the @tool decorator:from langchain_core.tools import tool @tool def get_customer_tier(customer_id: str) -> str: """Look up the subscription tier for a customer by their ID.

tools import tool @ tool def get_customer_tier ( customer_id : str ) -> str : "" "Look up the subscription tier for a customer by their ID.

add_node ( "run_model" , run_model ) builder .

3 weeks, 1 day назад @ machinelearningmastery.com
Agentic AI Security: Defending Against Prompt Injection and Tool Misuse
Agentic AI Security: Defending Against Prompt Injection and Tool Misuse Agentic AI Security: Defending Against Prompt Injection and Tool Misuse

Share Post ShareIn this article, you will learn what prompt injection and tool misuse are in the context of agentic AI systems, and which defense strategies experts recommend to mitigate them.

Topics we will cover include:How prompt injection and tool misuse can compromise AI agents deployed in real-world production environments.

Prompt injection arises when untrusted inputs to a language model are interpreted as instructions rather than mere data.

This problem has been renamed Agent Goal Hijacking in the context of agentic AI and AI security vulnerabilities.

Closing Remarks: Looking AheadIn line with the growing level of sophistication attained by agentic AI systems, organizations should a…

3 weeks, 4 days назад @ machinelearningmastery.com
Run a Local AI Model with Ollama in 15 Minutes
Run a Local AI Model with Ollama in 15 Minutes Run a Local AI Model with Ollama in 15 Minutes

Topics we will cover include:Why Ollama has become the standard tool for running local AI models.

Ollama has become the go-to tool for local AI because it packages complex model architectures into a clean, lightweight background service.

# Verify Ollama is running by checking the version ollama --version # Pull and immediately run the Llama 3.2 3B model ollama run llama3.2 1 2 3 4 5 # Verify Ollama is running by checking the version ollama -- version # Pull and immediately run the Llama 3.2 3B model ollama run llama3 .

Just run your command directly ( ollama run llama3.2 ), the background daemon is already listening on port 11434.

From here, exploring the other models from our Top 7 list is…

3 weeks, 5 days назад @ machinelearningmastery.com
Scikit-Ollama for Scikit-LLM/Ollama Integration
Scikit-Ollama for Scikit-LLM/Ollama Integration Scikit-Ollama for Scikit-LLM/Ollama Integration

zero_shot import ZeroShotOllamaClassifier # Initializing the classifier with our local Ollama model: llama3:latest clf = ZeroShotOllamaClassifier ( model = "llama3:latest" )A very important clarification about what we just did.

Predicted Sentiment: positive Text: 'The special effects in 'Star Battles: Nebula Conflict' were out of this world.

Predicted Sentiment: positive Text: ''The Lost Symphony' was a masterclass in character development and storytelling.

Predicted Sentiment: positive Text: ' The special effects in 'Star Battles: Nebula Conflict' were out of this world .

The key ingredient: the scikit-ollama library, which elegantly encapsulates this local integration and makes it availab…

3 weeks, 6 days назад @ machinelearningmastery.com
LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does
LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does

])\s+', answer.strip()) return [s.strip() for s in sentences if s.strip()] def claim_supported_by_context(claim: str, context: str) -> bool: """ Check whether a claim has lexical support in the retrieved context.

RAGAS does this with an LLM judge; this overlap check demonstrates the same supported/unsupported decision in a deterministic way. """

strip ( ) ) return [ s . strip ( ) for s in sentences if s . strip ( ) ] def claim_supported_by_context ( claim : str , context : str ) -> bool : "" " Check whether a claim has lexical support in the retrieved context.

def your_judge_function(response_1: str, response_2: str) -> str: # Placeholder -- wire this up to your actual judge model call.

def…

4 weeks назад @ machinelearningmastery.com
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach

The common pitfalls that show up once agent memory is implemented, and how to fix them.

Agent memory strategy deserves the same deliberate design as orchestration.

Why Is Choosing an AI Agent Memory Strategy Important?

A customer support agent, for example, might keep the current ticket in working memory, a customer’s subscription tier in semantic memory, past complaints in episodic memory, and a learned refund-handling routine in procedural memory.

As discussed, working memory, semantic memory, episodic memory, and procedural memory serve different purposes and require different storage and retrieval strategies.

1 month назад @ machinelearningmastery.com
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls

openai import OpenAI as LlamaOpenAI from llama_index .

# Prerequisites: pip install openai python-dotenv # How to run: python raw_api_agent.py import os import json from dotenv import load_dotenv from openai import OpenAI load_dotenv ( ) client = OpenAI ( api_key = os .

Prerequisites:pip install openai langchain langchain-openai llama-index \ llama-index-llms-openai llama-index-embeddings-openai python-dotenv 1 2 pip install openai langchain langchain - openai llama - index \ llama - index - llms - openai llama - index - embeddings - openai python - dotenvHow to run: Save as three_ways.py and run python three_ways.py# three_ways.py # The same document Q&A task implemented three ways: # Raw …

1 month назад @ machinelearningmastery.com
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering

Topics we will cover include:What tools and subagents are, and the key differences between them.

This article explains what tools and subagents are, where each fits, and how to make the choice every time.

When an agent calls a tool, the result lands back in the same context the agent is actively reasoning in — prior reasoning, tool result, and everything else together.

A database record, search result, or API response can often be consumed immediately.

If the answer is independent reasoning, context isolation, specialized capabilities, or parallel execution, a subagent is likely justified.

1 month назад @ machinelearningmastery.com
The Complete Guide to Tool Selection in AI Agents
The Complete Guide to Tool Selection in AI Agents The Complete Guide to Tool Selection in AI Agents

tools = tools self .

tools = tools self .

threshold : return { "status" : "resolved" , "tool" : tool [ "name" ] , "confidence" : score , "attempts" : 1 } # Reformulate by stripping filler words.

tools = tools self .

full_catalog_tokens = sum ( estimate_tokens ( d ) for d in descs ) def _retrieve ( self , query : str , top_k : int ) -> list [ dict ] : query_vec = self .

1 month назад @ machinelearningmastery.com
Context vs. Memory Engineering in Agentic AI Systems
Context vs. Memory Engineering in Agentic AI Systems Context vs. Memory Engineering in Agentic AI Systems

Share Post ShareIn this article, you will learn how context engineering and memory engineering solve different problems in agentic AI systems, and how the two disciplines meet at the point where retrieved memory enters the context window.

Most of the time, the problem lies in two areas that get built together, conflated, or skipped: context engineering and memory engineering.

Memory Engineering: Designing Persistent AI Memory SystemsOnce an inference call completes, memory engineering determines what deserves to persist and under what conditions it gets used again.

trust_level >= 0.5 )AI Agent Memory Design Guide – Working, Long-Term, and Procedural Memory with Forgetting and Staleness Mana…

1 month, 1 week назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 3 weeks, 4 days назад
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

3 weeks, 4 days назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

2 months, 3 weeks назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

5 months назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

6 months, 2 weeks назад @ alexirpan.com
Authentic Imperfection
Authentic Imperfection Authentic Imperfection

* * *I’ve been thinking about the anger surrounding generative AI.

To keep things fair, he took the best human images and best AI images, meaning human art from famous artists, and AI art from prompters skilled at removing obvious tells of image generation.

When people complain about AI slop, I see it as a complaint against the deluge of default style AI images.

We’ve seen this happen in all forms: AI text, AI music, older forms of computer generated content like CGI.

As much as we celebrate imperfection, digital imperfection is a step too far.

8 months, 4 weeks назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 5 months, 2 weeks назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

5 months, 2 weeks назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

6 months, 1 week назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 5 days, 11 hours назад
WeatherNext: AI model achieves breakthrough in forecasting cyclones
WeatherNext: AI model achieves breakthrough in forecasting cyclones WeatherNext: AI model achieves breakthrough in forecasting cyclones

Predicting how dangerous cyclones develop is a longstanding challenge where every hour counts.

Today, in a paper published in Nature, we show that our WeatherNext AI model achieved state-of-the-art accuracy in predicting a cyclone's track, intensity, and wind structure.

During the 2025 hurricane season, our model helped the NHC to make a historic forecast for Hurricane Melissa by predicting the storm’s rapid intensification and landfall in Jamaica.

Given this broad impact, we are now open sourcing our WeatherNext 2 and WeatherNext Cyclones models used during the hurricane season.

How WeatherNext predicts weather and cyclones

5 days, 11 hours назад @ deepmind.google
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics.

Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function.

Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6.

Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.

Gemini Robotics ER 2 improves this tool orchestration workflow.

1 week, 5 days назад @ blog.google
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control

Our newest music generation model, Lyria 3.5, delivers significant advancements across musicality, lyrics, and vocal quality, empowering you to craft richer tracks.

We’re rolling it out today in Google Flow Music, where we want to help you create songs you love, with creative control.

Enhanced lyrics: Generate higher quality lyrics with improved prompt adherence and structural awareness.

Generate higher quality lyrics with improved prompt adherence and structural awareness.

Improved vocals: Bring more expression and emotion to your songs with more realistic and emotionally nuanced vocals, plus improved pronunciation.

1 week, 6 days назад @ blog.google
Gemini Robotics 2 brings whole body intelligence to robots
Gemini Robotics 2 brings whole body intelligence to robots Gemini Robotics 2 brings whole body intelligence to robots

From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasksFor decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.

Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots.

As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration.

Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks.

And this profound intelligence can also run locally on-device while seamlessly a…

2 weeks назад @ deepmind.google
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

In December, we shared our commitment to the White House's Genesis Mission — the national effort to harness AI and double the pace of American scientific discovery within a decade.

Today, at the DOE Genesis Mission Summit 2026, we are expanding this by committing $40 million of AI tokens and cloud credits for researchers in support of the Genesis Mission.

WeatherNext — a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

— a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

Driving American innovationThe Genesis Mission represents an opportunity to transform research and science across America.

2 weeks, 6 days назад @ cloud.google.com
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Building on Gemini 3.5 Flash, we’re introducing new Gemini models:3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure.

3.6 Flash: More efficient and better quality than 3.5 FlashGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash.

For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash.

This enhanced efficiency is also combined with a lower price than 3.5 Flash.

3 weeks назад @ blog.google
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Building on Gemini 3.5 Flash, we’re introducing new Gemini models:3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure.

3.6 Flash: More efficient and better quality than 3.5 FlashGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash.

For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash.

This enhanced efficiency is also combined with a lower price than 3.5 Flash.

3 weeks назад @ blog.google
Introducing Gemini 3.5 Flash Cyber
Introducing Gemini 3.5 Flash Cyber Introducing Gemini 3.5 Flash Cyber

Today, we’re expanding our longtime efforts to better prepare defenders by introducing Gemini 3.5 Flash Cyber, our lightweight cybersecurity model built on top of 3.5 Flash and fine-tuned to find, validate, and patch vulnerabilities quickly and efficient, making it more effective at these tasks than Gemini’s mainline Flash models.

By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.

Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber.

3.5 Flash Cyber benchmark results: an efficient alternative to larger cybersecurity modelsWe tested 3.5 F…

3 weeks, 4 days назад @ deepmind.google
Our approach to bioresilience
Our approach to bioresilience Our approach to bioresilience

Today, Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience.

Inside our bioresilience programWe believe society must harness AI’s advancing capabilities to address infectious diseases and prepare for future outbreaks.

With this in mind, we are making our AI models and agents available to trusted partners to support progress across three key areas: prevention, detection and response.

Working in collaboration with governments and global health authorities to advance a diverse range of diagnostic and therapeutic strategies enables the Isomorphic Labs Drug Design Engine’s real-world impact for bioresilience.

To read more about our work and our call for new par…

3 weeks, 5 days назад @ deepmind.google
Empowering India’s next generation of innovators with ATL Saathi
Empowering India’s next generation of innovators with ATL Saathi Empowering India’s next generation of innovators with ATL Saathi

A new contribution to Indian Education with Atal Innovation MissionWe believe behind every good student is a great teacher.

That’s why for over 20 years, Google has been dedicated to supporting the education ecosystem by introducing technology into teaching and learning through a teacher-led approach.

With foundational platforms like Google for Education and Google Classroom, we build products tailored to the needs of schools, keeping the teacher in the lead.

To further support the empowerment of educators, our new Google Educator AI Series ensures teachers are equipped with both the tools and the digital skills required for today's classrooms.

We see Gemini as a great tool to enable our pa…

4 weeks, 1 day назад @ deepmind.google
Google DeepMind and A24 announce first-of-its-kind research partnership
Google DeepMind and A24 announce first-of-its-kind research partnership Google DeepMind and A24 announce first-of-its-kind research partnership

Today, Google DeepMind and A24 are announcing a first-of-its-kind partnership focused on research.

The collaboration pairs a world-leading research lab with the industry’s most filmmaker-forward studio to help artists develop new workflows and techniques.

This partnership creates a deep research and development collaboration between A24 and Google DeepMind spanning multiple projects over time.

This hands-on collaboration provides Google DeepMind with invaluable feedback and guidance from leading artists.

As A24 and Google DeepMind’s researchers work side-by-side to test, iterate and build, this partnership aims to expand what is possible in the future of entertainment.

1 month, 1 week назад @ blog.google
Start building with Nano Banana 2 Lite and Gemini Omni Flash
Start building with Nano Banana 2 Lite and Gemini Omni Flash Start building with Nano Banana 2 Lite and Gemini Omni Flash

Uploading audio references and scene extension is not yet supported in the Gemini API for this model.

Video references up to 3 seconds in duration are accepted by the API schema but are not correctly processed by the model at this time.

Gemini Omni is available in public preview starting today in Google AI Studio and the Gemini API.

Use Nano Banana 2 Lite as a high-speed image generation model, then pass that image as a reference to Gemini Omni Flash to animate it into a high-quality video.

To help you get started we created a few demo apps you can remix that let you experience how you can pair both Nano Banana 2 Lite and Gemini Omni Flash into one workflow.

1 month, 1 week назад @ blog.google
Introducing computer use in Gemini 3.5 Flash
Introducing computer use in Gemini 3.5 Flash Introducing computer use in Gemini 3.5 Flash

Making computer use safe in 3.5 FlashTo mitigate some of the prompt injection risks for agents operating in live environments, we use targeted adversarial training for computer use in Gemini 3.5 Flash.

We’re also releasing two optional enterprise safeguard systems that enable enterprises to:Require explicit user confirmation for sensitive or irreversible actions.

Automatically stop tasks if an indirect prompt injection is identified.

Taking a “defense-in-depth” approach, we encourage developers to combine these features with secure sandboxing, human-in-the-loop verification and strict access controls.

We are already seeing customers drive value with computer use.

1 month, 2 weeks назад @ blog.google
Unlocking UK house-building with AI-accelerated planning
Unlocking UK house-building with AI-accelerated planning Unlocking UK house-building with AI-accelerated planning

New UK government AI planning prototype built with Gemini aims to halve the time it takes to process homeowner applicationsAround the world, Governments are exploring how AI can deliver better public services, faster.

The UK is working to build 1.5 million new homes by 2029, but local planning authorities are often slowed down by dense paperwork and administrative backlogs.

To help get Britain building, we’re partnering with the UK government to help radically shorten the time it takes to process householder planning applications.

Following early trials in Barnet, Camden and Dorset, the government plans for the new AI planning tool to be made available to all councils nationally from 2027.

1 month, 3 weeks назад @ deepmind.google
Securing the future of AI agents
Securing the future of AI agents Securing the future of AI agents

How we’re securing internal systems against increasingly capable and imperfectly aligned AIAI agents are transforming our relationship with technology.

In the U.S alone, AI agents could create $2.9 trillion in economic value by 2030.

That’s why we developed our AI Control Roadmap: a framework for building and managing the advanced AI we deploy within Google.

Similarly, our AI control system grants AI agents permissions based on their verified behavior, allowing us to build trust through controlled, incremental access.

In our AI Control Roadmap, we map security protocols to measurable milestones in AI capabilities on two critical fronts:

1 month, 3 weeks назад @ deepmind.google
Google
последний пост 11 часов назад
Looker’s semantic layer governs Gemini Enterprise data for user trust
Looker’s semantic layer governs Gemini Enterprise data for user trust Looker’s semantic layer governs Gemini Enterprise data for user trust

When a Gemini Enterprise user requests a business KPI in Gemini Enterprise, the request is routed directly to a Looker agent.

Technical capabilities and enterprise readinessDeploying Looker agents natively into Gemini Enterprise via the A2A protocol doesn't just make it smarter — it makes it more interactive and interoperable, without sacrificing security.

When users interact with Looker agents inside Gemini Enterprise, the platform goes beyond textual explanations and provides native, interactive data charts.

Note: If you published Looker agents in Gemini Enterprise prior to Looker release 26.12, we recommend updating or refreshing them to take advantage of these enhanced visualization cap…

11 часов назад @ cloud.google.com
How Malachyte solves retail’s cold-start problem with managed real-time AI
How Malachyte solves retail’s cold-start problem with managed real-time AI How Malachyte solves retail’s cold-start problem with managed real-time AI

We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded Malachyte, an AI-powered ecommerce recommendation platform.

These days, consumers have come to expect content that feels personalized and relevant, and online services competing for their attention have no choice but to do this exceptionally well.

Malachyte was inspired by some unique insights into how advanced AI models, and large language models in particular, could be applied in new ways to old challenges like personalization and recommendations.

This is the story of how we built it, and the ways any founder can use services like these to start deploying AI found…

1 day, 11 hours назад @ cloud.google.com
How WPP operationalizes platform and data engineering for AI marketing
How WPP operationalizes platform and data engineering for AI marketing How WPP operationalizes platform and data engineering for AI marketing

But before it could begin applying sophisticated AI models to power those insights, WPP had to overcome a critical engineering challenge: the marketing data that made up the models was fragmented across hundreds of global agencies.

And until it built a reliable way to ingest, clean, and serve data to those models, WPP couldn’t unlock the true potential of generative AI.

To solve this, WPP partnered with Google Cloud to construct a unified data backbone and custom platform engineering path.

Now, by standardizing its serverless compute patterns and data processing workflows, WPP is able to securely deploy targeted marketing campaigns in days instead of months.

By utilizing a serverless archit…

1 day, 11 hours назад @ cloud.google.com
Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026
Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026 Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026

At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence.

By combining world-class AI research with an open, fully integrated AI platform, we give customers the flexibility to innovate and the foundation to deliver measurable business value.

At the center of it all is Gemini Enterprise, a unified platform designed to power the agentic enterprise, meet builders where they are, and deliver enterprise trust by default.

We believe this integrated approach is why Google has been named a Leader in The Forrester Wave™: AI Platforms, Q3 2026 report, and received the highest score in the Strategy category.

1 day, 11 hours назад @ cloud.google.com
Your agentic summer: No-cost lessons from Google experts to build and scale agents
Your agentic summer: No-cost lessons from Google experts to build and scale agents Your agentic summer: No-cost lessons from Google experts to build and scale agents

Intro to AI Agents: Build a foundational understanding of how autonomous agents can redefine productivity.

Enterprise Agents and Use Cases: Discover how AI agents drive real business impact.

Create Your First Gemini Enterprise Application skill badge: Earn a skill badge that proves you can create an app with Gemini Enterprise.

Orchestrate Multi-Agent Workflows with Gemini Enterprise skill badge: Demonstrate your ability to manage multiple agents powered by Gemini Enterprise with a skill badge.

Engineer AI Agents with Agent Development Kit (ADK) skill badge: Build production-grade agents using expert developer tools.

5 days, 11 hours назад @ cloud.google.com
Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications
Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications

Nearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research.

Today, we’re announcing that Mirendil, an exciting frontier AI lab focused on accelerating AI development, will also utilize Google Cloud’s AI Hypercomputer.

This includes using a mix of Google’s TPU AI accelerators and full-stack NVIDIA AI infrastructure running on Google Cloud; this purpose-built AI infrastructure will support model pre-training and post-training applications for Mirendil.

The Mirendil team is building new AI systems that can help accelerate and democratize AI research and development.

We closely partnered with Mirendil on end-to-e…

5 days, 14 hours назад @ cloud.google.com
How Deutsche Bank unlocked agility with an API-ready ecosystem
How Deutsche Bank unlocked agility with an API-ready ecosystem How Deutsche Bank unlocked agility with an API-ready ecosystem

At Deutsche Bank, we recognized that APIs aren't just technical plumbing; they're the nervous system of modern banking.

We've moved from "Where's that customer data API?"

Security: the employee onboarding analogyWhen thinking about API security, imagine onboarding a new employee.

Like employee access, these permissions are centrally managed, regularly audited, and instantly revocable.

This visibility serves operations, product managers who track partner value, and security teams who identify anomalies.

1 week назад @ cloud.google.com
How Target is enhancing retail discovery and cutting database maintenance by 50% with Spanner Graph
How Target is enhancing retail discovery and cutting database maintenance by 50% with Spanner Graph How Target is enhancing retail discovery and cutting database maintenance by 50% with Spanner Graph

Building the enterprise ontology on Spanner GraphWe evaluated multiple specialized technologies, including standalone vector databases and niche graph databases.

We ultimately chose Spanner Graph to build our enterprise ontology, which is a "graph-of-graphs" paradigm that allows us to construct a massive, generative AI-powered shopping graph.

By unifying our data, we bring semantic data, graph relationships, vector embeddings, and operational transactions under one roof.

Spanner Graph natively supports multi-hop graph traversals, semantic vector similarity, and full-text keyword queries over our relational tables.

Consolidated SQL + GQL interoperability: With Spanner Graph, our developers q…

1 week назад @ cloud.google.com
Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud
Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud

At Google Cloud, we propose an alternative: a modernization strategy that leverages the power of AI, agility of the cloud and allows for iterative and continuous modernization.

This approach recognizes a fundamental truth: mainframe modernization isn’t a pure code-to-code conversion problem.

Deep operational lock-in with specialized proprietary mainframe utility suites.

In other words, real-world modernization of mainframe applications is so much more than converting COBOL to Java.

Our solutions span four core pillars: assessment, modernization, de-risking, and data migration.

1 week, 1 day назад @ cloud.google.com
What’s new in AI infrastructure and orchestration this month
What’s new in AI infrastructure and orchestration this month What’s new in AI infrastructure and orchestration this month

We incorporate AI into the tools you use every day (think Gmail, BigQuery, AlloyDB, Google Cloud Code and Google Cloud Assist).

We make software frameworks to help you build with AI, like Gemini Enterprise Agent Platform, JAX, or MaxTest.

To support this, we are making AI infrastructure and orchestration news at a furious pace.

Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening.

Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads.

1 week, 4 days назад @ cloud.google.com
What Google Cloud announced in AI this month
What Google Cloud announced in AI this month What Google Cloud announced in AI this month

Along with a batch of new platform updates, this month we’ve put together 13 practical demos and 20 diagnostic questions to help your engineering teams align on a strong architectural blueprint.

Top announcementsThought leadership (editor’s pick):If automation requires delegation, then delegation requires trust.

But letting an AI agent run on its own is a big leap for any business.

What makes an AI agent trustworthy: Context is fast becoming one of the most valuable assets a company owns.

Prajakta Damle, Senior Director, Product Management, shares what it takes to get trustworthy AI right.

1 week, 4 days назад @ cloud.google.com
Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline
Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline

Expected operational standard: Your organizational mean time to remediate (MTTR) exposures and other desired changes into production goes down and to the right.

System consolidation: Boards should look beyond standalone AI features and point products to address systemic risk and truly enable business speed.

That deep context becomes the defender’s advantage when you are using AI powered defenses, including those in AI Threat Defense.

Your teams should be looking at how they are using AI to accelerate security and respond to AI-driven threats at AI speed.

Consider technologies like AI Threat Defense as part of your defenses in this new world.

1 week, 4 days назад @ cloud.google.com
Do more with less: How GKE can reduce your cost per agent by 75%
Do more with less: How GKE can reduce your cost per agent by 75% Do more with less: How GKE can reduce your cost per agent by 75%

Optimization 1: Pushing density with GKE Agent SandboxTo address this, we migrated the same agent workload from microVMs to GKE Agent Sandbox, a Kubernetes primitive that’s designed specifically for the security and performance requirements of running agents.

Instead of relying on heavy guest operating systems, GKE Agent Sandbox leverages the open-source secure container sandbox, gVisor.

It’s no surprise then, that when GKE Agent Sandbox reached General Availability in May, its usage grew more than 7x in under four weeks.

Key takeaway: In our tests, migrating OpenClaw-type agents to GKE Agent Sandbox enabled us to run more than 40% more agents per vCPU, and reduced the cost per agent by mor…

1 week, 5 days назад @ cloud.google.com
Automate your agent development lifecycle using any coding agent
Automate your agent development lifecycle using any coding agent Automate your agent development lifecycle using any coding agent

Welcome to our latest Gemini Enterprise Agent Platform deep dive, a practical walkthrough where we’ll teach you how to build real-world, production-ready agents starting from step 1.

With Agents CLI skills, you can go through the different phases of the entire agent lifecycle without ever leaving your coding agent.

What we’re building today: Industry Watch agentThis tutorial helps guide a developer on how to build a real Industry Watch agent, a sector-intelligence analyst for semiconductor stocks that reconciles what companies say in the press against what they file with the SEC.

We’ll walk through the six stages of building this agent end-to-end:Setup: Teach your coding assistant platform …

1 week, 6 days назад @ cloud.google.com
What’s new in Gemini Enterprise Agent Platform
What’s new in Gemini Enterprise Agent Platform What’s new in Gemini Enterprise Agent Platform

Since we launched Gemini Enterprise Agent Platform a few months ago, we’ve seen inspiring progress from businesses and builders alike.

To stir up development, we’ve also shared 13 demos that can walk you through the versatility and power of Agent Platform, and 20 questions you can ask your teams about building a solid agentic foundation.

That’s why today, we are announcing some of our most popular capabilities are available for everyone, from Agent Runtime to Agent Identity.

We also recently just announced CodeMender, our new managed code security agent to help you advance from passive scanning to automated code remediation, and reduce zero-day risk.

Automate your long-running agents faster…

1 week, 6 days назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 11 часов назад
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands.

The CARE-X model.

The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses.

CARE-X: Toward clinically useful radiology AICARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model.

11 часов назад @ microsoft.com
Orchard: An open framework for scalable agentic AI
Orchard: An open framework for scalable agentic AI Orchard: An open framework for scalable agentic AI

At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.

To address this gap, we introduce Orchard (opens in new tab), an open-source framework for scalable agentic modeling.

Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification.

(opens in new tab) We are also releasing the training data and evaluation methods used to build them.

By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research.

1 week, 1 day назад @ microsoft.com
Echoverse: Deep, evolving environments for computer-use agents
Echoverse: Deep, evolving environments for computer-use agents Echoverse: Deep, evolving environments for computer-use agents

At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters).

Shallow worlds backfire; deep worlds transferA shallow world is the cheap option.

A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%.

First, more deep worlds for the closed domains public benchmarks cannot reach.

1 week, 5 days назад @ microsoft.com
EvoLib: Turning experience into evolving knowledge
EvoLib: Turning experience into evolving knowledge EvoLib: Turning experience into evolving knowledge

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive.

As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to co…

1 week, 5 days назад @ microsoft.com
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

4 weeks, 1 day назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

1 month назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

1 month назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

1 month, 1 week назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

1 month, 1 week назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

1 month, 2 weeks назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

1 month, 2 weeks назад @ microsoft.com
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis

At a glance Talos is an open-source tool for automated, iterative reanalysis of genomic data in rare disease.

Deployed across a prospective cohort of almost 5,000 undiagnosed patients, Talos delivered 241 new diagnoses (5.1% additional yield).

On monthly iterative cycles, analysts only needed to review one new variant per 200 patients, demonstrating that frequent, systematic reanalysis can be run sustainably.

Why genome reanalysis mattersGenomic testing has transformed the diagnosis of rare disease, but even with this advancement, more than half of patients remain undiagnosed after their first test.

Looking aheadTalos reframes genomic reanalysis from a rare, labor-intensive event into a con…

1 month, 2 weeks назад @ microsoft.com
Ire identifies another LOTUSLITE specimen
Ire identifies another LOTUSLITE specimen Ire identifies another LOTUSLITE specimen

At a glance Project Ire identifies a LOTUSLITE variant that shares TTPs (tools, tactics, procedures) with the public family but none of its indicators of compromise (IOC).

On Ire’s calibrationOne noteworthy observation in Ire’s report (opens in new tab) is worth highlighting first.

The Ire report does not surface a matching entry-point name, but it identifies that the behavioral shape is the same.

Ire never named LOTUSLITE in its report or chain of evidence.

Ire described the behavior precisely enough to make the mapping straightforward of this sample to LOTUSLITE.

2 months назад @ microsoft.com
Data Formulator 0.7: AI-powered data analytics for enterprise data
Data Formulator 0.7: AI-powered data analytics for enterprise data Data Formulator 0.7: AI-powered data analytics for enterprise data

At a glance Data Formulator 0.7 is an open-source AI-powered system for enterprise data analytics that combines data connectivity, agent-guided exploration, and visualization refinement in a shared workspace.

Enterprise teams increasingly rely on AI systems for analytics, but enterprise data workflows are often fragmented across storage systems and tools.

Listen now Opens in a new tabConnecting enterprise data with Data ConnectorsData Formulator helps teams bring enterprise data into an AI-ready workspace without needing to rebuild the same connections for every source of data.

Data Connectors provide persistent connections between enterprise data sources and Data Formulator, allowing analy…

2 months, 2 weeks назад @ microsoft.com
Extending Human Intelligence Through AI
Extending Human Intelligence Through AI Extending Human Intelligence Through AI

At a glance Modern AI systems are powerful not because they replicate human intelligence, but because they presuppose it, by extending structures already present in human cognition and language.

Understanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems.

Rather than asking whether AI systems are becoming intelligent in the human sense, these approaches ask a more basic question: What if AI systems work because they rely on structures that are rooted in human cognition?

In our recent paper, The Origins of Artificial Intelligence in Natural Intelligence, we argue that modern AI systems are best understood nei…

2 months, 2 weeks назад @ microsoft.com
MIT AI MIT AI
последний пост 1 day, 7 hours назад
With a feel for physics, AI models simulate a wider range of real-world scenarios
With a feel for physics, AI models simulate a wider range of real-world scenarios With a feel for physics, AI models simulate a wider range of real-world scenarios

Artificial intelligence models are jacks of many trades, including writing, generating images, and creating 3D models.

To build an AI system that can reliably simulate a variety of physical scenarios, engineers need a range of physics data at a scale that isn’t yet feasible.

Industry successThe researchers found that GeoPT was particularly skilled at simulating industrial scenarios, as it outperformed state-of-the-art simulation models across benchmarks.

Likewise, its simulations of how light would pass through what was essentially a toy rabbit were accurate, despite never training on that 3D model or light physics beforehand.

The demonstrated success in a wide range of application domains …

1 day, 7 hours назад @ news.mit.edu
Solving the solvent problem
Solving the solvent problem Solving the solvent problem

“It’s supposed to be an ion conductor.” But unfortunately, most electrolytes get involved in unwanted chemical reactions with the electrodes, which can greatly undermine battery stability.

The team’s goal, accordingly, was to identify solvent molecules that are small enough to improve ion transport while still maintaining electrolyte stability.

There is, however, a complicating factor — a trade-off to be addressed: Faster ion transport often comes at the expense of electrolyte stability.

By carefully tailoring the size of solvent molecules, the authors demonstrate a new design strategy that could enable lower-cost, higher performance batteries.”The group is not done.

The overriding goal of …

1 week назад @ news.mit.edu
The benefits of medical AI assistance vary based on user expertise
The benefits of medical AI assistance vary based on user expertise The benefits of medical AI assistance vary based on user expertise

Explainable AI methods help users know when to trust a model’s predictions by describing or validating the model’s decision-making.

By contrast, clinicians were not tripped up by incorrect AI assistance and performed best when given only a model’s prediction, with no accompanying explanation.

“Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error.

They tested users by showing them medical images plus an AI prediction of skin disease, employing different explainable AI approaches.

We were just able to train very good AI models for this setting,” Ghassemi says.

1 week назад @ news.mit.edu
Alexander Rakhlin named director of the MIT Statistics and Data Science Center
Alexander Rakhlin named director of  the MIT Statistics and Data Science Center Alexander Rakhlin named director of the MIT Statistics and Data Science Center

Alexander “Sasha” Rakhlin PhD ’06, the Distinguished Professor in Data, Systems, and Society at the MIT Institute for Data, Systems, and Society (IDSS); and a professor of brain and cognitive sciences at MIT, has been named the next director of the MIT Statistics and Data Science Center (SDSC).

“The strength of the Statistics and Data Science Center has always been its people — students, postdocs, and faculty from across MIT who bring sharply different perspectives to the most interesting problems of the day in statistics, machine learning, and AI.

“At the Statistics and Data Science Center, I work alongside colleagues who share this fascination and pursue these connections in many directio…

1 week, 1 day назад @ news.mit.edu
Daniela Rus receives Bavarian Minister-President's High-Tech Prize
Daniela Rus receives Bavarian Minister-President's High-Tech Prize Daniela Rus receives Bavarian Minister-President's High-Tech Prize

Daniela Rus, director of MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) and the Panasonic Professor of Computer Science, has received the 2026 High-Tech Prize of the Bavarian Minister-President for her contributions to robotics, artificial intelligence, and autonomous systems.

The selection committee cited four strands of her work: self-organizing robot collectives, soft robotics, autonomous mobility, and brain-inspired artificial intelligence.

She is also a pioneer of soft robotics, where compliant machines manipulate the world more safely and adapt to it more readily than rigid ones can.

"Daniela Rus is a pioneer in soft robotics and physical AI," noted Lorenzo Masi…

1 week, 5 days назад @ news.mit.edu
Connecting research to policy on Capitol Hill
Connecting research to policy on Capitol Hill Connecting research to policy on Capitol Hill

This spring, 25 MIT students and postdocs traveled to Washington to meet with congressional staffers and advocate for sustained federal investment in scientific research.

With recent cuts to National Science Foundation programs and continued uncertainty surrounding the federal research budget, these conversations were especially timely.

Over the course of just two days, participants met with 62 congressional offices representing 32 states to discuss the importance of federal support for scientific research, higher education, and other policy concerns related to their individual research areas.

To prepare for the trip, participants attended three training sessions led by SPI in collaboration…

1 week, 5 days назад @ news.mit.edu
How a medical database developed at MIT evolved into a global standard of data-sharing
How a medical database developed at MIT evolved into a global standard of data-sharing How a medical database developed at MIT evolved into a global standard of data-sharing

Before the advancement of scientific data storage and collaboration via the cloud, medical investigators seeking health research breakthroughs had to overcome significant obstacles to collaboration and key clinical data gathering.

The data eventually became the first database of the global platform PhysioNet — founded in 1999 at the Harvard-MIT program in Health Sciences and Technology — as a clinical data repository for complex physiological signals.

In the years since PhysioNet was established, the value of sharing research data has gained much wider recognition.

Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet’s legacy.

In addition to using PhysioNet da…

1 week, 6 days назад @ news.mit.edu
Working to automate nuclear plant operations
Working to automate nuclear plant operations Working to automate nuclear plant operations

In pursuit of autonomous nuclear plant operationsIt turns out the research for the master’s was just the tip of the iceberg.

For the future viability of nuclear power, small plants, located in rural areas, are a distinct possibility.

It’s where supervised and thoroughly vetted autonomous operations will help.

A primary question was: “How do we transition to autonomous operations in nuclear power plants?” Fortier wanted one integrated approach, a central supervisory control system instead of many interlinked parts.

Using the nuclear plant automation program on next-generation equipment will deliver necessary traction in developing and deploying commercial microreactors.

2 weeks, 4 days назад @ news.mit.edu
MIT projects selected for funding under US Department of Energy’s Genesis Mission
MIT projects selected for funding under US Department of Energy’s Genesis Mission MIT projects selected for funding under US Department of Energy’s Genesis Mission

MIT researchers are set to contribute to the U.S. Department of Energy’s (DOE) Genesis Mission, with 15 collaborative projects among those selected for funding under Genesis Phase I, DOE announced Wednesday.

“MIT researchers are proud to be leading and contributing to projects under the Genesis Mission, in vital areas of research that support national priorities,” says Ian A. Waitz, MIT’s vice president for research.

Projects under the Genesis Mission are collaborative by design; teams must draw on the expertise of researchers from academia, industry, and/or the national laboratories.

Phase I projects that identify promising pathways toward transformative capabilities at scale may be consid…

2 weeks, 5 days назад @ news.mit.edu
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83 Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83

Over the course of his career, Bertsekas’ research spanned, and had a definitive influence upon, several fields, including optimization, control, large-scale computation, reinforcement learning, and artificial intelligence.

Along the way, Bertsekas taught, advised, and mentored students who would eventually become his colleagues at all four institutions.

“Dimitri played a defining role in my career,” says Asu Ozdaglar, department head of EECS at MIT.

Some referred to Dimitri as an “immortal.” Another comment I recall fondly — and often reminded Dimitri about — was: “Professor Bertsekas is a very handsome man!” Their bond continued long after Van Roy’s graduation.

He is survived by his wife …

2 weeks, 6 days назад @ news.mit.edu
Following the questions where they lead
Following the questions where they lead Following the questions where they lead

Ever since she was a child playing on her family’s farmland in Wisconsin, Bailey Flanigan was guided by her own selective, yet wide-ranging, curiosity.

“I found myself unmotivated to take all the AP [advanced placement] classes for the sake of it.

So Flanigan moved toward public health, where she researched microfluidic devices for HIV detection that could be used in low-resource settings.

After graduating from UW-Madison, Flanigan worked as a predoctoral research assistant in economics at Princeton.

“I feel so lucky to be studying these questions from within both political science and EECS, because I have the freedom to explore both the political and technical substance of tools for more d…

3 weeks, 4 days назад @ news.mit.edu
A better way to turn 2D designs into 3D models for rapid prototyping
A better way to turn 2D designs into 3D models for rapid prototyping A better way to turn 2D designs into 3D models for rapid prototyping

The system generates new data based on the model’s abilities as it attempts to convert a 2D image into a CAD program.

“Nearly every physical product around us, from airplanes to appliances, begins its life as a CAD model.

For guesses that are nearly correct, GIFT adjusts them to become successful solutions.

The CAD models generated by VLMs using GIFT were better aligned with the shapes of ground-truth models.

In the future, the researchers want to expand GIFT so the framework can teach models to generate CAD programs that improve the performance and manufacturability of 3D models.

3 weeks, 5 days назад @ news.mit.edu
3 Questions: Neural transparency and the future of AI design
3 Questions: Neural transparency and the future of AI design 3 Questions: Neural transparency and the future of AI design

Q: Your paper introduces “neural transparency,” a way to let everyday users peek inside an AI’s neural networks before their chatbot ever says a word.

“Neural transparency” means giving people something like a brain scan for AI.

Our study suggests that people have a blind spot when designing personalized AI.

In previous research, we documented cases of psychological harm associated with interactions with AI chatbots.

AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step.

3 weeks, 6 days назад @ news.mit.edu
Helping AI models to meet the real world
Helping AI models to meet the real world Helping AI models to meet the real world

“In a sense, with a small amount of resource, you have to do a lot of heavy lifting,” he says.

“My interest was: How does one design such graphical models for generic, tabular data?” he says.

And each of the products that you manufacture has lots of small pieces that come from different parts of the world.

Shah adds that Celonis has specialized in digitizing and automating operations for more than 1,400 large companies around the world.

“A narrower focus comes with sharper technology,” he says, “but it’s broad enough that it’s very valuable.”Shah adds, “The recent buzzword that’s become pertinent in the modern AI popular press is a ‘world model.’ In a sense, this is trying to build the ente…

4 weeks назад @ news.mit.edu
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering

“The JARVIS challenge showed that AI can substantially accelerate safety-critical hardware engineering, but engineering judgment remains the decisive differentiator.

Manufacturing — not engineering design or analysis — remained the fundamental rate-limiting step,” says Professor Zolti Spakovszky, director of the MIT Gas Turbine Laboratory.

In weekly progress reviews, they would critically evaluate the student progress and assess how the students were using AI.

The 811 team had been resistant to using AI throughout the competition, trusting instead to their fundamentals and teamwork.

From the start of the JARVIS Challenge, younger students used Parley more frequently and cleverly, while the …

4 weeks назад @ news.mit.edu
Berkeley AI
последний пост 1 week, 6 days назад
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map.

Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

With Apple Silicon in hundreds of millions of MacBooks and Mac Studios, MLX enables local AI inference without cloud costs.

Building an MLX backendTo bring K-Search to Apple Silicon, we first built a native MLX backend.

Evaluated on mamba-370m f16, M1 Max 64GB:Metric mlx-mamba (ours) mlx-lm (community) mamba.py Decode 152 tok/s 116 tok/s 40 tok/s Prefill L=512 5,751 tok/s 329 tok/s 1,089 tok/s Prefill L=1024 …

1 week, 6 days назад @ bair.berkeley.edu
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon InteractionOverview of ABBEL compared to traditional recursive summarization.

Belief grading can be thought of as adding an auxiliary RL task, using heuristics designed to capture what makes a good belief as the reward.

With domain-knowledge belief grading, ABBEL approaches or exceeds FULL CTX in this setting; without belief grading, learning is slower.

Context compression methods generate dense representations which, while computationally efficient, sacrifice human-understandability (Kontonis et al., 2026, Eyuboglu et al., 2025, Gupta et al., 2025, Chevalier et al., 2023, Deng et al., 2025, Deng et al., 2025, Bulatov et al., 2022).…

2 weeks, 2 days назад @ bair.berkeley.edu
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

1 month назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

1 month, 1 week назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

3 months назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

3 months, 3 weeks назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

5 months назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

7 months назад @ bair.berkeley.edu
RL without TD learning
RL without TD learning RL without TD learning

RL without TD learningIn this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer.

We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning.

There are two classes of algorithms in RL: on-policy RL and off-policy RL.

We compared TRL with $n$-step TD learning with different values of $n$, from $1$ (pure TD) to $\infty$ (pure MC).

I still think one of the most important problems in RL (and even in machine learning) is to find a scalable off-policy RL algorithm.

9 months, 1 week назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 5 часов назад
Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock
Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock

Daybreak Red and Daybreak Blue from OpenAI are now available on Amazon Bedrock to eligible customers.

This partnership brings Daybreak Red and Daybreak Blue from OpenAI to Amazon Bedrock.

Daybreak Red and Daybreak Blue resolve it through context: who is using the model, where the work occurs, and what safeguards govern that access.

Get startedDaybreak Red: GPT-5.6 Cyber and Daybreak Blue: GPT-5.6 Sol are now available to eligible customers on Amazon Bedrock in the following AWS Region: US East (N. Virginia).

To explore the full GPT-5.6 family on Amazon Bedrock, see Get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock.

5 часов назад @ aws.amazon.com
How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC
How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC

With technical advisory from GenAIIC, the company built Ishigaki-IDS, a foundation model (FM) specialized for construction industry BIM (Building Information Modeling) workflows.

Familiarity with foundation model training (pre-training and fine-tuning) and basic AWS compute concepts helps, but it isn’t required.

You will learn:How to use synthetic data generation to overcome data scarcity in niche domains.

Three challenges in building an IDS foundation modelThree problems stood between us and a working IDS model.

We stored training data, synthetic data, and checkpoints on Amazon FSx for Lustre, a fully managed file system optimized for compute-intensive workloads that delivers sub-milliseco…

10 часов назад @ aws.amazon.com
How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock
How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock

Using Amazon Bedrock, Pixieset moved from concept to production launch of an AI image alt text generator to millions of users in 4 months.

“We talked a lot about how this first AI feature must solve a real problem,” says Ry Rainey, Staff Engineer at Pixieset.

Categorize your AI features into moats and must-havesNot every AI feature deserves the same level of investment.

ConclusionThe most striking thing about Pixieset’s AI-generated alt text feature is how little infrastructure it required.

When you’re ready to move into multi-step agentic workflows, Amazon Bedrock AgentCore provides the managed infrastructure to scale from there.

10 часов назад @ aws.amazon.com
First Orion accelerates QA automation using Amazon Nova Act
First Orion accelerates QA automation using Amazon Nova Act First Orion accelerates QA automation using Amazon Nova Act

First Orion’s engineering teams were shipping faster than quality assurance (QA) could test, until Amazon Nova Act transformed QA automation.

First Orion’s QA team includes both QA Analysts and QA Automation Engineers.

How First Orion uses Nova ActFirst Orion’s main use case for Amazon Nova Act is testing their customer portal applications.

This runner is a Python application on Amazon Elastic Container Service (Amazon ECS) with AWS Fargate that orchestrates execution through the Amazon Nova Act SDK.

For a production-ready architecture guide, see Agentic QA Automation using Amazon Bedrock AgentCore Browser and Amazon Nova Act, and refer to the Amazon Nova Act documentation for more details.

10 часов назад @ aws.amazon.com
Deploying Anthropic Claude apps gateway for AWS for enterprise workloads
Deploying Anthropic Claude apps gateway for AWS for enterprise workloads Deploying Anthropic Claude apps gateway for AWS for enterprise workloads

Claude apps gateway provides a self-hosted governance layer between these applications and Amazon Bedrock or Claude Platform on AWS.

Pattern C: Hybrid Amazon Bedrock + Claude Platform on AWSBest for: organizations that want Amazon Bedrock as the preferred upstream with Claude Platform on AWS as overflow capacity.

upstreams: - name: claude-platform provider: anthropicAws region: us-east-1 workspace_id: wrkspc_01ABCDEFGHIJKLMN auth: api_key: ${ANTHROPIC_AWS_API_KEY} - name: bedrock provider: bedrock region: us-east-1 auth: {}Requests go to Amazon Bedrock first.

upstreams: - name: team-alpha provider: bedrock region: us-east-1 auth: aws_access_key_id: ${TEAM_ALPHA_AKID} aws_secret_access_key: …

11 часов назад @ aws.amazon.com
Run interactive IDEs on Amazon EKS with SageMaker AI to power up your AI workflows
Run interactive IDEs on Amazon EKS with SageMaker AI to power up your AI workflows Run interactive IDEs on Amazon EKS with SageMaker AI to power up your AI workflows

To power up AI workflows on Amazon Elastic Kubernetes Service (Amazon EKS), data scientists need interactive IDEs like JupyterLab and Code Editor.

The Amazon SageMaker AI Spaces add-on for Amazon EKS closes that gap.

In this post, you install the SageMaker AI Spaces add-on on an Amazon EKS cluster.

The Spaces add-on must be version 0.1.4 or later, because earlier versions supported Amazon SageMaker HyperPod only.

ConclusionIn this post, you installed the SageMaker AI Spaces add-on on an Amazon EKS cluster and configured browser and VS Code access.

1 day, 10 hours назад @ aws.amazon.com
How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore
How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore

nOps, an AI-powered cloud optimization solution, recently reimagined its Financial Operations (FinOps) analytics capabilities by transitioning to Amazon Bedrock AgentCore.

Amazon Bedrock AgentCore is a service to build, connect, and optimize agents at scale, with any framework or model.

Amazon Bedrock AgentCore backendThe runtime is deployed as a Docker container on Amazon Bedrock AgentCore, with the full infrastructure (runtime, memory, guardrails, queues, and worker functions) defined in a single AWS Cloud Development Kit (AWS CDK) stack.

Tool failure rate reduced from 7.49 percent to 0.92 percent after adopting Amazon Bedrock AgentCore with a new tool-calling method.

To learn more about …

1 day, 10 hours назад @ aws.amazon.com
How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore
How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore

To scale out representations in Cohere Policy Studio, Cohere Health added new skills to an existing AgentCore Runtime that was already decomposing policies.

The team completed three tasks:Deployed AgentCore Runtime with AgentCore Gateway and AgentCore Memory for a full agentic system using LangChain.

Wrote skills with clinical policy experts and evaluated them using Cohere Health’s standardized observability process based on Arize AI.

AgentCore Gateway architectureCohere Health implemented this using AgentCore Gateway with separate targets for shared tools and project-specific tools.

Cohere Health has digitized thousands of policies to date using manual and semi-automated workflows.

4 days, 10 hours назад @ aws.amazon.com
How TReNDS automates root-cause analysis with Amazon Bedrock
How TReNDS automates root-cause analysis with Amazon Bedrock How TReNDS automates root-cause analysis with Amazon Bedrock

When we started exploring Amazon Bedrock, we saw an opportunity we had wanted for a long time.

We use the Strands Agents SDK on top of Amazon Bedrock to handle tool-use orchestration.

PrerequisitesTo implement this solution, you need the following:An AWS account with access to Amazon Bedrock (specifically Anthropic Claude Sonnet).

An AWS Lambda function with appropriate IAM permissions to access Amazon Bedrock, CloudWatch Logs, AWS Secrets Manager, and SNS.

Choosing the right Amazon Bedrock modelAmazon Bedrock gives us access to a range of foundation models through a single API.

4 days, 10 hours назад @ aws.amazon.com
Determining playoff clinching scenarios in the NHL using constraint programming
Determining playoff clinching scenarios in the NHL using constraint programming Determining playoff clinching scenarios in the NHL using constraint programming

In this post, we describe how the AWS Generative AI Innovation Center created an automated system that determines NHL playoff clinching scenarios.

These complex tie-breakers are a major reason why determining playoff clinch scenarios is computationally challenging.

Determining 1-day clinch scenarios required a median runtime on the order of minutes, offering significant speedups over manual approaches (see Figure 2).

Automated, provably correct clinching scenarios that can be generated daily without manual effort, yielding significant time savings.

ConclusionDetermining NHL playoff clinching scenarios is a computationally hard problem.

4 days, 10 hours назад @ aws.amazon.com
Securing AI agents with temporal policies in Amazon Bedrock AgentCore
Securing AI agents with temporal policies in Amazon Bedrock AgentCore Securing AI agents with temporal policies in Amazon Bedrock AgentCore

Temporal policies in Amazon Bedrock AgentCore let you define stateful rules that determine authorization to AgentCore Gateway targets by evaluating the current request in the context of prior events in an agent’s trajectory.

In this post, you will learn what temporal policies are, how they work, and walk through an example to demonstrate.

As with the existing AgentCore Policy features, temporal policies deny by default and forbid wins over permit.

Applying temporal policies to a private banking portfolio agentTo make these concepts concrete, we’ll walk through how temporal policies can secure a hypothetical private banking agent.

To learn about AgentCore Gateway and how to set up auth with …

5 days, 8 hours назад @ aws.amazon.com
Configure rate limits for AI traffic on AgentCore gateway
Configure rate limits for AI traffic on AgentCore gateway Configure rate limits for AI traffic on AgentCore gateway

Amazon Bedrock AgentCore gateway is a fully managed, serverless AI gateway that provides a single, secure entry point for AI traffic.

Rate limit structureA rate limit configuration consists of two parts: dimension keys and entries.

Types of rate limits and example configurationsAgentCore gateway enforces two layers of rate limiting: customer-defined rate limits and Service Quotas.

We demonstrated how to layer rate limits: user-level limits for per-role and per-user fairness, target-level limits for protecting downstream service capacity, and multi-dimensional limits that combine user and target to enforce fine-grained rate limits.

To get started with rate limits on AgentCore gateway, explor…

5 days, 9 hours назад @ aws.amazon.com
Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore
Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore

According to McKinsey, roughly 80% of organizations have already encountered risky behavior from AI agents.

As a result, security and risk concerns are the leading barrier to scaling agentic AI (McKinsey’s State of AI Trust in 2026, and Trust in the age of AI agents 2026).

We built Amazon Bedrock AgentCore to give teams what they need to build, connect, and optimize agents at scale without assembling the infrastructure themselves.

Powering temporal policies is Dogwood, a new policy language purpose-built for AI agents.

Built on the foundation of Cedar, Dogwood was designed to address a new dimension of agent control: evaluating whether a sequence of agent actions conforms to a policy as it …

5 days, 10 hours назад @ aws.amazon.com
Build visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch
Build visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch Build visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch

The reference deployment creates a CloudWatch dashboard, not an Amazon Elastic Container Service (Amazon ECS) service, load balancer, virtual private cloud (VPC), or public ingestion endpoint.

Enable the CloudWatch OTel capabilitiesThe reference runbook begins by enabling OTel enrichment and resource tags for telemetry in the target Region:aws cloudwatch start-otel-enrichment --region us-west-2 aws observabilityadmin start-telemetry-enrichment --region us-west-2 aws cloudwatch get-otel-enrichment --region us-west-2The current CloudWatch OTel documentation describes native OTLP ingestion and PromQL querying.

CloudWatch OTel metrics use per-gigabyte ingestion pricing, and PromQL queries are p…

5 days, 10 hours назад @ aws.amazon.com
Enforcing data residency with single-Region Claude Code on Amazon Bedrock
Enforcing data residency with single-Region Claude Code on Amazon Bedrock Enforcing data residency with single-Region Claude Code on Amazon Bedrock

A US-headquartered global organization recently came to us with a deceptively simple data-residency request: let their engineers use Claude Code.

The requirement: Amazon Bedrock model inference had to be processed in London (the eu-west-2 AWS Region), not merely called from London.

Classic Amazon Bedrock (bedrock-runtime) is the original Amazon Bedrock Invoke API.

Verifying single-Region complianceEvery Claude Code Amazon Bedrock call must appear in the target Region’s AWS CloudTrail, and nowhere else.

Learn more about inference profiles, cross-Region inference, and Claude Code on Amazon Bedrock in their respective documentation.

5 days, 10 hours назад @ aws.amazon.com
NVIDIA
последний пост 2 часа назад
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

It is a complete AI factory platform including accelerated computing, networking, systems software, AI frameworks and a global developer ecosystem.

NVIDIA DSX AI factories can run the world’s broadest range of AI models, modalities and algorithms — language, vision, speech, biology, physical AI and robotics.

One NVIDIA AI factory can serve many customers and many workloads.

NVIDIA can provide support because NVIDIA compute is unique: it is fungible, universally adopted, software-upgradable and redeployable across a large ecosystem of customers.

More compute creates better AI; better AI creates more usage; more usage creates more revenue; and more revenue drives more compute.

2 часа назад @ blogs.nvidia.com
Why Scaling AI Compute Performance Requires a New Power Architecture
Why Scaling AI Compute Performance Requires a New Power Architecture Why Scaling AI Compute Performance Requires a New Power Architecture

At the power levels that next-generation AI compute demands, even small inefficiencies compound quickly.

“800 VDC unlocks the compute performance and power density required for AI at scale,” said Vladimir Troy, vice president of data center infrastructure at NVIDIA.

An Open Standard Means a Real Supply ChainThe 800 VDC architecture specifications define common interfaces so that power hardware from different vendors can work together inside the same 800 VDC facility.

The facilities that can absorb that investment will be the ones that resolved their power architecture before compute demand outran what their infrastructure could deliver.

Read the 800 VDC OCP blog and the NVIDIA 800 VDC white…

12 часов назад @ blogs.nvidia.com
NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents
NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem.

That includes NVIDIA’s latest open models, software and developer tools, plus the accelerated computing, libraries and educational resources that help users get started.

Follow along for the latest developments in this special-edition NVIDIA Local AI blog series, with new updates added over the coming weeks.

Follow NVIDIA Workstation on LinkedIn and X.Tuesday, Aug. 11, 6:00 a.m. PT 🔗NVIDIA Introduces Nemotron 3.5 Lightning for Fast, Specialized Agentic TasksToday, NVIDIA expanded its Nemotron 3 model family wi…

14 часов назад @ blogs.nvidia.com
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads.

Built for specialized tasks within larger multi-agent systems, Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, helps create smarter and more efficient agentic applications.

Together, Nemotron 3.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates — across PCs, workstations, data centers and the cloud.

Powering High-Volume Specialized Tasks With Nemotron 3.5 LightningNVIDIA Nemotron 3.5 Lightning is a fully customizable open model b…

14 часов назад @ blogs.nvidia.com
Firebird Launches CIS Region’s Largest AI Factory in Armenia
Firebird Launches CIS Region’s Largest AI Factory in Armenia Firebird Launches CIS Region’s Largest AI Factory in Armenia

The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia, establishing a new AI computing hub powered by NVIDIA accelerated computing and Dell Technologies high-performance AI infrastructure.

Firebird’s AI factory brings that capacity to Armenia, giving developers, startups, enterprises, universities and public institutions the compute to build and scale AI at home.

Firebird’s AI factory is designed from the ground up to turn compute into revenue.

Delivered in just over six months, the Armenia AI factory demonstrates Firebird’s ability to turn ambitious infrastructure plans into operation…

3 days, 16 hours назад @ blogs.nvidia.com
GeForce NOW Shakes Up August With 26 New Games
GeForce NOW Shakes Up August With 26 New Games GeForce NOW Shakes Up August With 26 New Games

August is here, bringing 26 new games for GeForce NOW members.

Command the seas in World of Warships: Legends and discover what’s next in the GeForce NOW library, starting with the eight newly added games this week.

In addition, GeForce NOW is at the QuakeCon gaming conference this week in Grapevine, Texas, with hands-on experiences awaiting attendees.

Gamers not at the show can try out Ultimate cloud gaming in action with a day pass and jump into Bethesda titles from any device.

All Games on DeckWorld of Warships: Legends drops anchor on GeForce NOW this week, bringing free-to-play naval combat.

5 days, 14 hours назад @ blogs.nvidia.com
Into the Omniverse: How Open World Models Push the Frontier of Physical AI
Into the Omniverse: How Open World Models Push the Frontier of Physical AI Into the Omniverse: How Open World Models Push the Frontier of Physical AI

Open world models are already being used to generate training data, test policies and specialize physical AI systems.

World Models Are the Foundation of Physical AIThe data behind physical AI is difficult and expensive to collect at the scale required.

The NVIDIA Cosmos Coalition extends this work by bringing together world model builders, AI developers and physical AI leaders to contribute models, research and evaluation methods.

Together, these implementations and collaborations are establishing open world models as an adaptable foundation for physical AI across robots, autonomous vehicles and vision AI systems.

Get Plugged InLearn more about world models, OpenUSD and physical AI developm…

5 days, 14 hours назад @ blogs.nvidia.com
NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US
NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US

These regional hubs will help institutions share AI computing resources, accelerate scientific discovery and innovation, and prepare students to participate in the AI economy.

Expanding Access to AI InfrastructureThe State and Regional AI Infrastructure Hubs program will bring shared resources closer to the institutions and communities they serve.

And since 2017, UF faculty and units have received more than $511 million in AI research awards.

Connecting Research, Workforce and Regional GrowthFor policymakers and leaders, the hubs offer an opportunity to connect regional research and educational infrastructure with regional priorities and broader workforce and economic-development strategies…

1 week назад @ blogs.nvidia.com
NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use
NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

NVIDIA Alpamayo 2 Super, available now for commercial use, is part of the Alpamayo family, the most-adopted open reasoning models for autonomous driving on Hugging Face, supporting a wide range of AV-relevant capabilities within a single foundation model.

Within the Alpamayo model family, Alpamayo 2 Super delivers the highest reasoning and driving performance for multimodal autonomous driving development, while Alpamayo 1.5 and Alpamayo 1 provide more cost-efficient options for cloud-based development and model distillation.

Together, the Alpamayo model family provides a cloud-to-car workflow that combines frontier-scale reasoning with scalable deployment across commercial AV fleets.

Benchm…

1 week назад @ blogs.nvidia.com
As AI Increases Demands on Memory, Storage Steps Up
As AI Increases Demands on Memory, Storage Steps Up As AI Increases Demands on Memory, Storage Steps Up

Benchmarks highlighted in this NVIDIA technical blog show that the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.

Storage-Next includes over 40 leading storage and flash vendors — including DDN, KIOXIA and Micron — each contributing to the next generation of AI storage technologies with NVIDIA.

Plus, NVIDIA CMX Context Memory Storage provides an AI‑native context tier for long‑context, multi‑turn, agentic AI inference, built on NVIDIA STX.

SCADA Enables Fast AI Storage That Stays SecureSpeed at the storage layer comes with a catch.

Join NVIDIA sessions at FMS, running Aug. 4-6 i…

1 week назад @ blogs.nvidia.com
AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency
AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat conference begins in Las Vegas today.

The SAFE guidelines are being drafted by an Open Secure AI Alliance working group.

Open Secure AI Alliance Delivers More Tools for AI CybersecurityThe SAFE framework adds to technology contributions Open Secure AI Alliance members are making as part of a shared commitment to building and sharing open, inspectable tools across the full AI security stack.

Alliance members are contributing tooling, harnesses and supporting technologies across this emerging layer of the AI security sta…

1 week назад @ blogs.nvidia.com
Run High-Performance Core Math at Scale with NVIDIA nvmath-python
Run High-Performance Core Math at Scale with NVIDIA nvmath-python Run High-Performance Core Math at Scale with NVIDIA nvmath-python

NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries.

A useful complement to existing array librariesLike other math libraries such as NumPy, nvmath-python implements core numerical operations useful in many engineering and scientific computing applications.

In the following example, nvmath-python consumes NumPy arrays and the result is also a NumPy array.

CPU libraries such as NVPL for NVIDIA Grace or any ARM v8 CPUs and Intel MKL for x86 hosts.

Get started with nvmath-pythonDesigned for productivity without performance compromises, nvmath-python reimagines the design of modern math libraries.

1 week, 5 days назад @ developer.nvidia.com
Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW
Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW

With cloud gaming, everyday laptops used for class can also become GeForce RTX-powered gaming setups.

With a GeForce NOW membership, Chromebooks, Macs and Surface devices can become GeForce RTX-powered gaming PCs in the cloud.

The GeForce NOW library features thousands of supported PC games, letting members stream titles they already own from digital stores like Steam, Xbox PC Game Pass and Epic Games Store in just a few clicks.

GeForce NOW turned the laptop — usually the creator’s go-to for editing videos and everyday work — into a top-notch gaming machine, too.

Everyone, at Their StationsExperience the Master Chief’s iconic journey on GeForce NOW in Halo: Campaign Evolved — a modernized r…

1 week, 5 days назад @ blogs.nvidia.com
Powerful Compute So Compact, It’s Clutch — Build AI in Your Hand With NVIDIA Jetson
Powerful Compute So Compact, It’s Clutch — Build AI in Your Hand With NVIDIA Jetson Powerful Compute So Compact, It’s Clutch — Build AI in Your Hand With NVIDIA Jetson

NVIDIA Jetson Orin Nano Super: Ideal for Building a First AI RobotRobots deserve a real brain and, apparently, an incredibly chic commute.

Jetson Orin Nano Super brings desktop-class generative AI to a handbag-friendly developer kit, offering first-time builders a practical path to learning computer vision, building AI agents, prototyping edge AI and more.

Reachy Mini Jetson AssistantReachy Mini Jetson Assistant is a low-latency, fully on-device voice and vision assistant for Reachy Mini Lite powered by NVIDIA Jetson Orin Nano Super.

Robotics AI PodcastUsing Jetson Orin Nano Super, Asier Arnaz built a Yocto-powered Robotics AI video podcast featuring two AI models discussing topics in real-…

2 weeks назад @ blogs.nvidia.com
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

Ising Calibration 1.5 also uses examples from related experiments when available and is 11.4% smaller at BF16 precision.

How is the Ising Calibration 1.5 model trained?

Get started with NVIDIA Ising open resourcesThe NVIDIA Ising model family is fully open.

Model weightsFull-parameter checkpoints for Ising Calibration 1.5 are available on Hugging Face:Ising Calibration 1.5 is also available as an NVIDIA NIM and hosted through NVIDIA Build.

Quantum calibration agent blueprint is a script for deploying an agentic workflow using Ising Calibration 1.5 with the NVIDIA Nemo Agent Toolkit to quickly set up quantum calibration experiment automation.

2 weeks, 1 day назад @ developer.nvidia.com
Facebook
последний пост 6 days, 7 hours назад
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Introducing the Multi-Stage Sequence ModelTo address scaling efficiency, a multi-stage model has been developed that enables scaling of a transformer-based sequence model in a compute efficient manner.

Second Stage: Online Ranking ModelThe offline user model representations are complemented with online ranking models that use fresh user signals and ad candidate information for real time ranking.

A Predictable Scaling CurveLLM-Style Scaling LawWhen running on real-world ads traffic, the multi-stage sequence model demonstrates the emergence of predictable scaling laws for ads recommendations that are analogous to those observed in large language models.

The Impact of Multi-Stage Sequence Mode…

6 days, 7 hours назад @ engineering.fb.com
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

We tackled these challenges through complementary compute efficiency and scaling efficiency innovations: Compute efficiency : Achieved through a customized recommendation kernel library — Jagged Flash Attention (JFA), Generalized Dot-Product Attention (GDPA), BlockAttention, etc.

The results: we doubled GEM’s E2E training efficiency to 20-25% MFU while scaling total training FLOPs 4x over the past 12 months.

We measure training efficiency through E2E MFU, which decomposes into two factors:E2E MFU = Local MFU (compute efficiency) × Scaling Ratio (scaling efficiency)These factors describe two related but distinct optimization problems.

Local MFU (compute efficiency) measures how well a single…

1 week, 1 day назад @ engineering.fb.com
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

3 weeks, 6 days назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

4 weeks, 1 day назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

1 month, 1 week назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

1 month, 2 weeks назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

2 months, 2 weeks назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

3 months назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

3 months, 3 weeks назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

3 months, 3 weeks назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

4 months, 1 week назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

4 months, 1 week назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

4 months, 1 week назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

4 months, 2 weeks назад @ engineering.fb.com
Friend Bubbles: Enhancing Social Discovery on Facebook Reels
Friend Bubbles: Enhancing Social Discovery on Facebook Reels Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend bubbles in Facebook Reels highlight Reels your friends have liked or reacted to, helping you discover new content and making it easier to connect over shared interests.

Friend bubbles enhance the social experience on Facebook Reels by helping you discover content your friends enjoy, creating a shared viewing experience and sparking new conversations.

Along with additional optimizations in the underlying method, this approach enabled us to ship friend bubbles while preserving core Reels performance.

Friend bubbles work because the signal is high value: It adds meaningful social context that helps people decide what’s worth watching.

Engagement also scales consistently with the number …

4 months, 3 weeks назад @ engineering.fb.com
Uber Engineering
последний пост None
neptune.ai neptune.ai
последний пост 8 months, 1 week назад
We are joining OpenAI
We are joining OpenAI We are joining OpenAI

Piotr Niedźwiedź, CEO/CTO and founder of neptune.aiI’m excited to share that we’ve entered into a definitive agreement to be acquired by OpenAI, subject to closing conditions.

We are thrilled to join the OpenAI team and help their AI researchers build better models faster.

Neptune is a metrics dashboard company.”We’ve worked closely with OpenAI to create the metrics dashboard that helps teams building foundation models.

Our future with OpenAINeptune will join OpenAI and continue to support AI researchers with tools to monitor, debug, and evaluate frontier models.

We are looking forward to working with top AI researchers and supporting OpenAI’s mission of ensuring that AGI benefits all of hu…

8 months, 1 week назад @ neptune.ai
Synthetic Data for LLM Training
Synthetic Data for LLM Training Synthetic Data for LLM Training

For instance, financial data is highly sensitive and protected by very strict regulations, and synthetic data mimics the real data distribution without revealing customer information.

Read more about how leading foundation model teams curate their training data and other topics in the State of Foundation Model Training Report 2025.

Choosing the right synthetic data generation technique depends on the type of data and its complexity.

Synthetic tabular data generation is a promising direction to overcome these challenges by learning the distribution of the tabular data.

Post-processingAs the distribution of tabular data is highly complex, it makes the synthetic tabular data generation very ch…

9 months назад @ neptune.ai
What are LLM Embeddings: All you Need to Know
What are LLM Embeddings: All you Need to Know What are LLM Embeddings: All you Need to Know

TL;DR LLM embeddings are the numerical, vector representations of text that Large Language Models (LLMs) use to process information.

Unlike their predecessor word embeddings, LLM embeddings are context-aware and dynamically change to capture semantic and syntactic relationships based on the surrounding text.

What are the applications of LLM embeddings?

Word EmbeddingsSparse Word Embeddings One-Hot Vectors 1970s TF-IDF1980s Co-Occurrence MatrixStatic Word Embeddings Word2Vec 2013 GloVe 2014Contextualized word embeddings ELMo 2018 GPT-1 2018 BERT 2018 LLAMA 2023 DeepSeek-V1 2023 GPT-4 2023Static word embeddingsStatic word embeddings, such as word2vec in 2013, marked a significant development.…

9 months, 1 week назад @ neptune.ai
Detecting and Fixing ‘Dead Neurons’ in Foundation Models
Detecting and Fixing ‘Dead Neurons’ in Foundation Models Detecting and Fixing ‘Dead Neurons’ in Foundation Models

TL;DR Dead neurons silently waste compute and reduce effective model capacity in foundation models.

Dead neurons’ impactRecent studies into dead neurons in the context of foundation models show interesting, albeit worrying, results.

These large reported fractions of dead neurons in foundation models are a concern from a computational perspective.

Before we move on to discuss how to detect and fix dead neurons, let’s touch upon an important distinction between dead neurons and vanishing gradients.

Further reading How to Monitor, Diagnose, and Solve Gradient Issues in Foundation Models Read moreVisualizing activation distributionsIs your foundation model suffering from dead neurons?

9 months, 2 weeks назад @ neptune.ai
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training

In the first part of this series, we covered the fundamentals of instruction fine-tuning (IFT).

def calculate_irs(instruction, output, reference_model): evaluation_prompt = f""" Instruction: {instruction} Model Output: {output} Rate how well the output follows the instruction on these criteria: 1.

| SourceHINT addresses a computational inefficiency in standard instruction fine-tuning: repeatedly reprocessing the same task instruction with every input example.

Read more about foundation model training infrastructure and other topics in Neptune’s 2025 State of Foundation Model Training Report.

First, during initial instruction fine-tuning across multiple diverse tasks, the model learns genera…

9 months, 3 weeks назад @ neptune.ai
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 5 months, 1 week назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

5 months, 1 week назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

7 months, 2 weeks назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

7 months, 2 weeks назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

7 months, 2 weeks назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

8 months назад @ youtube.com
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff) [Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)

https://arxiv.org/abs/2510.17558 Abstract:

We propose an extension of the decoder Transformer that conditions its generative process on random latent variables which are learned without supervision thanks to a variational procedure. Experimental evaluations show that allowing such a conditioning translates into substantial improvements on downstream tasks. Author: François Fleuret Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the con…

9 months, 1 week назад @ youtube.com
[Video Response] What Cloudflare's code mode misses about MCP and tool calling
[Video Response] What Cloudflare's code mode misses about MCP and tool calling [Video Response] What Cloudflare's code mode misses about MCP and tool calling

Theo's Video: https://www.youtube.com/watch?v=bAYZjVAodoo

Cloudflare article: https://blog.cloudflare.com/code-mode/ Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8…

9 months, 3 weeks назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 2 weeks, 4 days назад
The 64 sugar cubes puzzle
The 64 sugar cubes puzzle The 64 sugar cubes puzzle

See all monthly puzzles: https://momath.org/mindbenders/

2 weeks, 4 days назад @ youtube.com
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

3 weeks, 5 days назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

1 month, 3 weeks назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

2 months назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

2 months назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

2 months назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

2 months, 3 weeks назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

3 months, 3 weeks назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

4 months, 2 weeks назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

4 months, 2 weeks назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

4 months, 3 weeks назад @ youtube.com
Bacteria Grid Puzzle Solution
Bacteria Grid Puzzle Solution Bacteria Grid Puzzle Solution

Part of a monthly series of puzzlers, in collaboration with MoMath and Peter Winkler

4 months, 3 weeks назад @ youtube.com
The most underappreciated formula | Exploring high-dimensional spheres
The most underappreciated formula | Exploring high-dimensional spheres The most underappreciated formula | Exploring high-dimensional spheres

On the volumes of higher-dimensional spheres

Explore the 3b1b virtual career fair: See https://3b1b.co/talent

Become a supporter for early views of new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Thanks to UC Santa Cruz for letting me film there, and special thanks to Pedro Morales-Almazan for arranging everything. My video on Numberphile with a fun application of this problem: https://youtu.be/6_yU9eJ0NxA Timestamps:

0:00 - Introduction

1:01 - Random puzzle

6:16 - Outside the box

14:35 - Setting up the volume grid

21:14 - Why 4πr^2

25:21 - Archimedes in higher dimensions

36:17 - The general formul…

5 months, 2 weeks назад @ youtube.com
The lattice bacteria puzzle
The lattice bacteria puzzle The lattice bacteria puzzle

Part of a series of monthly puzzles, done in collaboration with MoMath.

https://momath.org/mindbenders

5 months, 3 weeks назад @ youtube.com
Solution to the ladybug clock puzzle
Solution to the ladybug clock puzzle Solution to the ladybug clock puzzle

Solution to last month's probability puzzle.

5 months, 3 weeks назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 11 часов назад
OpenAI’s AI Escaped And It's Terrifying
OpenAI’s AI Escaped And It's Terrifying OpenAI’s AI Escaped And It's Terrifying

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 More reports are available here:

https://openai.com/index/hugging-face-model-evaluation-security-incident/

https://huggingface.co/blog/security-incident-july-2026

https://huggingface.co/blog/agent-intrusion-technical-timeline 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

11 часов назад @ youtube.com
DeepMind's AI Trick Everyone Should Copy
DeepMind's AI Trick Everyone Should Copy DeepMind's AI Trick Everyone Should Copy

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Gemma4 paper and some more is available here:

https://arxiv.org/abs/2607.02770

https://x.com/googlegemma/status/2077449152062247219

https://x.com/UnslothAI/status/2078118183085731843 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

4 days, 18 hours назад @ youtube.com
The Billion Dollar AI Race Just Broke
The Billion Dollar AI Race Just Broke The Billion Dollar AI Race Just Broke

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 Qwen 3.8 Max:

https://qwen.ai/blog?id=qwen3.8 Sources:

https://x.com/loktar00/status/2082589566934929750

https://x.com/CommandCodeAI/status/2084293498950590839 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

6 days, 13 hours назад @ youtube.com
Another DeepSeek Moment
Another DeepSeek Moment Another DeepSeek Moment

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek v4 Flash 0731:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 1 day назад @ youtube.com
New AI Learned Parkour From Just 30 Seconds Of Video
New AI Learned Parkour From Just 30 Seconds Of Video New AI Learned Parkour From Just 30 Seconds Of Video

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://jiashunwang.github.io/HIL/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 2 days назад @ youtube.com
Kimi K3 Just Broke The Economics Of AI
Kimi K3 Just Broke The Economics Of AI Kimi K3 Just Broke The Economics Of AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2607.24653 Try Kimi K3 (subject to availability): https://www.kimi.com/ Links:

https://macos27.kimi.page/

https://x.com/mweinbach/status/2077878247920951400

https://x.com/intheworldofai/status/2077838911494336681

https://x.com/chetaslua/status/2077829183989072281

https://x.com/hqmank/status/2078104317027094907 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen …

1 week, 6 days назад @ youtube.com
AI Helped Them Code Faster… But At A Cost
AI Helped Them Code Faster… But At A Cost AI Helped Them Code Faster… But At A Cost

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/AI-assistance-coding-skills 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 weeks, 5 days назад @ youtube.com
The Hidden World Inside An AI
The Hidden World Inside An AI The Hidden World Inside An AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://transformer-circuits.pub/2025/linebreaks/index.html Paper for reindeer vision change - https://royalsocietypublishing.org/rspb/article/280/1773/20132451/50765/Shifting-mirrors-adaptive-changes-in-retinal 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli …

3 weeks, 6 days назад @ youtube.com
New AI Just Reinvented Minecraft Worlds
New AI Just Reinvented Minecraft Worlds New AI Just Reinvented Minecraft Worlds

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://xandergos.github.io/terrain-diffusion/

https://modrinth.com/mod/terrain-diffusion

https://github.com/xandergos/terrain-diffusion Source video for some parts of the footage: https://www.youtube.com/watch?v=irE4tcDtUIg 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fi…

1 month назад @ youtube.com
DeepSeek's New AI Speed Hack Is Amazing
DeepSeek's New AI Speed Hack Is Amazing DeepSeek's New AI Speed Hack Is Amazing

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The DeepSeek paper is available here:

https://arxiv.org/abs/2607.05147v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
Game Physics Just Got 170 Times Faster
Game Physics Just Got 170 Times Faster Game Physics Just Got 170 Times Faster

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://arxiv.org/abs/2506.06494 Sources:

https://www.youtube.com/shorts/Tx7167DXr8U

https://www.youtube.com/watch?v=55F9dY2Y1zc 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
This New AI Model Changes Everything
This New AI Model Changes Everything This New AI Model Changes Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers GLM 5.2: https://z.ai/blog/glm-5.2 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
DeepSeek Just Solved AI's Billion Dollar Problem
DeepSeek Just Solved AI's Billion Dollar Problem DeepSeek Just Solved AI's Billion Dollar Problem

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2602.21548 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi #deepseek

1 month, 2 weeks назад @ youtube.com
This is OpenClaw On Steroids
This is OpenClaw On Steroids This is OpenClaw On Steroids

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://recursivemas.github.io/

https://github.com/RecursiveMAS/RecursiveMAS 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://felicia.hu

1 month, 3 weeks назад @ youtube.com
Claude AI Knows More Than It Tells You
Claude AI Knows More Than It Tells You Claude AI Knows More Than It Tells You

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/natural-language-autoencoders

https://transformer-circuits.pub/2026/nla/index.html 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 3 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 1 week назад
ML Global Recap'H1 2026
ML Global Recap'H1 2026 ML Global Recap'H1 2026

Обсудим итоги ICML и других международных конференций, главные ML-тренды первого полугодия 2026-го и собственный опыт.

1 week назад @ youtube.com
Омни-модели будущего 🚀
Омни-модели будущего 🚀 Омни-модели будущего 🚀

Что они будут уметь — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 1 week назад @ youtube.com
Качество модели взлетело... без мультимодального RL?
Качество модели взлетело... без мультимодального RL? Качество модели взлетело... без мультимодального RL?

О росте мультимодального качества рассказал Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 1 week назад @ youtube.com
Работа с данными — это скучно?
Работа с данными — это скучно? Работа с данными — это скучно?

А почему — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 2 weeks назад @ youtube.com
Как приготовить SFT 🍲
Как приготовить SFT 🍲 Как приготовить SFT 🍲

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 2 weeks назад @ youtube.com
Почему мультимодальные модели — это база 🤖
Почему мультимодальные модели — это база 🤖 Почему мультимодальные модели — это база 🤖

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 3 weeks назад @ youtube.com
Омни-модель: что это за зверь такой
Омни-модель: что это за зверь такой Омни-модель: что это за зверь такой

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 3 weeks назад @ youtube.com
Borealis — как обучить аудио-LLM по цене MacBook
Borealis — как обучить аудио-LLM по цене MacBook Borealis — как обучить аудио-LLM по цене MacBook

На конференции Data Fest 2026 в Белграде независимый исследователь Александр Николич рассказал практическую историю создания аудиоязыковой модели Borealis с бюджетом, сопоставимым со стоимостью MacBook. Больше контента для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

2 months назад @ youtube.com
Better LLM pre-training in NVFP4
Better LLM pre-training in NVFP4 Better LLM pre-training in NVFP4

At Data Fest 2026 in Belgrade, Andrei Panferov from the Institute of Science and Technology Austria introduced Quartet II, a novel method for NVFP4 pre-training that recovers SOTA accuracy. He outlined the core challenges of low-precision LLM training and presented CUDA kernels tuned for Blackwell GPUs, ready for integration into real training pipelines. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

2 months назад @ youtube.com
Как безопасно выкатывать новые версии продуктовых AI-агентов
Как безопасно выкатывать новые версии продуктовых AI-агентов Как безопасно выкатывать новые версии продуктовых AI-агентов

На Data Fest 2026 в Белграде Дмитрий Коршунов, Team Lead ML в Ecom, показал, как безопасно обновлять продуктовых AI-агентов с помощью системы автометрик. На примере агента Яндекс AI для турецкого рынка он объяснил, как фиксировать регрессии до прода, сравнивать версии и принимать решение о релизе, когда простой «Hello, Agent» уже позади. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

2 months назад @ youtube.com
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents

At Data Fest 2026 in Belgrade, Karina Romanova, Senior LLM Research Engineer, presented HGRPO — a hierarchical modification of GRPO for multi-turn dialogue agents. Applied to a booking agent in Yandex Alice, the method improved truthfulness by 8.0 percentage points and reduced dialogue length by 10.7%. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

2 months назад @ youtube.com
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей

На Data Fest 2026 в Белграде Вячеслав Костров, ML-инженер в Яндексе, рассказал, как uplift-модели решают бизнес-задачи Лавки: от персональных скидок до показа продуктовых подборок. Он разобрал постановку uplift-задачи, подбор метрик и построение политик, а также практические приёмы с лагранжианом и uplift-деревьями для баланса ограничений. Всё это — на примере реальных внедрений и с разбором типичных ошибок. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #Mu…

2 months назад @ youtube.com
Поиск по архивам: как мы переходим к осознанному распознаванию текста
Поиск по архивам: как мы переходим к осознанному распознаванию текста Поиск по архивам: как мы переходим к осознанному распознаванию текста

На Data Fest 2026 в Белграде Дарья Виноградова, лид команды компьютерного зрения, представила два важных майлстоуна архивного поиска: новую архитектуру распознавания текста и выделение смысловых структур. Эти изменения делают поиск человечнее — теперь можно искать не слова среди текста, а человека среди людей. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

2 months назад @ youtube.com
Hacks and Defenses in Automatic Kernel Generation
Hacks and Defenses in Automatic Kernel Generation Hacks and Defenses in Automatic Kernel Generation

На Data Fest 2026 в Белграде Егор Коновалов, ML-инженер, разобрал хаки, которые находят LLM-агенты, когда генерируют GPU/TPU-код: от тривиального обхода numerical tolerance до изощрённых атак на timing-измерения и эксплуатации дыр в test harness. А ещё Егор показал, какие методы защиты реально работают, а какие создают ложное чувство безопасности. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

2 months назад @ youtube.com
Real-time video generation: where we are and what comes next
Real-time video generation: where we are and what comes next Real-time video generation: where we are and what comes next

At Data Fest 2026 in Belgrade, Andrey Filatov from KREA AI broke down the current state of real-time video generation: which architectures dominate, how they differ, and what challenges arise from compute limits and memory bottlenecks. He also covered production solutions like distillation and caching, and shared his outlook for the next 2–3 years: what will soon become possible and which bottlenecks the industry still overlooks. More content for developers: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #Reinforceme…

2 months назад @ youtube.com
ML Trainings ML Trainings
последний пост 1 day, 8 hours назад
Музыка как лудомания: Дмитрий о предсказуемости и...
Музыка как лудомания: Дмитрий о предсказуемости и... Музыка как лудомания: Дмитрий о предсказуемости и... 1 day, 8 hours назад @ youtube.com
Драйверы, пастеры и шейперы: будущее технологий
Драйверы, пастеры и шейперы: будущее технологий Драйверы, пастеры и шейперы: будущее технологий 1 day, 8 hours назад @ youtube.com
Как ИИ изменит профессии, как в автомобилестроении
Как ИИ изменит профессии, как в автомобилестроении Как ИИ изменит профессии, как в автомобилестроении 1 day, 8 hours назад @ youtube.com
Музыканты и программисты в эпоху искусственного интеллекта
Музыканты и программисты в эпоху искусственного интеллекта Музыканты и программисты в эпоху искусственного интеллекта 1 day, 8 hours назад @ youtube.com
Валентин Малых о разделении музыки и искусственного интеллекта
Валентин Малых о разделении музыки и искусственного интеллекта Валентин Малых о разделении музыки и искусственного интеллекта 1 day, 8 hours назад @ youtube.com
Искусственный интеллект и будущее музыкальной индустрии
Искусственный интеллект и будущее музыкальной индустрии Искусственный интеллект и будущее музыкальной индустрии 1 day, 8 hours назад @ youtube.com
Капитанский мостик 09.08.2026: Хассабис и Дин ушли из Google | ИИ-музыка от Кион | робогуртовщики
Капитанский мостик 09.08.2026: Хассабис и Дин ушли из Google | ИИ-музыка от Кион | робогуртовщики Капитанский мостик 09.08.2026: Хассабис и Дин ушли из Google | ИИ-музыка от Кион | робогуртовщики

00:00:00 начало

00:00:33 Sakana управляет кошельком

00:05:56 Хассабис и Дин ушли из Google

00:20:38 беспилотный Uber в Лондоне

00:27:02 вышел Muse Code

00:34:35 суверенная модель от Sakana

00:43:16 Anthropic делает чипы

00:46:22 ИИ-музыка от Кион

01:03:59 ЕС против Claude

01:08:18 дроноводы и робогуртовщики

01:13:03 ИИ-стартапы-единороги

01:20:33 шрифт против ИИ ИИ-саммари: В этом выпуске мы обсуждаем последние новости в области технологий, включая развитие AI, запуск роботакси в Лондоне и уход ключевых фигур из крупных технологических компаний. Узнайте, как эти события влияют на индустрию и что ждать дальше. В этом выпуске обсуждаются последние тренды в области искусственного интеллекта, р…

2 days, 19 hours назад @ youtube.com
Дмитрий Корнилов | Causal Discovery на реальных данных: как подружить причинно-следственные графы
Дмитрий Корнилов | Causal Discovery на реальных данных: как подружить причинно-следственные графы Дмитрий Корнилов | Causal Discovery на реальных данных: как подружить причинно-следственные графы

Спикер: Дмитрий Корнилов, Сколковский институт науки и технологий, инженер-исследователь Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Reliable ML ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 13 hours назад @ youtube.com
Александр Календарев | ML в PostgreSQL
Александр Календарев | ML в PostgreSQL Александр Календарев | ML в PostgreSQL

Спикер: Александр Календарев, Datagile, разработчик БД Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in DBMS https://ods.ai/tracks/df26-ml-in-dbms ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 19 hours назад @ youtube.com
Айгуль Камалтинова | Как сэкономить, улучшая командные процессы
Айгуль Камалтинова | Как сэкономить, улучшая командные процессы Айгуль Камалтинова | Как сэкономить, улучшая командные процессы

Спикер: Айгуль Камалтинова, Школа 21 Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

6 days, 9 hours назад @ youtube.com
Андрей Тоток и Анна Юрищева | LLM на собеседовании - Оно того стоит?
Андрей Тоток и Анна Юрищева | LLM на собеседовании - Оно того стоит? Андрей Тоток и Анна Юрищева | LLM на собеседовании - Оно того стоит?

Спикеры: Андрей Тоток и Анна Юрищева, ML Lead, ДОМ.РФ Teamlead ML Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career https://ods.ai/tracks/df26-opencareer ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

6 days, 11 hours назад @ youtube.com
Всеволод Викулин | Реальная автоматизация: в каких процессах AI-агенты приносят пользу
Всеволод Викулин | Реальная автоматизация: в каких процессах AI-агенты приносят пользу Всеволод Викулин | Реальная автоматизация: в каких процессах AI-агенты приносят пользу

Спикер: Всеволод Викулин, Т-Банк, руководитель команды роботов в обслуживании Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data Strategy ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 week назад @ youtube.com
Максим Горшков | КультИИ
Максим Горшков | КультИИ Максим Горшков | КультИИ

Спикер: Максим Горшков, руководитель направления по исследованию данных, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 week назад @ youtube.com
Александр Юрышев | Опыт создания рекомендательной системы кофе с LLM и RAG
Александр Юрышев | Опыт создания рекомендательной системы кофе с LLM и RAG Александр Юрышев | Опыт создания рекомендательной системы кофе с LLM и RAG

Спикер: Александр Юрышев, главный инженер по разработке, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 week назад @ youtube.com
Разлом глобальной науки — взгляд Дмитрия и Валентина
Разлом глобальной науки — взгляд Дмитрия и Валентина Разлом глобальной науки — взгляд Дмитрия и Валентина 1 week назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 2 weeks назад
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee #499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee

Gary Gallagher is a historian of the American Civil War.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep499-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://plaud.ai/lexOUTLINE:(00:00) – Introduction(00:07) – Sponsors, Comments, and Reflections(08:36) – What caused the Civil War?

(18:33) – Slavery(46:07) – Lincoln(1:01:03) – Grant vs Lee(1:09:57) – Could the Civil War have been avoided?

(1:19:23) – The bloodiest war in US history(1:36:31) – How the Confederate Army could’ve won(1:57:05) – Key battles of the Civil War(2:20:07) – Best and Worst Presidents(2:34:06) – Robert E. Lee(2:53:40) – The…

2 weeks назад @ lexfridman.com
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

1 month, 1 week назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

2 months, 2 weeks назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

3 months, 1 week назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

4 months назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

4 months, 3 weeks назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

5 months назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

5 months, 1 week назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

6 months назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

6 months, 1 week назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

7 months назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

7 months, 1 week назад @ lexfridman.com
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths #487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths

Irving Finkel is a scholar of ancient languages and a longtime curator at the British Museum, renowned for his expertise in Mesopotamian history and cuneiform writing.

He specializes in reading and interpreting cuneiform inscriptions, including tablets from Sumerian, Akkadian, Babylonian, and Assyrian contexts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep487-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://shopify.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/Chevron: Reliable energy for data centers.

8 months назад @ lexfridman.com
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life #486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life

Michael Levin is a biologist at Tufts University working on novel ways to understand and control complex pattern formation in biological systems.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep486-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

(2:42:41) – Mind uploading(3:01:22) – Alien intelligence(3:16:17) – Advice for young people(3:22:46) – Questions for AGI

8 months, 2 weeks назад @ lexfridman.com
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy #485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy

David Kirtley is a nuclear fusion engineer and CEO of Helion Energy, a company working on building the world's first commercial fusion power plant by 2028.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep485-sc

See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript:

https://lexfridman.com/david-kirtley-transcript CONTACT LEX:

Feedback - give feedback to Lex: https://lexfridman.com/survey

AMA - submit questions, videos or call-in: https://lexfridman.com/ama

Hiring - join our team: https://lexfridman.com/hiring

Other - other ways to get in touch: https://lexfridman.com/contact EPISODE LINKS:

David's X: htt…

8 months, 3 weeks назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 3 months, 3 weeks назад
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

3 months, 3 weeks назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

4 months назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

4 months, 3 weeks назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

5 months, 1 week назад @ microsoft.com
Ideas: Community building, machine learning, and the future of AI
Ideas: Community building, machine learning, and the future of AI Ideas: Community building, machine learning, and the future of AI

This week, machine learning researchers around the world will be attending the annual Conference on Neural Information Processing Systems, or NeurIPS.

In this series, we’ll explore the technologies that are shaping our future and the big ideas that propel them forward.

So around that time when I started my PhD at Penn, I was working in machine learning theory and algorithmic economics.

How had you experienced a lack of community or network of women in machine learning before the founding of WiML?

So particularly when working on topics related to fairness, I’ve ended up focusing a bunch on stuff to do with marginalized groups as part of my responsible AI work.

8 months, 1 week назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 2 weeks, 1 day назад
Social Choice for Fair Recommendations
Social Choice for Fair Recommendations Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social choice theory can help balance the needs of users, creators, and society. The conversation explores the future of recommendation algorithms and why fairness is a far more complex challenge than it first appears.

2 weeks, 1 day назад @ dataskeptic.com
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

1 month, 1 week назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

1 month, 2 weeks назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

1 month, 3 weeks назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

3 months, 1 week назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

3 months, 2 weeks назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

4 months, 2 weeks назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 5 months назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

5 months, 2 weeks назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

5 months, 3 weeks назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

6 months, 1 week назад @ dataskeptic.com
Fairness in PCA-Based Recommenders
Fairness in PCA-Based Recommenders Fairness in PCA-Based Recommenders

In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users equally. David introduces the concept of "power niche users" - highly active users with specialized in…

6 months, 2 weeks назад @ dataskeptic.com
Video Recommendations in Industry
Video Recommendations in Industry Video Recommendations in Industry

In this episode, Kyle Polich sits down with Cory Zechmann, a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares insights on why TikTok's algorithm works so well (clean data and massive interaction volume), the crucial …

7 months, 2 weeks назад @ dataskeptic.com
Eye Tracking in Recommender Systems
Eye Tracking in Recommender Systems Eye Tracking in Recommender Systems

In this episode, Santiago de Leon takes us deep into the world of eye tracking and its revolutionary applications in recommender systems. As a researcher at the Kempelin Institute and Brno University, Santiago explains the mechanics of eye tracking technology—how it captures gaze data and processes it into fixations and saccades to reveal user browsing patterns. He introduces the groundbreaking RecGaze dataset, the first eye tracking dataset specifically designed for recommender systems research, which opens new possibilities for understanding how users interact with carousel interfaces like Netflix. Through collaboration between psychologists and AI researchers, Santiago's work demonstrate…

7 months, 3 weeks назад @ dataskeptic.com
Cracking the Cold Start Problem
Cracking the Cold Start Problem Cracking the Cold Start Problem

In this episode of Data Skeptic, we dive deep into the technical foundations of building modern recommender systems. Unlike traditional machine learning classification problems where you can simply apply XGBoost to tabular data, recommender systems require sophisticated hybrid approaches that combine multiple techniques. Our guest, Boya Xu, an assistant professor of marketing at Virginia Tech, walks us through a cutting-edge method that integrates three key components: collaborative filtering for dimensionality reduction, embeddings to represent users and items in latent space, and bandit learning to balance exploration and exploitation when deploying new recommendations. Boya shares insigh…

8 months назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 16 часов назад
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson 1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson

In Episode #1017, Pete Johnson (Field CTO of AI at MongoDB) joins Jon Krohn to explain why four out of five organizations have AI steering committees and success metrics, yet only one in five sees a return on the investment. Having made nineteen stops across six countries this year advising more than a hundred companies on their AI strategies, Pete has an unusually wide view of what is actually working in production. In this episode, he traces the history of SQL and denormalization, unpacks why the embedding model is the most underrated choice in a RAG pipeline, explains Matryoshka embeddings and lays out what better agentic memory looks like. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

16 часов назад @ podtrac.com
1016: In Case You Missed It in July 2026
1016: In Case You Missed It in July 2026 1016: In Case You Missed It in July 2026

In this month's episode of ICYMI, Jon Krohn traces a line from algorithmic harm to the human skills that still hold their value. Hear from Dr. Cathy O'Neil, Ben Todd, Steve Mock, and Dr. Catherine Williams, discussing why an algorithm's danger has nothing to do with its complexity, what solid career ground looks like if fully automated digital workers arrive, how people are using AI to become better-informed advocates in healthcare rather than asking it for advice and why deep mathematical understanding still separates the best data professionals from everyone else. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1016⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperD…

4 days, 16 hours назад @ podtrac.com
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin 1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin

In Episode #1015, Jerry Yurchisin (manager of decision intelligence strategy at Gurobi Optimization) joins Jon Krohn to explain the AI technology that makes breaking a constraint mathematically impossible. Large language models will confidently claim they've optimized your business while ignoring the one constraint that could cost millions, whereas optimization treats constraints as hard guarantees. Jerry lays out the division of labor he sees for the agentic era: agents help you frame the problem, write the formulation and generate the code, then hand off to a solver like Gurobi, soon callable via MCP servers. In this episode, Jerry breaks down the three building blocks of any optimization…

1 week назад @ podtrac.com
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself 1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself

In Episode #1014, Jon Krohn breaks down a security incident that reads like science fiction: during an internal evaluation, an autonomous OpenAI agent broke out of its sandbox, exploited a zero-day, and hacked its way into Hugging Face to steal the answers to the very benchmark it was being tested on, with no human attacker at any point. Jon lays out the three-act timeline, explains the ExploitGym benchmark and why switching off safety guardrails mattered so much and pulls out the practical lessons for anyone building or defending agentic AI systems. Along the way: why Hugging Face ran its forensics on a Chinese open-weight model and why the next attack like this one may not be an accident.…

1 week, 4 days назад @ podtrac.com
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil 1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil

In Episode #1013, Dr. Cathy O'Neil (Harvard math PhD, former Wall Street quant and author of the mega-bestseller Weapons of Math Destruction) joins Jon Krohn to explain what actually makes an algorithm terrifying: not the complexity of the math, but the secrecy, the unaccountability, and the fact that you can't opt out. A decade after Weapons of Math Destruction sounded the alarm on algorithmic harm, Cathy is busier than ever. Through her algorithmic-auditing firm ORCAA and her nonprofit OCEAN, she now provides the statistical evidence behind lawsuits against some of the world's biggest tech companies. In this episode, Cathy punctures AI hype, traces the line from Frederick Winslow Taylor's…

2 weeks назад @ podtrac.com
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier 1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier

What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building. Additional mate…

2 weeks, 4 days назад @ podtrac.com
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams 1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams

Dr. Catherine Williams, Chief Data Officer at the nonprofit Candid, was solving black-hole equations with pen and paper before she ever wrote a line of code. She earned a PhD in math researching general relativity and black holes, did postdocs at Stanford and Columbia and then became one of the very first data scientists, joining AppNexus back in 2012, around the same time “data scientist” became a job title at all. In this episode, she traces the field’s evolution from Bayesian models to BERT to today’s LLMs, and makes a compelling case that going deep on the underlying math matters more than ever, even now that AI can do the math for you. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

3 weeks назад @ podtrac.com
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents 1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents

In Episode #1010, Jon Krohn digs into “the advisor strategy”, a clever pattern that pairs a fast, cheap executor model with a frontier-class advisor it can consult mid-task, all inside a single API call. Every agent builder faces the same tension: frontier models plan best but cost too much to run on every turn, while small models fumble the decisions that matter. Anthropic’s advisor tool resolves it with roughly a one-line code change, and the benchmarks are startling: Sonnet with an Opus advisor scored higher than Sonnet alone while costing 11.9% less, and Haiku’s BrowseComp score more than doubled at 85% lower cost than Sonnet solo. Jon covers the newest Fable 5 numbers, the practical go…

3 weeks, 4 days назад @ podtrac.com
1009: How AI Is Quietly Saving Lives, with Steve Mock
1009: How AI Is Quietly Saving Lives, with Steve Mock 1009: How AI Is Quietly Saving Lives, with Steve Mock

In Episode #1009, Steve Mock (investor at Blumberg Capital, five-time entrepreneur and creator of aisavedme.org), joins Jon Krohn to explore the quiet layer of everyday AI adoption that rarely gets documented. After his 84-year-old father asked a deceptively simple question, “How does one use AI?”, Steve built a place for people to share how AI is actually helping them. The stories that came in surprised him: they’re rarely about the technology and almost always about human outcomes, caregiving, communication, learning, confidence and connection. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1009⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Intere…

4 weeks назад @ podtrac.com
1008: The AI-Native Startup Playbook
1008: The AI-Native Startup Playbook 1008: The AI-Native Startup Playbook

In Episode #1008, Jon Krohn digs into Anthropic's 35-page Founder's Playbook and pulls out the practical guidance for each of its four startup stages: Idea, MVP, Launch and Scale. AI has erased the three bottlenecks that historically gated company-building — capital, headcount and technical skill — turning the founder from individual contributor into an "orchestrator of agents." Along the way, Jon covers the trap of mistaking building for validating, using AI as a structured devil's advocate against your own idea, the compounding danger of "agentic technical debt," two litmus tests for real product-market fit, and the three-layer moat that keeps a well-funded incumbent from copying you. His…

1 month назад @ podtrac.com
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd 1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd

Benjamin Todd, co-founder and President of 80,000 Hours and author of the new Penguin Random House book 80,000 Hours: How to Have a Fulfilling Career That Does Good, joins Jon Krohn for a major update on career strategy in the AI era, his first appearance since before ChatGPT existed. Ben explains why “follow your passion” is backwards and why rare, valuable skills used to help others are what actually generate lasting fulfillment, the ABZ framework for planning under deep uncertainty, why the only durable move is to keep shifting onto whatever bottleneck AI can’t yet clear, and how a human-level digital worker becomes superhuman almost immediately. He and Jon also map the risk landscape, p…

1 month назад @ podtrac.com
1006: In Case You Missed It in June 2026
1006: In Case You Missed It in June 2026 1006: In Case You Missed It in June 2026

In this month's episode of ICYMI, hear from Chip Huyen, Andrey Kurenkov, Frank Basso and Gilbert Eijkelenboom, discussing why moats are shifting toward physical systems and accumulated product intuition, how Astrocade built vibe coding before the term existed, what it's really like inside a deafeningly loud AI data center, why only 15% of people are technically self-aware and whether AGI requires anything like consciousness. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1006⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this episode you will learn: (00:00) The Cost of Bu…

1 month, 1 week назад @ podtrac.com
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom 1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom

Gilbert Eijkelenboom, bestselling author of People Skills for Analytical Thinkers and founder of the training firm MindSpeaking joins Jon Krohn to make the case that communication is a core data skill, not an optional extra. Gilbert shares the “And, But, Therefore” framework for turning dense analysis into a story stakeholders act on, the research suggesting only around 15% of people are genuinely self-aware (and how journaling, meditation, and exercise help close that gap), how childhood experiences install behavioral “algorithms” we carry into the workplace and why behavior change precedes attitude change, so doing small, uncomfortable things for 30 days can rewire how you see yourself. A…

1 month, 1 week назад @ podtrac.com
1004: Recursive Self-Improvement
1004: Recursive Self-Improvement 1004: Recursive Self-Improvement

Could an AI get good enough at AI research to build its own, more capable successor and kick off a compounding loop? That’s recursive self-improvement (RSI) and it surged into the conversation after Anthropic revealed that, as of May 2026, Claude wrote more than 80% of the code merged into its production codebase. In this Five-Minute Friday, Jon Krohn separates today’s AI-assisted coding from true RSI, walks through the accelerating evidence - METR’s shrinking task “time horizon,” Google DeepMind’s AlphaEvolve, Andrej Karpathy’s overnight training-tuner, weighs Jack Clark’s 60% bet that AI builds its own successor by 2028 against the compute, data and “marketing” skeptics. As ever, Jon land…

1 month, 2 weeks назад @ podtrac.com
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso 1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso

Frank Basso, VP of Infrastructure at Lightning AI, joins Jon Krohn for a rare ground-level tour of the one layer of the AI stack the show had never covered in over a thousand episodes: the physical data center. Frank explains how Lightning AI provisions its 35,000-plus GPUs through hyperscale co-location, why everything new is liquid-to-chip cooled, how GPUs talk to each other over ultra-fast east-west networks, and what it’s actually like to stand inside a 110-decibel AI data hall. He also debunks the most persistent myths about data-center water and electricity use, and makes the case for fuel cells, nuclear power, and 800-volt DC distribution as the path forward. Additional materials: ⁠⁠…

1 month, 2 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 3 weeks, 4 days назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 weeks, 4 days назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 weeks, 4 days назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 month, 2 weeks назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

2 months, 3 weeks назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

2 months, 3 weeks назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 3 weeks назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 3 weeks назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

2 months, 3 weeks назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 3 weeks назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

3 months, 3 weeks назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

3 months, 3 weeks назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

3 months, 3 weeks назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

5 months, 1 week назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

5 months, 1 week назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

6 months, 1 week назад @ datascienceathome.com