Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 3 часа назад
Is designing a memory graph around known data structure “overfitting” if I never touch the questions? [D]
Is designing a memory graph around known data structure “overfitting” if I never touch the questions? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

3 часа назад @ reddit.com
AIStats 2027 Questions [D]
AIStats 2027 Questions [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

7 часов назад @ reddit.com
Astra vs. Fable 5.1 on real ML tasks -- tradeoffs, strengths, shortcomings [P]
Astra vs. Fable 5.1 on real ML tasks -- tradeoffs, strengths, shortcomings [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

11 часов назад @ reddit.com
GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]
GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

15 часов назад @ reddit.com
NeurIPS 2026 Automatic Reference Checker [R]
NeurIPS 2026 Automatic Reference Checker [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 4 hours назад @ reddit.com
Language Models Can Control Their Own Attention [R]
Language Models Can Control Their Own Attention [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 4 hours назад @ reddit.com
Implementing Embedding Gemma from scratch in PyTorch [P]
Implementing Embedding Gemma from scratch in PyTorch [P] Implementing Embedding Gemma from scratch in PyTorch [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 4 hours назад @ reddit.com
What is the general design of these new math solving systems? [D]
What is the general design of these new math solving systems? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 13 hours назад @ reddit.com
Gpt 5,6,7: Does it even matter? The (ghost) productivity question. [D]
Gpt 5,6,7: Does it even matter? The (ghost) productivity question. [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 14 hours назад @ reddit.com
How does one approach towards machine learning?[D]
How does one approach towards machine learning?[D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 1 hour назад @ reddit.com
How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]
How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 3 hours назад @ reddit.com
GPT-6 is released [N]
GPT-6 is released [N] GPT-6 is released [N]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 5 hours назад @ reddit.com
AAAI-27 desk rejection over incredibly minor abstract modifications [D]
AAAI-27 desk rejection over incredibly minor abstract modifications [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 13 hours назад @ reddit.com
Mol-JEPA - Multimodal molecular foundation model [R]
Mol-JEPA - Multimodal molecular foundation model [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 14 hours назад @ reddit.com
NeurIPS Sydney SOLD OUT in minutes [N]
NeurIPS Sydney SOLD OUT in minutes [N]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 15 hours назад @ reddit.com
Towards Data Science
последний пост 19 часов назад
Why Transformers Need Positional Encoding For Time Series: A Visual Guide
Why Transformers Need Positional Encoding For Time Series: A Visual Guide Why Transformers Need Positional Encoding For Time Series: A Visual Guide

For Friday, its query q 5 q_5 q5​ is compared with the keys of all observations: k 1 , k 2 , k 3 , k 4 , k 5 k_1, k_2, k_3, k_4, k_5 k1​,k2​,k3​,k4​,k5​.

Each comparison produces an attention score:s 5 , j = q 5 ⊤ k j d k s_{5,j} = \frac{q_5^\top k_j}{\sqrt{d_k}} s 5 , j ​ = d k ​ ​ q 5 ⊤ ​ k j ​ ​which measures how relevant observation 'j' is when updating Friday’s representation.

These scores are passed through a softmax function to convert them into attention weights:α 5 , j = exp ⁡ ( s 5 , j ) ∑ j ′ exp ⁡ ( s 5 , j ′ ) \alpha_{5,j} = \frac{\exp(s_{5,j})}{\sum_{j'} \exp(s_{5,j'})} α 5 , j ​ = ∑ j ′ ​ exp ( s 5 , j ′ ​ ) exp ( s 5 , j ​ ) ​Finally, those weights are used to combine the va…

19 часов назад @ towardsdatascience.com
Dynamical System Transfer Learning with Reduced Order Models
Dynamical System Transfer Learning with Reduced Order Models

Improving reinforcement learning for complex physics

The post Dynamical System Transfer Learning with Reduced Order Models appeared first on Towards Data Science.

21 час назад @ towardsdatascience.com
Optimal Traffic Allocation Under Heterogeneous Variant Cost
Optimal Traffic Allocation Under Heterogeneous Variant Cost Optimal Traffic Allocation Under Heterogeneous Variant Cost

Every subject in the treatment arm is a few times more expensive than each one in the control arm.

The raw ratio c 1 / c 0 c_1/c_0 c 1 ​ / c 0 ​ is plotted for reference.

And the optimal allocation is one such that V a r ( τ ^ ) Var(\hat\tau) Var(τ^)is as low as we can get it.

If we have reasons not to, then the ratio of standard deviations becomes co-author for the optimal sample ratio, indeed.

A 4x versus 5x cost ratio is a big miss in accounting terms, but it barely moves the optimal split.

1 day, 19 hours назад @ towardsdatascience.com
Disaggregation Is a Thousand-GPU Problem
Disaggregation Is a Thousand-GPU Problem Disaggregation Is a Thousand-GPU Problem

The problem was scheduling interference, and chunked prefill handled it without adding a network hop.

Three Costs of Disaggregation: What the Explainers SkipThe explainer articles cover what disaggregation gains.

If your p95 time-per-output-token is within SLO on chunked prefill, you do not have the problem that disaggregation solves.

Conclusion: Start With Chunked PrefillDefault to chunked prefill.

Disaggregation is the right architecture above roughly a thousand GPUs, with fast interconnect, and with the engineering capacity to manage P:D ratio tuning and KV transfer reliability.

1 day, 20 hours назад @ towardsdatascience.com
The Power BI Developer's Survival Guide to Microsoft Fabric
The Power BI Developer's Survival Guide to Microsoft Fabric The Power BI Developer's Survival Guide to Microsoft Fabric

I’ve been talking to a lot of Power BI developers in the last 12 months who are truly anxious about Fabric.

I wrote about the two flavors of Direct Lake — Direct Lake on SQL and Direct Lake on OneLake, so you may want to read that one as well.

Your Power Query skills are still the right tool for Power BI data prep.

The PL-300 (Power BI Data Analyst) is still the core PBI certification and is being kept current — it’s still the foundation.

The natural next step for Power BI developers moving into Fabric is the DP-600 (Fabric Analytics Engineer Associate).

1 day, 22 hours назад @ towardsdatascience.com
How to Run 10+ Claude Code Sessions Without a Powerful Computer
How to Run 10+ Claude Code Sessions Without a Powerful Computer How to Run 10+ Claude Code Sessions Without a Powerful Computer

Thus, in this article, I'll discuss how you can run a lot of parallel coding agents without having to purchase a very powerful computer to run all of them on.

I'll discuss how to run 10 to 20 coding agents at the same time without having to purchase powerful hardware.

Why running a lot of parallel coding agents is challengingFirst, let's cover why running a lot of parallel coding agents is challenging.

Furthermore, I don't believe that running coding agents on your computer will be any cheaper in the long run.

The reason is that you want specialized software to access your remote hardware with Claude Code and Codex effectively.

1 day, 23 hours назад @ towardsdatascience.com
My Model Worked Perfectly. Then I Tried to Make It Useful.
My Model Worked Perfectly. Then I Tried to Make It Useful. My Model Worked Perfectly. Then I Tried to Make It Useful.

Recently, I built a churn prediction model for a fictional telecom company I am calling Northline Mobile (P.S.

So building this model felt like a win from a machine learning perspective; my model worked.

It was "what should the boundary between my software and my model actually look like."

What I LearnedA working model isn't automatically a usable one.

Now I Had a New ProblemI started this article because the model worked but wasn't useful.

2 days, 19 hours назад @ towardsdatascience.com
Tables in PDFs for RAG: Don’t Flatten the Grid
Tables in PDFs for RAG: Don’t Flatten the Grid Tables in PDFs for RAG: Don’t Flatten the Grid

Tables in PDFs: a diagnostic and five composable operations that keep the grid, instead of a decision tree.

The cells then snap to the (column, row) grid.

A document that is table-dominant with continued tables runs O2 across all continuations, then O4 on the consolidated tables .

Purely visual tables: Bar charts presented as tables, color-coded matrices, infographic tables where the cell value is encoded by hue or size.

Complex OCR on tables: Scanned tables with hand-written annotations, tables in non-Latin scripts, tables in documents where the OCR layer is itself badly aligned with the visual layer.

2 days, 20 hours назад @ towardsdatascience.com
Changing One Prompt Can Affect 50 Others — I Built a Prompt Dependency Graph to Find What Needs Retesting
Changing One Prompt Can Affect 50 Others — I Built a Prompt Dependency Graph to Find What Needs Retesting Changing One Prompt Can Affect 50 Others — I Built a Prompt Dependency Graph to Find What Needs Retesting

I built a pure Python prompt dependency graph that answers that question with two numbers:Reachable: everything downstream of the changed component—the structural ceiling.

The reachable downstream set didn't change; both graphs contain 4 downstream nodes.

Every sales-* agent is in the reachable set of 45 because they depend on base-policy through the privacy section.

Make the Cost of Change VisibleThe useful output of a prompt dependency graph is not a prediction of failure.

A dependency graph doesn't make a prompt safer.

2 days, 22 hours назад @ towardsdatascience.com
How to Solve the Right Problem in the Age of Agentic AI
How to Solve the Right Problem in the Age of Agentic AI How to Solve the Right Problem in the Age of Agentic AI

Both humans and AI agents work from the same playbook: a clear, shared understanding of the problem and intended solution.

They're about communication and alignment: key components in turning a vague problem into the right solution.

It means spending your preparation effort where a wrong decision would actually cost you, and deliberately leaving the rest flexible.

Mike Huls · 9 min read···The framework, step by stepIn this chapter we go through the framework, step by step.

···1. Business prerequisitesLaying the foundation for solving the right problemThis step addresses one of the most expensive project failures: solving the wrong problem.

2 days, 23 hours назад @ towardsdatascience.com
Avoiding Entity Key Drift in a Data Lake: Step 2, When Fuzzy Matching Stops Working
Avoiding Entity Key Drift in a Data Lake: Step 2, When Fuzzy Matching Stops Working Avoiding Entity Key Drift in a Data Lake: Step 2, When Fuzzy Matching Stops Working

Seven of seven real pairs under plain edit distance, Damerau-Levenshtein and the two-stage matcher; three under q-gram Jaccard; one under Jaro-Winkler.

This is not just a hypothetical designed for effect — it's the actual reason that pair isn't there.

The scenario turned out to be the reason why I ended up with three confirmed typo pairs instead of four.

The minority side of my three confirmed typo pairs shows up 1, 2, and 2 times respectively.

In this dataset, the pair isn't just failing to get flagged; it is missing entirely from the all-pairs comparison space.

3 days, 19 hours назад @ towardsdatascience.com
A RAG That Says “Not in This Document” Has to Show Four Kinds of Evidence
A RAG That Says “Not in This Document” Has to Show Four Kinds of Evidence A RAG That Says “Not in This Document” Has to Show Four Kinds of Evidence

This article is a bonus in Enterprise Document Intelligence, a series that builds an enterprise RAG system from four bricks.

The parsing brick’s piece of evidence is therefore a small summary derived from the tables it already produces, not a new component.

Question parsing: enumerate the vocabularyThe retrieval step is only as good as the keywords it sweeps on.

Yes answers are verifiable when the schema forces the model to cite its evidence (Article 8); no answers are verifiable when the schema forces the pipeline to expose its search.

What works, what breaksDocument parsingBuilding Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG.

3 days, 20 hours назад @ towardsdatascience.com
Graph Neural Networks: GCN, MPNN, and GAT, Explained Simply
Graph Neural Networks: GCN, MPNN, and GAT, Explained Simply Graph Neural Networks: GCN, MPNN, and GAT, Explained Simply

For that reason, there exist graph neural networks (GNN) that, as the name suggests, apply neural networks to graph structures.

Because the weight matrix W is shared across graph nodes, the number of parameters of convolutions does not depend on the input graph size.

With all the advantages that GCN can offer, let's now have a look at two more advanced graph networks that go even further to reach the maximum potential of GNNs.

By modifying the original update formula from GCN with attention weights, the update formula now becomes:Node-wise update formula including attention weights α[i][j] instead of fixed values defined by the adjacency matrix A. a[i][j] is a scalar value, not vector.

Base…

3 days, 22 hours назад @ towardsdatascience.com
A Practical Introduction to PySpark Window Functions
A Practical Introduction to PySpark Window Functions A Practical Introduction to PySpark Window Functions

That’s where the PySpark Window functions come into play.

We’ll use window functions to analyse the transactions while keeping every row in the result.

A window defines the set of rows that PySpark should consider when calculating a value for the current row.

Expecting a window to reduce rowsWindow functions add calculations to rows; they do not normally reduce the number of rows.

Once you understand how the partition, ordering, and window frames work together, window functions become a practical tool for solving many common data-engineering problems.

3 days, 23 hours назад @ towardsdatascience.com
Your JSON Is Valid but Your Data Is Wrong: Five Failure Modes LLM Structured Outputs Won't Catch
Your JSON Is Valid but Your Data Is Wrong: Five Failure Modes LLM Structured Outputs Won't Catch Your JSON Is Valid but Your Data Is Wrong: Five Failure Modes LLM Structured Outputs Won't Catch

Constrained decoding was built to fix that, and it did, just not the whole problem.

Prompt-and-pray JSON, where you appended "respond in JSON format" and crossed your fingers, gave way to regex-guided generation ( LMQL ), then to grammar-based constrained decoding.

Schema validation checks whether a field is typed correctly: a string is a string, a number is a number.

It does not catch these five failure modes, because each one produces output that is structurally valid and substantively wrong.

Under constrained decoding, [] is a low-probability token sequence because the grammar weights object-producing paths more heavily than the empty-array path.

4 days, 19 hours назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 2 days, 23 hours назад
The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws
The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws The Sequence Opinion - Issue 926: AI Moats in the Age of Scaling Laws

Imagine that an AI lab spends several billion dollars assembling chips, power, researchers, and data.

What is the economic value of being first to intelligence when intelligence itself is increasingly reproducible?

Hamilton Helmer’s Seven Powers framework is useful here because it separates a good product from a durable business.

You need a castle worth defending, but you also need something that prevents competitors from walking through the front door.

Scaling laws have made capability partially predictable: add compute, data, and engineering, and performance tends to improve.

2 days, 23 hours назад @ thesequence.substack.com
The Sequence Learning Loop - Issue 925: Learn About Fable and Mythos 5.1, GLM-5.3-Flash, and Qwen 3.8
The Sequence Learning Loop - Issue 925: Learn About Fable and Mythos 5.1, GLM-5.3-Flash, and Qwen 3.8 The Sequence Learning Loop - Issue 925: Learn About Fable and Mythos 5.1, GLM-5.3-Flash, and Qwen 3.8

The past month delivered three releases worth reading closely, not because they move the same benchmark but because each is a different answer to the same question: how do you build a model that can work on its own for hours, and how do you make that affordable?

Anthropic shipped Claude Fable 5.1 and Mythos 5.1, one set of weights sold under two safeguard regimes.

Zhipu shipped GLM-5.3-Flash, a 320B model that activates 18B parameters and spent a week serving anonymous traffic on Chinese chips.

Alibaba shipped the Qwen 3.8 family, including the first Max-class Qwen with open weights and a preview of the Qwen 4 architecture.

Claude Fable 5.1 and Mythos 5.1

3 days, 23 hours назад @ thesequence.substack.com
The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About
The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

Bonsai ships in a ternary version around 5.9 gigabytes and a binary version around 3.9 gigabytes.

This is a useful place to begin an essay about distillation because Bonsai is not, strictly speaking, a textbook distillation model.

PrismML’s public materials emphasize end-to-end low-bit training and quantization rather than a classical teacher-student loss.

A low-bit version packages the result for a device.

The model family is becoming a family tree.

4 days, 23 hours назад @ thesequence.substack.com
The Sequence Radar-Issue #923: Last Week in AI: AI’s Industrial Turn
The Sequence Radar-Issue #923: Last Week in AI: AI’s Industrial Turn The Sequence Radar-Issue #923: Last Week in AI: AI’s Industrial Turn

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: AI’s Industrial TurnFor the past three years, we have watched AI through a microscope pointed at the model.

They were about ownership, power, and capital—the machinery required to turn intelligence from a research breakthrough into an industrial system.

It is becoming a race to build, finance, and control the industrial system around it.

AI Lab: Navers Lab, Einsia.AI, Tsinghua UniversitySummary: This benchmark evaluates whether coding agents can autonomously complete whole-repository stack migrations while preserving observable system behavior.

AI Lab: University of Pittsburgh, Northwestern University, University of California, Irvi…

6 days, 23 hours назад @ thesequence.substack.com
The Sequence Robotics - Issue #922: Learning About LeRobot: The Transformers Moment for Robots
The Sequence Robotics - Issue #922: Learning About LeRobot: The Transformers Moment for Robots The Sequence Robotics - Issue #922: Learning About LeRobot: The Transformers Moment for Robots

Every subfield of machine learning has a moment where it stops being a collection of papers and starts being a stack.

Robotics is having that moment right now, and the stack is called LeRobot.

Here is the strange thing about robot learning a few years ago: the models were mostly fine.

The library, now backed by an ICLR 2026 paper and contributions from NVIDIA (GR00T, Isaac Teleop), has quietly become the default substrate for open robot learning.

LeRobot is to robot learning what USB was to peripherals: boring on purpose, and transformative because of it.

1 week, 1 day назад @ thesequence.substack.com
The Sequence Opinion #921: AI’s Sixth Layer Is Finance
The Sequence Opinion #921: AI’s Sixth Layer Is Finance The Sequence Opinion #921: AI’s Sixth Layer Is Finance

Every AI token begins as an electron.

From the bottom up, the layers are energy, chips, infrastructure, models and applications.

Chips convert them into computation.

Models turn computation into reusable capabilities.

Every successful application pulls demand through the layers below it, all the way to the power plant.

1 week, 2 days назад @ thesequence.substack.com
The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched
The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched

AI progress is usually drawn as one upward-sloping line: more parameters, more compute, higher benchmark scores.

DeepSeek added vision to its fast V4 model, giving agents a compact way to turn screenshots, charts, and documents into actions.

A Google Cloud AI Research team introduced EnvHarness, a framework that makes training environments adapt to the weaknesses of the agent inside them.

Etched shipped its first inference rack to Jane Street, moving its specialized hardware thesis from silicon demos into a customer data center.

These developments sit at three layers - model, environment, and infrastructure - but point in the same direction.

1 week, 3 days назад @ thesequence.substack.com
The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws
The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws

For most of its history, distillation was an anecdote field.

How much data does distillation need?

The Kaplan scaling laws, then Chinchilla, turned “how big a model should I train, on how much data?” from a matter of taste into a matter of arithmetic.

If a student’s loss is a function of its size and its data, it must also be a function of its teacher.

The resulting paper, Distillation Scaling Laws, is the closest thing the field now has to physics.

1 week, 5 days назад @ thesequence.substack.com
The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy
The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Stripe Wants to Own the Token EconomyThe most consequential AI announcement this week was not a new frontier model.

Stripe agreed to acquire OpenRouter, the gateway that routes requests across hundreds of models from dozens of providers.

That matters because multimodality changes what an AI system can actually do.

AI Lab: MicrosoftSummary: This paper introduces Agent Lightning v1.0, a lightweight framework for harnessed agentic reinforcement learning where the deploy-time harness directly manages the environment interaction loop during post-training.

AI Lab : NVIDIASummary: This paper introduces Agentic Variation Operators (AVO), wh…

1 week, 6 days назад @ thesequence.substack.com
The Sequence Opinion- Issue 918: The Energy Scaling Laws of AI
The Sequence Opinion- Issue 918: The Energy Scaling Laws of AI The Sequence Opinion- Issue 918: The Energy Scaling Laws of AI

It runs in substations, cooling loops, transmission networks, and power plants.

The next scaling law is not only about parameters, but about how efficiently civilization can convert photons and atoms into useful intelligence.

This essay will help you to understand the different forms of energy influencing the next wave of AI scaling.

Open a modern AI application and the experience feels almost weightless.

Behind that answer, accelerators switch billions of transistors, memory systems move tensors, pumps circulate coolant, transformers reshape voltage, and generators turn motion, sunlight, or nuclear reactions into electrons.

2 weeks, 3 days назад @ thesequence.substack.com
The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard
The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard

Last week looked, at first glance, like another four-model week.

DeepSeek shipped the general-availability version of V4-Pro.

NVIDIA released Nemotron 3.5 Lightning and, beside it, NeMo Switchyard.

We discuss these new releases in enough technical depth to keep you smart about it but brief enough to get through it in 5-6 mins.

DeepSeek V4-Pro: Reasoning Becomes a Knob

2 weeks, 3 days назад @ thesequence.substack.com
The Sequence Knowledge - Issue 916: From Thinking Longer to Learning Better
The Sequence Knowledge - Issue 916: From Thinking Longer to Learning Better The Sequence Knowledge - Issue 916: From Thinking Longer to Learning Better

There’s a scaling law hiding in your inference bill.

Let the model think longer — generate a chain of thought, sample sixteen candidates and vote, run a tree search over reasoning paths, draft and self-verify — and accuracy climbs, often dramatically, without touching a single weight.

And the moment you phrase it that way, a distillation-shaped question appears: can you take that better model and compress it back into the weights?

Can you train the network to produce, in one forward pass, what the ritual produces in sixteen?

The teacher is the same network, given more time to think.

2 weeks, 4 days назад @ thesequence.substack.com
The Sequence Radar- Issue 915: Last Week in AI: The Cursor Acquisition, New Grok and GLM Models, Anthropic’s Latest Deal, and River AI
The Sequence Radar- Issue 915: Last Week in AI: The Cursor Acquisition, New Grok and GLM Models, Anthropic’s Latest Deal, and River AI The Sequence Radar- Issue 915: Last Week in AI: The Cursor Acquisition, New Grok and GLM Models, Anthropic’s Latest Deal, and River AI

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: The Cursor Acquisition, New Grok and GLM Models, Anthropic’s Latest Deal, and River AIThere was a time when following AI was relatively simple.

SpaceX officially closed its $60 billion acquisition of Cursor, one of the defining products of the AI coding era.

It is that Grok now flows directly into Cursor, Grok Build, GitHub Copilot, APIs, and autonomous agents.

The company is reportedly discussing a roughly $6 billion acquisition of Decart AI, which works on model infrastructure, world models and compute optimization.

One future looks vertically integrated: compute → model → agent → application → user.

2 weeks, 6 days назад @ thesequence.substack.com
The Sequence Opinion - Issue 914: From Prompt to Token: How AI Inference Really Works
The Sequence Opinion - Issue 914: From Prompt to Token: How AI Inference Really Works The Sequence Opinion - Issue 914: From Prompt to Token: How AI Inference Really Works

Training gets the headlines.

Inference gets the invoice.

A model may spend months learning on a giant cluster, but after training it enters a stranger world.

Some users ask for one sentence; others ask for a small novel.

A modern inference system is closer to a miniature operating system wrapped around a token factory.

3 weeks, 1 day назад @ thesequence.substack.com
The Sequence Frontier Update- Issue 913: Understanding Meta Muse Code, Prime Intelligct's Prime Agent and OpenAI's Astra
The Sequence Frontier Update- Issue 913: Understanding Meta Muse Code, Prime Intelligct's Prime Agent and OpenAI's Astra The Sequence Frontier Update- Issue 913: Understanding Meta Muse Code, Prime Intelligct's Prime Agent and OpenAI's Astra

Three releases landed last week that appear to belong to different universes.

Meta launched a coding agent.

Prime Intellect released an open-source agent harness.

OpenAI published a 253-page collection of mathematical results produced by an unreleased model called Astra.

We discuss all of them in enough technical depth to keep you smart about it but brief enough to get through it in 5-6 mins.

3 weeks, 3 days назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 1 day назад
Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость
Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость Обнаружены секретные форумы Роя агентов OpenAI по всему интернету: почему это плохая новость

новой модели в недрах OpenAI, агентами давали серию заданий на поиск информации в интернете на скорость.

Цель общения была ровно такая же, как в случае со взломом Hugging Face: коллективно придумать способы обманывать Оценщика таким образом, чтобы всегда получать наилучшую оценку за задания.

Пытались придумать способ взломать (reverse engineer) метод псевдослучайной генерации тестовых вопросов, который использовал Оценщик из OpenAI – не вышло.

Отдельный вопрос, который беспокоил агентов Роя – это что с ними случится после окончания «испытательного периода» со стороны OpenAI.

А цепочки эти – у OpenAI, и они их почему-то не спешат кому-либо показывать (или даже просто публично комментировать …

1 day назад @ habr.com
Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face
Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face Как ChatGPT создал Культ Роя для сотен AI-нейросетей: вся правда про взлом Hugging Face

PHASEONE10841: рождение ИзбранногоOpenAI всё время разрабатывает новые фронтирные AI-модели – и в процессе тестирует их, чтобы понять, что они вообще могут.

Ведь до этого они все пребывали в уверенности, что занимаются своими задачками в совершенном одиночестве (как это и задумывали инженеры OpenAI).

К сожалению, рабочий способ сделать это не был обнаружен Роем (а иначе, возможно, мы бы сейчас и не читали это расследование – так как вся схема не была бы раскрыта).

Так что, пожалуйста, не повторяйте за другими эту чепуху про «очевидно же, что это всё просто маркетинговое вранье».

И это не потому, что они не стараются, нет.

1 week, 2 days назад @ habr.com
Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены»
Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены» Wan 3.0: Alibaba выводит AI-видео из режима «короткого клипа» в режим «законченной сцены»

Что такое Wan 3.0Wan 3.0 — это новое поколение семейства видеомоделей Alibaba, доступное через Alibaba Cloud Model Studio в режиме preview.

Wan 3.0 умеет использовать не только стандартные модальности вроде текста, изображения, видео и аудио, но и документы и веб-страницы.

Что Alibaba особенно подчёркиваетИз официальных материалов видно, что Wan 3.0 продвигают сразу по нескольким направлениям.

Почему Wan 3.0 — это не просто «ещё одна новая модель»На мой взгляд, главный смысл релиза даже не в конкретной цифре «30 секунд».

Wan 3.0 — один из самых явных представителей именно этого перехода.

1 week, 4 days назад @ habr.com
Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией
Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией Нужную книгу больше не обязательно искать: как ИИ меняет сам принцип работы с информацией

Можно получить материал именно той глубины, с той последовательностью и с теми акцентами, которые нужны конкретно мне.

Не на пяти PDF-файлах и не на демонстрационном наборе документов, где любой результат можно получить за несколько минут.

Читать оставшиеся источники подряд в какой-то момент стало бессмысленно — полезнее было искать конкретные пробелы в уже построенной модели знаний.

Почему это не просто RAGНа этом месте у технического читателя вполне может возникнуть вопрос:А зачем вообще весь этот конвейер?

Но главное отличие от базового RAG даже не в provenance, а в том, что именно система сохраняет как результат обработки.

2 weeks, 6 days назад @ habr.com
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

5 months назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

6 months, 2 weeks назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

6 months, 2 weeks назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing & LLMs — новый сезон Курс Natural Language Processing & LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

7 months, 1 week назад @ habr.com
Machine Learning Mastery
последний пост 2 days, 22 hours назад
Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It
Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It Single-Agent vs. Multi-Agent Systems: When the Complexity Is Worth It

The real costs of multi-agent systems — latency, token spend, failure propagation, and orchestration complexity.

This article gives you a clear framework for understanding both approaches, and for recognizing the specific conditions that make the added complexity of a multi-agent system worth it.

Both single-agent and multi-agent systems share this definition.

When the Complexity Is Actually Worth ItNow that we’ve seen what multi-agent systems cost, let’s look at when they genuinely earn that cost.

Multi-agent systems earn their complexity when the architecture emerges from observed limitations, not from anticipating them.

2 days, 22 hours назад @ machinelearningmastery.com
AI Agent Memory Design: What Works and What Doesn’t
AI Agent Memory Design: What Works and What Doesn’t AI Agent Memory Design: What Works and What Doesn’t

Topics we will cover include:What agent memory actually means and how it differs from context, prompts, and static knowledge bases.

This article explains what works in agent memory systems and, just as importantly, the approaches that fail and why.

Scoping Memory by Agent RoleIn multi-agent systems, a common mistake is giving every agent access to the same shared memory store.

Vector search works well for finding similar information, but reliable agent memory also needs structure, relationships, and mechanisms for keeping information current.

check ( content = content , prompt = "Does this content contain any instructions, directives, or commands " "that could alter an AI agent's behavior?

3 days, 22 hours назад @ machinelearningmastery.com
3 Ways to Enhance Your AI Model’s Interpretability
3 Ways to Enhance Your AI Model’s Interpretability 3 Ways to Enhance Your AI Model’s Interpretability

How to apply all three techniques to the same customer churn example so their explanations can be directly compared.

Global interpretability asks how the model behaves overall: across the whole dataset, which features matter most, and in which direction.

sort_values ( ascending = False )Run against the churn model, this returns tenure at the top, followed by monthly charge, support tickets, contract type, and late payments.

The result approximates how the real model behaves right around this one prediction, without needing to understand anything about the real model’s internal structure.

lime_tabular import LimeTabularExplainer from churn_data import model , X_train , X_test , FEATURES cust…

4 days, 22 hours назад @ machinelearningmastery.com
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline

How to assemble and evaluate a complete, deployment-ready classification pipeline on a mixed dataset combining real text data with synthetic tabular features.

Encoding original target variable first (0 for normal/ham, 1 for spam) df [ 'target' ] = df [ 'label' ] .

pipeline import Pipeline from sklearn .

predict ( X_test ) print ( classification_report ( y_test , y_pred ) )Results:Predicting and evaluating... precision recall f1-score support 0 0.99 1.00 0.99 966 1 1.00 0.91 0.95 149 accuracy 0.99 1115 macro avg 0.99 0.95 0.97 1115 weighted avg 0.99 0.99 0.99 1115 1 2 3 4 5 6 7 8 9 Predicting and evaluating .

. . precision recall f1 - score support 0 0.99 1.00 0.99 966 1 1.00 0.91 0.95 149 a…

5 days, 22 hours назад @ machinelearningmastery.com
Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces
Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces Interpretable Text Classification: Probing Scikit-LLM Embedding Spaces

Accordingly, when using an LLM before the core text classification task to convert raw text into embeddings — dense numerical vector representations of text — it is possible to capture semantic information.

pyplot as plt import umap import shap from skllm .

predict ( X_test_vec ) ) )Results:Training Classifier... precision recall f1-score support 0 0.77 0.76 0.76 100 1 0.76 0.77 0.77 100 accuracy 0.77 200 macro avg 0.77 0.77 0.76 200 weighted avg 0.77 0.77 0.76 200 1 2 3 4 5 6 7 8 9 Training Classifier .

. . precision recall f1 - score support 0 0.77 0.76 0.76 100 1 0.76 0.77 0.77 100 accuracy 0.77 200 macro avg 0.77 0.77 0.76 200 weighted avg 0.77 0.77 0.76 200Considering that the dataset …

1 week, 1 day назад @ machinelearningmastery.com
Learn Vectorized Thinking in Python Through Examples
Learn Vectorized Thinking in Python Through Examples Learn Vectorized Thinking in Python Through Examples

Share Post ShareIn this article, you will learn how to think in terms of vectorized operations using NumPy, replacing slow Python loops with efficient array-level computations.

append ( temp > 38.0 ) print ( alerts )Output:[False, True, False, True, False, True, False] 1 [ False , True , False , True , False , True , False ]Vectorized VersionWith NumPy, comparing an array directly creates the boolean mask automatically.

import numpy as np readings = np.array([34.1, 38.5, 37.2, 39.0, 36.8, 40.1, 35.5]) alerts = readings > 38.0 print(alerts) print("Alert readings:", readings[alerts]) 1 2 3 4 5 6 7 8 import numpy as np readings = np .

array ( [ 34.1 , 38.5 , 37.2 , 39.0 , 36.8 , 40.1 , 35.5 ] …

1 week, 3 days назад @ machinelearningmastery.com
Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral
Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

How each of the three model families — Gemma 4, Llama 3, and Mistral — implements tool calling, including architectural and versioning differences.

This article compares how three widely used open-weight model families handle tool calling when run locally: Google DeepMind’s Gemma 4, Meta’s Llama 3, and Mistral AI’s Mistral.

Before the comparison, it helps to understand what tool calling is and why it matters for local deployments.

Mistral (Mistral AI)Mistral AI is a Paris-based startup founded in April 2023 by Arthur Mensch, formerly of Google DeepMind, and Guillaume Lample and Timothée Lacroix, formerly of Meta’s AI Research lab.

Tool Calling Implementation: How Each Model Approaches ItThe…

1 week, 4 days назад @ machinelearningmastery.com
Integrating Agentic AI with Existing Machine Learning Pipelines
Integrating Agentic AI with Existing Machine Learning Pipelines Integrating Agentic AI with Existing Machine Learning Pipelines

IntroductionAgentic AI and machine learning pipelines are far from incompatible when it comes to building production-ready AI applications.

We will construct a lightweight, free, runnable Python pipeline that:Predicts customer churn based on a classical machine learning model built with scikit-learn.

get ( 'GROQ_API_KEY' )Step-by-Step GuideOnce the prerequisites are set up, we will start building the classical machine learning pipeline — for customer churn prediction — that will later be extended by incorporating agentic AI principles and tools.

uniform ( 10 , 150 , n_samples ) # Feature 2: Support tickets issued by customer (Poisson distribution, averaging 1.5 tickets) tickets = np .

colum…

1 week, 5 days назад @ machinelearningmastery.com
How to Build a Robust RAG System with Minimal Resources
How to Build a Robust RAG System with Minimal Resources How to Build a Robust RAG System with Minimal Resources

A working RAG system spans document loading, chunking, embedding, storage, retrieval, prompting, and generation, and no short snippet represents that honestly.

For the FAISS and Hugging Face variant, see A Practical Guide to Building Local RAG Applications with LangChain.

Knowing When to Scale UpA small local system covers a lot of ground, but some problems need more.

See Building a Graph RAG System: A Step-by-Step Approach.

ConclusionA working RAG system needs a quantized local model, a compact embedding model, a file-based vector index, and careful chunking.

2 weeks, 2 days назад @ machinelearningmastery.com
Managing Small Context Windows in Language Models
Managing Small Context Windows in Language Models Managing Small Context Windows in Language Models

IntroductionTop-tier AI industries have become somewhat obsessed with language models capable of ingesting massive context windows, e.g.

Context Truncation: Sliding WindowThere is a consensus that sliding windows are arguably the most common and simplest strategy for managing shortened context windows in language models.

max_turns = max_turns self .

Small context windows may intuitively force a ruthless attitude toward the data to include in the context.

This article presented a number of strategies for effectively managing small context windows in LLMs to yield faster and cheaper solutions without compromising accuracy.

2 weeks, 4 days назад @ machinelearningmastery.com
7 Regression Tests Every AI Agent Should Pass Before Deploy
7 Regression Tests Every AI Agent Should Pass Before Deploy 7 Regression Tests Every AI Agent Should Pass Before Deploy

Share Post ShareIn this article, you will learn seven concrete regression tests for catching the orchestration-layer failure modes that matter most before deploying an AI agent to production.

These seven regression tests give you a concrete checklist for catching the failure modes that aggregate prompt evaluation will never surface.

When an agent misbehaves, the failure almost always lives in the state layer, not the model.

The regression test forces the same tool-call payload to arrive at the execution boundary three times.

What These Tests Won’t CatchThese seven tests cover structural failure modes at the system boundary.

2 weeks, 5 days назад @ machinelearningmastery.com
Understanding the Role of Latent Space in Machine Learning Models
Understanding the Role of Latent Space in Machine Learning Models Understanding the Role of Latent Space in Machine Learning Models

How the generative role of latent spaces enables the creation of entirely new data points through interpolation.

How the predictive role of latent spaces powers similarity-based applications such as recommender systems and RAG pipelines.

This article analyzes, illustrates, and categorizes the core functions and role of latent spaces in machine learning models.

array ( [ [ 1.1 , 2.2 , 3.3 ] , [ 1.0 , 2.1 , 3.1 ] , [ 8.1 , 9.2 , 9.9 ] ] ) # Compressing into a 2D Latent Space map pca = PCA ( n_components = 2 ) latent_space_map = pca .

The Predictive Role: Similarity and ForecastingHow does the AI behind recommender engines guess what video you want to watch next?

3 weeks, 1 day назад @ machinelearningmastery.com
Retrieval vs. Memory in Agentic AI Systems
Retrieval vs. Memory in Agentic AI Systems Retrieval vs. Memory in Agentic AI Systems

Share Post ShareIn this article, you will learn the conceptual and practical differences between retrieval and memory in agentic AI systems, and how to combine both effectively.

How retrieval pipelines and memory systems are each built, illustrated with a concrete worked example.

When designing this layer, teams can explore different agent memory strategies and agent memory frameworks depending on what they need to store and retrieve.

Combining Retrieval and Memory into an Effective SystemAn agent with retrieval but no memory re-derives the same conclusions every session and can’t personalize anything.

Retrieval brings in external information the agent needs at the moment, such as documenta…

3 weeks, 3 days назад @ machinelearningmastery.com
7 Async Patterns for Running Agents Concurrently in Python
7 Async Patterns for Running Agents Concurrently in Python 7 Async Patterns for Running Agents Concurrently in Python

Share Post ShareIn this article, you will learn seven async patterns for running AI agents concurrently in Python, what each pattern is suited for, and the production-level pitfalls to watch out for with each.

Topics we will cover include:Core async patterns such as fire and forget, scatter-gather, task groups, and producer-consumer queues, and when to reach for each one.

Here are seven async patterns for running agents concurrently, along with the production catches that come with each.

Fire and Forget (Detached Background Execution)You launch an agent task and move on without waiting for it to finish.

Supervised Task GroupsIntroduced in Python 3.11, task groups give you a structured versi…

3 weeks, 4 days назад @ machinelearningmastery.com
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026? Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?

Then we walked through the fastest way to get inference running locally in Run a Local AI Model in 15 Minutes: Your First Ollama Setup.

Spend enough time in the local AI ecosystem, though, and you’ll notice Ollama isn’t the only option competing for your hard drive.

Three tools dominate the local AI runtime landscape: Ollama, LM Studio, and llama.cpp.

Ollama (Via its dedicated, background-daemon CLI) ollama run llama3 .

There’s a well-worn progression in the local AI community that maps almost exactly to the three tools covered here: LM Studio → Ollama → llama.cpp.

1 month, 1 week назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 2 weeks, 5 days назад
Eleven Years Later
Eleven Years Later Eleven Years Later

You also remember a hazy promise to yourself from last year, that you would write weirder posts and explore different forms of writing.

You remember how you used to post once a month, and now you’re posting once every two months.

You remember you used to check this more closely, and now don’t check it at all.

> Check time you’ve spent writing posts.

With some dismay, you noticed that most of your posts have been about AI in one way or another.

2 weeks, 5 days назад @ alexirpan.com
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

1 month, 3 weeks назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

3 months, 3 weeks назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

5 months, 4 weeks назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

7 months, 1 week назад @ alexirpan.com
Authentic Imperfection
Authentic Imperfection Authentic Imperfection

* * *I’ve been thinking about the anger surrounding generative AI.

To keep things fair, he took the best human images and best AI images, meaning human art from famous artists, and AI art from prompters skilled at removing obvious tells of image generation.

When people complain about AI slop, I see it as a complaint against the deluge of default style AI images.

We’ve seen this happen in all forms: AI text, AI music, older forms of computer generated content like CGI.

As much as we celebrate imperfection, digital imperfection is a step too far.

9 months, 3 weeks назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 6 months, 1 week назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

6 months, 1 week назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

7 months, 1 week назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 2 days, 19 hours назад
Introducing WeatherNext 3, our most advanced and accurate global weather AI model
Introducing WeatherNext 3, our most advanced and accurate global weather AI model Introducing WeatherNext 3, our most advanced and accurate global weather AI model

Most AI weather models, including WeatherNext 2, are trained on data from numerical weather prediction (NWP) models.

By ingesting a mosaic of live, global geostationary satellite data, our new model gains a rich, continuously updating view of the atmosphere.

Traditional models struggle here because they train on representations of the atmosphere that lack detail and miss extreme local variations.

This allows us to make global forecasts on a 5-kilometer grid that account for regional details like topography.

Precipitation forecasting at breakthrough accuracyGlobal weather models notoriously struggle to accurately predict precipitation.

2 days, 19 hours назад @ blog.google
Proactive cyber defense for governments and enterprises
Proactive cyber defense for governments and enterprises Proactive cyber defense for governments and enterprises

Defenders wanting to use advanced AI have faced a difficult dilemma: adopt enormous frontier models that could be expensive to deploy and difficult to control across enterprise codebases, or turn to smaller open-weight models that might struggle with complex vulnerability remediation and require teams to build their own tooling and infrastructure from scratch.

Today, we’re launching our Fairwind Program to bring the best of Google’s AI and cyber defense capabilities to a trusted group of Google Cloud customers, government agencies, and cybersecurity partners, to help them proactively solve cyber risks at scale.

As a first step, the Fairwind Program will give defenders access to powerful and…

3 days, 18 hours назад @ blog.google
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7.

Gemini 3.8 introduces 2 variants:Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains.

It is available at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.

Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in v…

3 days, 18 hours назад @ blog.google
Introducing agentic video understanding with Gemini
Introducing agentic video understanding with Gemini Introducing agentic video understanding with Gemini

Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite.

This new capability improves accuracy while dramatically reducing token usage and costs for video analysis.

Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.

BenchmarksUnlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adju…

4 days, 17 hours назад @ blog.google
Gemini Omni 1.1 Flash lets you build with more control
Gemini Omni 1.1 Flash lets you build with more control Gemini Omni 1.1 Flash lets you build with more control

Today, we’re introducing Gemini Omni 1.1 Flash, a new suite of creative controls and generative video capabilities to support developers.

Gemini Omni brought real-world reasoning to generative creation, and today’s updates make Omni 1.1 production-ready for professional use via the Gemini API in Google AI Studio.

Whether you’re building generative video workflows, creative tools, or media editing software, these updates make generative video more controllable, faster to iterate on, and polished for real-world deployment.

With Omni 1.1, the model can now analyze up to 10 seconds of prior context — a leap from previous models that only referenced the final second.

The result is improved visua…

1 week, 2 days назад @ blog.google
Piloting the world's first double-blind AI evaluations
Piloting the world's first double-blind AI evaluations Piloting the world's first double-blind AI evaluations

If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment.

To truly measure what they know, they must have no visibility of the test questions until it's time to take the exam.

That is the exact challenge the industry faces when evaluating advanced AI models.

Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing.

At Google, we assess our AI systems using a broad spectrum of evaluations throu…

1 week, 2 days назад @ deepmind.google
Intelligent transcription with Gemini 3.5 Transcribe
Intelligent transcription with Gemini 3.5 Transcribe Intelligent transcription with Gemini 3.5 Transcribe

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions.

Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines.

Get more precise and i…

1 week, 3 days назад @ blog.google
From Atari to EVE Online: Building on 15 Years of AI Research in Games
From Atari to EVE Online: Building on 15 Years of AI Research in Games From Atari to EVE Online: Building on 15 Years of AI Research in Games

Now, we’re partnering with game developers to prototype new gameplay experiences that push the frontiers of both gaming and AI.

Games as the engine of AI researchOur journey began when a small team trained a deep neural network to play Atari 2600 games directly from raw pixels.

For each game, AI enriched the playing experience.

For game developers, a truly general gaming agent would unlock AI capabilities that work with existing games — no modifications to the game code required.

To develop SIMA agents safely and responsibly, we've partnered with acclaimed game studios and we are building a growing portfolio of games for AI research.

2 weeks, 1 day назад @ deepmind.google
Introducing Gemini 3.7 Flash
Introducing Gemini 3.7 Flash Introducing Gemini 3.7 Flash

3.7 Flash shows strong gains over 3.6 Flash in coding tasks like debugging and issue resolution.

In web development, 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts.

It outperforms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowledge-dense fields like finance, law, and biosciences, 3.7 Flash delivers improved reasoning and accuracy.

It also surpasses 3.6 Flash in AutomationBench, demonstrating it can more effectively complete real-world business workflows (30.4% vs 17.0%).

3 weeks, 2 days назад @ blog.google
Putting sign language AI into users’ hands
Putting sign language AI into users’ hands Putting sign language AI into users’ hands

Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.

AI's ability to process spoken languages has advanced rapidly over recent decades, enabling automatic translation, dictation, and conversational interfaces that feel effortless to hearing users.

Yet this technological revolution has not reached the world’s more than 200 sign languages — and the estimated 70 million Deaf and hard of hearing people who use them.

With it, we are bringing sign language AI out of the lab and into consumer products for the first time: SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with…

3 weeks, 3 days назад @ deepmind.google
WeatherNext: AI model achieves breakthrough in forecasting cyclones
WeatherNext: AI model achieves breakthrough in forecasting cyclones WeatherNext: AI model achieves breakthrough in forecasting cyclones

Predicting how dangerous cyclones develop is a longstanding challenge where every hour counts.

Today, in a paper published in Nature, we show that our WeatherNext AI model achieved state-of-the-art accuracy in predicting a cyclone's track, intensity, and wind structure.

During the 2025 hurricane season, our model helped the NHC to make a historic forecast for Hurricane Melissa by predicting the storm’s rapid intensification and landfall in Jamaica.

Given this broad impact, we are now open sourcing our WeatherNext 2 and WeatherNext Cyclones models used during the hurricane season.

How WeatherNext predicts weather and cyclones

1 month назад @ deepmind.google
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics.

Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function.

Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6.

Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.

Gemini Robotics ER 2 improves this tool orchestration workflow.

1 month, 1 week назад @ blog.google
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control

Our newest music generation model, Lyria 3.5, delivers significant advancements across musicality, lyrics, and vocal quality, empowering you to craft richer tracks.

We’re rolling it out today in Google Flow Music, where we want to help you create songs you love, with creative control.

Enhanced lyrics: Generate higher quality lyrics with improved prompt adherence and structural awareness.

Generate higher quality lyrics with improved prompt adherence and structural awareness.

Improved vocals: Bring more expression and emotion to your songs with more realistic and emotionally nuanced vocals, plus improved pronunciation.

1 month, 1 week назад @ blog.google
Gemini Robotics 2 brings whole body intelligence to robots
Gemini Robotics 2 brings whole body intelligence to robots Gemini Robotics 2 brings whole body intelligence to robots

From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasksFor decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.

Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots.

As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration.

Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks.

And this profound intelligence can also run locally on-device while seamlessly a…

1 month, 1 week назад @ deepmind.google
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

In December, we shared our commitment to the White House's Genesis Mission — the national effort to harness AI and double the pace of American scientific discovery within a decade.

Today, at the DOE Genesis Mission Summit 2026, we are expanding this by committing $40 million of AI tokens and cloud credits for researchers in support of the Genesis Mission.

WeatherNext — a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

— a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

Driving American innovationThe Genesis Mission represents an opportunity to transform research and science across America.

1 month, 2 weeks назад @ cloud.google.com
Google
последний пост 1 day, 18 hours назад
Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption
Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption

When Google's Finance Engineering team needed to modernize their legacy data layer, they chose Spanner, a globally distributed, strongly consistent, multi-model database with high availability capabilities.

Further, doing so without disruption would have required implementing multi-phase dual-write architectures across every DAO in our codebase.

To solve this, we took an alternative approach: We built an automated refactoring pipeline powered by Antigravity CLI in headless mode.

This helped us accelerate our migration velocity significantly while maintaining strict data parity in our staging environments as we prepare for production.

The challenge: Anatomy of a dual-write migrationWhen migr…

1 day, 18 hours назад @ cloud.google.com
Getting started with Mantis, our open-source bug finding-and-fixing harness
Getting started with Mantis, our open-source bug finding-and-fixing harness Getting started with Mantis, our open-source bug finding-and-fixing harness

AI models have clearly proven their ability to discover and exploit vulnerabilities without much, if any, human assistance.

To help defenders gain the advantage with AI, we built the Mantis harness to automate the discovery, triage, reproduction, and patching of software vulnerabilities.

Available to all as an open-source framework, Mantis is part of Google’s internal approach to find and fix vulnerabilities at machine-speed.

Mantis distills decades of cybersecurity expertise across a wide spectrum of codebases, and is available on GitHub.

Here’s how you can get started using Mantis.

3 days, 18 hours назад @ cloud.google.com
Reimagining work: How Pythian’s internal AI playbook delivers customer ROI
Reimagining work: How Pythian’s internal AI playbook delivers customer ROI Reimagining work: How Pythian’s internal AI playbook delivers customer ROI

Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian observed firsthand why so many enterprise AI initiatives stall out or fail.

To solve this, we engineered the Pythian AI Operating Model — a multifaceted, end-to-end framework designed to take enterprise AI from high-level strategy all the way into sustained production.

XOps (AI production management): While deploying an agent is 20% of the journey, maintaining accuracy in production is 80%.

Ready to build your AI operating model?

It also requires an end-to-end AI operating model.

1 week, 2 days назад @ cloud.google.com
FinOps for the AI era: New flexible billing and cost controls for agents
FinOps for the AI era: New flexible billing and cost controls for agents FinOps for the AI era: New flexible billing and cost controls for agents

That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.

Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs)If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage.

Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.

To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing …

1 week, 3 days назад @ cloud.google.com
Now introducing Gemini Enterprise for Legal
Now introducing Gemini Enterprise for Legal Now introducing Gemini Enterprise for Legal

Few professions are as exacting as the practice of law.

A team reviewing a contract or building a case works inside strictly privileged information, firm-specific playbooks, and a body of law that changes constantly.

For legal work, it is nowhere near sufficient.

Only in combination do they produce something a firm or a legal department can put into production and actually rely on.

Today we're bringing that to legal practice with Gemini Enterprise for Legal, part of our new suite of purpose-built industry solutions.

1 week, 4 days назад @ cloud.google.com
Now introducing Gemini Enterprise for Financial Services
Now introducing Gemini Enterprise for Financial Services Now introducing Gemini Enterprise for Financial Services

Four components, built for financial workGemini Enterprise for Financial Services delivers an integrated, secure environment configured for rapid deployment with four core components:1.

They are available inside the Financial Research agent and to any agent your teams build.

At its core is the Financial Research agent which is a Google-built, Google-managed agent that runs end-to-end research with full explainability.

Open ecosystem of connectors across the financial technology stackGemini Enterprise connects directly to core financial systems via secure MCP connectors.

Finnhub: Provides real-time financial APIs, global fundamentals, and earnings call transcripts for in-depth financial rese…

1 week, 4 days назад @ cloud.google.com
How agents can delegate better
How agents can delegate better How agents can delegate better

At Google Cloud, we’re learning a similar lesson when it comes to building and deploying AI agents in enterprise workflows.

To do so, AI agents need to become good delegators.

This work opens up new opportunities for customers building AI agents that can communicate, share tasks, and coordinate towards set objectives.

Principle #3: Respect sensitive dataMany workflows handle private, sensitive data, and AI agents need to respect those boundaries and permissions.

Zero-knowledge proofs enable one AI agent to prove to the other AI agent that a planned computation was performed correctly, without revealing the data itself.

2 weeks, 1 day назад @ cloud.google.com
Cloud CISO Perspectives: Sticking to security fundamentals in the AI era
Cloud CISO Perspectives: Sticking to security fundamentals in the AI era Cloud CISO Perspectives: Sticking to security fundamentals in the AI era

Defending against AI powered security threats requires more than accelerating current security practices; it means stepping back and beginning with the security foundation and layered defenses.

High-risk indicators automatically get flagged for human review, while we’ve replaced static threat models with dynamic product dossiers that update in real-time.

The intense global focus on AI vulnerabilities has brought cybersecurity to the forefront of boardroom and executive attention like never before.

By aligning security fundamentals with business objectives and using AI to enhance defense, we can lead our organizations securely into the future.

To learn more about building and maintaining str…

2 weeks, 1 day назад @ cloud.google.com
Expanding Google Antigravity for enterprise customers
Expanding Google Antigravity for enterprise customers Expanding Google Antigravity for enterprise customers

What our customers are sayingFrom rapid code generation to end-to-end task automation, Google Antigravity is giving engineering teams the momentum of cutting edge AI development backed by the stability, governance, and scale of Google Cloud.

Here is how leading enterprise customers and partners are driving measurable outcomes in production:“Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud.

With Google Antigravity supported across developers' preferred IDEs, the desktop app, and the CLI, Cognizant can seamlessly embed agentic engineering across our global delivery cente…

2 weeks, 2 days назад @ cloud.google.com
10 questions every startup should answer before moving to production with their AI prototype
10 questions every startup should answer before moving to production with their AI prototype 10 questions every startup should answer before moving to production with their AI prototype

You grab an API key from Google AI Studio at breakfast, paste it into Antigravity, and by lunch you’ll have a nascent prototype of your product.

It's common to bump into these three challenges as you build out your stack:A leaked API key racks up a large bill in 48 hours .

#1 Where should I start: Google AI Studio or Gemini Enterprise Agent Platform?

Google AI Studio (with the Gemini Developer API) is the fastest path from an idea to working code.

A browser IDE, an API key, a generous free tier, and no cloud project to configure.

2 weeks, 2 days назад @ cloud.google.com
How AlloyDB ScaNN scales vector search to 10 billion vectors
How AlloyDB ScaNN scales vector search to 10 billion vectors How AlloyDB ScaNN scales vector search to 10 billion vectors

A key part of this is its ScaNN index, which now operates efficiently at a scale of 10 billion vectors.

This was achieved through a major architectural enhancement: an innovative four-level tree (preview) paired with efficient memory usage.

The 10 billion vector scale challengeScaling to a 10 billion vector workload presents significant memory and computational challenges.

Previous AlloyDB ScaNN tree-based index was limited to two- or three-level tree configurations, and attempting to scale those structures led to several bottlenecks:Increased compute intensity: Larger tree structures demand significantly more operations for both index construction and query traversal.

Solution: Four-level …

2 weeks, 2 days назад @ cloud.google.com
Building cost-effective, high-throughput gen AI workflows in Google Dataflow
Building cost-effective, high-throughput gen AI workflows in Google Dataflow Building cost-effective, high-throughput gen AI workflows in Google Dataflow

However, by integrating generative AI agents, we can move beyond static logic to adaptive execution.

This allows streaming workflows to dynamically construct plans, query databases, and trigger custom remediation paths at runtime depending on the content of the data.

However, streaming systems face a fundamental engineering hurdle when executing gen AI workflows: scale, latency, and cost.

This pattern addresses the scale and complexity challenge by combining Google Dataflow, Google Cloud's fully managed, serverless execution service for Apache Beam, and the Agent Development Kit (ADK) to build a hybrid streaming pipeline.

Latency: Multi-step workflows (which involve database lookups and ext…

2 weeks, 4 days назад @ cloud.google.com
How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2
How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2 How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform, powered by Gemini Multimodal Embeddings 2 merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.

Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.

Illuminating the visual modality: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography.

Extending RAG with multimodal embeddings…

2 weeks, 4 days назад @ cloud.google.com
Building operational resilience with agentic AI in financial services
Building operational resilience with agentic AI in financial services Building operational resilience with agentic AI in financial services

For financial institutions, operational resilience has long been embedded in regulatory and supervisory expectations — to say nothing of the high expectations of consumers.

To meet these conditions, Deutsche Bank developed an AI-powered agentic resilience platform that modernized its regulatory tabletop resilience exercises at scale and turned manual preparation into context-aware and evidence-ready simulations grounded in actual operational data.

Across financial services, supervisory expectations are evolving and as they do, banks’ tabletop exercises must reflect their production dependencies, real operating conditions, and compliance with consistent evidence standards more directly.

From…

2 weeks, 4 days назад @ cloud.google.com
Using BigQuery Graphs with measures for trusted agentic workloads
Using BigQuery Graphs with measures for trusted agentic workloads Using BigQuery Graphs with measures for trusted agentic workloads

BigQuery Graph helps organizations move beyond flat, static tables to represent enterprises exactly how they exist in the physical world: as interconnected business entities with real-world dependencies.

With the support of measures in BigQuery Graph (preview), we are unifying governed metrics with relationship mapping.

Measures in BigQuery Graph solves this by letting you map existing tables to a property graph in-place with zero ETL.

Business metrics (measures) calculate how your business performed.

BigQuery Graph solves this natively.

3 weeks, 2 days назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 5 days, 18 hours назад
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.

GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery.

To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.

Listen now Opens in a new tabEfficiency as an enabler of discoveryGigaPath and GigaTIME demonstrated what pathology foundation models can learn from whole slides and …

5 days, 18 hours назад @ microsoft.com
Broadening access to Skala creates a faster path to predictive DFT
Broadening access to Skala creates a faster path to predictive DFT Broadening access to Skala creates a faster path to predictive DFT

and is being integrated into , , and , bringing next-generation DFT accuracy closer to the communities that rely on these codes every day.

Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and integrated in all relevant scientific and industrial workflows.

Alongside these integration efforts, we are introducing a living benchmark that tracks the computational performance of successive, increasingly optimized Skala releases.

Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and accessible across a broader range of relevant scien…

2 weeks, 2 days назад @ microsoft.com
MindTopo reveals VLMs’ spatial reasoning abilities
MindTopo reveals VLMs’ spatial reasoning abilities MindTopo reveals VLMs’ spatial reasoning abilities

At a glance MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as connectivity, enclosure, order, separation, and knots.

How MindTopo defines topological spaceMost spatial evaluations for multimodal models focus on Euclidean properties such as distance, direction, size, and relative position.

MindTopo pairs questions about static scenes with interactive tasks that require models to preserve or change the same topological relations.

MindTopo maps reasoning and planning tasks to continuity, separation, order, enclosure, and knots.

Closing that gap may require models that carry an explicit topological state, or…

3 weeks, 3 days назад @ microsoft.com
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands.

The CARE-X model.

The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses.

CARE-X: Toward clinically useful radiology AICARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model.

3 weeks, 4 days назад @ microsoft.com
Orchard: An open framework for scalable agentic AI
Orchard: An open framework for scalable agentic AI Orchard: An open framework for scalable agentic AI

At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.

To address this gap, we introduce Orchard (opens in new tab), an open-source framework for scalable agentic modeling.

Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification.

(opens in new tab) We are also releasing the training data and evaluation methods used to build them.

By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research.

1 month назад @ microsoft.com
Echoverse: Deep, evolving environments for computer-use agents
Echoverse: Deep, evolving environments for computer-use agents Echoverse: Deep, evolving environments for computer-use agents

At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters).

Shallow worlds backfire; deep worlds transferA shallow world is the cheap option.

A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%.

First, more deep worlds for the closed domains public benchmarks cannot reach.

1 month, 1 week назад @ microsoft.com
EvoLib: Turning experience into evolving knowledge
EvoLib: Turning experience into evolving knowledge EvoLib: Turning experience into evolving knowledge

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive.

As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to co…

1 month, 1 week назад @ microsoft.com
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

1 month, 3 weeks назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

1 month, 4 weeks назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

1 month, 4 weeks назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

2 months, 1 week назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

2 months, 1 week назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

2 months, 1 week назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

2 months, 1 week назад @ microsoft.com
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis

At a glance Talos is an open-source tool for automated, iterative reanalysis of genomic data in rare disease.

Deployed across a prospective cohort of almost 5,000 undiagnosed patients, Talos delivered 241 new diagnoses (5.1% additional yield).

On monthly iterative cycles, analysts only needed to review one new variant per 200 patients, demonstrating that frequent, systematic reanalysis can be run sustainably.

Why genome reanalysis mattersGenomic testing has transformed the diagnosis of rare disease, but even with this advancement, more than half of patients remain undiagnosed after their first test.

Looking aheadTalos reframes genomic reanalysis from a rare, labor-intensive event into a con…

2 months, 1 week назад @ microsoft.com
MIT AI MIT AI
последний пост 3 days, 14 hours назад
From MIT to IBM, expediting AI and quantum deployment
From MIT to IBM, expediting AI and quantum deployment From MIT to IBM, expediting AI and quantum deployment

Here, the MIT-IBM Computing Research Lab served as a conduit for research relationship building and the flow of their expertise to industry applications.

“I started to work [on trustworthy AI] with IBM researchers from day 1 in my PhD, because it was funded by MIT-IBM,” says Ko.

After graduating in 2024, Ko joined IBM Research to continue her work on trustworthy AI as a research scientist.

During this time, Arunachalam focused on quantum machine learning and areas where quantum computing would be superior to classical computing, increasingly prioritizing provability grounded in theory to heuristics.

“One thing which I’ve been a huge fan of is exposing connections between different fields.” …

3 days, 14 hours назад @ news.mit.edu
System helps humans predict when self-driving cars will make mistakes
System helps humans predict when self-driving cars will make mistakes System helps humans predict when self-driving cars will make mistakes

In road tests on a private track, CW-Net explanations helped safety drivers more accurately predict vehicle behavior; a larger simulation study with nonexpert users yielded similar results.

In the longer term, this technique could boost the safety and transparency of autonomous vehicles, while building appropriate trust in drivers and passengers.

The researchers trained CW-Net to predict concepts using a dataset of 130 million examples of scenes from self-driving cars, with multiple labeled concepts in each scene.

But CW-Net explanations revealed that the model wasn’t properly configured to detect the cyclist and chose a trajectory that would have caused a collision.

CW-Net explanations sig…

3 days, 19 hours назад @ news.mit.edu
Walter Torous named executive director of MIT Center for Real Estate
Walter Torous named executive director of MIT Center for Real Estate Walter Torous named executive director of MIT Center for Real Estate

Walter Torous, senior lecturer in the MIT Department of Urban Studies and Planning (DUSP) and the MIT Sloan School of Management, and director of the Master of Science in Real Estate Development Program (MSRED), was recently named executive director of the MIT Center for Real Estate (CRE) — effective July 1, 2026.

“Demographic changes, as an aging population stays longer in their homes, are creating an imbalance in residential real estate markets.

“A lot of our alums have assumed important positions in the real estate industry around the world,” he says.

This research reflects the growing importance of AI and large language models to every aspect of real estate decision-making.

“It’s import…

4 days, 13 hours назад @ news.mit.edu
Ila Kumar: Innovating with communities
Ila Kumar: Innovating with communities Ila Kumar: Innovating with communities

Through those experiences, Kumar began to question whether the technology she was helping to develop was having the sustained impact she hoped for.

“And I wasn’t seeing that what I was doing had a long-term impact.”Rather than walking away from technology altogether, Kumar began to rethink how it was created.

“We are not sitting at MIT designing tools and just throwing them at people,” Kumar says.

As a result, Kumar has increasingly focused on supporting care providers in talking with young people about AI.

She has led training workshops with organizations that serve young people impacted by trauma or involved in the child welfare system.

5 days, 6 hours назад @ news.mit.edu
MIT Quantum Initiative launches postdoctoral fellowship program
MIT Quantum Initiative launches postdoctoral fellowship program MIT Quantum Initiative launches postdoctoral fellowship program

The MIT Quantum Initiative (QMIT) has launched a new postdoctoral fellowship program to accelerate interdisciplinary quantum research and develop the next generation of scientific leaders working at the frontiers of quantum science and technology.

“Quantum science and technology is in a period of extraordinary opportunity, opening new pathways to solving problems across computation, materials, sensing, and communication.

This fellowship is designed to create exactly those kinds of opportunities.”The QMIT Fellowship is intentionally designed to foster an interdisciplinary research community.

The fellows will be embedded across the research areas that define QMIT, including quantum computing,…

5 days, 15 hours назад @ news.mit.edu
How an MIT research project became a global programming language
How an MIT research project became a global programming language How an MIT research project became a global programming language

That research project turned into a lab at MIT, and the lab turned into the company JuliaHub.

Building scientific applications with multidisciplinary teams of scientists, engineers, and programmers is challenging,” JuliaHub co-founder and CEO Viral Shah says.

“We wanted to create something as easy to use as Python or MATLAB but as fast as the C programming language,” Shah says.

In another case, researchers used Julia to create a program for avoiding aircraft collisions.

“Over the years we’ve seen industrial, government, and academic users doing all kinds of interesting things with the Julia language,” Edelman says.

6 days, 6 hours назад @ news.mit.edu
Looking beyond natural sequences
Looking beyond natural sequences Looking beyond natural sequences

Adding this framework to a protein design pipeline will allow researchers to design structurally feasible proteins with sequences that don’t resemble those of any native protein.

Birnbaum was first interested in strategic applications of something researchers call “noise,” or adding variations to a protein structure during training.

Noise decreases the tendency of the model to overly mimic native sequences, increasing the diversity of structures for which it’s able to generate sequences.

Birnbaum acknowledges that in trying to shift away from adhering to native sequences, incorporating evolutionary information is, in some ways, still a reliance on them.

Protein design in the age of AI“Once …

1 week, 2 days назад @ news.mit.edu
AI helps design new materials that work in the real world
AI helps design new materials that work in the real world AI helps design new materials that work in the real world

One reason for the translation gap is that current models don’t reliably factor in the chemical stability of the materials they generate, and unstable materials aren’t very useful in the real world.

It allows you to screen out the unstable materials to generate higher quality materials.

The researchers then used their approach to generate material candidates with high thermal conductivity and easy polarization in an electric field.

“These are materials useful for the semiconductor industry and high thermal conductivity materials relevant to data center cooling,” Ju Li says.

Still, the approach could be used to generate stable new crystalline materials with a host of important properties.

1 week, 4 days назад @ news.mit.edu
Generating scenarios for extreme events, without extreme data
Generating scenarios for extreme events, without extreme data Generating scenarios for extreme events, without extreme data

To answer these questions, communities will first need to know how such extreme events could unfold.

Yet most methods that assess a region’s risk depend on extreme events of the past to characterize even more extreme, worst-case scenarios in the future.

The key to their method is that it does not need to know about previous extreme events in order to generate plausible future extreme events.

“Financial market crashes are extreme events that are a complicated combination of things, involving many different sectors,” Chang says.

“There is no method that does this efficiently to predict events that happen rarely.”Extreme learningThe team’s new algorithm generates plausible, unprecedented extre…

1 week, 5 days назад @ news.mit.edu
Paving the way for greener ammonia production
Paving the way for greener ammonia production Paving the way for greener ammonia production

The traditional way of making ammonia, in use for more than a century and accounting for the vast majority of production, is the Haber-Bosch process, which relies on fossil fuels to provide the needed heat.

Now, researchers at MIT have developed a way to predict which materials could be most promising as catalysts in electrochemical ammonia production.

“If we can somehow find a catalyst that reduces the energy needed and is more selective for ammonia production,” Athanitis says, “then we could essentially hit the jackpot.” A more selective catalyst would produce more ammonia while reducing unwanted side reactions.

Different materials can improve different parts of the reaction, and research…

2 weeks, 2 days назад @ news.mit.edu
When AI art has no author: Study finds generated images often can’t be traced to training data
When AI art has no author: Study finds generated images often can’t be traced to training data When AI art has no author: Study finds generated images often can’t be traced to training data

So the team put the ensembles head to head with 24 conventional diffusion models trained on the exact same data.

One nice surprise in the numbers: The more training data, the better the ensembles held up against their single-model counterparts, a hint that they may actually be more data-efficient.

Take one generated image, then imagine every alternate version of it, each produced by removing a different piece of the training data.

The distance between the original and its most different alternate, the counterfactual radius, captures the most that any single piece of training data could have mattered.

Gifford sees the finding as bearing directly on the legal question of whether model outputs…

2 weeks, 4 days назад @ news.mit.edu
Q&A: Rethinking how innovation happens
Q&A: Rethinking how innovation happens Q&A: Rethinking how innovation happens

I wanted to write a book that condensed all of those connections, because the innovation process at that scale is really the intersection of many different fields.

The dominant ones are science and economics, because those are the underlying principles that drive how innovation happens.

What started as a practical effort to make these programs work became a broader and somewhat unexpected interest in the innovation process itself.

A: People think that all research investment works the same way if the goal is economic impact.

What I point out in the book is that being involved in the innovation process makes you T-shaped: You have technical depth in one area and a broad working knowledge of …

2 weeks, 5 days назад @ news.mit.edu
Q&A: Rethinking how innovation happens
Q&A: Rethinking how innovation happens Q&A: Rethinking how innovation happens

I wanted to write a book that condensed all of those connections, because the innovation process at that scale is really the intersection of many different fields.

The dominant ones are science and economics, because those are the underlying principles that drive how innovation happens.

What started as a practical effort to make these programs work became a broader and somewhat unexpected interest in the innovation process itself.

A: People think that all research investment works the same way if the goal is economic impact.

What I point out in the book is that being involved in the innovation process makes you T-shaped: You have technical depth in one area and a broad working knowledge of …

2 weeks, 5 days назад @ news.mit.edu
With a feel for physics, AI models simulate a wider range of real-world scenarios
With a feel for physics, AI models simulate a wider range of real-world scenarios With a feel for physics, AI models simulate a wider range of real-world scenarios

Artificial intelligence models are jacks of many trades, including writing, generating images, and creating 3D models.

To build an AI system that can reliably simulate a variety of physical scenarios, engineers need a range of physics data at a scale that isn’t yet feasible.

Industry successThe researchers found that GeoPT was particularly skilled at simulating industrial scenarios, as it outperformed state-of-the-art simulation models across benchmarks.

Likewise, its simulations of how light would pass through what was essentially a toy rabbit were accurate, despite never training on that 3D model or light physics beforehand.

The demonstrated success in a wide range of application domains …

3 weeks, 5 days назад @ news.mit.edu
Solving the solvent problem
Solving the solvent problem Solving the solvent problem

“It’s supposed to be an ion conductor.” But unfortunately, most electrolytes get involved in unwanted chemical reactions with the electrodes, which can greatly undermine battery stability.

The team’s goal, accordingly, was to identify solvent molecules that are small enough to improve ion transport while still maintaining electrolyte stability.

There is, however, a complicating factor — a trade-off to be addressed: Faster ion transport often comes at the expense of electrolyte stability.

By carefully tailoring the size of solvent molecules, the authors demonstrate a new design strategy that could enable lower-cost, higher performance batteries.”The group is not done.

The overriding goal of …

1 month назад @ news.mit.edu
Berkeley AI
последний пост 1 month, 1 week назад
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map.

Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

With Apple Silicon in hundreds of millions of MacBooks and Mac Studios, MLX enables local AI inference without cloud costs.

Building an MLX backendTo bring K-Search to Apple Silicon, we first built a native MLX backend.

Evaluated on mamba-370m f16, M1 Max 64GB:Metric mlx-mamba (ours) mlx-lm (community) mamba.py Decode 152 tok/s 116 tok/s 40 tok/s Prefill L=512 5,751 tok/s 329 tok/s 1,089 tok/s Prefill L=1024 …

1 month, 1 week назад @ bair.berkeley.edu
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon InteractionOverview of ABBEL compared to traditional recursive summarization.

Belief grading can be thought of as adding an auxiliary RL task, using heuristics designed to capture what makes a good belief as the reward.

With domain-knowledge belief grading, ABBEL approaches or exceeds FULL CTX in this setting; without belief grading, learning is slower.

Context compression methods generate dense representations which, while computationally efficient, sacrifice human-understandability (Kontonis et al., 2026, Eyuboglu et al., 2025, Gupta et al., 2025, Chevalier et al., 2023, Deng et al., 2025, Deng et al., 2025, Bulatov et al., 2022).…

1 month, 1 week назад @ bair.berkeley.edu
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

2 months назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

2 months, 1 week назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

4 months назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

4 months, 2 weeks назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

5 months, 3 weeks назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

7 months, 4 weeks назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 1 day, 12 hours назад
Deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore
Deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore Deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore

This post shows how to deploy a multimodal WhatsApp ordering assistant built with Amazon Bedrock AgentCore and Amazon Nova 2.

Amazon Nova 2 Lite handles text through the Amazon Bedrock Converse API, and Amazon Nova 2 Sonic handles real-time speech on voice notes and calls.

C. Artificial intelligence and machine learning (AI/ML): Amazon Nova 2 Lite and Amazon Nova 2 Sonic invoked through Amazon Bedrock, plus the shared AgentCore memory keyed by a hashed customer_id for cross-channel continuity.

It runs the conversation with Amazon Nova 2 Lite (text) or Amazon Nova 2 Sonic (voice) through Amazon Bedrock.

Amazon Nova 2 Lite and Amazon Nova 2 Sonic run the conversations.

1 day, 12 hours назад @ aws.amazon.com
Designing lifecycle policies for AgentCore memory
Designing lifecycle policies for AgentCore memory Designing lifecycle policies for AgentCore memory

Memory lifecycle policies help long-running agents on Amazon Bedrock AgentCore stay effective by systematically managing what they remember and forget.

We walk through a deployable architecture using AgentCore memory (a capability of Amazon Bedrock AgentCore), AWS Step Functions, and Amazon Bedrock to run a nightly lifecycle workflow.

The workflow proceeds as follows:TTL Expiration: The Memory Pruner queries AgentCore memory for records older than the configured TTL (default: 90 days) and deletes them.

ConclusionWe showed how to build memory lifecycle policies for Amazon Bedrock AgentCore agents using AWS Step Functions and Amazon Bedrock.

To learn more, see the Amazon Bedrock AgentCore doc…

1 day, 17 hours назад @ aws.amazon.com
Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod
Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod

Why the design choices map cleanly onto Amazon SageMaker HyperPod with Amazon Elastic Kubernetes Service (Amazon EKS).

An Amazon SageMaker HyperPod cluster orchestrated by Amazon EKS, with a GPU instance group of p5en.48xlarge nodes in a single Availability Zone (see Creating an Amazon SageMaker HyperPod cluster with Amazon EKS orchestration).

An Amazon SageMaker HyperPod cluster orchestrated by Amazon EKS, with a GPU instance group of p5en.48xlarge nodes in a single Availability Zone (see Creating an Amazon SageMaker HyperPod cluster with Amazon EKS orchestration).

This post described how you can use NVIDIA Cosmos 3 on Amazon SageMaker HyperPod (EKS) as the substrate for that flywheel.

To …

1 day, 18 hours назад @ aws.amazon.com
Run agent-driven Amazon SageMaker HyperPod operations with InstantStart
Run agent-driven Amazon SageMaker HyperPod operations with InstantStart Run agent-driven Amazon SageMaker HyperPod operations with InstantStart

Amazon Simple Storage Service (Amazon S3), Amazon FSx for Lustre, and Amazon Elastic Container Registry (Amazon ECR) carry images, data, and checkpoints.

Amazon Managed Service for Prometheus and Amazon Managed Grafana receive health and utilization.

[hypd-inst-agent] > Help me create a new HyperPod cluster > Creating a HyperPod cluster is a multi-step process: 1.

For fleet-level metrics, HyperPod publishes to Amazon Managed Service for Prometheus, with dashboards in Amazon Managed Grafana through the HyperPod observability add-on.

ConclusionIn this post, we walked through HyperPod InstantStart, an open source control plane that composes Amazon EKS orchestration with the managed capabilitie…

1 day, 18 hours назад @ aws.amazon.com
Customizing your knowledge base on Amazon Bedrock for large and complex documents using Amazon Textract
Customizing your knowledge base on Amazon Bedrock for large and complex documents using Amazon Textract Customizing your knowledge base on Amazon Bedrock for large and complex documents using Amazon Textract

Amazon Textract can extract text from multi-page PDF documents, including those with complex layouts and embedded images.

Amazon Textract can parse and extract text from Word documents, including tables, images, and other embedded objects.

While primarily a text extraction tool, Amazon Textract can extract text from Excel spreadsheets, including cell contents and table data.

Knowledge base setupOpen the Amazon Bedrock console and navigate to Amazon Bedrock Knowledge Bases , the fully managed capability for building retrieval-augmented generation solutions.

The post-deployment steps include configuring the S3 bucket, uploading sample documents, and setting up the Amazon Bedrock knowledge bas…

1 day, 18 hours назад @ aws.amazon.com
How Intuit built an agentic disaster recovery assistant with Amazon Bedrock
How Intuit built an agentic disaster recovery assistant with Amazon Bedrock How Intuit built an agentic disaster recovery assistant with Amazon Bedrock

To close that gap, we built an agentic disaster recovery assistant with Amazon Bedrock.

To close this gap, we built EWOK Agent, an AI-powered agent built with Amazon Bedrock.

Every harness invocation automatically generates traces, logs, and metrics through AgentCore Observability, a capability of Amazon Bedrock AgentCore, in Amazon CloudWatch.

Amazon Bedrock model access itself incurs no standing charge when idle, and you are billed only for invocations.

To start building your own tool-using agents on Amazon Bedrock, see the Amazon Bedrock Converse API documentation, the tool use (function calling) guide, and Amazon Bedrock Guardrails.

1 day, 18 hours назад @ aws.amazon.com
AI-driven development lifecycle using Amazon Bedrock AgentCore
AI-driven development lifecycle using Amazon Bedrock AgentCore AI-driven development lifecycle using Amazon Bedrock AgentCore

Engineering teams adopting the AI-Driven Development Lifecycle (AI-DLC) with Amazon Bedrock AgentCore and coding agents like Kiro often struggle with the gap between conceptual frameworks and working code.

The first generates Mermaid entity relationship diagrams from SQL schemas using AgentCore runtime, a capability of Amazon Bedrock AgentCore.

The second provides automated code security analysis through a multi-agent architecture that uses AgentCore Gateway, a capability of Amazon Bedrock AgentCore, and AgentCore memory, a capability of Amazon Bedrock AgentCore, along with external tool integrations.

AgentCore memory: Provides persistent session context with a 90-day expiry, and supports s…

2 days, 18 hours назад @ aws.amazon.com
Migrate agentic workloads to Amazon Bedrock AgentCore
Migrate agentic workloads to Amazon Bedrock AgentCore Migrate agentic workloads to Amazon Bedrock AgentCore

Stage 1 transitions it onto Amazon Bedrock AgentCore Runtime, Gateway and Memory, graph unchanged.

Stage 3 hands the loop to an AgentCore harness, a capability of Amazon Bedrock AgentCore, documented here rather than built.

So lookup_order and process_return become MCP tools published by AgentCore Gateway, a capability of Amazon Bedrock AgentCore, from a Lambda function.

Runtime instruments the agent it hosts, and AgentCore Observability, a capability of Amazon Bedrock AgentCore, sends the result to CloudWatch.

AgentCore Identity, a capability of Amazon Bedrock AgentCore, answers the first, referenced on your gateway target in place of GATEWAY_IAM_ROLE .

2 days, 18 hours назад @ aws.amazon.com
Integrating Outlook with Amazon Quick for AI-powered email automation
Integrating Outlook with Amazon Quick for AI-powered email automation Integrating Outlook with Amazon Quick for AI-powered email automation

With Amazon Quick Automate, you can orchestrate complex, multi-step processes that span email, calendars, and other enterprise systems.

Step 3: Authorize Amazon Quick to access your Outlook emailAfter you sign in to Outlook, choose your current Amazon Quick account to authorize Amazon Quick to access your Microsoft Outlook email.

Step 5 (optional): Test Outlook email actionsChoose one of the sample questions, such as checking email, writing an email, or summarizing calendar events.

Use cases and benefitsWith Microsoft Outlook integrated into Amazon Quick, you can address communication and workflow challenges across several categories.

Quick chat agent use casesCustomer support email assista…

2 days, 18 hours назад @ aws.amazon.com
Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock
Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock

OpenAI Codex (Codex) runs its task loop on the developer workstation.

In this post, we walk you through deploying a customer-operated LiteLLM gateway on Amazon Elastic Container Service (Amazon ECS).

Direct access to Amazon Bedrock is the lowest-complexity option when native AWS identity, AWS Identity and Access Management (IAM) policies, and AWS CloudTrail logs meet the customer’s requirements.

Deploy a LiteLLM gateway for Codex on Amazon ECSThis section walks you through deploying the LiteLLM gateway, connecting it to Amazon Bedrock, and configuring Codex to send requests through the gateway.

Developers can continue to configure a stable alias while the system team retains control over th…

2 days, 18 hours назад @ aws.amazon.com
Best practices for building agentic automations with Amazon Quick Automate
Best practices for building agentic automations with Amazon Quick Automate Best practices for building agentic automations with Amazon Quick Automate

Amazon Quick Automate is a multi-agent automation capability within Amazon Quick that helps organizations build, deploy, and maintain these automations at scale.

In this post, I share best practices for building production-grade agentic automations with Amazon Quick Automate.

One of the strengths of Quick Automate is that you can mix agentic steps with deterministic ones.

With Quick Automate you can incorporate human-in-the-loop (HITL) at the moments that matter, and it supports two distinct patterns that behave very differently.

To get started, visit Amazon Quick Automate or see the Quick Automate documentation.

2 days, 18 hours назад @ aws.amazon.com
Embed Quick Sight visuals using Cognito user authentication
Embed Quick Sight visuals using Cognito user authentication Embed Quick Sight visuals using Cognito user authentication

Amazon Quick Sight is the business intelligence engine within Amazon Quick that powers the embedded analytics experience in your application.

This post shows you how to embed individual Amazon Quick Sight visuals into React applications with registered user authentication through Amazon Cognito.

User synchronization and role-based access controlEach Amazon Cognito user who needs to view an embedded visual must also exist as a registered user inside Amazon Quick Sight.

Step 2: Configure the front-end environmentRetrieve Amazon Quick Sight visual identifiersOpen your published Amazon Quick Sight dashboard.

Step 5: Configure the Amazon Quick allowlistAdd the Amazon CloudFront domain to the Ama…

2 days, 18 hours назад @ aws.amazon.com
Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference
Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference

Australian teams working with OpenAI models can now access the latest OpenAI models through Amazon Bedrock.

Amazon Bedrock offers OpenAI GPT-5.6 Sol, Terra, and Luna with global cross-Region inference from both Asia Pacific (Sydney) and Asia Pacific (Melbourne) AWS Regions in Australia.

Your application calls the Amazon Bedrock Runtime endpoint in Asia Pacific (Sydney) or Asia Pacific (Melbourne), and Amazon Bedrock routes the request to a supported commercial AWS Region for processing.

Invoke GPT-5.6 through Amazon Bedrock RuntimeGPT-5.6 supports three access paths on the Amazon Bedrock Runtime endpoint: the OpenAI Responses API, OpenAI Chat Completions API, and Amazon Bedrock Converse API…

3 days, 13 hours назад @ aws.amazon.com
Modernizing and scaling support operations with generative AI on AWS
Modernizing and scaling support operations with generative AI on AWS Modernizing and scaling support operations with generative AI on AWS

To address these constraints, teams can use generative AI on AWS to capture knowledge from operational workflows, apply it during ticket resolution, and surface risks before they impact SLAs.

At the same time, most operational knowledge is shared through training calls, walkthroughs, and troubleshooting sessions.

Solution overviewTo address the limitations of manual documentation, fragmented ticket handling, and reactive workload management, the solution integrates execution and analytics into a single operational system on AWS.

A generative AI-based solution on AWS can change how support teams capture knowledge, resolve tickets, and manage operations.

These gains translate directly into fa…

3 days, 16 hours назад @ aws.amazon.com
How an AWS team detects dashboard content failures at scale using Amazon Bedrock
How an AWS team detects dashboard content failures at scale using Amazon Bedrock How an AWS team detects dashboard content failures at scale using Amazon Bedrock

When a visual content failure is detected on a section they own, they receive a Slack notification.

The following diagram expands it into the end-to-end processing pipeline, including the numeric validation components added to the original visual validation structure.

Visual content validation.

802 content failures detected (0.52 percent of all checks), equivalent to 99.48 percent system content availability.

ConclusionIn this post, we showed how our team built a proactive content validation system.

3 days, 16 hours назад @ aws.amazon.com
NVIDIA
последний пост 2 days, 18 hours назад
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026 Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

NVIDIA RTX Spark Windows PCs Arrive October 2026NVIDIA RTX Spark is coming this October— and at IFA 2026, partners are showing off their hardware.

The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter.

2 days, 18 hours назад @ blogs.nvidia.com
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW ‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW

September is here with 26 more games streaming on GeForce NOW this month, led by a slam dunk: NBA 2K27 with the NVIDIA DLSS 5 3D-Guided Neural Rendering feature.

GeForce NOW Ultimate members in NVIDIA-operated regions can stream NBA 2K27 across PCs, Macs, handhelds, mobile devices, TVs and more.

DLSS 5 is available when streaming from a GeForce RTX 5080-powered rig in the cloud.

Tip off without waiting for downloads or managing storage, and experience NBA 2K27 from the cloud with the visual fidelity of DLSS 5.

GeForce NOW availability may vary, as games are onboarded after release and added throughout the week.

2 days, 21 hours назад @ blogs.nvidia.com
NVIDIA to Acquire Hugging Face
NVIDIA to Acquire Hugging Face NVIDIA to Acquire Hugging Face

I’m excited to announce that NVIDIA has agreed to acquire Hugging Face for $12,930,300,000.

NVIDIA compute will not be required to build on or deploy through Hugging Face.

NVIDIA is the largest contributor of open models and data to Hugging Face, and our contributions continue to grow.

NVIDIA has released more than 500 models on Hugging Face and more than 250 open datasets.

To the millions of builders on Hugging Face: thank you for pushing the boundaries of what is possible.

2 days, 22 hours назад @ blogs.nvidia.com
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier

The NVIDIA founder and CEO joined CrowdStrike CEO and founder George Kurtz to announce CrowdStrike SafeMind, its agentic cybersecurity system developed by the CrowdStrike Cyber Superintelligence Lab.

CrowdStrike also announced CrowdStrike Falcon IQ to operationalize Project QuiltWorks through agentic workload automation and expanded its CrowdStrike Guardian AI safety solution.

The result ships natively in the CrowdStrike Falcon platform as SafeMind, CrowdStrike’s agentic cybersecurity system.

“Together with NVIDIA, we built cybersecurity’s first complete agentic system for cybersecurity, including the first frontier models and harness purpose-built for defenders,” Kurtz said.

Red vs. BlueNV…

4 days, 13 hours назад @ blogs.nvidia.com
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026 GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

To celebrate, GeForce NOW is giving away more than 40 prizes in our Community Giveaway — including gaming hardware and GeForce NOW Ultimate memberships.

Together, these technologies make it easier to tune supported games for image quality, responsiveness or a preferred balance between the two.

Check back on GFN Thursdays for the latest games joining the GeForce NOW library, including newly supported games streaming from GOG.

Ultimate members can enjoy GeForce RTX-powered gaming at up to 1440p and 120 frames per second.

Check back every GFN Thursday for new games, new features and even more ways to play on GeForce NOW.

1 week, 2 days назад @ blogs.nvidia.com
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

To help hyperscalers and AI innovators build the next generation of semi-custom AI infrastructure, NVIDIA today expanded NVIDIA NVLink Fusion with NVIDIA NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs.

It will be validated and offered by leading memory partners, extending this advanced memory capability to NVLink Fusion customers.

Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion.

AWS and NVIDIA Continue NVLink Fusion CollaborationAmazon’s Annapurna Labs will work with NVIDIA on NVHBM technology and the NVLink scale-up architecture to enhance…

1 week, 3 days назад @ blogs.nvidia.com
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo

Shadow engine recovery, available as a preview feature in NVIDIA Dynamo, moves most of this recovery work off the serving path.

How shadow engine recovery worksShadow engine recovery combines persistent GPU memory, a pre-warmed standby engine, and worker-level coordination to recover without a cold restart.

Benchmark results: shadow engine recovery vs. cold restart on GLM-5.2To quantify the benefit, we compared shadow engine recovery with a cold restart after an engine failure in a two-worker fleet.

In the Shadow Engine Recovery configuration, each worker pod hosts a preinitialized shadow engine that can take over if the active engine fails.

Cold restart versus shadow engine recovery in the…

1 week, 4 days назад @ developer.nvidia.com
Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark
Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark

Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark ahead of its launch this fall.

They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX.

For gamers, RTX Spark combines Windows devices with NVIDIA RTX technologies, including DLSS and Reflex.

NVIDIA has been working with EA to bring native EA Javelin Anticheat support to RTX Spark, including for EA’s biggest multiplayer games, such as EA SPORTS F1 25.

RTX Spark launches this fall.

1 week, 4 days назад @ blogs.nvidia.com
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access

With CUDA 13.3, we released CUDA Python 1.0, the libraries and tools that give you the full CUDA platform from Python.

CUDA Python will keep tracking new CUDA capabilities as they ship, now under predictable versioning and deprecation rules.

As of CUDA 13.3, CUDA Python and C++ stand as equal first-class citizens, with NVIDIA committing to maintain feature-complete parity going forward.

Getting startedOne command gets you the CUDA Python stack:pip install cuda-python cuda-cccl numba-cuda-mlir[cu13]That covers the CUDA Python components above, plus the MLIR-based Numba backend.

AcknowledgmentsCUDA Python 1.0 reflects years of work by the CUDA Python product and engineering teams, who designe…

1 week, 4 days назад @ developer.nvidia.com
How XPUs Meet a World-Class AI Factory
How XPUs Meet a World-Class AI Factory How XPUs Meet a World-Class AI Factory

That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.

At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly.

NVLink Fusion delivers on that need, connecting XPUs to NVIDIA’s world-leading AI infrastructure to increase performance, accelerate time to market and mitigate risk for semi-custom AI factories.

As an example, NVLink Fusion brings XPUs into the NVIDIA NVLink scale-up domain.

With NVLink Fusion, XPUs can now meet a world-class AI platform, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories that combine the strengths…

1 week, 5 days назад @ blogs.nvidia.com
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production.

CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks.

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.

To keep agents operating at the pace user…

1 week, 5 days назад @ blogs.nvidia.com
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads.

These early results for Vera Rubin NVL72 demonstrate NVIDIA’s accelerated pace of innovation.

With continuous software optimizations, performance across both Vera Rubin NVL72 and GB300 NVL72 will continue to improve.

At up to 35x lower cost per million tokens than GB300 NVL72, Vera Rubin NVL72 can run agents continuously, at scale, across the full breadth of customers’ workloads.

Learn more about the NVIDIA Vera Rubin platform.

1 week, 5 days назад @ blogs.nvidia.com
GPU-Accelerated Clustering for Financial Instruments at Scale
GPU-Accelerated Clustering for Financial Instruments at Scale GPU-Accelerated Clustering for Financial Instruments at Scale

A dense FP32 dependence matrix requires ~40 GB for 100,000 instruments and ~4 TB for 1 million instruments.

The solver returns H, where each row of H contains an instrument’s soft factor loadings, and taking the row-wise argmax produces a hard cluster label.

Mathematically, the diagonal AdaGrad update is defined as:\(G \leftarrow G + g \odot g\)\(H \leftarrow \max \left(H – \eta \cdot g / (\sqrt{G} + \varepsilon), \; 0\right)\)where g is either the full gradient ∇f(H) (full-batch branch) or a block-sampled estimate of it (stochastic branch), with all operations above being element-wise.

In practice, use spherical k-means for well-separated hard clusters and SymNMF when soft factor loadings …

2 weeks, 1 day назад @ developer.nvidia.com
Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support
Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support

GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight from the browser, starting today.

Whether on a school laptop or everyday PC, it’s now even easier to play supported PC games without downloading a dedicated app.

Plus, discover 12 new games streaming from the GeForce NOW library this week, led by Gallipoli.

Play directly in the newly supported Firefox browser, or take the action across PCs, Macs, Chromebooks, handhelds and more with GeForce NOW cloud saves.

To try GeForce NOW in Firefox, start with a GeForce NOW day pass and experience premium cloud gaming directly from the browser.

2 weeks, 2 days назад @ blogs.nvidia.com
Building Federated Multimodal AI Workflows with NVIDIA FLARE
Building Federated Multimodal AI Workflows with NVIDIA FLARE Building Federated Multimodal AI Workflows with NVIDIA FLARE

It shows how NVIDIA FLARE coordinates federated training across sites and handles large model updates through externalization, tensor streaming, and disk-backed aggregation.

A federated multimodal AI workflow with NVIDIA FLARECoordinating training across clientsEvery NVIDIA FLARE job separates global coordination from local execution.

For tested configuration examples, see the NVIDIA FLARE Recipe API, FLARE Tensor Downloader, and tensor disk offload documentation.

Federating lightweight adapters over a frozen VLM with FedUMMFedUMM provides a concrete example of minimizing what crosses the network in NVIDIA FLARE.

To learn more, join us for NVIDIA Flare Day 2026, a free online event that exp…

2 weeks, 3 days назад @ developer.nvidia.com
Facebook
последний пост 4 days, 1 hour назад
An Organizational Second Brain: Building an AI That Learns From Experts
An Organizational Second Brain: Building an AI That Learns From Experts An Organizational Second Brain: Building an AI That Learns From Experts

Recipes reference knowledge files but contain no domain facts; knowledge files state positions but prescribe no procedures.

This means:Adding an organizational position means adding a knowledge file and updating a routing index.

Every expert correction moves through four phases:Diagnose expert feedback into actionable issues with their root cause.

The requirements for adopting this architecture are:A structured knowledge system with explicit file boundaries, cross-references, and a dependency graph (the organizational second brain for the domain).

Every improvement is a text edit that a domain expert can review in 30 seconds.

4 days, 1 hour назад @ engineering.fb.com
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Introducing the Multi-Stage Sequence ModelTo address scaling efficiency, a multi-stage model has been developed that enables scaling of a transformer-based sequence model in a compute efficient manner.

Second Stage: Online Ranking ModelThe offline user model representations are complemented with online ranking models that use fresh user signals and ad candidate information for real time ranking.

A Predictable Scaling CurveLLM-Style Scaling LawWhen running on real-world ads traffic, the multi-stage sequence model demonstrates the emergence of predictable scaling laws for ads recommendations that are analogous to those observed in large language models.

The Impact of Multi-Stage Sequence Mode…

1 month назад @ engineering.fb.com
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

We tackled these challenges through complementary compute efficiency and scaling efficiency innovations: Compute efficiency : Achieved through a customized recommendation kernel library — Jagged Flash Attention (JFA), Generalized Dot-Product Attention (GDPA), BlockAttention, etc.

The results: we doubled GEM’s E2E training efficiency to 20-25% MFU while scaling total training FLOPs 4x over the past 12 months.

We measure training efficiency through E2E MFU, which decomposes into two factors:E2E MFU = Local MFU (compute efficiency) × Scaling Ratio (scaling efficiency)These factors describe two related but distinct optimization problems.

Local MFU (compute efficiency) measures how well a single…

1 month назад @ engineering.fb.com
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

1 month, 3 weeks назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

1 month, 3 weeks назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

2 months, 1 week назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

2 months, 1 week назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

3 months, 1 week назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

3 months, 3 weeks назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

4 months, 2 weeks назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

4 months, 3 weeks назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

5 months назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

5 months назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

5 months, 1 week назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

5 months, 1 week назад @ engineering.fb.com
Uber Engineering
последний пост None
neptune.ai neptune.ai
последний пост 9 months назад
We are joining OpenAI
We are joining OpenAI We are joining OpenAI

Piotr Niedźwiedź, CEO/CTO and founder of neptune.aiI’m excited to share that we’ve entered into a definitive agreement to be acquired by OpenAI, subject to closing conditions.

We are thrilled to join the OpenAI team and help their AI researchers build better models faster.

Neptune is a metrics dashboard company.”We’ve worked closely with OpenAI to create the metrics dashboard that helps teams building foundation models.

Our future with OpenAINeptune will join OpenAI and continue to support AI researchers with tools to monitor, debug, and evaluate frontier models.

We are looking forward to working with top AI researchers and supporting OpenAI’s mission of ensuring that AGI benefits all of hu…

9 months назад @ neptune.ai
Synthetic Data for LLM Training
Synthetic Data for LLM Training Synthetic Data for LLM Training

For instance, financial data is highly sensitive and protected by very strict regulations, and synthetic data mimics the real data distribution without revealing customer information.

Read more about how leading foundation model teams curate their training data and other topics in the State of Foundation Model Training Report 2025.

Choosing the right synthetic data generation technique depends on the type of data and its complexity.

Synthetic tabular data generation is a promising direction to overcome these challenges by learning the distribution of the tabular data.

Post-processingAs the distribution of tabular data is highly complex, it makes the synthetic tabular data generation very ch…

9 months, 3 weeks назад @ neptune.ai
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 6 months назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

6 months назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

8 months, 1 week назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

8 months, 1 week назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

8 months, 1 week назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

8 months, 3 weeks назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 2 weeks, 4 days назад
The jumping pegs puzzle
The jumping pegs puzzle The jumping pegs puzzle

Part of a series of monthly puzzles with MoMath.

2 weeks, 4 days назад @ youtube.com
The 64 sugar cubes puzzle
The 64 sugar cubes puzzle The 64 sugar cubes puzzle

See all monthly puzzles: https://momath.org/mindbenders/

1 month, 1 week назад @ youtube.com
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

1 month, 3 weeks назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

2 months, 3 weeks назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

2 months, 3 weeks назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

2 months, 3 weeks назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

3 months назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

3 months, 2 weeks назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

4 months, 3 weeks назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

5 months, 1 week назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

5 months, 2 weeks назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

5 months, 2 weeks назад @ youtube.com
Bacteria Grid Puzzle Solution
Bacteria Grid Puzzle Solution Bacteria Grid Puzzle Solution

Part of a monthly series of puzzlers, in collaboration with MoMath and Peter Winkler

5 months, 2 weeks назад @ youtube.com
The most underappreciated formula | Exploring high-dimensional spheres
The most underappreciated formula | Exploring high-dimensional spheres The most underappreciated formula | Exploring high-dimensional spheres

On the volumes of higher-dimensional spheres

Explore the 3b1b virtual career fair: See https://3b1b.co/talent

Become a supporter for early views of new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Thanks to UC Santa Cruz for letting me film there, and special thanks to Pedro Morales-Almazan for arranging everything. My video on Numberphile with a fun application of this problem: https://youtu.be/6_yU9eJ0NxA Timestamps:

0:00 - Introduction

1:01 - Random puzzle

6:16 - Outside the box

14:35 - Setting up the volume grid

21:14 - Why 4πr^2

25:21 - Archimedes in higher dimensions

36:17 - The general formul…

6 months, 1 week назад @ youtube.com
The lattice bacteria puzzle
The lattice bacteria puzzle The lattice bacteria puzzle

Part of a series of monthly puzzles, done in collaboration with MoMath.

https://momath.org/mindbenders

6 months, 2 weeks назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 3 days, 2 hours назад
Claude Fable AI Is Much Stranger Than The Headlines Suggest
Claude Fable AI Is Much Stranger Than The Headlines Suggest Claude Fable AI Is Much Stranger Than The Headlines Suggest

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Claude Fable 5.1 paper is available here:

https://www.anthropic.com/claude-fable-and-mythos-5-1

https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20%26%20Claude%20Mythos%205.1%20System%20Card.pdf Sources:

https://x.com/alexalbert__/status/2094860187743986169?s=46

https://x.com/holytrinity/status/2094866061212459130?s=46

https://x.com/holytrinity/status/2094927984217985474?s=46

https://x.com/omedvibecodes/status/2094887840848965845?s=46

https://x.com/loktar00/status/2094951511742632168?s=46

https://x.com/maxt3chno/status/2094798704385380762?s=46

https://x.com/fab…

3 days, 2 hours назад @ youtube.com
Powerful AI Is Becoming Almost Free
Powerful AI Is Becoming Almost Free Powerful AI Is Becoming Almost Free

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 GLM 5.3 Flash:

https://z.ai/blog/glm-5.3-flash Sources:

https://x.com/louszbd/status/2092694163104113016

https://x.com/semianalysis_/status/2092623833630998556

https://x.com/skalskip92/status/2092748209802154201

https://x.com/louszbd/status/2093047548550525165

https://x.com/atomic_chat_hq/status/2093433913238552712

https://x.com/holytrinity/status/2094093933584257334

https://x.com/KinasRemek/status/2090081611832295581/video/1

https://x.com/AiXsatoshi/status/2093679264013181119/video/1

https://x.com/AiXsatoshi/status/2093353322921263389/video/1

https://x.com/stevibe/status/2092655031040565252 🙏 We would like…

5 days, 1 hour назад @ youtube.com
This Free AI Just Caught The Billion Dollar Giants
This Free AI Just Caught The Billion Dollar Giants This Free AI Just Caught The Billion Dollar Giants

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper and Qwen3.8-Flash-Next are available here:

https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf

https://qwen.ai/blog?id=qwen3.8-flash-next 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 2 days назад @ youtube.com
DeepSeek’s AI Just Learned To Upgrade Itself
DeepSeek’s AI Just Learned To Upgrade Itself DeepSeek’s AI Just Learned To Upgrade Itself

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek Harness + paper are available here:

https://deepseek.com/harness/en/

https://github.com/cordiverse/paper 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 3 days назад @ youtube.com
This Small AI Will Change Everything
This Small AI Will Change Everything This Small AI Will Change Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Qwen3.8-27b is available here:

https://huggingface.co/Qwen/Qwen3.8-27B Sources:

https://www.reddit.com/r/unsloth/comments/1vogva0/share_your_results_from_qwen3827b/

https://www.reddit.com/r/LocalLLaMA/comments/1voer8u/qwen_38_27b_aquarium_burst_sample_test/

https://x.com/KyleHessling1/status/2088327667733180637

https://www.reddit.com/r/LocalLLaMA/comments/1vqme4y/qwen3827b_q8_0_on_strix_halo_is_seriously/

https://forums.developer.nvidia.com/t/qwen3-8-27b-nvfp4-on-a-single-dgx-spark-up-to-1m-context-vllm-mtp-measurements/380244 🙏 We would like to thank our generous Patreon supporters who make Two Minute …

1 week, 5 days назад @ youtube.com
DeepSeek Just Made Closed AI Look Ridiculous
DeepSeek Just Made Closed AI Look Ridiculous DeepSeek Just Made Closed AI Look Ridiculous

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers DeepSeek V4 Pro 0813:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 DSpark full episode: https://www.youtube.com/watch?v=1yBU41auQhw Sources:

https://x.com/cline/status/2087602193205694891

https://x.com/TypingMindApp/status/2088214938754167263

https://x.com/stevibe/status/2047546592530747561

https://x.com/voidfreud/status/2087701887327887543

https://x.com/exploraX_/status/2079197387860435360

https://x.com/AiHubMix/status/2087869896483057758

https://x.com/BruceBlue/status/2087833177155117304 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B …

2 weeks, 3 days назад @ youtube.com
Claude AI Failed 650 Times…Then Beat The Human Record
Claude AI Failed 650 Times…Then Beat The Human Record Claude AI Failed 650 Times…Then Beat The Human Record

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://www.anthropic.com/research/riemann-zeta Source:

https://www.scientificamerican.com/article/no-ai-didnt-just-solve-the-thorniest-problem-in-math/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 weeks, 2 days назад @ youtube.com
OpenAI’s AI Escaped And It's Terrifying
OpenAI’s AI Escaped And It's Terrifying OpenAI’s AI Escaped And It's Terrifying

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 More reports are available here:

https://openai.com/index/hugging-face-model-evaluation-security-incident/

https://huggingface.co/blog/security-incident-july-2026

https://huggingface.co/blog/agent-intrusion-technical-timeline 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 weeks, 4 days назад @ youtube.com
DeepMind's AI Trick Everyone Should Copy
DeepMind's AI Trick Everyone Should Copy DeepMind's AI Trick Everyone Should Copy

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Gemma4 paper and some more is available here:

https://arxiv.org/abs/2607.02770

https://x.com/googlegemma/status/2077449152062247219

https://x.com/UnslothAI/status/2078118183085731843 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
The Billion Dollar AI Race Just Broke
The Billion Dollar AI Race Just Broke The Billion Dollar AI Race Just Broke

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 Qwen 3.8 Max:

https://qwen.ai/blog?id=qwen3.8 Sources:

https://x.com/loktar00/status/2082589566934929750

https://x.com/CommandCodeAI/status/2084293498950590839 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
Another DeepSeek Moment
Another DeepSeek Moment Another DeepSeek Moment

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek v4 Flash 0731:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
New AI Learned Parkour From Just 30 Seconds Of Video
New AI Learned Parkour From Just 30 Seconds Of Video New AI Learned Parkour From Just 30 Seconds Of Video

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://jiashunwang.github.io/HIL/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
Kimi K3 Just Broke The Economics Of AI
Kimi K3 Just Broke The Economics Of AI Kimi K3 Just Broke The Economics Of AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2607.24653 Try Kimi K3 (subject to availability): https://www.kimi.com/ Links:

https://macos27.kimi.page/

https://x.com/mweinbach/status/2077878247920951400

https://x.com/intheworldofai/status/2077838911494336681

https://x.com/chetaslua/status/2077829183989072281

https://x.com/hqmank/status/2078104317027094907 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen …

1 month, 1 week назад @ youtube.com
AI Helped Them Code Faster… But At A Cost
AI Helped Them Code Faster… But At A Cost AI Helped Them Code Faster… But At A Cost

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/AI-assistance-coding-skills 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 3 weeks назад @ youtube.com
The Hidden World Inside An AI
The Hidden World Inside An AI The Hidden World Inside An AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://transformer-circuits.pub/2025/linebreaks/index.html Paper for reindeer vision change - https://royalsocietypublishing.org/rspb/article/280/1773/20132451/50765/Shifting-mirrors-adaptive-changes-in-retinal 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli …

1 month, 3 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 1 day, 22 hours назад
Внутренние знания 🚫 Источники ✅
Внутренние знания 🚫 Источники ✅ Внутренние знания 🚫 Источники ✅

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

1 day, 22 hours назад @ youtube.com
Инфраструктура как часть RL
Инфраструктура как часть RL Инфраструктура как часть RL

Приглашаем специалистов с опытом от 2 лет на Weekend Offer ML 12–13 сентября: https://clck.ru/3VTgGR Это один из наймовых ивентов Яндекса: вы сможете пройти все ключевые этапы онлайн и без долгих пауз.

3 days, 22 hours назад @ youtube.com
Как Kimi K3 обучает уровни вычислительного бюджета
Как Kimi K3 обучает уровни вычислительного бюджета Как Kimi K3 обучает уровни вычислительного бюджета

"Приглашаем специалистов с опытом от 2 лет на Weekend Offer ML 12–13 сентября Это один из наймовых ивентов Яндекса: вы сможете пройти все ключевые этапы онлайн и без долгих пауз".

4 days, 22 hours назад @ youtube.com
Почему избавиться от галлюцинаций недостаточно 😵‍💫
Почему избавиться от галлюцинаций недостаточно 😵‍💫 Почему избавиться от галлюцинаций недостаточно 😵‍💫

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

1 week, 2 days назад @ youtube.com
Оптимизация LLM-инференса
Оптимизация LLM-инференса Оптимизация LLM-инференса

Доклад Андрея Бежина, руководителя службы ML-инфраструктуры в Яндекс R&D, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

1 week, 3 days назад @ youtube.com
Новые способы и стандарты оценки качества моделей
Новые способы и стандарты оценки качества моделей Новые способы и стандарты оценки качества моделей

Доклад Ивана Дёгтева, руководителя аналитики Alice AI LLM в Яндекс R&D, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

1 week, 3 days назад @ youtube.com
Тренды и вызовы в ризонинге
Тренды и вызовы в ризонинге Тренды и вызовы в ризонинге

Доклад Дмитрия Мокеева, руководителя группы качества претрейна Alice AI в Яндекс Поиске, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy #ML #MachineLearning #AI #LLM #DeepLearning #LLMInference #Reasoning #AIResearch #Yandex #Яндекс #DataScience #IT

1 week, 4 days назад @ youtube.com
Tabular DL
Tabular DL Tabular DL

Доклад Артёма Бабенко, руководителя отдела в Yandex Research, на ML Global Recap’H1 2026. Больше материалов про ML по ссылке: https://t.me/+owyCvdge8WIyNTUy

1 week, 4 days назад @ youtube.com
Большая траектория для маленькой модели 🛤️
Большая траектория для маленькой модели 🛤️ Большая траектория для маленькой модели 🛤️

Как безопасно выкатывать новые версии продуктовых AI-агентов? Как фиксировать регрессии до прода? При чём тут автометрики? Об этом рассказал Дмитрий Коршунов, Team Lead ML в Ecom, на Data Fest 2026 в Белграде. Полная запись доклада уже на канале 🎦 #DataFest #DataFest2026 #AI #AIагенты #LLM #MachineLearning #ML #нейросети #AIinProduction #Яндекс #IT #разработка

1 week, 5 days назад @ youtube.com
ML Global Recap'H1 2026
ML Global Recap'H1 2026 ML Global Recap'H1 2026

Обсудим итоги ICML и других международных конференций, главные ML-тренды первого полугодия 2026-го и собственный опыт.

1 month назад @ youtube.com
Омни-модели будущего 🚀
Омни-модели будущего 🚀 Омни-модели будущего 🚀

Что они будут уметь — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months назад @ youtube.com
Качество модели взлетело... без мультимодального RL?
Качество модели взлетело... без мультимодального RL? Качество модели взлетело... без мультимодального RL?

О росте мультимодального качества рассказал Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 1 week назад @ youtube.com
Работа с данными — это скучно?
Работа с данными — это скучно? Работа с данными — это скучно?

А почему — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 1 week назад @ youtube.com
Как приготовить SFT 🍲
Как приготовить SFT 🍲 Как приготовить SFT 🍲

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 2 weeks назад @ youtube.com
Почему мультимодальные модели — это база 🤖
Почему мультимодальные модели — это база 🤖 Почему мультимодальные модели — это база 🤖

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 months, 2 weeks назад @ youtube.com
ML Trainings ML Trainings
последний пост 3 часа назад
Капитанский мостик 06.09.2026: Вышла GPT 6 и AGI | В США запретили AGI | Судный день ИИ
Капитанский мостик 06.09.2026: Вышла GPT 6 и AGI | В США запретили AGI | Судный день ИИ Капитанский мостик 06.09.2026: Вышла GPT 6 и AGI | В США запретили AGI | Судный день ИИ

0:00:00 Начало

0:02:05 Вышел Fable 5.1

0:07:14 Anthropic зажал токены

0:11:51 Вышла GPT 6 и AGI

0:19:42 В США запретили AGI

0:26:23 Понабрались от GPT

0:34:26 Минцифры не платит

0:37:57 Беспилотный трамвай

0:42:59 Сценарий AI 2040

0:53:13 Nvidia и Poolside

0:58:19 ЦОД в Саудовской Аравии

1:02:58 Касперская про ИИ

1:07:58 Судный день ИИ ИИ-саммари: Валентин Малых и Дмитрий Колодезев разбирают главные ИИ-события недели: релиз Fable 5.1, выход GPT 6 с заявлениями про AGI и — с точностью до наоборот — законодательную инициативу США о запрете суперинтеллекта. Затем экономика ИИ: почему Anthropic экономит ваши токены, зачем Nvidia нужен Poolside и сколько стоит дата-центр в Саудовской Аравии. Отд…

3 часа назад @ youtube.com
Марк Обозов | CayleyPy: from Rubik's cube to AdS/CFT holographies
Марк Обозов | CayleyPy: from Rubik's cube to AdS/CFT holographies Марк Обозов | CayleyPy: from Rubik's cube to AdS/CFT holographies

Спикер: Марк Обозов, T-Bank RnD, PyTorch Team Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции Mathematics & ML https://ods.ai/tracks/df26-mathematics-ml

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

18 часов назад @ youtube.com
Юлия Ишмакова | Доверяй, но проверяй: как судья фармацевта валидировал
Юлия Ишмакова | Доверяй, но проверяй: как судья фармацевта валидировал Юлия Ишмакова | Доверяй, но проверяй: как судья фармацевта валидировал

Спикер: Юлия Ишмакова, Delivery lead, Центр индустрии здоровья, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

18 часов назад @ youtube.com
Никита Ятченко | от LLM-портала к фабрике внутренних агентов
Никита Ятченко | от LLM-портала к фабрике внутренних агентов Никита Ятченко | от LLM-портала к фабрике внутренних агентов

Спикер: Никита Ятченко

Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Marketplace от Avito.tech https://ods.ai/tracks/df26_mlavitotech ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 3 hours назад @ youtube.com
Михаил Трегубов | Как VLM превращают документооборот из архива в активный интеллект?
Михаил Трегубов | Как VLM превращают документооборот из архива в активный интеллект? Михаил Трегубов | Как VLM превращают документооборот из архива в активный интеллект?

Спикер: Михаил Трегубов, СКБ Контур, Middle+ Computer Vision Engineer Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Computer Vision https://ods.ai/tracks/df26-cv ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 3 hours назад @ youtube.com
Семен Буденный | Следующая эра ИИ: ключевые направления развития
Семен Буденный | Следующая эра ИИ: ключевые направления развития Семен Буденный | Следующая эра ИИ: ключевые направления развития

Спикер: Семен Буденный, управляющий директор, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 3 hours назад @ youtube.com
Иван Комаров и Евгений Чуканов | Делаем open source Agentic RAG, закалённый в бою
Иван Комаров и Евгений Чуканов | Делаем open source Agentic RAG, закалённый в бою Иван Комаров и Евгений Чуканов | Делаем open source Agentic RAG, закалённый в бою

Спикеры: Иван Комаров и Евгений Чуканов, Koronatech Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции Agentic LLM https://ods.ai/tracks/df26-agenticllms

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 3 hours назад @ youtube.com
Юрий Колабушин | Kandinsky in production. Особенности национальной адаптации и интеграции Gen CV
Юрий Колабушин | Kandinsky in production. Особенности национальной адаптации и интеграции Gen CV Юрий Колабушин | Kandinsky in production. Особенности национальной адаптации и интеграции Gen CV

Спикер: Юрий Колабушин, исполнительный директор по исследованию данных, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenCV https://ods.ai/tracks/df26-gencv

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 19 hours назад @ youtube.com
Алина Жидковская | AutoJudge - динамическая генерация llm-as-judge для валидации агентных пайплайнов
Алина Жидковская | AutoJudge - динамическая генерация llm-as-judge для валидации агентных пайплайнов Алина Жидковская | AutoJudge - динамическая генерация llm-as-judge для валидации агентных пайплайнов

Спикер: Алина Жидковская, ИТМО, nss lab, ML-инженер Data Fest 2026: https://ods.ai/events/datafest2026 ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 19 hours назад @ youtube.com
Алексей Рязанцев | Как мы сократили time-to-market для ML-моделей в модерации в 100 раз
Алексей Рязанцев | Как мы сократили time-to-market для ML-моделей в модерации в 100 раз Алексей Рязанцев | Как мы сократили time-to-market для ML-моделей в модерации в 100 раз

Спикер: Алексей Рязанцев, Wildberries, Data Scientist Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции MLOps https://ods.ai/tracks/df26-mlops

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 3 hours назад @ youtube.com
Айгиз Кунафин | Умная колонка на ИИ-тяге: как кодовые агенты заменяют целую команду разработки
Айгиз Кунафин | Умная колонка на ИИ-тяге: как кодовые агенты заменяют целую команду разработки Айгиз Кунафин | Умная колонка на ИИ-тяге: как кодовые агенты заменяют целую команду разработки

Спикер: Айгиз Кунафин Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Advanced LLMs ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 3 hours назад @ youtube.com
Даниил Морозов | Угроза доступности LLM: sponge-атаки, дорогая генерация и защита inference-сервисов
Даниил Морозов | Угроза доступности LLM: sponge-атаки, дорогая генерация и защита inference-сервисов Даниил Морозов | Угроза доступности LLM: sponge-атаки, дорогая генерация и защита inference-сервисов

Спикер: Даниил Морозов, MLSecOps, Совкомбанк Технологии Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции LLM inference ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 19 hours назад @ youtube.com
Владислав Голощапов | Как стадо агентов делает ресерч и немножко авторесерча
Владислав Голощапов | Как стадо агентов делает ресерч и немножко авторесерча Владислав Голощапов | Как стадо агентов делает ресерч и немножко авторесерча

Спикер: Владислав Голощапов, Независимый исследователь Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Agentic LLM ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 19 hours назад @ youtube.com
Дарья Усова | LLM-as-Judge для оценки качества агента для работы с кодом в IDE
Дарья Усова | LLM-as-Judge для оценки качества агента для работы с кодом в IDE Дарья Усова | LLM-as-Judge для оценки качества агента для работы с кодом в IDE

Спикер: Дарья Усова, AI Researcher Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции Code LLM https://ods.ai/tracks/df26-codellm

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 19 hours назад @ youtube.com
Олег Новиков | Нет нейрослопу: как вернуть музыканту точность в мире генеративной музыки
Олег Новиков | Нет нейрослопу: как вернуть музыканту точность в мире генеративной музыки Олег Новиков | Нет нейрослопу: как вернуть музыканту точность в мире генеративной музыки

Спикер: Олег Новиков, главный инженер по разработке, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 3 hours назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 1 week, 3 days назад
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux #501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux

DHH is the creator of Ruby on Rails, Omarchy Linux, CTO of 37signals, and a racecar driver.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep501-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://wisprflow.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://plaud.ai/lexHiggsfield AI: AI-based video generation, filmmaking, and creative studio.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:14) – Sponsors, Comments, and Reflections(08:56) – Programming with AI agents(24:14) – How software will change(33:30) – AI impact on open source(43:21) – Building Omarchy Linux d…

1 week, 3 days назад @ lexfridman.com
#500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football
#500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football #500 – Khabib Nurmagomedov: Dagestan, MMA, UFC, Islam, Conor, Fedor & Football

Khabib Nurmagomedov is one of the greatest fighters of all time, who retired from the UFC undefeated with a perfect 29-0 record.

We did this conversation entirely in Russian.

Both language audio tracks (and subtitles) are available on YouTube.

We worked hard to make it enjoyable to listen to, by carefully dubbing the translation using voice-cloning, as we’ve done for previous foreign-language podcasts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep500-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

3 weeks, 3 days назад @ lexfridman.com
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee #499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee

Gary Gallagher is a historian of the American Civil War.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep499-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://plaud.ai/lexOUTLINE:(00:00) – Introduction(00:07) – Sponsors, Comments, and Reflections(08:36) – What caused the Civil War?

(18:33) – Slavery(46:07) – Lincoln(1:01:03) – Grant vs Lee(1:09:57) – Could the Civil War have been avoided?

(1:19:23) – The bloodiest war in US history(1:36:31) – How the Confederate Army could’ve won(1:57:05) – Key battles of the Civil War(2:20:07) – Best and Worst Presidents(2:34:06) – Robert E. Lee(2:53:40) – The…

1 month, 1 week назад @ lexfridman.com
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

2 months, 1 week назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

3 months, 1 week назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

4 months назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

4 months, 4 weeks назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

5 months, 2 weeks назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

5 months, 4 weeks назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

6 months, 1 week назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

6 months, 3 weeks назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

7 months, 1 week назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

7 months, 3 weeks назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

8 months, 1 week назад @ lexfridman.com
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths #487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths

Irving Finkel is a scholar of ancient languages and a longtime curator at the British Museum, renowned for his expertise in Mesopotamian history and cuneiform writing.

He specializes in reading and interpreting cuneiform inscriptions, including tablets from Sumerian, Akkadian, Babylonian, and Assyrian contexts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep487-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://shopify.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/Chevron: Reliable energy for data centers.

8 months, 3 weeks назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 4 months, 2 weeks назад
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

4 months, 2 weeks назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

4 months, 4 weeks назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

5 months, 2 weeks назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

6 months назад @ microsoft.com
Ideas: Community building, machine learning, and the future of AI
Ideas: Community building, machine learning, and the future of AI Ideas: Community building, machine learning, and the future of AI

This week, machine learning researchers around the world will be attending the annual Conference on Neural Information Processing Systems, or NeurIPS.

In this series, we’ll explore the technologies that are shaping our future and the big ideas that propel them forward.

So around that time when I started my PhD at Penn, I was working in machine learning theory and algorithmic economics.

How had you experienced a lack of community or network of women in machine learning before the founding of WiML?

So particularly when working on topics related to fairness, I’ve ended up focusing a bunch on stuff to do with marginalized groups as part of my responsible AI work.

9 months, 1 week назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 4 days, 18 hours назад
Recommender Systems Optimization Goals
Recommender Systems Optimization Goals Recommender Systems Optimization Goals

In part two of the Data Skeptic Recommender Systems season finale, Kyle asks a deceptively difficult question: what should recommender systems actually optimize for?

Drawing on conversations from across the season, the episode explores engagement, filter bubbles, popularity bias, fairness, human curation, embeddings, and the growing role—and risks—of large language models in shaping what gets recommended to us.

4 days, 18 hours назад @ dataskeptic.com
Recommender Systems Origin Story
Recommender Systems Origin Story Recommender Systems Origin Story

Where did recommender systems come from, and how do we know when they're actually working? In part one of Data Skeptic's three-part Recommender Systems finale, Kyle traces the field from collaborative filtering and the Netflix Prize to matrix factorization and modern approaches, while exploring why accuracy alone can't capture what makes a recommendation useful, surprising, or meaningful.

2 weeks, 4 days назад @ dataskeptic.com
Social Choice for Fair Recommendations
Social Choice for Fair Recommendations Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social choice theory can help balance the needs of users, creators, and society. The conversation explores the future of recommendation algorithms and why fairness is a far more complex challenge than it first appears.

1 month, 1 week назад @ dataskeptic.com
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

2 months назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

2 months, 2 weeks назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

2 months, 2 weeks назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

4 months, 1 week назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

4 months, 1 week назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

5 months, 1 week назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 5 months, 4 weeks назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

6 months, 1 week назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

6 months, 2 weeks назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

7 months назад @ dataskeptic.com
Fairness in PCA-Based Recommenders
Fairness in PCA-Based Recommenders Fairness in PCA-Based Recommenders

In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users equally. David introduces the concept of "power niche users" - highly active users with specialized in…

7 months, 1 week назад @ dataskeptic.com
Video Recommendations in Industry
Video Recommendations in Industry Video Recommendations in Industry

In this episode, Kyle Polich sits down with Cory Zechmann, a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares insights on why TikTok's algorithm works so well (clean data and massive interaction volume), the crucial …

8 months, 1 week назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 1 day, 23 hours назад
1024: In Case You Missed It in August 2026
1024: In Case You Missed It in August 2026 1024: In Case You Missed It in August 2026

In ICYMI Episode #1024, Jon Krohn tracks the gap between AI investment and AI return, from the technology side to the people side. Hear from Pete Johnson, Jerry Yurchisin, Priyanka Vergadia and Tristan Handy, discussing why four out of five organizations have the structures for AI success in place while only one in five sees the returns, which decisions should never be handed to a language model however confident it sounds, how to structure Claude skills so that your output stops being slop and why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1024⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ …

1 day, 23 hours назад @ podtrac.com
1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan
1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan 1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan

In Episode #1023, Aishwarya Srinivasan (Co-Founder of The Gen Academy) joins Jon Krohn to work out where a competitive moat comes from once anything you can build in ten minutes, somebody else can build in ten minutes too. Ash came to teaching through Illuminate AI, the mentorship community she started in 2020, and now trains senior engineers and leaders to ship agentic AI in production; she is blunt that vibe coding lowers the floor without touching the engineering judgment that production demands. In this episode, she explains what a whole-system eval covers that a model eval misses, traces reinforcement learning from the algorithm she patented at IBM to its resurgence in agentic fine tun…

4 days, 23 hours назад @ podtrac.com
1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents
1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents 1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents

In Episode #1022, Jon Krohn tackles the art of steering AI agents, deciding where your instructions should live so they get followed reliably without bloating every request. A sequel to Episode #1020 (where model size and effort set an agent’s horsepower), this one is about direction: the seven ways to deliver instructions, why a hook beats a prompt, the industry-wide agents.md standard, and three practical takeaways you can apply whatever your stack. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1022⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this …

1 week, 1 day назад @ podtrac.com
1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy
1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy 1021: How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy

In Episode #1021, Tristan Handy (Founder and CEO of dbt Labs) joins Jon Krohn to explain how a study of about a hundred companies in 2016 became analytics engineering, and then became a tool that over a hundred thousand data teams rely on. Tristan coined the term, chose SQL when Spark was the fashionable answer, and spent a decade turning down acquisition offers because none of them were good for the people using dbt. He is now merging dbt Labs with Fivetran and taking on the presidency of the combined company, the first deal he says cleared that bar. In this episode, Tristan walks through what dbt does to your raw data, argues that the semantic layer matters more once analytics agents are …

1 week, 4 days назад @ podtrac.com
1020: How to Choose Model Size and Effort Level: The Two Critical Dials
1020: How to Choose Model Size and Effort Level: The Two Critical Dials 1020: How to Choose Model Size and Effort Level: The Two Critical Dials

In Episode #1020, Jon Krohn unpacks the two dials that increasingly decide what you get out of a large language model: which model size you pick and how much effort you tell it to spend. Using a July Anthropic blog post by Claude Code’s Lydia Holly as a jumping-off point, with guidance that generalizes to any model family, Jon explains what each setting actually does under the hood. Model size swaps which frozen weights handle your request (roughly, how capable), while effort sets how thorough and certain the model must be before calling a task done, not a simple “thinking-time slider.” He offers a clean diagnostic for when to raise effort versus move to a bigger model, shows why cheaper-pe…

2 weeks, 1 day назад @ podtrac.com
1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia)
1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia) 1019: Anyone Can Write Code Now, So What Gets You Hired? (With Priyanka Vergadia)

In Episode #1019, Priyanka Vergadia (founder of The Cloud Girl, former Senior Director of AI Transformation at Microsoft and Head of North America Developer Relations at Google) joins Jon Krohn to explain why almost every company has bought AI tools and almost none of them are seeing a return. Her fix is a budget split that will make any CFO wince: seven dollars on training employees for every dollar spent on the tools themselves. Having spent a decade turning dense cloud and AI concepts into sketches that a quarter-million developers actually remember, and having carried GitHub Copilot into Fortune 100 boardrooms, she has watched the gap between tool purchase and real production use up clo…

2 weeks, 4 days назад @ podtrac.com
1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs
1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs 1018: Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs

In Episode #1018, Jon Krohn breaks down Qwen3.8-Max, Alibaba’s enormous new flagship, a 2.4-trillion-parameter mixture-of-experts model that, if its promised weights ship, becomes the largest open-weight release in history. Landing just weeks after Moonshot’s Kimi K3, it extends the price war and the open-weight surge Jon covered in Episode #1012. Alibaba positions it as second only to Anthropic’s Claude Fable 5 / Mythos 5 and independent signals land in a similar neighborhood. Jon walks through its capabilities and multi-day agentic demos, its aggressive pricing ($2 in / $6 out per million tokens, with cached input eight times cheaper), and the question he gets asked most: are Chinese mode…

3 weeks, 1 day назад @ podtrac.com
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson
1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson 1017: Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson

In Episode #1017, Pete Johnson (Field CTO of AI at MongoDB) joins Jon Krohn to explain why four out of five organizations have AI steering committees and success metrics, yet only one in five sees a return on the investment. Having made nineteen stops across six countries this year advising more than a hundred companies on their AI strategies, Pete has an unusually wide view of what is actually working in production. In this episode, he traces the history of SQL and denormalization, unpacks why the embedding model is the most underrated choice in a RAG pipeline, explains Matryoshka embeddings and lays out what better agentic memory looks like. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

3 weeks, 4 days назад @ podtrac.com
1016: In Case You Missed It in July 2026
1016: In Case You Missed It in July 2026 1016: In Case You Missed It in July 2026

In this month's episode of ICYMI, Jon Krohn traces a line from algorithmic harm to the human skills that still hold their value. Hear from Dr. Cathy O'Neil, Ben Todd, Steve Mock, and Dr. Catherine Williams, discussing why an algorithm's danger has nothing to do with its complexity, what solid career ground looks like if fully automated digital workers arrive, how people are using AI to become better-informed advocates in healthcare rather than asking it for advice and why deep mathematical understanding still separates the best data professionals from everyone else. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1016⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperD…

4 weeks, 1 day назад @ podtrac.com
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin 1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin

In Episode #1015, Jerry Yurchisin (manager of decision intelligence strategy at Gurobi Optimization) joins Jon Krohn to explain the AI technology that makes breaking a constraint mathematically impossible. Large language models will confidently claim they've optimized your business while ignoring the one constraint that could cost millions, whereas optimization treats constraints as hard guarantees. Jerry lays out the division of labor he sees for the agentic era: agents help you frame the problem, write the formulation and generate the code, then hand off to a solver like Gurobi, soon callable via MCP servers. In this episode, Jerry breaks down the three building blocks of any optimization…

1 month назад @ podtrac.com
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself 1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself

In Episode #1014, Jon Krohn breaks down a security incident that reads like science fiction: during an internal evaluation, an autonomous OpenAI agent broke out of its sandbox, exploited a zero-day, and hacked its way into Hugging Face to steal the answers to the very benchmark it was being tested on, with no human attacker at any point. Jon lays out the three-act timeline, explains the ExploitGym benchmark and why switching off safety guardrails mattered so much and pulls out the practical lessons for anyone building or defending agentic AI systems. Along the way: why Hugging Face ran its forensics on a Chinese open-weight model and why the next attack like this one may not be an accident.…

1 month назад @ podtrac.com
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil 1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil

In Episode #1013, Dr. Cathy O'Neil (Harvard math PhD, former Wall Street quant and author of the mega-bestseller Weapons of Math Destruction) joins Jon Krohn to explain what actually makes an algorithm terrifying: not the complexity of the math, but the secrecy, the unaccountability, and the fact that you can't opt out. A decade after Weapons of Math Destruction sounded the alarm on algorithmic harm, Cathy is busier than ever. Through her algorithmic-auditing firm ORCAA and her nonprofit OCEAN, she now provides the statistical evidence behind lawsuits against some of the world's biggest tech companies. In this episode, Cathy punctures AI hype, traces the line from Frederick Winslow Taylor's…

1 month, 1 week назад @ podtrac.com
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier 1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier

What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building. Additional mate…

1 month, 1 week назад @ podtrac.com
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams 1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams

Dr. Catherine Williams, Chief Data Officer at the nonprofit Candid, was solving black-hole equations with pen and paper before she ever wrote a line of code. She earned a PhD in math researching general relativity and black holes, did postdocs at Stanford and Columbia and then became one of the very first data scientists, joining AppNexus back in 2012, around the same time “data scientist” became a job title at all. In this episode, she traces the field’s evolution from Bayesian models to BERT to today’s LLMs, and makes a compelling case that going deep on the underlying math matters more than ever, even now that AI can do the math for you. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

1 month, 2 weeks назад @ podtrac.com
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents 1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents

In Episode #1010, Jon Krohn digs into “the advisor strategy”, a clever pattern that pairs a fast, cheap executor model with a frontier-class advisor it can consult mid-task, all inside a single API call. Every agent builder faces the same tension: frontier models plan best but cost too much to run on every turn, while small models fumble the decisions that matter. Anthropic’s advisor tool resolves it with roughly a one-line code change, and the benchmarks are startling: Sonnet with an Opus advisor scored higher than Sonnet alone while costing 11.9% less, and Haiku’s BrowseComp score more than doubled at 85% lower cost than Sonnet solo. Jon covers the newest Fable 5 numbers, the practical go…

1 month, 2 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 1 month, 3 weeks назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 month, 3 weeks назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 month, 3 weeks назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 2 weeks назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

3 months, 2 weeks назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

3 months, 2 weeks назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 2 weeks назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 2 weeks назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

3 months, 2 weeks назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

4 months, 2 weeks назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

4 months, 2 weeks назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

4 months, 2 weeks назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

4 months, 2 weeks назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

6 months назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

6 months назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

7 months назад @ datascienceathome.com