Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 7 часов назад
Made a small model that extracts text from a white background [P]
Made a small model that extracts text from a white background [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

7 часов назад @ reddit.com
Recent project I worked on: End to End Edge ML platform [D]
Recent project I worked on: End to End Edge ML platform [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

18 часов назад @ reddit.com
CICD / KAFKA / KUBERNETES / Interview questions (MLE) [R]
CICD / KAFKA / KUBERNETES / Interview questions (MLE) [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

21 час назад @ reddit.com
Missed AAAI reciprocal reviewer nomination deadline — risk of desk rejection? [D]
Missed AAAI reciprocal reviewer nomination deadline — risk of desk rejection? [D] Missed AAAI reciprocal reviewer nomination deadline — risk of desk rejection? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

23 часа назад @ reddit.com
Multi-Tenant SaaS: Which Architecture Would You Choose? [D]
Multi-Tenant SaaS: Which Architecture Would You Choose? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 4 hours назад @ reddit.com
Neurips 2026 Main Track Theory Paper Tracker- Discussion Thread [D]
Neurips 2026 Main Track Theory Paper Tracker- Discussion Thread [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 5 hours назад @ reddit.com
I want to use AI coding agents for machine learning projects [D]
I want to use AI coding agents for machine learning projects [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 6 hours назад @ reddit.com
Open-weight 4B models approach o3-level medical question answering in Swedish [P]
Open-weight 4B models approach o3-level medical question answering in Swedish [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 9 hours назад @ reddit.com
We compared different LLMs on IMO 2026 [R]
We compared different LLMs on IMO 2026 [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 13 hours назад @ reddit.com
I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]
I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P] I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 14 hours назад @ reddit.com
Understanding GPU Inference Workloads [D]
Understanding GPU Inference Workloads [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 17 hours назад @ reddit.com
Link plots/figures in NeurIPS rebuttal [R]
Link plots/figures in NeurIPS rebuttal [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 18 hours назад @ reddit.com
Need Arxiv endorsement for CS.LG to publish a preprint about Learning Stable Latent Manifolds from noisy sensor data /Observation space [P]
Need Arxiv endorsement for CS.LG to publish a preprint about Learning Stable Latent Manifolds from noisy sensor data /Observation space [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 22 hours назад @ reddit.com
Paper lengths, and reasonable assumptions in ML conferences. [D]
Paper lengths, and reasonable assumptions in ML conferences. [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 2 hours назад @ reddit.com
Why first person video may matter for robot learning[D]
Why first person video may matter for robot learning[D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 4 hours назад @ reddit.com
Towards Data Science
последний пост 4 часа назад
“Los Movimientos”: The Routing Problem That Nearly Broke My Spirit
“Los Movimientos”: The Routing Problem That Nearly Broke My Spirit “Los Movimientos”: The Routing Problem That Nearly Broke My Spirit

In this article, we will:Translate the real-world nightmare of Los Movimientos into a formal pickup-and-delivery routing problem.

Understand the main elements of the problem, including vehicles, pickup locations, delivery locations, capacities, travel times, and time windows.

A household may request delivery between 8 a.m. and noon.

f"{minutes_to_clock(pyo.value(model.S[k, node]))}" f" | {node_location[node]:<18}" f" | {node_type[node].title():<8}" f"{request:<16}" f" | Passengers onboard: {onboard}" )The routes dictionary stores the ordered sequence needed for the plots.

The printed output then gives operators the planned time, location, activity, associated request, and number of passenge…

4 часа назад @ towardsdatascience.com
Reducing Human Annotation with ML Active Learning
Reducing Human Annotation with ML Active Learning Reducing Human Annotation with ML Active Learning

Using active learning, you can often achieve comparable performance to normal supervised learning with many fewer annotated training examples.

Learning From Uncertain or Edge-Case SamplesOne major benefit of active learning is that your model can focus on learning from the most difficult examples.

With active learning these rare cases can be presented early on and delivered to human annotators for review.

Active Learning TechniquesThe intuition behind active learning is using a query strategy to determine which unlabeled samples should be labeled first.

Concluding, Active Learning helps you prioritize what samples to request human annotation, and it’s specially useful when you have many unl…

6 часов назад @ towardsdatascience.com
The Most Beautiful Statistic: The History and the Science of the Humble Mean
The Most Beautiful Statistic: The History and the Science of the Humble Mean The Most Beautiful Statistic: The History and the Science of the Humble Mean

The sample mean serves as your working estimate of the population mean.

In other words, your sample mean will generally be an estimate of the population mean.

Suppose you wish to estimate an unknown quantity x. x could be the position of a new star you discovered in the night sky.

Let x 1 , x 2 , x 3 , …, x n be several observations of a random variable X.

Then, x 1 , x 2 , x 3 , …, x n are the readings taken on July 1 in n randomly selected years.

7 часов назад @ towardsdatascience.com
How I Reproduced BM25, Dense Retrieval, and SPLADE on a 16GB MacBook
How I Reproduced BM25, Dense Retrieval, and SPLADE on a 16GB MacBook How I Reproduced BM25, Dense Retrieval, and SPLADE on a 16GB MacBook

Diagram comparing BM25, dense retrieval, and SPLADE.

It is used here for a direct, same collection comparison of BM25, dense retrieval, and SPLADE side by side.

BM25, dense retrieval, and SPLADE side by side on NFCorpusNFCorpus is small enough to compare all three methods on the same laptop CPU.

What this actually takes, in practiceReproducing BM25, dense retrieval, and SPLADE correctly, at real scale, on a normal laptop, is genuinely doable.

BM25 is unsupervised sparse retrieval; SPLADE is learned sparse retrieval.

9 часов назад @ towardsdatascience.com
How to Efficiently Prompt Claude Code
How to Efficiently Prompt Claude Code How to Efficiently Prompt Claude Code

It just requires a different skill set and a different way of thinking to effectively prompt agents nowadays.

I’ll discuss how to work with and prompt your coding agents effectively.

How I prompt Claude CodeNow I’ll move on to how I prompt Claude Code specifically.

You basically have this discussion with your coding agent to align the implementation in your head with what the coding agent will actually implement.

ConclusionIn this article, I discussed how to prompt Claude Code and other coding agents effectively.

1 day, 6 hours назад @ towardsdatascience.com
How to Give an LLM Agent a Browser
How to Give an LLM Agent a Browser How to Give an LLM Agent a Browser

The Mental ModelAt a high level, a browser-using agent is best understood as an LLM placed in an interaction loop with a browser.

For the tech stack, here we’ll use the OpenAI Agents SDK to power the agent runtime and Playwright MCP to connect the agent to the browser.

Our final, configured agent looks like this:# pip install openai-agents from agents import Agent, ModelSettings from openai.types.shared import Reasoning agent = Agent( name="Support Console Browser Agent", model="gpt-5.4", model_settings=ModelSettings( reasoning=Reasoning(effort="medium"), ), instructions=AGENT_INSTRUCTIONS, mcp_servers=[playwright_server], )There are three pieces we need to unpack here, i.e., the LLM client…

1 day, 8 hours назад @ towardsdatascience.com
How to Optimize Vector Search When RAM Gets Too Expensive: On-Disk vs. In-Memory ANN Indexes
How to Optimize Vector Search When RAM Gets Too Expensive: On-Disk vs. In-Memory ANN Indexes How to Optimize Vector Search When RAM Gets Too Expensive: On-Disk vs. In-Memory ANN Indexes

At this scale, storing indexes and associated data in RAM will cost thousands of dollars per month, and HNSW can become a scalability bottleneck.

The implementations vary, although many modern ANN algorithms (for example, HNSW and DiskANN) rely on a graph structure to provide low query latency.

It’s a great choice if search latency is not critical and the index size is expected to be large.

The trade-offAs with everything in engineering, the cost reduction provided by on-disk ANN algorithms is not free.

General guidance is that disk-based ANN algorithms provide slower latency than HNSW (in RAM) just because RAM access is much faster.

2 days, 6 hours назад @ towardsdatascience.com
The Fluid Simulator That Doesn’t Solve the Fluid Equations
The Fluid Simulator That Doesn’t Solve the Fluid Equations The Fluid Simulator That Doesn’t Solve the Fluid Equations

Quick take: LBM replaces the intractable Navier-Stokes PDE with a much simpler update rule on a discrete velocity lattice.

Why the obvious approach is painfulThe standard route to fluid simulation goes through the Navier-Stokes equations, the macroscopic conservation laws that govern virtually everything we care about in fluid dynamics.

The physical wall sits halfway between the last fluid node and the first solid node, giving the scheme second-order spatial accuracy.

A perfectly equilibrium fluid would show no such structure.

The Navier-Stokes equations feel fundamental: they’re the equations of fluid dynamics, written at the scale we can see and touch.

2 days, 8 hours назад @ towardsdatascience.com
Tabular LLMs: An Introduction to the Foundation Models That Predict Your Spreadsheet
Tabular LLMs: An Introduction to the Foundation Models That Predict Your Spreadsheet Tabular LLMs: An Introduction to the Foundation Models That Predict Your Spreadsheet

What a tabular foundation model isA tabular foundation model is a single pretrained model that predicts on any table, zero-shot.

The contrast with the time-series foundation models from the t0 post comes down to order.

First, every single model above the best GBDT configuration is a tabular foundation model; the only non-FM entries above it are the AutoGluon ensemble pipelines.

(In time series I found the opposite — proper tuning halved the foundation models’ apparent edge.

Hatched bars are oracle ceilings; the blue bar is the deployable validation-picked router between the two foundation models, above both of its members.

3 days, 4 hours назад @ towardsdatascience.com
Build and Run an Intelligent Document Processing (IDP) System in the Cloud
Build and Run an Intelligent Document Processing (IDP) System in the Cloud Build and Run an Intelligent Document Processing (IDP) System in the Cloud

, I built a fully automated intelligent document processing (IDP) system that ran in the cloud on Amazon Web Services.

Click the Next button and select “AWS Step Functions” as your target.

SummaryThis article walks through the process of building a simplified Intelligent Document Processing system on AWS.

This work is carried out by AWS Lambda functions.

For our test cases, we could probably have bypassed the AWS Textract step altogether and just had Bedrock classify and extract the information directly from the images.

3 days, 6 hours назад @ towardsdatascience.com
Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship
Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship

So the answer is not “use the big model” and it is not “use the small model”.

Flagship models cost an order of magnitude more per call than a small local one.

If speed is the constraint, the cheap local model is often the wrong default, and that belongs in the dispatcher’s criteria.

The small local model goes from 38% to 62% correct.

That forces a local model regardless of what would be cheapest or strongest in the cloud, and the starting choice becomes the best local model that fits the machine.

3 days, 7 hours назад @ towardsdatascience.com
Context Windows Forget What Matters — I Built a Usage-Reinforced Decay Engine for AI Agent Memory
Context Windows Forget What Matters — I Built a Usage-Reinforced Decay Engine for AI Agent Memory Context Windows Forget What Matters — I Built a Usage-Reinforced Decay Engine for AI Agent Memory

TL;DR:AI agent memory systems evict old context using a sliding window: if something hasn’t been touched in N turns, it’s gone.

The sliding window baseline is that vertical line cutting straight through.

It completely ignores both of these curves and drops the data at a fixed cutoff, no matter how stable either memory actually is.

In my first version of the synthetic session generator, I scheduled recall turns using pure uniform-random sampling across the entire session length.

Sensitivity Sweep (N=20 seeds per configuration, 15 threshold × window combinations tested)Threshold Window Ebbinghaus FRR Baseline FRR Result holds?

3 days, 9 hours назад @ towardsdatascience.com
When Data Science Makes Us Sad: The Story of an Overbooked Flight
When Data Science Makes Us Sad: The Story of an Overbooked Flight When Data Science Makes Us Sad: The Story of an Overbooked Flight

The probability of a passenger showing up for a flight is 95% based on historical flight data.

There will be at least one overbooked passengers if 301, 302, 303, or 304 passengers show up for the flight.

We can use the probabilities we already calculated for these events:If 301 passengers show up, there is 1 overbooked passenger.

If 302 passengers show up, there are 2 overbooked passengers, and so on.

Out of those 10,000 flights, we expect to have about 1.66 overbooked passengers.

4 days, 4 hours назад @ towardsdatascience.com
Most RAG Hallucinations Are Extraction Errors: Seven Patterns for a Typed Generation Contract
Most RAG Hallucinations Are Extraction Errors: Seven Patterns for a Typed Generation Contract Most RAG Hallucinations Are Extraction Errors: Seven Patterns for a Typed Generation Contract

The fix is a small set of generation patterns, one per answer shape.

In RAG the model reads the context; when the answer is wrong, the cause is upstream, in the extraction chain (parsing, question, retrieval, or the generation contract itself).

A typed contract is three regions, each with its own purpose and validator – Image by authorBelow are the seven patterns that keep generation on the typed-contract side.

→ Article 8A (the answer contract) develops the typed contract in full.

The seven patterns share one move: refuse the LLM-as-oracle frame, and treat generation as filling a typed contract.

4 days, 6 hours назад @ towardsdatascience.com
Lessons Learned After 8.5 Years of ML
Lessons Learned After 8.5 Years of ML Lessons Learned After 8.5 Years of ML

When I started my studies of machine learning, the only thing I back then knew about it was the classic machine learning techniques, such as KNN or clustering algorithms.

However, I think that there are still some lessons that are applicable regardless of the field one actually is doing machine learning research or machine learning practices in.

In this edition of my lessons learned articles I look back at the previous eight years and try to distill lessons that are persisting in coming up over and over again.

I frequently like to compare the progress in machine learning with the progress one makes in sportive activities.

Only afterwards can you advance to the high-output zone, where you pr…

4 days, 7 hours назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 1 day, 9 hours назад
The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack
The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack

Next Week in The Sequence:Our series about AI model distillation continues with another exciting technique.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI StackWhen I started The Sequence years ago, AI was still a relatively niche field, followed closely by researchers, a small group of builders, and a few overly enthusiastic people like me.

AI Lab: Meta AISummary: This paper introduces GAMUT, a multimodal benchmark designed to evaluate the factual completeness of long-form generations rather than just their factual precision.

AI Lab: Microsoft ResearchSummary: This paper presents Experiential Learning (EL)…

1 day, 9 hours назад @ thesequence.substack.com
The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA?
The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA? The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA?

THESIS Google is the closest strategic mirror of NVIDIA’s full-stack method, but not a universal drop-in replacement; AMD and AWS make a literal “only” claim too strong.

This is useful, but incomplete in the same way that comparing airlines by engine thrust is incomplete.

It is an industrial system that turns models into running software with unusually little friction.

That is the strongest form of the case for Google as NVIDIA’s only viable competitor.

Google is the closest full-stack strategic rival, not a universal drop-in replacement, and AWS and AMD make the word “only” uncomfortable.

4 days, 10 hours назад @ thesequence.substack.com
The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time
The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time

Inkling is best understood not as a single 975-billion-parameter brain that fires all at once, but as a giant warehouse of specialist capacity.

The more interesting number is the one beside it: 41 billion parameters active per token.

A good mental picture is a university with 256 specialist departments on each relevant floor.

It selects six departments that appear useful for this token, adds two general-purpose departments that always attend, combines their work, and moves on.

A line of Python might summon one set of specialists; a phrase in Greek, a diagram label, or a piece of audio may summon another.

5 days, 10 hours назад @ thesequence.substack.com
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models

The distilled 32B model started solving competition math it had no business solving.

The 7B model began verifying its own work and branching its reasoning mid-stream — emergent behaviors nobody trained into it directly.

A grab-bag of small dense models suddenly reasoned like something ten times their size.

We spent an entire installment establishing why naive sequence-level imitation is the wrong tool, and then the single most important reasoning-distillation result of the decade is naive sequence-level imitation.

The answer is the whole story of this installment, and it turns out to be more interesting than either “imitation works” or “imitation doesn’t.”

6 days, 10 hours назад @ thesequence.substack.com
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race

In the AI of the Week , we discuss Thinking Machine first open weights model.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: China, Compression and the Open-Model RaceFor years, AI progress has been narrated as a horse race: larger models, higher benchmark scores, more expensive clusters.

Moonshot describes it as the first open model in the three-trillion-parameter class, although the full weights are not due until later this month.

This week, AI stopped looking like a single race.

AI Lab: Shanghai AI LaboratorySummary: ADVANCED MATHBENCH introduces a rigorous evaluation suite focusing on the generation and process-level verification of advanced, natural-language mathematical pr…

1 week, 1 day назад @ thesequence.substack.com
The Sequence Opinion #896: Spark, Compute, and the Two Metas
The Sequence Opinion #896: Spark, Compute, and the Two Metas The Sequence Opinion #896: Spark, Compute, and the Two Metas

The occasion was the launch of Muse Spark 1.1, the second model out of Meta Superintelligence Labs and the first Meta model ever to ship with a price tag.

Two days earlier Meta shipped Muse Image, its first image generation model from the new lab.

Chips, datacenters, cloud, models, API, apps, devices.

At the layer where models meet users, the app and agent layer, Meta might be the favorite.

At the layer where models get made, the evidence is thin and the structural arguments cut against it.

1 week, 4 days назад @ thesequence.substack.com
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break

OpenAI’s audit of SWE-Bench Pro shows why a precise score can still be a poor measure - and why coding agents may become essential tools for auditing the benchmarks that grade them.

A frontier coding score can look wonderfully precise: 80.3 percent, one decimal place, clean enough to rank models and anchor product claims.

That is the uncomfortable conclusion of OpenAI’s audit of SWE-Bench Pro.

OpenAI estimates that roughly 30 percent of the public benchmark is broken.

OpenAI has withdrawn its earlier recommendation that the field adopt SWE-Bench Pro.

1 week, 5 days назад @ thesequence.substack.com
The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era
The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era

Looking back at the 2015 distillation paper, what’s striking isn’t the temperature trick or the dark-knowledge framing --- it’s the world the paper quietly assumed.

There was a fixed input distribution.

There was a teacher that produced a probability vector over a closed set of classes.

There was a student trained to match that vector.

Then language models arrived, and one by one, every assumption broke.

2 weeks назад @ thesequence.substack.com
The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack

In the opinion section, we are going to debate Meta’s opportunities and tremendous challenges to catch up with the AI frontier labs.

This week’s GPT-5.6, GPT-Live, ChatGPT Work, Grok 4.5, and Muse Spark 1.1 reveal a shift.

Meta’s Muse Spark 1.1 makes the race more crowded and cheaper.

AI Lab: LMMS-Lab, NTU MMLab, and MicrosoftSummary: This paper introduces SkillOpt-Lite, a minimal viable pipeline for autonomous agent skill optimization that replaces complex algorithmic architectures with a file-system-based trajectory exploration approach.

Muse Spark 1.1Meta introduced Muse Spark 1.1, the second version of its multimodal reasoning model.

2 weeks, 1 day назад @ thesequence.substack.com
The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough
The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough

Quick note: For over two years now, we’ve been running The Sequence without sponsors.

Literally every week we receive tens of inquires to sponsor The Sequence.

The question is deceptively simple: what makes a domain a good domain for AI?

Not good in the sense of commercially interesting, but good in the sense that if you point a modern training pipeline at it, capability actually compounds.

Verifiability

2 weeks, 4 days назад @ thesequence.substack.com
The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI
The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI

Quick note: For over two years now, we’ve been running The Sequence without sponsors.

Literally every week we receive tens of inquires to sponsor The Sequence.

There’s a running joke in machine learning that the field’s most valuable model isn’t a transformer at all — it’s gradient-boosted trees fit on a CSV.

You hand it the whole problem — training rows, test rows, all of it — as one giant prompt, and it answers.

Understanding TabFM really requires understanding that lineage, so let’s start there.

2 weeks, 5 days назад @ thesequence.substack.com
The Sequence Knowledge #890: A Brief History of Model Distillation
The Sequence Knowledge #890: A Brief History of Model Distillation The Sequence Knowledge #890: A Brief History of Model Distillation

The story most people tell about knowledge distillation starts in 2015, with Geoffrey Hinton, Oriol Vinyals, and Jeff Dean introducing a clever softmax temperature trick and a phrase — “dark knowledge” — that immediately lodged itself in the field’s vocabulary.

The real history is quieter, more pragmatic, and worth recovering, because the conceptual moves the field made between 2006 and 2015 still define how we think about distillation today.

The vocabulary changed.

The diagrams changed.

The underlying question — what exactly is being transferred from a teacher to a student?

2 weeks, 6 days назад @ thesequence.substack.com
The Sequence Radar #889: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land Grab
The Sequence Radar #889: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land Grab The Sequence Radar #889: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land Grab

The opinion section, discusses that domains are a good fit for rapid progress of AI models vs which ones are challenging.

Subscribe and don’t miss out:📝 Editorial: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land GrabThis week’s developments in AI provide a clear direction where the space is going: more capable models and the imperative of capable delivery capabilities.

Start at the model layer, where we just watched the first frontier model get hot-patched by a government.

One layer up, Anthropic launched what I’d call claude --science .

Claude ScienceAnthropic announced Claude Science, a workbench for scientists.

3 weeks, 1 day назад @ thesequence.substack.com
The Sequence Opinion #888: Everything You Need to Know About the AI in Space Race
The Sequence Opinion #888: Everything You Need to Know About the AI in Space Race The Sequence Opinion #888: Everything You Need to Know About the AI in Space Race

The core thesis of this essay is simple to state: space is becoming a new frontier for AI — and one of the most competitive ones.

When the scarce thing was ideas, the frontier was architectures; when it was data, the frontier was the open web; when it was FLOPs, the frontier was the fab.

Today the scarce thing is energy — grid capacity, cooling water, land, permits — and orbit is the one place in reach where energy is effectively unmetered and no zoning board has jurisdiction.

This essay discusses the core thesis of AI in space: value proposition, key players, architecture differences and much more.

The core thesis: compute is now an energy problem, and space is an energy solution

3 weeks, 4 days назад @ thesequence.substack.com
The Sequence AI of the Week #887: Meta's Autodata: When Models Learn to Make Their Own Lessons
The Sequence AI of the Week #887: Meta's Autodata: When Models Learn to Make Their Own Lessons The Sequence AI of the Week #887: Meta's Autodata: When Models Learn to Make Their Own Lessons

Today, we are covering an amazing paper published by Meta last week: https://arxiv.org/abs/2606.25996There is a quiet shift happening in AI training.

For years, the center of gravity was the model: more parameters, more GPUs, better architectures, longer context windows, better optimizers.

Meta’s new Autodata work flips that perspective.

Not “ask a strong model to generate a million examples and hope the distribution is useful.” Instead, Autodata treats data generation like a miniature research loop.

An AI agent creates examples, tests them, studies the failures, updates its recipe, and tries again.

3 weeks, 5 days назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 3 months, 3 weeks назад
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

3 months, 3 weeks назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером &amp; Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером &amp; Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

5 months, 1 week назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

5 months, 1 week назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing &amp; LLMs — новый сезон Курс Natural Language Processing &amp; LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

5 months, 4 weeks назад @ habr.com
Machine Learning Mastery
последний пост 2 weeks, 3 days назад
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach

The common pitfalls that show up once agent memory is implemented, and how to fix them.

Agent memory strategy deserves the same deliberate design as orchestration.

Why Is Choosing an AI Agent Memory Strategy Important?

A customer support agent, for example, might keep the current ticket in working memory, a customer’s subscription tier in semantic memory, past complaints in episodic memory, and a learned refund-handling routine in procedural memory.

As discussed, working memory, semantic memory, episodic memory, and procedural memory serve different purposes and require different storage and retrieval strategies.

2 weeks, 3 days назад @ machinelearningmastery.com
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls

openai import OpenAI as LlamaOpenAI from llama_index .

# Prerequisites: pip install openai python-dotenv # How to run: python raw_api_agent.py import os import json from dotenv import load_dotenv from openai import OpenAI load_dotenv ( ) client = OpenAI ( api_key = os .

Prerequisites:pip install openai langchain langchain-openai llama-index \ llama-index-llms-openai llama-index-embeddings-openai python-dotenv 1 2 pip install openai langchain langchain - openai llama - index \ llama - index - llms - openai llama - index - embeddings - openai python - dotenvHow to run: Save as three_ways.py and run python three_ways.py# three_ways.py # The same document Q&A task implemented three ways: # Raw …

2 weeks, 4 days назад @ machinelearningmastery.com
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering

Topics we will cover include:What tools and subagents are, and the key differences between them.

This article explains what tools and subagents are, where each fits, and how to make the choice every time.

When an agent calls a tool, the result lands back in the same context the agent is actively reasoning in — prior reasoning, tool result, and everything else together.

A database record, search result, or API response can often be consumed immediately.

If the answer is independent reasoning, context isolation, specialized capabilities, or parallel execution, a subagent is likely justified.

2 weeks, 6 days назад @ machinelearningmastery.com
The Complete Guide to Tool Selection in AI Agents
The Complete Guide to Tool Selection in AI Agents The Complete Guide to Tool Selection in AI Agents

tools = tools self .

tools = tools self .

threshold : return { "status" : "resolved" , "tool" : tool [ "name" ] , "confidence" : score , "attempts" : 1 } # Reformulate by stripping filler words.

tools = tools self .

full_catalog_tokens = sum ( estimate_tokens ( d ) for d in descs ) def _retrieve ( self , query : str , top_k : int ) -> list [ dict ] : query_vec = self .

3 weeks назад @ machinelearningmastery.com
Context vs. Memory Engineering in Agentic AI Systems
Context vs. Memory Engineering in Agentic AI Systems Context vs. Memory Engineering in Agentic AI Systems

Share Post ShareIn this article, you will learn how context engineering and memory engineering solve different problems in agentic AI systems, and how the two disciplines meet at the point where retrieved memory enters the context window.

Most of the time, the problem lies in two areas that get built together, conflated, or skipped: context engineering and memory engineering.

Memory Engineering: Designing Persistent AI Memory SystemsOnce an inference call completes, memory engineering determines what deserves to persist and under what conditions it gets used again.

trust_level >= 0.5 )AI Agent Memory Design Guide – Working, Long-Term, and Procedural Memory with Forgetting and Staleness Mana…

3 weeks, 4 days назад @ machinelearningmastery.com
Context Window Management for Long-Running Agents: Strategies and Tradeoffs
Context Window Management for Long-Running Agents: Strategies and Tradeoffs Context Window Management for Long-Running Agents: Strategies and Tradeoffs

Share Post ShareIn this article, you will learn five practical strategies for managing context windows in long-running AI agent applications, along with the key tradeoffs each approach introduces.

Five distinct context management strategies: sliding windows, recursive summarization, structured state management, ephemeral context via RAG, and dynamic context routing.

Accordingly, shifting from “LLMs as prompt-response engines” to “(agent-endowed) LLMs as long-running background processes” turns context windows into a major AI engineering bottleneck.

For all these reasons, managing context windows in the long run requires specific strategies like sliding windows, tiered memory, and dynamic su…

3 weeks, 6 days назад @ machinelearningmastery.com
Model Context Protocol Explained in 3 Levels of Difficulty
Model Context Protocol Explained in 3 Levels of Difficulty Model Context Protocol Explained in 3 Levels of Difficulty

Share Post ShareIn this article, you will learn how the Model Context Protocol (MCP) standardizes the way AI applications connect to external tools and data sources, broken down across three levels of depth.

How the host, client, and server work together, and what happens when a model’s request flows through an MCP server.

The transport options, security risks, and deployment choices that matter once an MCP server is running in production.

Accessing information outside that context requires external tools.

Because both sides follow the same protocol, an MCP server can be used by any compatible MCP client without requiring a custom integration for that specific client.

4 weeks назад @ machinelearningmastery.com
The AI Agent Tech Stack Explained
The AI Agent Tech Stack Explained The AI Agent Tech Stack Explained

According to Atlan’s research on AI agent memory, 95% of enterprise generative AI pilots delivered zero measurable ROI in 2025, with failure attributed to context readiness rather than model quality.

isoformat ( ) , "user" : user_input , "agent" : agent _ response } ) def load_recent_episodes ( n : int = 5 ) -> str : "" "Retrieve the last N episodes as a formatted string for injection into context."

Prerequisites:pip install langchain langchain-openai langchain-chroma langchain-text-splitters chromadb python-dotenv 1 pip install langchain langchain - openai langchain - chroma langchain - text - splitters chromadb python - dotenvHow to run: Save as rag_pipeline.py, add OPENAI_API_KEY to your…

1 month назад @ machinelearningmastery.com
Agentic Workflow vs. Autonomous Agent: What’s the Difference?
Agentic Workflow vs. Autonomous Agent: What’s the Difference? Agentic Workflow vs. Autonomous Agent: What’s the Difference?

lower ( ) : return "billing" return "general" def extract ( raw_input : str ) -> str : "" "Step 1 -- always runs, always leads to step 2.

def handle_billing(text: str) -> str: return f"[BILLING TEAM] Routed: {text[:50]}" def handle_technical(text: str) -> str: return f"[TECH SUPPORT] Routed: {text[:50]}" def handle_general(text: str) -> str: return f"[GENERAL QUEUE] Routed: {text[:50]}" # The branch map IS the entire decision space.

def handle_billing ( text : str ) -> str : return f "[BILLING TEAM] Routed: {text[:50]}" def handle_technical ( text : str ) -> str : return f "[TECH SUPPORT] Routed: {text[:50]}" def handle_general ( text : str ) -> str : return f "[GENERAL QUEUE] Routed: {text…

1 month назад @ machinelearningmastery.com
Context Windows Are Not Memory: What AI Agent Developers Need to Understand
Context Windows Are Not Memory: What AI Agent Developers Need to Understand Context Windows Are Not Memory: What AI Agent Developers Need to Understand

Topics we will cover include:Why a context window behaves like a stateless scratchpad rather than persistent memory.

Deeming a huge context window as “memory” is, in architectural terms, similar to buying a 25-foot-wide office desk because you are reluctant to acquire a filing cabinet.

Context WindowA context window in an AI model, particularly agent-based ones with underlying language models, is like a desk surface or a stateless scratchpad.

When passing an agent a conversation history spanning over 200K tokens (large context window), it isn’t remembering what happened at a previous step in time.

Generate a compact summary for the active prompt summary = summarizer_model.generate(raw_trans…

1 month назад @ machinelearningmastery.com
Clustering Unstructured Text with LLM Embeddings and HDBSCAN
Clustering Unstructured Text with LLM Embeddings and HDBSCAN Clustering Unstructured Text with LLM Embeddings and HDBSCAN

Share Post ShareIn this article, you will learn how to build a text clustering pipeline by combining large language model embeddings with HDBSCAN, a density-based clustering algorithm, to automatically discover topics in unlabeled text data.

target } ) df = df [ df [ 'text' ] .

) print ( "Sample document:" ) print ( df [ 'text' ] .

cluster import HDBSCAN # Initializing HDBSCAN # min_cluster_size=8: we specified that each cluster must have at least 8 documents clusterer = HDBSCAN ( min_cluster_size = 8 , min_samples = 3 , store_centers = 'centroid' ) df [ 'cluster' ] = clusterer .

shape [ 1 ] ) ] ) reduced_df [ 'cluster' ] = df [ 'cluster' ] # Getting all unique pairwise combinations of the …

1 month назад @ machinelearningmastery.com
Building Browser-Using AI Agents in Python
Building Browser-Using AI Agents in Python Building Browser-Using AI Agents in Python

# Prerequisites: pip install playwright && playwright install chromium # How to run: python scrape_books.py import asyncio import json from playwright .

"" page = await get_page ( ) await page .

"" page = await get_page ( ) await page .

"" page = await get_page ( ) await page .

"" page = await get_page ( ) try : await page .

1 month назад @ machinelearningmastery.com
The Roadmap to Mastering AI Agent Evaluation
The Roadmap to Mastering AI Agent Evaluation The Roadmap to Mastering AI Agent Evaluation

Step 1: Understanding Why Agent Evaluation Is ImportantThe instinct when an agent fails is to treat it as a prompting problem: the system prompt needs to be clearer.

Step 2: Defining What Agent Evaluation Success Looks LikeEvaluation is only as good as its success criteria.

Step 4: Grading Agent Reasoning and Output Quality with Model-Based JudgesSome agent evaluation dimensions resist deterministic checking — output quality, tone, faithfulness to retrieved context, appropriate empathy.

Step 5: Matching Agent Evaluation Strategy to Agent TypeGrading strategies apply broadly, but agent type determines which graders carry the most weight and which failure modes to prioritize.

Summary of Key A…

1 month, 1 week назад @ machinelearningmastery.com
Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM
Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM

How to build, run, and evaluate a zero-shot sentiment classification pipeline using scikit-learn-compatible syntax.

from sklearn.pipeline import Pipeline from skllm.models.gpt.classification.zero_shot import ZeroShotGPTClassifier # Define the end-to-end pipeline sentiment_pipeline = Pipeline([ ("cleaner", text_cleaner), # Updated to use Groq's active Llama 3.1 8B model ("llm_classifier", ZeroShotGPTClassifier(model="custom_url::llama-3.1-8b-instant")) ]) # Fit the pipeline # Note: For Zero-Shot classification, fit() doesn't train the LLM.

pipeline import Pipeline from skllm .

. . Actual : negative | Predicted : negative Review : This entry is certainly interesting for series fans ( like mys…

1 month, 1 week назад @ machinelearningmastery.com
AI Agent Tool Design: What Works and What Doesn’t
AI Agent Tool Design: What Works and What Doesn’t AI Agent Tool Design: What Works and What Doesn’t

What Works in AI Agent Tool Design1.

The difference becomes clearer when comparing a multi-action tool against dedicated single-purpose tools:# Avoid: action-based multi-behavior tool @tool def manage_customer( action: str, customer_id: str | None = None, data: dict | None = None ): """ action: create | get | update | delete | suspend """ ... # Prefer: single-responsibility tools @tool def create_customer(data: CustomerInput) -> Customer: """Create a new customer record."""

... @tool def suspend_customer(customer_id: str, reason: str) -> SuspensionResult: """Suspend a customer account."""

. . # Prefer: single-responsibility tools @ tool def create_customer ( data : CustomerInput ) -> Custom…

1 month, 1 week назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 1 week, 3 days назад
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

1 week, 3 days назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

2 months, 1 week назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

4 months, 2 weeks назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

5 months, 4 weeks назад @ alexirpan.com
Authentic Imperfection
Authentic Imperfection Authentic Imperfection

* * *I’ve been thinking about the anger surrounding generative AI.

To keep things fair, he took the best human images and best AI images, meaning human art from famous artists, and AI art from prompters skilled at removing obvious tells of image generation.

When people complain about AI slop, I see it as a complaint against the deluge of default style AI images.

We’ve seen this happen in all forms: AI text, AI music, older forms of computer generated content like CGI.

As much as we celebrate imperfection, digital imperfection is a step too far.

8 months, 1 week назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 5 months назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

5 months назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

5 months, 3 weeks назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 5 days, 7 hours назад
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

In December, we shared our commitment to the White House's Genesis Mission — the national effort to harness AI and double the pace of American scientific discovery within a decade.

Today, at the DOE Genesis Mission Summit 2026, we are expanding this by committing $40 million of AI tokens and cloud credits for researchers in support of the Genesis Mission.

WeatherNext — a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

— a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

Driving American innovationThe Genesis Mission represents an opportunity to transform research and science across America.

5 days, 7 hours назад @ cloud.google.com
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Building on Gemini 3.5 Flash, we’re introducing new Gemini models:3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure.

3.6 Flash: More efficient and better quality than 3.5 FlashGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash.

For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash.

This enhanced efficiency is also combined with a lower price than 3.5 Flash.

6 days, 5 hours назад @ blog.google
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Building on Gemini 3.5 Flash, we’re introducing new Gemini models:3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure.

3.6 Flash: More efficient and better quality than 3.5 FlashGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash.

For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash.

This enhanced efficiency is also combined with a lower price than 3.5 Flash.

6 days, 5 hours назад @ blog.google
Introducing Gemini 3.5 Flash Cyber
Introducing Gemini 3.5 Flash Cyber Introducing Gemini 3.5 Flash Cyber

Today, we’re expanding our longtime efforts to better prepare defenders by introducing Gemini 3.5 Flash Cyber, our lightweight cybersecurity model built on top of 3.5 Flash and fine-tuned to find, validate, and patch vulnerabilities quickly and efficient, making it more effective at these tasks than Gemini’s mainline Flash models.

By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.

Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber.

3.5 Flash Cyber benchmark results: an efficient alternative to larger cybersecurity modelsWe tested 3.5 F…

1 week, 3 days назад @ deepmind.google
Our approach to bioresilience
Our approach to bioresilience Our approach to bioresilience

Today, Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience.

Inside our bioresilience programWe believe society must harness AI’s advancing capabilities to address infectious diseases and prepare for future outbreaks.

With this in mind, we are making our AI models and agents available to trusted partners to support progress across three key areas: prevention, detection and response.

Working in collaboration with governments and global health authorities to advance a diverse range of diagnostic and therapeutic strategies enables the Isomorphic Labs Drug Design Engine’s real-world impact for bioresilience.

To read more about our work and our call for new par…

1 week, 4 days назад @ deepmind.google
Empowering India’s next generation of innovators with ATL Saathi
Empowering India’s next generation of innovators with ATL Saathi Empowering India’s next generation of innovators with ATL Saathi

A new contribution to Indian Education with Atal Innovation MissionWe believe behind every good student is a great teacher.

That’s why for over 20 years, Google has been dedicated to supporting the education ecosystem by introducing technology into teaching and learning through a teacher-led approach.

With foundational platforms like Google for Education and Google Classroom, we build products tailored to the needs of schools, keeping the teacher in the lead.

To further support the empowerment of educators, our new Google Educator AI Series ensures teachers are equipped with both the tools and the digital skills required for today's classrooms.

We see Gemini as a great tool to enable our pa…

2 weeks назад @ deepmind.google
Google DeepMind and A24 announce first-of-its-kind research partnership
Google DeepMind and A24 announce first-of-its-kind research partnership Google DeepMind and A24 announce first-of-its-kind research partnership

Today, Google DeepMind and A24 are announcing a first-of-its-kind partnership focused on research.

The collaboration pairs a world-leading research lab with the industry’s most filmmaker-forward studio to help artists develop new workflows and techniques.

This partnership creates a deep research and development collaboration between A24 and Google DeepMind spanning multiple projects over time.

This hands-on collaboration provides Google DeepMind with invaluable feedback and guidance from leading artists.

As A24 and Google DeepMind’s researchers work side-by-side to test, iterate and build, this partnership aims to expand what is possible in the future of entertainment.

3 weeks, 3 days назад @ blog.google
Start building with Nano Banana 2 Lite and Gemini Omni Flash
Start building with Nano Banana 2 Lite and Gemini Omni Flash Start building with Nano Banana 2 Lite and Gemini Omni Flash

Uploading audio references and scene extension is not yet supported in the Gemini API for this model.

Video references up to 3 seconds in duration are accepted by the API schema but are not correctly processed by the model at this time.

Gemini Omni is available in public preview starting today in Google AI Studio and the Gemini API.

Use Nano Banana 2 Lite as a high-speed image generation model, then pass that image as a reference to Gemini Omni Flash to animate it into a high-quality video.

To help you get started we created a few demo apps you can remix that let you experience how you can pair both Nano Banana 2 Lite and Gemini Omni Flash into one workflow.

3 weeks, 6 days назад @ blog.google
Introducing computer use in Gemini 3.5 Flash
Introducing computer use in Gemini 3.5 Flash Introducing computer use in Gemini 3.5 Flash

Making computer use safe in 3.5 FlashTo mitigate some of the prompt injection risks for agents operating in live environments, we use targeted adversarial training for computer use in Gemini 3.5 Flash.

We’re also releasing two optional enterprise safeguard systems that enable enterprises to:Require explicit user confirmation for sensitive or irreversible actions.

Automatically stop tasks if an indirect prompt injection is identified.

Taking a “defense-in-depth” approach, we encourage developers to combine these features with secure sandboxing, human-in-the-loop verification and strict access controls.

We are already seeing customers drive value with computer use.

1 month назад @ blog.google
Unlocking UK house-building with AI-accelerated planning
Unlocking UK house-building with AI-accelerated planning Unlocking UK house-building with AI-accelerated planning

New UK government AI planning prototype built with Gemini aims to halve the time it takes to process homeowner applicationsAround the world, Governments are exploring how AI can deliver better public services, faster.

The UK is working to build 1.5 million new homes by 2029, but local planning authorities are often slowed down by dense paperwork and administrative backlogs.

To help get Britain building, we’re partnering with the UK government to help radically shorten the time it takes to process householder planning applications.

Following early trials in Barnet, Camden and Dorset, the government plans for the new AI planning tool to be made available to all councils nationally from 2027.

1 month, 1 week назад @ deepmind.google
Securing the future of AI agents
Securing the future of AI agents Securing the future of AI agents

How we’re securing internal systems against increasingly capable and imperfectly aligned AIAI agents are transforming our relationship with technology.

In the U.S alone, AI agents could create $2.9 trillion in economic value by 2030.

That’s why we developed our AI Control Roadmap: a framework for building and managing the advanced AI we deploy within Google.

Similarly, our AI control system grants AI agents permissions based on their verified behavior, allowing us to build trust through controlled, incremental access.

In our AI Control Roadmap, we map security protocols to measurable milestones in AI capabilities on two critical fronts:

1 month, 1 week назад @ deepmind.google
DiffusionGemma: 4x faster text generation
DiffusionGemma: 4x faster text generation DiffusionGemma: 4x faster text generation

While the AI research community has explored diffusion-based text generation for years, applying it to large models has remained a challenge.

DiffusionGemma changes this by shifting how models use hardware.

But when run locally for a single user, this word-by-word process leaves your dedicated GPU or TPU underutilized — it spends most of its time simply waiting for the next "keystroke."

By giving the computer's processor a larger chunk of work at once, DiffusionGemma utilizes your hardware to its full potential.

It upgrades your model inference from a single, sequential typewriter to a massive printing press that stamps the entire block of text simultaneously.

1 month, 2 weeks назад @ blog.google
Investing in multi-agent AI safety research
Investing in multi-agent AI safety research Investing in multi-agent AI safety research

Scaling AI Safety Research for a Multi-Agent WorldFor the past decade, we’ve focused on making individual AI models more capable, helpful and safe.

The funding call focuses on the study of how large-scale multi-agent AI systems behave as a group, and how we can provide frameworks to understand and mitigate against potential risks.

Scaling the frontier of multi-agent safety researchAlthough foundational frameworks for multi-agent safety exist, the rapid evolution of these systems requires an immediate, large-scale expansion of research.

A collaborative call to actionNo single lab can solve multi-agent safety alone.

Building realistic, reproducible environments to evaluate, compare and accele…

1 month, 2 weeks назад @ deepmind.google
Fluid, natural voice translation with Gemini 3.5 Live Translate
Fluid, natural voice translation with Gemini 3.5 Live Translate Fluid, natural voice translation with Gemini 3.5 Live Translate

Today, we’re taking our next step with the release of Gemini 3.5 Live Translate, our latest audio model for live speech-to-speech translation.

The model automatically detects 70+ languages and generates smooth, natural-sounding translated speech that preserves the speakers' intonation, pacing and pitch.

Unlike turn by turn systems that wait for the speaker to finish speaking before responding, 3.5 Live Translate generates speech continuously, balancing the trade-off between waiting for context to improve quality and translating immediately to stay in sync with the speaker.

Gemini 3.5 Live Translate is rolling out starting today across Google products:For developers in public preview via the…

1 month, 2 weeks назад @ blog.google
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Introducing Gemma 4 12B: a unified, encoder-free multimodal model Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Today, we are introducing Gemma 4 12B, our latest model designed to bring agentic multimodal intelligence directly to laptops.

Here’s an overview of what makes Gemma 4 12B unique:Novel unified architecture: No multimodal encoders.

Advanced reasoning: Benchmark performance nearing our 26B model, unlocking powerful multi-step reasoning and agentic workflows.

Benchmark performance nearing our 26B model, unlocking powerful multi-step reasoning and agentic workflows.

Small enough to run locally on consumer laptops with 16GB of RAM, it unlocks powerful multimodal and agentic experiences right on your machine.

1 month, 2 weeks назад @ blog.google
Google
последний пост 4 days, 5 hours назад
The Blueprint: How Voicify makes AI-enabled ordering a delight for customers
The Blueprint: How Voicify makes AI-enabled ordering a delight for customers The Blueprint: How Voicify makes AI-enabled ordering a delight for customers

Founded in 2018, Voicify reimagines the traditional phone call with the goal of transforming every call into a seamless and engaging experience.

We realized that specialized, purpose-driven AI assistants were the key to businesses maintaining excellent service at scale.

Under the hood, Gemini Flash, served via Gemini Enterprise Agent Platform, vastly improves latency, minimizing user wait times and preventing hang-ups.

To grow the business — and call volume — and to handle traffic spikes, we switched from Google AI Studio to Vertex AI and its current incarnation in Gemini Enterprise.

Specific Gemini Enterprise Agent Platform features help us manage high call volumes without experiencing ser…

4 days, 5 hours назад @ cloud.google.com
Now in preview: Find and fix software vulnerabilities with CodeMender
Now in preview: Find and fix software vulnerabilities with CodeMender Now in preview: Find and fix software vulnerabilities with CodeMender

It examines and remediates existing code security issues without sacrificing development velocity by:Deploying the best-fit model .

Find and fix vulnerabilities with AIBorn from Google DeepMind's pioneering AI research, CodeMender transforms vulnerability management from a manual bottleneck into an autonomous, high-speed system.

CodeMender brings AI into a critical part of the security lifecycle by accelerating the path from validated vulnerability to tested fix.

"CodeMender consistently identified critical vulnerabilities that our other AI-enabled tools completely missed.

How the CodeMender agent worksWe’ve fine-tuned CodeMender’s harness to be continuously updated with the latest Google D…

6 days, 6 hours назад @ cloud.google.com
13 hands-on demos to build on Gemini Enterprise Agent Platform
13 hands-on demos to build on Gemini Enterprise Agent Platform 13 hands-on demos to build on Gemini Enterprise Agent Platform

Earlier this year, we introduced Gemini Enterprise Agent Platform, where you can build, scale, govern, and optimize agents.

Today, we’re sharing 13 demos that walk you through what Agent Platform can do.

The ambient expense agent codelab is the most complete "Agent Platform in action" demo in the set.

Deploy a stateful data science agent to Agent Runtime (formerly known as Agent Engine).

You deploy a multi-tool ADK agent on Agent Runtime that calls MCP servers on Cloud Run through Agent Gateway.

1 week, 3 days назад @ cloud.google.com
Google is a Leader and positioned furthest in Vision and highest in Execution in the 2026 Gartner® Magic Quadrant™ for Conversational AI Platforms
Google is a Leader and positioned furthest in Vision and highest in Execution in the 2026 Gartner® Magic Quadrant™ for Conversational AI Platforms Google is a Leader and positioned furthest in Vision and highest in Execution in the 2026 Gartner® Magic Quadrant™ for Conversational AI Platforms

Building the next generation of customer experiences with Gemini Enterprise for Customer ExperienceEnterprise customer experiences are entering a new era.

Today, Gemini Enterprise for Customer Experience brings these capabilities together to give your customers a frictionless experience.

Built for production AIAt the center of Gemini Enterprise for Customer Experience is CX Agent Studio, Google’s platform for building intelligent customer experience agents.

This is why Gemini Enterprise for Customer Experience and CX Agent Studio run natively on Google Cloud’s complete, first-party AI stack.

Gartner research publications consist of the opinions of Gartner's research organization and should …

1 week, 4 days назад @ cloud.google.com
Three lessons in accelerating foundation model upgrades
Three lessons in accelerating foundation model upgrades Three lessons in accelerating foundation model upgrades

For most engineering teams, upgrading to a new model checkpoint means months of manual toil to verify performance.

And the industry is moving at breakneck pace – since 2023, we’ve announced six major model evolutions, bringing us to Gemini 3.5 today.

As part of that, our team built an agentic workflow that completes model upgrades in hours instead of months.

Their goal was to migrate to the latest out-of-the-box foundation model, guided purely by prompt engineering.

Build an agentic loop: You can use the Agent Development Kit within Gemini Enterprise Agent Platform to create your agent.

1 week, 4 days назад @ cloud.google.com
Cloud CISO Perspectives: How AI leverages deep context as the defender’s advantage
Cloud CISO Perspectives: How AI leverages deep context as the defender’s advantage Cloud CISO Perspectives: How AI leverages deep context as the defender’s advantage

This ensures autonomy under human supervision, empowering engineering and security teams to eliminate backlogs and secure the software development lifecycle without sacrificing speed.

The key to countering this is enforcing Zero Trust for AI, and directing teams toward approved architectures with proper governance.

That means securing AI infrastructure requires building from the ground up, and not bolting on.

Fight AI with AI.

Learn more about how to secure your software lifecycle with Google AI Threat Defense.

1 week, 4 days назад @ cloud.google.com
Why AI apps fail in production (And how Google solved it)
Why AI apps fail in production (And how Google solved it) Why AI apps fail in production (And how Google solved it)

We are living in the golden age of the weekend AI side project.

Thanks to agentic engineering and LLMs, the time to go from a blank IDE to a functional local application has dropped from quarters to hours.

The data is sobering: only 5% of AI prototypes make it to production; the other 95% fall into the validation abyss.

For developers, watching people on social media ship lightning-fast AI deployments while you’re stuck in endless validation loops is maddening.

But as AI engineering leader Addy Osmani points out in our premiere of Emergent, unconstrained agentic orchestration inside an enterprise introduces an unpredictable blast radius.

1 week, 5 days назад @ cloud.google.com
IDC: Why the right networking approach is foundational to agentic AI
IDC: Why the right networking approach is foundational to agentic AI IDC: Why the right networking approach is foundational to agentic AI

Editor’s note: Today we hear from IDC on the results of its 2026 AI in Networking Special Report Survey exploring the enterprises' concerns about networking infrastructure to support the rise of agentic AI in their organizations.

Agentic AI specifically heightens these concerns by introducing more distributed and dynamic interactions across applications, services, APIs, tools, and data sources.

The right platform for agentic AI should be open, flexible, and able to evolve.

Businesses must meet their AI objectives while carefully managing dynamic agentic AI systems.

As agentic AI systems continue to evolve, the demands they place are unlikely to be addressed through best-of-breed point solut…

1 week, 5 days назад @ cloud.google.com
Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software
Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software

Gemini Enterprise: A unified system for the agentic eraA great foundation model is only as valuable as an organization's ability to safely put it to work.

Developers can build agents using Gemini 3.5 Flash on the Gemini Enterprise Agent Platform, or use it in your projects in Google AI Studio and Antigravity.

Get startedDownload the IDC MarketScape: Worldwide Foundation Model Software 2026 Vendor Assessment excerpt to learn why organizations are choosing Google Cloud.

Explore Gemini Enterprise today, or speak to your Google Cloud account representative to schedule a hands-on technical workshop.

IDC MarketScape: Worldwide Foundation Model Software 2026 Vendor Assessment, Doc #US54427726, Jul…

1 week, 6 days назад @ cloud.google.com
Claude at scale on Google Cloud: Frontier AI, built for enterprise production
Claude at scale on Google Cloud: Frontier AI, built for enterprise production Claude at scale on Google Cloud: Frontier AI, built for enterprise production

Running frontier AI in production is demanding — accelerators to manage, latency to hold steady across continents, regulated data to keep in-region, and long-context requests to serve reliably.

Claude on Google Cloud is built for exactly this.

In our case, Claude brings the reasoning, and Google Cloud brings the managed infrastructure, global reach, and compliance posture that enterprises already run on.

Managed infrastructure that frees engineering timeClaude on Google Cloud runs on fully managed infrastructure, so enterprise teams ship features instead of building clusters.

Invoking Claude is operationally identical to invoking any other Google Cloud service: the same IAM policies, the sa…

1 week, 6 days назад @ cloud.google.com
Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs
Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs

Furthermore, k8s-aibom treats the Kubernetes cluster state as a pure functional input: Identical cluster inputs produce byte-identical ML-BOM documents.

Commercial AI security platforms extend the picture with cloud-native posture management, but typically through external scanning shaped around vendor-specific data models.

Standard monitoring tools indicate that a container is running, but can’t prove whether an AI model was explicitly configured by a platform engineer or dynamically pulled by an autonomous script at runtime.

Getting startedIt’s rare that a technical solution like k8s-aibom can help mitigate the multi-faceted problem of shadow AI, impacting CISOs, governance, risk, and com…

2 weeks назад @ cloud.google.com
Frontier and Center: Who evaluates the evaluations?
Frontier and Center: Who evaluates the evaluations? Frontier and Center: Who evaluates the evaluations?

Yet this is exactly how we evaluate AI agents.

Along the way, the added fidelity exposed some deeper issues with the quality of emergent evaluation cases themselves.

Difficulty, measuredWhen it comes to retrieval, evaluation cases are often stratified into tiers of difficulty.

What we need is a rigorous approach that can modulate the difficulty of evaluation cases.

Therefore, we can adjust the difficulty of evaluation cases by adding or removing terms with varying levels of informative power.

2 weeks, 3 days назад @ cloud.google.com
Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud
Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud

Visit the blog to read more how BASF used AlphaEvolve to improve their existing planning and forecasting models by over 80%.

Visit the blog to read more about how Jetbrains used AlphaEvolve to improve their IDE performance by over 15-20%.

Kinaxis: Improving optimization and forecasting systems"Kinaxis researchers have used AlphaEvolve to materially improve both the speed and quality of highly mature forecasting and optimization algorithms.

Visit the blog to read more about how Klarna used AlphaEvolve to double Training Speed and improve performance for their foundational models.

Visit the blog to read more about how Schröedinger used AlphaEvolve to quadruple the speed of molecular discovery.

2 weeks, 4 days назад @ cloud.google.com
A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace
A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace

Here’s an overview of these architectural elements:Customer project: Where users discover agents via the dedicated Agent Marketplace category within Google Cloud Marketplace and interact with these agents through the Gemini Enterprise app.

Step 2: Review the organizational requirements to sell on MarketplaceJoin the Google Cloud Partner Network : If you're new to offering your solutions on Marketplace, join the Google Cloud Partner Network.

Nominate your agent for Google Cloud Marketplace by contacting your Google Cloud representative.

A2A Agent Card: Create an Agent Card, a JSON file declaring capabilities (skills), authentication methods, and service endpoints.

A2A agent cardTo list your …

2 weeks, 6 days назад @ cloud.google.com
Drive proactive security, prioritize risks with Google Threat Intelligence and Wiz ASM
Drive proactive security, prioritize risks with Google Threat Intelligence and Wiz ASM Drive proactive security, prioritize risks with Google Threat Intelligence and Wiz ASM

To help you be more proactive by matching your real-world exposures with real-time adversary activity, we’ve begun integration efforts between Google Threat Intelligence and Wiz Attack Surface Management (ASM).

By connecting exposure and validated exploitable risks directly to real-time threat intelligence, we can help you detect and prioritize external-facing exploitable issues and uncover logic-driven vulnerabilities with AI scanning at the speed needed for today’s defenses.

This allows you to shift to a strategy that prioritizes actions based on the real-world threats that pose the greatest risks to your organization.

Combining these two perspectives on threats can help you move from rea…

2 weeks, 6 days назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 2 weeks назад
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

2 weeks назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

2 weeks, 4 days назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

2 weeks, 5 days назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

3 weeks, 6 days назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

3 weeks, 6 days назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

1 month назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

1 month назад @ microsoft.com
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis

At a glance Talos is an open-source tool for automated, iterative reanalysis of genomic data in rare disease.

Deployed across a prospective cohort of almost 5,000 undiagnosed patients, Talos delivered 241 new diagnoses (5.1% additional yield).

On monthly iterative cycles, analysts only needed to review one new variant per 200 patients, demonstrating that frequent, systematic reanalysis can be run sustainably.

Why genome reanalysis mattersGenomic testing has transformed the diagnosis of rare disease, but even with this advancement, more than half of patients remain undiagnosed after their first test.

Looking aheadTalos reframes genomic reanalysis from a rare, labor-intensive event into a con…

1 month назад @ microsoft.com
Ire identifies another LOTUSLITE specimen
Ire identifies another LOTUSLITE specimen Ire identifies another LOTUSLITE specimen

At a glance Project Ire identifies a LOTUSLITE variant that shares TTPs (tools, tactics, procedures) with the public family but none of its indicators of compromise (IOC).

On Ire’s calibrationOne noteworthy observation in Ire’s report (opens in new tab) is worth highlighting first.

The Ire report does not surface a matching entry-point name, but it identifies that the behavioral shape is the same.

Ire never named LOTUSLITE in its report or chain of evidence.

Ire described the behavior precisely enough to make the mapping straightforward of this sample to LOTUSLITE.

1 month, 2 weeks назад @ microsoft.com
Data Formulator 0.7: AI-powered data analytics for enterprise data
Data Formulator 0.7: AI-powered data analytics for enterprise data Data Formulator 0.7: AI-powered data analytics for enterprise data

At a glance Data Formulator 0.7 is an open-source AI-powered system for enterprise data analytics that combines data connectivity, agent-guided exploration, and visualization refinement in a shared workspace.

Enterprise teams increasingly rely on AI systems for analytics, but enterprise data workflows are often fragmented across storage systems and tools.

Listen now Opens in a new tabConnecting enterprise data with Data ConnectorsData Formulator helps teams bring enterprise data into an AI-ready workspace without needing to rebuild the same connections for every source of data.

Data Connectors provide persistent connections between enterprise data sources and Data Formulator, allowing analy…

2 months назад @ microsoft.com
Extending Human Intelligence Through AI
Extending Human Intelligence Through AI Extending Human Intelligence Through AI

At a glance Modern AI systems are powerful not because they replicate human intelligence, but because they presuppose it, by extending structures already present in human cognition and language.

Understanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems.

Rather than asking whether AI systems are becoming intelligent in the human sense, these approaches ask a more basic question: What if AI systems work because they rely on structures that are rooted in human cognition?

In our recent paper, The Origins of Artificial Intelligence in Natural Intelligence, we argue that modern AI systems are best understood nei…

2 months назад @ microsoft.com
MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models

Built as the next generation of Magentic-UI, it combines a redesigned app with a harness optimized for small models.

MagenticBrain and Fara1.5 are small models designed for orchestration and computer-use tasks, respectively.

Together, these releases explore how far agentic performance can be pushed with smaller models, codesigned tools, and an optimized execution harness.

Today, Microsoft Research AI Frontiers releases MagenticLite (opens in new tab), an experimental agentic application designed for small models.

The result is an agent that runs efficiently, keeps data on the user’s machine, and supports a broad range of agentic tasks.

2 months, 1 week назад @ microsoft.com
Vega: Zero-knowledge proofs for digital identity in the age of AI
Vega: Zero-knowledge proofs for digital identity in the age of AI Vega: Zero-knowledge proofs for digital identity in the age of AI

Vega puts these building blocks together into a single proof system.

The hashing problem, and how folding solves itA credential proof must do two expensive things: hash the credential bytes with SHA-256 and verify the issuer’s digital signature.

Making it zero-knowledge, cheaplyA proof system needs to be zero-knowledge: the verifier should learn nothing beyond the claim being proved.

Device bindingA zero-knowledge credential proof is only useful if it is tied to the person holding the credential.

The proof system powering Vega is already available as the open-source spartan2 (opens in new tab) project on GitHub.

2 months, 1 week назад @ microsoft.com
Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability
Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability

Our recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated workflows.

Using a controlled evaluation methodology, we examine how well information is preserved across these extended workflows.

We use chained transformation-and-inversion tasks that evaluate whether semantic content is preserved accurately across extended delegated workflows.

Azure AI Foundry Labs Get a glimpse of potential future directions for AI, with these experimental technologies from Microsoft Research.

At the same time, the findings should not be interpreted as evidence that AI systems lack practical value in real-world work today.

2 months, 1 week назад @ microsoft.com
mimalloc: A new, high-performance, scalable memory allocator for the modern era
mimalloc: A new, high-performance, scalable memory allocator for the modern era mimalloc: A new, high-performance, scalable memory allocator for the modern era

mimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free.

The mimalloc memory allocator was initially designed in 2020 as a fast allocator for the state-of-the-art Lean (opens in new tab) and Koka (opens in new tab) programming languages developed at RiSE, both of which use novel compiler-guided reference counting (see Perceus).

ja .LBB0_generic leaq 7 ( %rsi ), %rax ; round to sizeof(void*) andq $-8 , %rax movq 232 ( %rdi , %rax ), %rcx ; rcx = heap->small_pages[index] movq 8 ( %rcx ), %rax ; block = rax = page->free testq %rax , %rax ; block == NULL?

Thus, mimalloc has three free lists per (64 KiB) mimalloc page, and effectively that …

2 months, 2 weeks назад @ microsoft.com
MIT AI MIT AI
последний пост 3 days, 17 hours назад
Working to automate nuclear plant operations
Working to automate nuclear plant operations Working to automate nuclear plant operations

In pursuit of autonomous nuclear plant operationsIt turns out the research for the master’s was just the tip of the iceberg.

For the future viability of nuclear power, small plants, located in rural areas, are a distinct possibility.

It’s where supervised and thoroughly vetted autonomous operations will help.

A primary question was: “How do we transition to autonomous operations in nuclear power plants?” Fortier wanted one integrated approach, a central supervisory control system instead of many interlinked parts.

Using the nuclear plant automation program on next-generation equipment will deliver necessary traction in developing and deploying commercial microreactors.

3 days, 17 hours назад @ news.mit.edu
MIT projects selected for funding under US Department of Energy’s Genesis Mission
MIT projects selected for funding under US Department of Energy’s Genesis Mission MIT projects selected for funding under US Department of Energy’s Genesis Mission

MIT researchers are set to contribute to the U.S. Department of Energy’s (DOE) Genesis Mission, with 15 collaborative projects among those selected for funding under Genesis Phase I, DOE announced Wednesday.

“MIT researchers are proud to be leading and contributing to projects under the Genesis Mission, in vital areas of research that support national priorities,” says Ian A. Waitz, MIT’s vice president for research.

Projects under the Genesis Mission are collaborative by design; teams must draw on the expertise of researchers from academia, industry, and/or the national laboratories.

Phase I projects that identify promising pathways toward transformative capabilities at scale may be consid…

4 days, 9 hours назад @ news.mit.edu
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83 Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83

Over the course of his career, Bertsekas’ research spanned, and had a definitive influence upon, several fields, including optimization, control, large-scale computation, reinforcement learning, and artificial intelligence.

Along the way, Bertsekas taught, advised, and mentored students who would eventually become his colleagues at all four institutions.

“Dimitri played a defining role in my career,” says Asu Ozdaglar, department head of EECS at MIT.

Some referred to Dimitri as an “immortal.” Another comment I recall fondly — and often reminded Dimitri about — was: “Professor Bertsekas is a very handsome man!” Their bond continued long after Van Roy’s graduation.

He is survived by his wife …

5 days, 4 hours назад @ news.mit.edu
Following the questions where they lead
Following the questions where they lead Following the questions where they lead

Ever since she was a child playing on her family’s farmland in Wisconsin, Bailey Flanigan was guided by her own selective, yet wide-ranging, curiosity.

“I found myself unmotivated to take all the AP [advanced placement] classes for the sake of it.

So Flanigan moved toward public health, where she researched microfluidic devices for HIV detection that could be used in low-resource settings.

After graduating from UW-Madison, Flanigan worked as a predoctoral research assistant in economics at Princeton.

“I feel so lucky to be studying these questions from within both political science and EECS, because I have the freedom to explore both the political and technical substance of tools for more d…

1 week, 3 days назад @ news.mit.edu
A better way to turn 2D designs into 3D models for rapid prototyping
A better way to turn 2D designs into 3D models for rapid prototyping A better way to turn 2D designs into 3D models for rapid prototyping

The system generates new data based on the model’s abilities as it attempts to convert a 2D image into a CAD program.

“Nearly every physical product around us, from airplanes to appliances, begins its life as a CAD model.

For guesses that are nearly correct, GIFT adjusts them to become successful solutions.

The CAD models generated by VLMs using GIFT were better aligned with the shapes of ground-truth models.

In the future, the researchers want to expand GIFT so the framework can teach models to generate CAD programs that improve the performance and manufacturability of 3D models.

1 week, 4 days назад @ news.mit.edu
3 Questions: Neural transparency and the future of AI design
3 Questions: Neural transparency and the future of AI design 3 Questions: Neural transparency and the future of AI design

Q: Your paper introduces “neural transparency,” a way to let everyday users peek inside an AI’s neural networks before their chatbot ever says a word.

“Neural transparency” means giving people something like a brain scan for AI.

Our study suggests that people have a blind spot when designing personalized AI.

In previous research, we documented cases of psychological harm associated with interactions with AI chatbots.

AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step.

1 week, 5 days назад @ news.mit.edu
Helping AI models to meet the real world
Helping AI models to meet the real world Helping AI models to meet the real world

“In a sense, with a small amount of resource, you have to do a lot of heavy lifting,” he says.

“My interest was: How does one design such graphical models for generic, tabular data?” he says.

And each of the products that you manufacture has lots of small pieces that come from different parts of the world.

Shah adds that Celonis has specialized in digitizing and automating operations for more than 1,400 large companies around the world.

“A narrower focus comes with sharper technology,” he says, “but it’s broad enough that it’s very valuable.”Shah adds, “The recent buzzword that’s become pertinent in the modern AI popular press is a ‘world model.’ In a sense, this is trying to build the ente…

1 week, 6 days назад @ news.mit.edu
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering

“The JARVIS challenge showed that AI can substantially accelerate safety-critical hardware engineering, but engineering judgment remains the decisive differentiator.

Manufacturing — not engineering design or analysis — remained the fundamental rate-limiting step,” says Professor Zolti Spakovszky, director of the MIT Gas Turbine Laboratory.

In weekly progress reviews, they would critically evaluate the student progress and assess how the students were using AI.

The 811 team had been resistant to using AI throughout the competition, trusting instead to their fundamentals and teamwork.

From the start of the JARVIS Challenge, younger students used Parley more frequently and cleverly, while the …

1 week, 6 days назад @ news.mit.edu
How MIT students are helping to prevent cyberattacks
How MIT students are helping to prevent cyberattacks How MIT students are helping to prevent cyberattacks

To counter such threats, Lecturer Jungwoo Chun and Ford Professor of Urban and Environmental Planning Lawrence Susskind launched the MIT Cybersecurity Clinic in 2019.

Much like a legal or medical clinic, the course doubles as hands-on training for students and a pro-bono service to at-risk communities.

After completing instructional modules and passing a certification exam, students are assigned in teams to a client.

The Cybersecurity Clinic aims to round out the knowledge of students from every discipline.

In either case, Susskind and Chun check in periodically with clients for at least two years following each engagement.

2 weeks назад @ news.mit.edu
AI agents create virtual playgrounds to help robots get crucial training data
AI agents create virtual playgrounds to help robots get crucial training data AI agents create virtual playgrounds to help robots get crucial training data

It turns out that AI agents, or semi-autonomous programs that “think” and complete well-defined tasks, could help produce the lifelike virtual settings that robots need.

The most telling test: they dropped a pretrained robot policy — an AI controller trained largely on real-world data, which had never seen a SceneSmith scene — into the generated environments.

The team also teleoperated robots through the virtual spaces, guiding them to open cabinets, put away bottles, and navigate between rooms.

Behind the scenesThe agents that SceneSmith uses each have a well-defined role in the generative process, fleshing out scenes in stages.

It can take multiple hours to produce a single scene because …

2 weeks назад @ news.mit.edu
New method aims to keep kids safe from illegal AI-generated content
New method aims to keep kids safe from illegal AI-generated content New method aims to keep kids safe from illegal AI-generated content

When tested, the auditing procedure identified model variations that had been specialized to generate CSAM with 100 percent accuracy.

Auditing adaptationsRecent techniques have made it easier for users to specialize a generative AI model for their task through a process known as fine-tuning.

This has led to a wave of new generative AI model variants for a variety of purposes, like producing watercolor images that mimic an artistic movement.

They tested their method on variations of three types of models, comparing the results to ground-truth data from LoRA adaptors known for generating CSAM, other harmful images, and safe content.

Their method was 100 percent accurate in identifying models …

2 weeks назад @ news.mit.edu
Tiny robot boats build floating structures
Tiny robot boats build floating structures Tiny robot boats build floating structures

A team of MIT researchers sees it as a dynamic, Lego-like construction site.

Each robot, about the size of a dinner plate at 21 centimeters square, is a self-contained vessel with its own thrusters, sensors, and magnetic latches.

In that final mode, called collective transport, a planner charts a trajectory for the whole structure and each robot computes its own contribution.

“Our boats become more stable by joining together, like the ant raft, if you have waves or currents,” Hagemann says.

The team thanks MIT Sea Grant and Professor Michael Triantafyllou for providing the test tank.

2 weeks, 4 days назад @ news.mit.edu
How novice coders can develop AI programs for military applications
How novice coders can develop AI programs for military applications How novice coders can develop AI programs for military applications

We both wanted to understand better where and how AI could be used by nontechnical users in the military."

During the project, Lynch completed several professional development courses in AI and familiarized himself with both military and nonmilitary uses of the technology.

For the basis for his code generation, he used the paid models of three AI chatbots: Anthropic's Claude, OpenAI's ChatGPT, and Google's Gemini.

For example, he often encountered difficulties with the AI chatbots lacking hierarchical focus and modifying unrelated code sections.

Although AI can generate significant amounts of functional code, code review remains a bottleneck in this space.

2 weeks, 6 days назад @ news.mit.edu
Jesse Thaler named director of the Laboratory for Nuclear Science
Jesse Thaler named director of the Laboratory for Nuclear Science Jesse Thaler named director of the Laboratory for Nuclear Science

Professor Jesse Thaler has been named director of the MIT Laboratory for Nuclear Science (LNS), effective Aug. 1.

Thaler is a theoretical particle physicist who combines techniques from quantum field theory and machine learning to address outstanding questions in fundamental physics.

Mike Williams, professor of physics, will succeed Thaler as IAIFI director.

Established in 1946 to support nuclear and particle physics, LNS now encompasses research spanning cosmology, gravity, field theory, and quantum information science.

As head of LNS, Thaler will also oversee his home center of CTP-LI, which last year received a donation from the Leinweber Foundation to establish a network of theoretical …

2 weeks, 6 days назад @ news.mit.edu
Toward a future that preserves benefits of neurotechnology for all
Toward a future that preserves benefits of neurotechnology for all Toward a future that preserves benefits of neurotechnology for all

Sava’s concept was inspired by an internship at IBM, where she worked on a project with the PACE Center in London.

As advanced medical technology gets closer to hitting consumer markets, the need for guardrails on protected usage should increase.

What might begin as a neural implant to aid in communication could become a device used to police one’s innermost thoughts.

From its inception, the competition has consistently attracted undergraduate and graduate students from across a wide range of disciplines.

The judges also named four honorable mentions, each of whom received a $500 cash prize.

3 weeks назад @ news.mit.edu
Berkeley AI
последний пост 1 day, 12 hours назад
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon InteractionOverview of ABBEL compared to traditional recursive summarization.

Belief grading can be thought of as adding an auxiliary RL task, using heuristics designed to capture what makes a good belief as the reward.

With domain-knowledge belief grading, ABBEL approaches or exceeds FULL CTX in this setting; without belief grading, learning is slower.

Context compression methods generate dense representations which, while computationally efficient, sacrifice human-understandability (Kontonis et al., 2026, Eyuboglu et al., 2025, Gupta et al., 2025, Chevalier et al., 2023, Deng et al., 2025, Deng et al., 2025, Bulatov et al., 2022).…

1 day, 12 hours назад @ bair.berkeley.edu
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

2 weeks, 6 days назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

3 weeks, 5 days назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

2 months, 2 weeks назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

3 months, 1 week назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

4 months, 2 weeks назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

6 months, 2 weeks назад @ bair.berkeley.edu
RL without TD learning
RL without TD learning RL without TD learning

RL without TD learningIn this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer.

We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning.

There are two classes of algorithms in RL: on-policy RL and off-policy RL.

We compared TRL with $n$-step TD learning with different values of $n$, from $1$ (pure TD) to $\infty$ (pure MC).

I still think one of the most important problems in RL (and even in machine learning) is to find a scalable off-policy RL algorithm.

8 months, 4 weeks назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 4 часа назад
Beyond RAG: Task-aware knowledge compression for enterprise AI on AWS
Beyond RAG: Task-aware knowledge compression for enterprise AI on AWS Beyond RAG: Task-aware knowledge compression for enterprise AI on AWS

This post shows you how to address that gap using task-aware knowledge compression (TAKC), a technique that pre-compresses entire knowledge bases into task-specific representations deployed on AWS.

Straightforward factual questions hit the ultra-compressed cache while complex analytical questions use the lightly compressed cache.

Amazon ElastiCache Serverless stores the compressed cache, while Amazon S3 holds raw data, chunks, and cache backups (KMS-encrypted).

Define the infrastructure as a single AWS Cloud Development Kit (AWS CDK) stack and deploy it with a single command.

AWS CDK CLI.

4 часа назад @ aws.amazon.com
Deepgram enhances Amazon SageMaker AI support with AWS IAM Temporary Delegation
Deepgram enhances Amazon SageMaker AI support with AWS IAM Temporary Delegation Deepgram enhances Amazon SageMaker AI support with AWS IAM Temporary Delegation

Amazon SageMaker AI delivers a managed deployment option for self-hosted Deepgram speech AI models.

With this integration, Deepgram has reduced the time for initial investigation on a SageMaker AI support ticket from days to minutes.

An active Deepgram support contract or account that includes access to the Deepgram support ticketing system.

Delete the SageMaker AI endpoint: aws sagemaker delete-endpoint --endpoint-name .. Delete the endpoint configuration: aws sagemaker delete-endpoint-config --endpoint-config-name .. Delete the model: aws sagemaker delete-model --model-name .. Delete the associated Amazon CloudWatch log group to stop log storage charges: aws logs delete-log-group --log…

4 часа назад @ aws.amazon.com
How Guardoc transforms medical document processing with Amazon Nova models
How Guardoc transforms medical document processing with Amazon Nova models How Guardoc transforms medical document processing with Amazon Nova models

The solution: Amazon Nova models on Amazon BedrockGuardoc Health built its pipeline on the Amazon Nova family of models and Amazon Bedrock, combining several AWS services to handle each stage of document processing.

Amazon Nova 2 Lite is a low-cost, multimodal foundation model (FM) in the Amazon Nova family, designed for fast, affordable processing of text, images, and video inputs.

Amazon Textract processes high-volume, structured extraction while the Amazon Nova family of models applies multimodal reasoning to handle the more complex cases.

By combining Amazon Textract, Amazon Titan Text Embeddings V2, Amazon DynamoDB, and the Amazon Nova family of models available through Amazon Bedrock,…

5 часов назад @ aws.amazon.com
Introducing Claude Opus 5 on AWS: Anthropic’s most capable Opus model
Introducing Claude Opus 5 on AWS: Anthropic’s most capable Opus model Introducing Claude Opus 5 on AWS: Anthropic’s most capable Opus model

Claude Opus 5 is Anthropic’s most advanced Opus model and the first in the fifth generation.

According to Anthropic, Claude Opus 5 matches Claude Fable 5’s top-tier intelligence in many domains at Opus-tier pricing.

What makes Claude Opus 5 differentAccording to Anthropic, Claude Opus 5 delivers a step-change in coding.

Claude Opus 5 also improves on Claude Opus 4.8’s cyber capabilities across the board, from coding to cyber security.

Getting started with Claude Opus 5 on Amazon BedrockYou can get started with Claude Opus 5 in the Amazon Bedrock console.

3 days, 3 hours назад @ aws.amazon.com
Build an explainable next-best-product recommendation system for banking on AWS
Build an explainable next-best-product recommendation system for banking on AWS Build an explainable next-best-product recommendation system for banking on AWS

Building a deep learning-based explainable next-best-product recommendation system helps banking institutions predict which product a customer needs next.

Traditional rule-based systems and collaborative filtering approaches often fail to capture the complex temporal patterns in customer product adoption journeys.

In this post, we present the architecture and design decisions behind a Next-Best-Product (NBP) recommendation system using Amazon SageMaker AI and PyTorch.

For guidance on writing least-privilege IAM policies for SageMaker AI, see Identity-based policy examples for SageMaker AI.

ConclusionThis post showed how to build a Next-Best-Product recommendation system for banking using Py…

3 days, 5 hours назад @ aws.amazon.com
Get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock
Get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock Get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock

OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock.

To run an existing OpenAI SDK application on Amazon Bedrock, replace the OpenAI base URL with the bedrock-mantle endpoint, use the corresponding Amazon Bedrock model ID, and authenticate with an Amazon Bedrock API key or AWS credentials.

GPT-5.6 Sol, Terra, and Luna are third-party models from OpenAI, made available on Amazon Bedrock and subject to the OpenAI terms.

Get started with GPT-5.6 on Amazon BedrockComplete the following steps to start using GPT-5.6 on Amazon Bedrock.

ConclusionIn this post, we showed how to get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock.

3 days, 5 hours назад @ aws.amazon.com
Best practices for applying Amazon Bedrock Guardrails to code generation workflows
Best practices for applying Amazon Bedrock Guardrails to code generation workflows Best practices for applying Amazon Bedrock Guardrails to code generation workflows

For the previous post, see Build safe generative AI applications like a pro: best practices with Amazon Bedrock Guardrails.

In this post, we explain how Amazon Bedrock Guardrails can be configured for code generation workflows with coding assistants to overcome these constraints.

Best Practices: Architecture Patterns for Code Generation WorkflowsTo address these challenges, we recommend a set of architecture patterns that optimize guardrail usage for code generation workflows.

High-risk code: Full evaluation with all safeguards Standard code: Lightweight evaluation (sensitive info only) Low-risk code: Skip intermediate evaluation; validate at commit only """ # Patterns that indicate high-ri…

3 days, 22 hours назад @ aws.amazon.com
Evaluating AI Agents: A production blueprint with Strands and AgentCore
Evaluating AI Agents: A production blueprint with Strands and AgentCore Evaluating AI Agents: A production blueprint with Strands and AgentCore

The pipeline combines the Strands Agents SDK with Amazon Bedrock AgentCore, a fully managed service for deploying and operating AI agents at scale.

spanning build-time testing with strands-agents-evals (the open source evaluation library for Strands Agents) and production monitoring with Amazon Bedrock AgentCore Evaluations.

The worked example: A dealer stock search agentMotorway built the dealer stock search agent on the Strands Agents SDK and Amazon Bedrock AgentCore.

The HelpfulnessEvaluator and TrajectoryEvaluator from strands-agents-evals use LLM-as-judge scoring to assess whether the agent’s reasoning holds together.

Monitor production behavior with AgentCore EvaluationsAfter you depl…

4 days, 4 hours назад @ aws.amazon.com
Building trade assistant: How Jefferies optimized front office trading operations with AI
Building trade assistant: How Jefferies optimized front office trading operations with AI Building trade assistant: How Jefferies optimized front office trading operations with AI

Solution overviewThe Front Office trade assistant represents a shift in how Capital Markets’ Front Office Equity traders interact with data.

The user interaction begins with the Jefferies front-end trading interface, which now includes an embedded AI assistant widget.

Once a trader logs into GFM, they have a UI widget to interact with the Trade Assistant Agent.

The team also chose Amazon Bedrock for its flexibility to choose different LLMs as the trade assistant evolves.

Lessons learnedThroughout the journey to deliver the Front Office trade assistant agent, the Jefferies team uncovered several key lessons that shaped its enterprise AI strategy and implementation roadmap.

4 days, 4 hours назад @ aws.amazon.com
Building multi-Region visualizations with Highcharts in Amazon Quick
Building multi-Region visualizations with Highcharts in Amazon Quick Building multi-Region visualizations with Highcharts in Amazon Quick

When your carrier performance data spans multiple regions, your dashboard must reconcile fundamentally different competitive structures within a single view.

This post shows you how to build multi-Region carrier performance dashboards in Quick Sight using Highcharts custom visualizations to overcome native chart limitations.

Solution architectureThe following diagram shows how the multi-Region architecture federates carrier performance data while maintaining data sovereignty.

US carriers (Carriers 1–3) and UK carriers (Carriers 4–7) appear with different color schemes, making it easy to compare cross-region performance at a glance.

ConclusionIn this post, you learned how to extend Amazon Qu…

4 days, 4 hours назад @ aws.amazon.com
Detecting silent agent failures with Amazon Bedrock AgentCore optimization
Detecting silent agent failures with Amazon Bedrock AgentCore optimization Detecting silent agent failures with Amazon Bedrock AgentCore optimization

Amazon Bedrock AgentCore optimization provides insights that help you discover, explain, and prioritize behavioral failures in your deployed AI agents, including the silent ones that never generate an error signal.

What insights in Amazon Bedrock AgentCore optimization deliversInsights operate one layer above your existing observability stack.

Each session analysis method extracts one or more attributes from each session.

One aggregate explanation per cluster covering affected sessions, specific enough to act on without opening individual traces.

Enable insights in Amazon Bedrock AgentCore optimization for your agent and see what your dashboards are missing.

4 days, 4 hours назад @ aws.amazon.com
Agentic retrieval for Amazon Bedrock Managed Knowledge Base
Agentic retrieval for Amazon Bedrock Managed Knowledge Base Agentic retrieval for Amazon Bedrock Managed Knowledge Base

Agentic retrieval for Amazon Bedrock Managed Knowledge Bases is designed for these questions.

Agentic retrieval plans and iterates over retrieval and can generate a response in the same call.

What agentic retrieval isAgentic retrieval is a new retrieval mode in Amazon Bedrock Managed Knowledge Bases, available through the AgenticRetrieveStream API.

Question Complexity Δ (Gain) with Agentic Retrieve Agentic Hops (avg) 2-hop questions +22.8 1.98 3-hop questions +31.9 3.50 4-hop questions +37.3 4.78Two patterns stand out.

Guardrails are configured through policyConfiguration.bedrockGuardrailConfiguration , and only the BLOCK action is supported; the MASK action is not supported with agentic re…

4 days, 4 hours назад @ aws.amazon.com
AI Teammates: how monday.com runs production AI agents on Amazon Bedrock
AI Teammates: how monday.com runs production AI agents on Amazon Bedrock AI Teammates: how monday.com runs production AI agents on Amazon Bedrock

AI Teammates are agentic AI on Amazon Bedrock, and few engineering organizations run them in production at the scale that monday.com does.

The architecture in one diagramThe seven AWS services that we used are Amazon Simple Notification Service (Amazon SNS), Amazon Simple Queue Service (Amazon SQS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Relational Database Service (Amazon RDS), Amazon ElastiCache, Amazon Elastic File System (Amazon EFS), and Amazon Simple Storage Service (Amazon S3).

Amazon Bedrock as the model fabricAmazon Bedrock is more than a place to get tokens.

Talk to us: the monday AI Engineering team, or the Amazon Bedrock team, will be happy to share our insights.…

5 days, 5 hours назад @ aws.amazon.com
Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova
Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

We first examine the reasoning suppression problem, then introduce Self-Distilled Reasoning (SDR), validate it across three benchmarks, and provide practical recommendations.

Self-distilled reasoningWe propose an effective approach that uses self-distilled reasoning (SDR) for SFT customization on datasets without reasoning traces.

Constructing reasoning traces for SFTWe query Amazon Bedrock to obtain reasoning traces for a given dataset.

Training Reasoning specifies whether reasoning mode was enabled during the SFT training process.

4 Dataset has < 50% reasoning traces If latency constraints allow for reasoning during inference, pre-fill missing traces with Nova 2 Lite reasoning (one-time, …

6 days, 4 hours назад @ aws.amazon.com
Custom OS installation now available on AWS DeepRacer devices
Custom OS installation now available on AWS DeepRacer devices Custom OS installation now available on AWS DeepRacer devices

AWS DeepRacer devices are fully autonomous 1/18th scale race cars driven by models trained through reinforcement learning.

As shipped, secure firmware on the AWS DeepRacer devices boots AWS-signed operating systems, including versions of Ubuntu 16.04 and 20.04.

It’s particularly valuable for community members who want to create and distribute custom AWS DeepRacer installation distributions and media.

On bootup, the developer shim blinks “DEVELOPER MODE” in Morse code on the AWS DeepRacer device’s built-in lights.

Option C: Custom OSIf you’re building a custom OS for the AWS DeepRacer device, create a boot volume on the device.

1 week назад @ aws.amazon.com
NVIDIA
последний пост 5 часов назад
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

Ising Calibration 1.5 also uses examples from related experiments when available and is 11.4% smaller at BF16 precision.

How is the Ising Calibration 1.5 model trained?

Get started with NVIDIA Ising open resourcesThe NVIDIA Ising model family is fully open.

Model weightsFull-parameter checkpoints for Ising Calibration 1.5 are available on Hugging Face:Ising Calibration 1.5 is also available as an NVIDIA NIM and hosted through NVIDIA Build.

Quantum calibration agent blueprint is a script for deploying an agentic workflow using Ising Calibration 1.5 with the NVIDIA Nemo Agent Toolkit to quickly set up quantum calibration experiment automation.

5 часов назад @ developer.nvidia.com
Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

The Open Secure AI Alliance — building on the leadership of the Linux Foundation’s Akrites initiative and OpenSSF community work — will work to remediate and disclose vulnerabilities using open technologies.

That is the mission of the Open Secure AI Alliance: to ensure defenders everywhere have open, frontier tools they can trust and control.

NVIDIA is contributing open models, model weights, data and new agent harness research to the Open Secure AI Alliance to speed the development of new cybersecurity tools and techniques.

That future is worth building — and the Open Secure AI Alliance invites governments, industry and researchers to join in the work of defending the AI era together.

Lear…

12 часов назад @ blogs.nvidia.com
NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs
NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs

The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems.

To help meet that challenge, NVIDIA is collaborating with industry leaders Cadence and Synopsys to optimize critical electronic design automation (EDA) applications for the NVIDIA Vera CPU.

While GPUs and AI have accelerated many aspects of chip design, several critical EDA workloads remain heavily dependent on CPU performance.

The results highlight Vera’s ability to accelerate two of the most compute-intensive stages of modern chip design.

Bringing Vera to the Design ProcessNVIDIA is deploying Vera throughout the EDA workflows used to create futu…

20 часов назад @ blogs.nvidia.com
At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners
At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners

At this week’s AI Summit in San Francisco, South Korean President Jae Myung Lee and some of the country’s top business leaders and researchers are meeting with NVIDIA and ecosystem partners to chart Korea’s AI progress.

To start, NVIDIA and the Korea Advanced Institute of Science and Technology (KAIST) today announced a joint AI research lab at the KAIST Kim Jaechul Graduate School of AI in Seoul, dedicated to advancing agentic AI for South Korea.

The collaboration will establish a robust academic AI research program, bringing together NVIDIA full-stack AI expertise, NVIDIA Nemotron open models and NVIDIA AI Cloud partner computing with the world-class scientific talent at KAIST, one of Asi…

3 days, 16 hours назад @ blogs.nvidia.com
GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud
GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud

Set sail in Path of Exile: Curse of the Allflame and charge in Battlefield 6 Season 4 both launching major content for members this week.

Then revisit Capcom legends like Breath of Fire IV, Dino Crisis and Dino Crisis 2, jump into Halo: Campaign Evolved Advanced Access and discover nine titles arriving on GeForce NOW.

Path of Exile: Curse of the Allflame launches Friday, July 24, sending Exiles into the perilous Frozen Seas.

Raw instinct takes over in Dino Crisis and Dino Crisis 2, the survival-horror series that combines pulse-pounding action, resource management and thrilling dinosaur encounters.

One GeForce NOW member recently shared they found themselves back on GeForce NOW Ultimate bec…

4 days, 8 hours назад @ blogs.nvidia.com
NVIDIA AI Supercomputer Comes Online at Naval Postgraduate School
NVIDIA AI Supercomputer Comes Online at Naval Postgraduate School NVIDIA AI Supercomputer Comes Online at Naval Postgraduate School

NVIDIA founder and CEO Jensen Huang today visited the Naval Postgraduate School in Monterey, California, to commission an NVIDIA DGX GB300 system — bringing one of the world’s most powerful AI platforms fully online for the students, researchers and faculty at the U.S. military’s flagship graduate university.

It’s built based on an NVIDIA AI Technology Center on the Monterey campus — a dedicated hub for AI research and graduate instruction, now anchored by the DGX GB300.

“The most important advice that I would give to someone is to engage the technology,” Huang said.

To learn more, watch NPS and NVIDIA researchers present on AI modeling and simulation at NVIDIA GTC Washington, D.C.NVIDIA Pa…

4 days, 19 hours назад @ blogs.nvidia.com
NVIDIA Open Sources First GPU-Accelerated Medical Physics Simulation Framework
NVIDIA Open Sources First GPU-Accelerated Medical Physics Simulation Framework NVIDIA Open Sources First GPU-Accelerated Medical Physics Simulation Framework

NVIDIA Medical Physics Simulation framework — a new open source, GPU-accelerated capability within NVIDIA Isaac for Healthcare — announced today, helps medical robotics developers model anatomy-device interaction, generate hard-to-capture scenarios, test in silico, and train or evaluate robot policies before hardware-heavy testing.

Medical Physics Simulation brings together classical physics simulation and generative AI physics simulation.

NVIDIA Cosmos-H Dreams, the real-time generative AI physics simulation capability within Medical Physics Simulation, helps model visual scene dynamics learned from procedural data.

Medtronic Structural Heart is exploring applying Medical Physics Simulatio…

5 days, 8 hours назад @ blogs.nvidia.com
Built in Fort Worth: Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems
Built in Fort Worth: Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems Built in Fort Worth: Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems

The AI era runs on AI infrastructure.

Wistron opened its first U.S. manufacturing facility today in Fort Worth — a 324,000-square-foot greenfield plant producing superchips at the heart of some of the world’s most capable AI systems.

The Fort Worth plant is where that capacity takes shape.

The Wistron Fort Worth factory is one of the projects making that number real.

A new chapter of American manufacturing is taking shape in Fort Worth.

5 days, 22 hours назад @ blogs.nvidia.com
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

NVIDIA Vera Rubin is here, and it’s going gigascale.

Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure.

The Vera Rubin platform is built from chip to grid to deliver the highest performance per watt and the lowest token cost.

Tuesday, July 21, 8:00 a.m. PT 🔗NVIDIA Vera Rubin Powers Europe’s Open-Model EraVera Rubin is delivering next-generation performance to Europe’s AI infrastructure.

Tuesday, July 21, 8:00 a.m. PT 🔗NVIDIA Vera Rubin NVL72 on CoreWeave Demonstrates 10x More Tokens Per Megawatt Than Blackwell in BenchmarkEmbracing extreme co-design with NVIDIA, Coreweave is delivering an order o…

6 days, 5 hours назад @ blogs.nvidia.com
Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories
Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories

Marking a networking milestone, NVIDIA Spectrum-6 — a 102.4-terabit-per-second Ethernet switch system delivering 2x the capacity of previous-generation systems and built as part of the NVIDIA Vera Rubin platform — is arriving across the world’s gigascale AI factories.

Spectrum-6 anchors the next generation of the NVIDIA Spectrum-X Ethernet platform, delivering the bandwidth, scale and intelligence needed to operate an AI factory as one end-to-end computing system.

AI Performance Is a Network ProblemPeak GPU performance alone no longer predicts the performance of an AI factory.

Unlike off-the-shelf Ethernet, the NVIDIA Spectrum-X Ethernet networking platform for AI factory scale-out combines…

6 days, 6 hours назад @ blogs.nvidia.com
At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI

The NVIDIA keynote, taking place today, July 20, at 3:45 p.m. PT, will feature NVIDIA AI research and engineering leaders Neil Ashton, Edward Liu and Ming-Yu Liu discussing neural rendering techniques, world models and simulation methods for AI, built by AI.

At SIGGRAPH, NVIDIA announced the Synthetic Video Detector NVIDIA NIM microservice, part of the NVIDIA AI for Media platform, to bring an AI-assisted detection signal into editorial and media workflows.

The 4-billion-parameter omnimodel is optimized for memory-efficient deployment and high throughput on NVIDIA Jetson, NVIDIA RTX PRO and NVIDIA DGX systems, as well as GeForce RTX GPUs.

AI Agents Made Easy: Build and Run Personal AI Agent…

1 week назад @ blogs.nvidia.com
Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin
Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin

BMS announced today it is deploying its second NVIDIA DGX SuperPOD, this one built on eight DGX Vera Rubin NVL72 systems — the most powerful and energy-efficient AI cluster in life sciences.

It will give researchers at the global pharmaceutical giant access to a unified AI platform — including NVIDIA BioNeMo Agent Toolkit for biological AI — for running predictions, training models and powering agentic workflows across the full drug discovery pipeline.

AI is also applied in lead optimization stages of drug discovery using a methodology Sheth calls “Predict First,” which informs experimental gating based on design predictions.

These research AI applications have significant impact on compute…

1 week назад @ blogs.nvidia.com
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI

Unlike a generative model responding to a prompt, an agentic model must plan, use different tools and recover from problems it encounters mid-run.

Agentic AI introduces a new compute pattern for post-training, making it the central workload of the agentic era and the primary driver of intelligence per dollar.

That means that every improvement to cost per token flows directly into intelligence per dollar.

Agentic Post-Training DemystifiedPost-training is where intelligence is built.

NVIDIA NeMo open libraries, such as NeMo Gym for training environments and NeMo RL for distributed post-training, turn post-training from bespoke research code into repeatable infrastructure.

1 week, 3 days назад @ blogs.nvidia.com
Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW
Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW

Onimusha: Way of the Sword is coming to GeForce NOW at launch, with the playable demo available this week.

Plus, GeForce NOW officially launches in India, moving from beta to public availability — meaning gamers can sign up without a waitlist.

Onimusha: Way of the Sword will arrive on the cloud at launch on Thursday, Sept. 3.

GeForce NOW now supports UPI payments, giving gamers across India a fast, secure and convenient way to purchase memberships and day passes.

— the off-the-rails, train‑wrecking action game — is now available to stream on GeForce NOW, letting members embrace pure chaos from almost any device.

1 week, 4 days назад @ blogs.nvidia.com
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI

To meet that need, NVIDIA today introduced the T3000 and T2000, new modules based on the NVIDIA Thor architecture that enable mass-market robotics and edge AI applications at scale.

Going Wide on Edge AI With T2000The Jetson T2000 brings Thor architecture to a broader range of edge AI systems.

With the introduction of the new NVIDIA Jetson modules, NVIDIA now offers a scalable edge AI platform spanning performance from 70 TOPS to 2,000 teraflops, enabling developers to address virtually any edge AI workload.

These skills support the entire Jetson portfolio, including Jetson Thor and Jetson Orin, enabling developers to run more capable workloads on lower-memory configurations.

Delivering Cos…

1 week, 4 days назад @ blogs.nvidia.com
Facebook
последний пост 1 week, 5 days назад
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

1 week, 5 days назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

2 weeks назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

3 weeks, 6 days назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

1 month назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

2 months назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

2 months, 2 weeks назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

3 months, 1 week назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

3 months, 1 week назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

3 months, 3 weeks назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

3 months, 3 weeks назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

3 months, 4 weeks назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

3 months, 4 weeks назад @ engineering.fb.com
Friend Bubbles: Enhancing Social Discovery on Facebook Reels
Friend Bubbles: Enhancing Social Discovery on Facebook Reels Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend bubbles in Facebook Reels highlight Reels your friends have liked or reacted to, helping you discover new content and making it easier to connect over shared interests.

Friend bubbles enhance the social experience on Facebook Reels by helping you discover content your friends enjoy, creating a shared viewing experience and sparking new conversations.

Along with additional optimizations in the underlying method, this approach enabled us to ship friend bubbles while preserving core Reels performance.

Friend bubbles work because the signal is high value: It adds meaningful social context that helps people decide what’s worth watching.

Engagement also scales consistently with the number …

4 months, 1 week назад @ engineering.fb.com
Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation
Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation

Meta’s Ranking Engineer Agent (REA) autonomously executes key steps across the end-to-end machine learning (ML) lifecycle for ads ranking models.

Powering these interactions are highly sophisticated, complex and massively distributed machine learning (ML) models that continuously evolve to serve both advertisers and people who use the platforms.

Optimizing these ML models has traditionally been time-consuming.

To address this, Meta built the Ranking Engineer Agent, an autonomous AI agent designed to drive the end-to-end ML lifecycle and iteratively evolve Meta’s ads ranking models at scale.

ML training jobs run for hours or days, far beyond what any session-bound assistant can manage.

4 months, 1 week назад @ engineering.fb.com
Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps
Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps

Nowhere is this more apparent than in mobile security, where a single class of vulnerability can be replicated across hundreds of call sites scattered throughout a sprawling, multi-app codebase serving billions of users.

Meta’s Product Security team has developed a two-pronged strategy to address this:Designing secure-by-default frameworks that wrap potentially unsafe Android OS APIs and make the secure path the easiest path for developers, andLeveraging generative AI to automate the migration of existing code to those frameworks at scale.

The result is a system that can propose, validate, and submit security patches across millions of lines of code with minimal friction for the engineers w…

4 months, 2 weeks назад @ engineering.fb.com
Uber Engineering
последний пост None
neptune.ai neptune.ai
последний пост 7 months, 3 weeks назад
We are joining OpenAI
We are joining OpenAI We are joining OpenAI

Piotr Niedźwiedź, CEO/CTO and founder of neptune.aiI’m excited to share that we’ve entered into a definitive agreement to be acquired by OpenAI, subject to closing conditions.

We are thrilled to join the OpenAI team and help their AI researchers build better models faster.

Neptune is a metrics dashboard company.”We’ve worked closely with OpenAI to create the metrics dashboard that helps teams building foundation models.

Our future with OpenAINeptune will join OpenAI and continue to support AI researchers with tools to monitor, debug, and evaluate frontier models.

We are looking forward to working with top AI researchers and supporting OpenAI’s mission of ensuring that AGI benefits all of hu…

7 months, 3 weeks назад @ neptune.ai
Synthetic Data for LLM Training
Synthetic Data for LLM Training Synthetic Data for LLM Training

For instance, financial data is highly sensitive and protected by very strict regulations, and synthetic data mimics the real data distribution without revealing customer information.

Read more about how leading foundation model teams curate their training data and other topics in the State of Foundation Model Training Report 2025.

Choosing the right synthetic data generation technique depends on the type of data and its complexity.

Synthetic tabular data generation is a promising direction to overcome these challenges by learning the distribution of the tabular data.

Post-processingAs the distribution of tabular data is highly complex, it makes the synthetic tabular data generation very ch…

8 months, 2 weeks назад @ neptune.ai
What are LLM Embeddings: All you Need to Know
What are LLM Embeddings: All you Need to Know What are LLM Embeddings: All you Need to Know

TL;DR LLM embeddings are the numerical, vector representations of text that Large Language Models (LLMs) use to process information.

Unlike their predecessor word embeddings, LLM embeddings are context-aware and dynamically change to capture semantic and syntactic relationships based on the surrounding text.

What are the applications of LLM embeddings?

Word EmbeddingsSparse Word Embeddings One-Hot Vectors 1970s TF-IDF1980s Co-Occurrence MatrixStatic Word Embeddings Word2Vec 2013 GloVe 2014Contextualized word embeddings ELMo 2018 GPT-1 2018 BERT 2018 LLAMA 2023 DeepSeek-V1 2023 GPT-4 2023Static word embeddingsStatic word embeddings, such as word2vec in 2013, marked a significant development.…

8 months, 3 weeks назад @ neptune.ai
Detecting and Fixing ‘Dead Neurons’ in Foundation Models
Detecting and Fixing ‘Dead Neurons’ in Foundation Models Detecting and Fixing ‘Dead Neurons’ in Foundation Models

TL;DR Dead neurons silently waste compute and reduce effective model capacity in foundation models.

Dead neurons’ impactRecent studies into dead neurons in the context of foundation models show interesting, albeit worrying, results.

These large reported fractions of dead neurons in foundation models are a concern from a computational perspective.

Before we move on to discuss how to detect and fix dead neurons, let’s touch upon an important distinction between dead neurons and vanishing gradients.

Further reading How to Monitor, Diagnose, and Solve Gradient Issues in Foundation Models Read moreVisualizing activation distributionsIs your foundation model suffering from dead neurons?

9 months назад @ neptune.ai
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training

In the first part of this series, we covered the fundamentals of instruction fine-tuning (IFT).

def calculate_irs(instruction, output, reference_model): evaluation_prompt = f""" Instruction: {instruction} Model Output: {output} Rate how well the output follows the instruction on these criteria: 1.

| SourceHINT addresses a computational inefficiency in standard instruction fine-tuning: repeatedly reprocessing the same task instruction with every input example.

Read more about foundation model training infrastructure and other topics in Neptune’s 2025 State of Foundation Model Training Report.

First, during initial instruction fine-tuning across multiple diverse tasks, the model learns genera…

9 months, 1 week назад @ neptune.ai
How to Optimize LLM Inference
How to Optimize LLM Inference How to Optimize LLM Inference

Large Language Model (LLM) inference at scale is challenging as it involves transferring massive amounts of model parameters and data and performing computations on large tensors.

In the following, we’ll use the Llama model family architecture as a specific example to understand the LLM workload at inference.

For a far more detailed analysis of the LLM workload at inference, see the chapter All About Transformer Inference in the book How to Scale Your Model, published by Google DeepMind.

See also How to Run LLMs Locally Read moreA quick primer on hardware for LLM inferenceA typical LLM inference cluster consists of several nodes, each with a multi-core CPU and multiple accelerator devices, …

9 months, 2 weeks назад @ neptune.ai
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 4 months, 3 weeks назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

4 months, 3 weeks назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

7 months назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

7 months назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

7 months назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

7 months, 2 weeks назад @ youtube.com
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff) [Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)

https://arxiv.org/abs/2510.17558 Abstract:

We propose an extension of the decoder Transformer that conditions its generative process on random latent variables which are learned without supervision thanks to a variational procedure. Experimental evaluations show that allowing such a conditioning translates into substantial improvements on downstream tasks. Author: François Fleuret Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the con…

8 months, 4 weeks назад @ youtube.com
[Video Response] What Cloudflare's code mode misses about MCP and tool calling
[Video Response] What Cloudflare's code mode misses about MCP and tool calling [Video Response] What Cloudflare's code mode misses about MCP and tool calling

Theo's Video: https://www.youtube.com/watch?v=bAYZjVAodoo

Cloudflare article: https://blog.cloudflare.com/code-mode/ Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8…

9 months, 1 week назад @ youtube.com
[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)
[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant) [Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)

Paper: https://arxiv.org/abs/2508.21038 Abstract:

Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-following, coding, and more. These new benchmarks push embeddings to work for any query and any notion of relevance that could be given. While prior works have pointed out theoretical limitations of vector embeddings, there is a common assumption that these difficulties are exclusively due to unrealistic queries, and those that are not can be overcome with better training data and larger models. In this work, we demonstrate that we may encounter these theoretical limitations in realist…

9 months, 2 weeks назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 3 days, 6 hours назад
The 64 sugar cubes puzzle
The 64 sugar cubes puzzle The 64 sugar cubes puzzle

See all monthly puzzles: https://momath.org/mindbenders/

3 days, 6 hours назад @ youtube.com
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

1 week, 4 days назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

1 month, 1 week назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

1 month, 2 weeks назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

1 month, 2 weeks назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

1 month, 2 weeks назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

2 months назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

3 months, 1 week назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

4 months назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

4 months назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

4 months, 1 week назад @ youtube.com
Bacteria Grid Puzzle Solution
Bacteria Grid Puzzle Solution Bacteria Grid Puzzle Solution

Part of a monthly series of puzzlers, in collaboration with MoMath and Peter Winkler

4 months, 1 week назад @ youtube.com
The most underappreciated formula | Exploring high-dimensional spheres
The most underappreciated formula | Exploring high-dimensional spheres The most underappreciated formula | Exploring high-dimensional spheres

On the volumes of higher-dimensional spheres

Explore the 3b1b virtual career fair: See https://3b1b.co/talent

Become a supporter for early views of new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Thanks to UC Santa Cruz for letting me film there, and special thanks to Pedro Morales-Almazan for arranging everything. My video on Numberphile with a fun application of this problem: https://youtu.be/6_yU9eJ0NxA Timestamps:

0:00 - Introduction

1:01 - Random puzzle

6:16 - Outside the box

14:35 - Setting up the volume grid

21:14 - Why 4πr^2

25:21 - Archimedes in higher dimensions

36:17 - The general formul…

5 months назад @ youtube.com
The lattice bacteria puzzle
The lattice bacteria puzzle The lattice bacteria puzzle

Part of a series of monthly puzzles, done in collaboration with MoMath.

https://momath.org/mindbenders

5 months, 1 week назад @ youtube.com
Solution to the ladybug clock puzzle
Solution to the ladybug clock puzzle Solution to the ladybug clock puzzle

Solution to last month's probability puzzle.

5 months, 1 week назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 1 week, 4 days назад
AI Helped Them Code Faster… But At A Cost
AI Helped Them Code Faster… But At A Cost AI Helped Them Code Faster… But At A Cost

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/AI-assistance-coding-skills 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 4 days назад @ youtube.com
The Hidden World Inside An AI
The Hidden World Inside An AI The Hidden World Inside An AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://transformer-circuits.pub/2025/linebreaks/index.html Paper for reindeer vision change - https://royalsocietypublishing.org/rspb/article/280/1773/20132451/50765/Shifting-mirrors-adaptive-changes-in-retinal 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli …

1 week, 5 days назад @ youtube.com
New AI Just Reinvented Minecraft Worlds
New AI Just Reinvented Minecraft Worlds New AI Just Reinvented Minecraft Worlds

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://xandergos.github.io/terrain-diffusion/

https://modrinth.com/mod/terrain-diffusion

https://github.com/xandergos/terrain-diffusion Source video for some parts of the footage: https://www.youtube.com/watch?v=irE4tcDtUIg 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fi…

2 weeks, 1 day назад @ youtube.com
DeepSeek's New AI Speed Hack Is Amazing
DeepSeek's New AI Speed Hack Is Amazing DeepSeek's New AI Speed Hack Is Amazing

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The DeepSeek paper is available here:

https://arxiv.org/abs/2607.05147v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

2 weeks, 6 days назад @ youtube.com
Game Physics Just Got 170 Times Faster
Game Physics Just Got 170 Times Faster Game Physics Just Got 170 Times Faster

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://arxiv.org/abs/2506.06494 Sources:

https://www.youtube.com/shorts/Tx7167DXr8U

https://www.youtube.com/watch?v=55F9dY2Y1zc 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 weeks, 3 days назад @ youtube.com
This New AI Model Changes Everything
This New AI Model Changes Everything This New AI Model Changes Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers GLM 5.2: https://z.ai/blog/glm-5.2 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 weeks, 5 days назад @ youtube.com
DeepSeek Just Solved AI's Billion Dollar Problem
DeepSeek Just Solved AI's Billion Dollar Problem DeepSeek Just Solved AI's Billion Dollar Problem

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2602.21548 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi #deepseek

1 month назад @ youtube.com
This is OpenClaw On Steroids
This is OpenClaw On Steroids This is OpenClaw On Steroids

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://recursivemas.github.io/

https://github.com/RecursiveMAS/RecursiveMAS 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://felicia.hu

1 month, 1 week назад @ youtube.com
Claude AI Knows More Than It Tells You
Claude AI Knows More Than It Tells You Claude AI Knows More Than It Tells You

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/natural-language-autoencoders

https://transformer-circuits.pub/2026/nla/index.html 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 1 week назад @ youtube.com
NVIDIA's New Free AI - A Gift To All of Us
NVIDIA's New Free AI - A Gift To All of Us NVIDIA's New Free AI - A Gift To All of Us

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Nemotron 3 Ultra paper is available here:

https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/ Free Rendering course and source code:

https://users.cg.tuwien.ac.at/zsolnai/gfx/rendering-course/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://f…

1 month, 1 week назад @ youtube.com
AI Agents as "Games Masters"? 🎮🔥
AI Agents as "Games Masters"? 🎮🔥 AI Agents as "Games Masters"? 🎮🔥

Check the pinned comment for the link to the full interview. Could AI agents eventually become the "Games Master" driving your gaming storylines? We explore the concept of AI assisting players or creating dynamic, non-scripted narratives. Discover how AI is currently being tested inside immersive game environments to change how we play. 🧠 Hashtags: #aiingames #gaming #ai #gamedev #futuretech

1 month, 3 weeks назад @ youtube.com
DeepMind’s New AI Found A Strange New Way To Think
DeepMind’s New AI Found A Strange New Way To Think DeepMind’s New AI Found A Strange New Way To Think

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://github.com/google-deepmind/alphaproof-nexus-results

https://arxiv.org/html/2605.22763v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 3 weeks назад @ youtube.com
Meet the AI "Co-Scientist" Changing Everything 🤖🧪 #ai
Meet the AI "Co-Scientist" Changing Everything 🤖🧪 #ai Meet the AI "Co-Scientist" Changing Everything 🤖🧪 #ai 1 month, 3 weeks назад @ youtube.com
Claude Opus 4.8: Lying Machine No More
Claude Opus 4.8: Lying Machine No More Claude Opus 4.8: Lying Machine No More

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers Anthropic's Opus 4.8: https://www.anthropic.com/news/claude-opus-4-8 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 3 weeks назад @ youtube.com
A Second Nobel Prize for AlphaFold? 🧬🏆 #alphafold #deepmind #nobelprize #science #ai
A Second Nobel Prize for AlphaFold? 🧬🏆 #alphafold #deepmind #nobelprize #science #ai A Second Nobel Prize for AlphaFold? 🧬🏆 #alphafold #deepmind #nobelprize #science #ai

Check the pinned comment for the link to the full interview. We're discussing whether a "second order Nobel" prize is on the horizon for AI-driven science. With over 3 million researchers already using AlphaFold, the real-world impact is already historic. Hear what the experts think about what comes next for scientific discovery! 🔬

1 month, 3 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 3 weeks, 4 days назад
Омни-модели будущего 🚀
Омни-модели будущего 🚀 Омни-модели будущего 🚀

Что они будут уметь — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

3 weeks, 4 days назад @ youtube.com
Качество модели взлетело... без мультимодального RL?
Качество модели взлетело... без мультимодального RL? Качество модели взлетело... без мультимодального RL?

О росте мультимодального качества рассказал Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

4 weeks назад @ youtube.com
Работа с данными — это скучно?
Работа с данными — это скучно? Работа с данными — это скучно?

А почему — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month назад @ youtube.com
Как приготовить SFT 🍲
Как приготовить SFT 🍲 Как приготовить SFT 🍲

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month назад @ youtube.com
Почему мультимодальные модели — это база 🤖
Почему мультимодальные модели — это база 🤖 Почему мультимодальные модели — это база 🤖

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 1 week назад @ youtube.com
Омни-модель: что это за зверь такой
Омни-модель: что это за зверь такой Омни-модель: что это за зверь такой

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 1 week назад @ youtube.com
Borealis — как обучить аудио-LLM по цене MacBook
Borealis — как обучить аудио-LLM по цене MacBook Borealis — как обучить аудио-LLM по цене MacBook

На конференции Data Fest 2026 в Белграде независимый исследователь Александр Николич рассказал практическую историю создания аудиоязыковой модели Borealis с бюджетом, сопоставимым со стоимостью MacBook. Больше контента для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
Better LLM pre-training in NVFP4
Better LLM pre-training in NVFP4 Better LLM pre-training in NVFP4

At Data Fest 2026 in Belgrade, Andrei Panferov from the Institute of Science and Technology Austria introduced Quartet II, a novel method for NVFP4 pre-training that recovers SOTA accuracy. He outlined the core challenges of low-precision LLM training and presented CUDA kernels tuned for Blackwell GPUs, ready for integration into real training pipelines. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
Как безопасно выкатывать новые версии продуктовых AI-агентов
Как безопасно выкатывать новые версии продуктовых AI-агентов Как безопасно выкатывать новые версии продуктовых AI-агентов

На Data Fest 2026 в Белграде Дмитрий Коршунов, Team Lead ML в Ecom, показал, как безопасно обновлять продуктовых AI-агентов с помощью системы автометрик. На примере агента Яндекс AI для турецкого рынка он объяснил, как фиксировать регрессии до прода, сравнивать версии и принимать решение о релизе, когда простой «Hello, Agent» уже позади. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents

At Data Fest 2026 in Belgrade, Karina Romanova, Senior LLM Research Engineer, presented HGRPO — a hierarchical modification of GRPO for multi-turn dialogue agents. Applied to a booking agent in Yandex Alice, the method improved truthfulness by 8.0 percentage points and reduced dialogue length by 10.7%. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей

На Data Fest 2026 в Белграде Вячеслав Костров, ML-инженер в Яндексе, рассказал, как uplift-модели решают бизнес-задачи Лавки: от персональных скидок до показа продуктовых подборок. Он разобрал постановку uplift-задачи, подбор метрик и построение политик, а также практические приёмы с лагранжианом и uplift-деревьями для баланса ограничений. Всё это — на примере реальных внедрений и с разбором типичных ошибок. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #Mu…

1 month, 2 weeks назад @ youtube.com
Поиск по архивам: как мы переходим к осознанному распознаванию текста
Поиск по архивам: как мы переходим к осознанному распознаванию текста Поиск по архивам: как мы переходим к осознанному распознаванию текста

На Data Fest 2026 в Белграде Дарья Виноградова, лид команды компьютерного зрения, представила два важных майлстоуна архивного поиска: новую архитектуру распознавания текста и выделение смысловых структур. Эти изменения делают поиск человечнее — теперь можно искать не слова среди текста, а человека среди людей. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
Hacks and Defenses in Automatic Kernel Generation
Hacks and Defenses in Automatic Kernel Generation Hacks and Defenses in Automatic Kernel Generation

На Data Fest 2026 в Белграде Егор Коновалов, ML-инженер, разобрал хаки, которые находят LLM-агенты, когда генерируют GPU/TPU-код: от тривиального обхода numerical tolerance до изощрённых атак на timing-измерения и эксплуатации дыр в test harness. А ещё Егор показал, какие методы защиты реально работают, а какие создают ложное чувство безопасности. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
Real-time video generation: where we are and what comes next
Real-time video generation: where we are and what comes next Real-time video generation: where we are and what comes next

At Data Fest 2026 in Belgrade, Andrey Filatov from KREA AI broke down the current state of real-time video generation: which architectures dominate, how they differ, and what challenges arise from compute limits and memory bottlenecks. He also covered production solutions like distillation and caching, and shared his outlook for the next 2–3 years: what will soon become possible and which bottlenecks the industry still overlooks. More content for developers: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #Reinforceme…

1 month, 2 weeks назад @ youtube.com
AI-генерация учебного контента и проверка открытых ответов студентов
AI-генерация учебного контента и проверка открытых ответов студентов AI-генерация учебного контента и проверка открытых ответов студентов

Доклад из секции ML & Education конференции Data Fest 2026 в гостях у Яндекса «AI-генерация учебного контента и проверка открытых ответов студентов». Спикер — Денис Королёв, доцент, МИЭМ НИУ ВШЭ. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 2 weeks назад @ youtube.com
ML Trainings ML Trainings
последний пост 5 часов назад
Алексей Козлов | Как устроено ранжирование рекламных объявлений в Авито
Алексей Козлов | Как устроено ранжирование рекламных объявлений в Авито Алексей Козлов | Как устроено ранжирование рекламных объявлений в Авито

Спикер: Алексей Козлов, Авито Тех, Senior DS Engineer Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Advertising https://ods.ai/tracks/df26-ml-in-advertising

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 часов назад @ youtube.com
Аркадий Аврамчук, Мария Тимонина | AI в спорте:Аналитика по видео в командных и индивидуальных видах
Аркадий Аврамчук, Мария Тимонина | AI в спорте:Аналитика по видео в командных и индивидуальных видах Аркадий Аврамчук, Мария Тимонина | AI в спорте:Аналитика по видео в командных и индивидуальных видах

Спикеры: Аркадий Аврамчук, Мария Тимонина, руководитель направления, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 часов назад @ youtube.com
Светлана Ладанова | Camera-Control в моделях Text/Image-to-Video
Светлана Ладанова | Camera-Control в моделях Text/Image-to-Video Светлана Ладанова | Camera-Control в моделях Text/Image-to-Video

Спикер: Светлана Ладанова, руководитель направления, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 часов назад @ youtube.com
Капитанский мостик 26.07.2026: Amazon против AGI | Китайские модели в США | HuggingFace взломали
Капитанский мостик 26.07.2026: Amazon против AGI | Китайские модели в США | HuggingFace взломали Капитанский мостик 26.07.2026: Amazon против AGI | Китайские модели в США | HuggingFace взломали

0:00:00 Начало

0:00:59 Anaconda и Kilo Code

0:07:07 Stripe и OpenRouter

0:12:13 Amazon против AGI

0:15:49 Вышел KodaCode 1.0

0:20:18 Cerebras и AMD

0:24:06 Microsoft и Kimi K3

0:26:17 Илон Маск и будущее

0:39:04 Claude, Sakana и уязвимости

0:50:27 Cisco и уязвимости

0:53:13 ЦОДы в РФ для Китая

0:58:27 Китайские модели в США

1:01:33 Google запечет модели

1:06:04 Z.ai и китайский ЦОД

1:09:55 HuggingFace взломали ИИ-саммари: Обсуждение последних новостей в сфере технологий, включая AI, контейнеризацию, рынок чипов и изменения в индустрии. Узнайте о новых проектах, стратегиях компаний и трендах развития. В этом выпуске мы обсуждаем последние новости в области технологий, регуляции рынка и искус…

1 day, 14 hours назад @ youtube.com
Анастасия Рысьмятова | LLM в Авито
Анастасия Рысьмятова | LLM в Авито Анастасия Рысьмятова | LLM в Авито

Спикер: Анастасия Рысьмятова Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Marketplace от Avito.tech

https://ods.ai/tracks/df26_mlavitotech ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 5 hours назад @ youtube.com
Анатолий Мастрюков | Как трансформеры повысили разнообразие кандидатов на Главной Авито
Анатолий Мастрюков | Как трансформеры повысили разнообразие кандидатов на Главной Авито Анатолий Мастрюков | Как трансформеры повысили разнообразие кандидатов на Главной Авито

Спикер: Анатолий Мастрюков Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Marketplace от Avito.tech

https://ods.ai/tracks/df26_mlavitotech ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 5 hours назад @ youtube.com
Алина Бабенко | Модели в монетизации
Алина Бабенко | Модели в  монетизации Алина Бабенко | Модели в монетизации

Спикер: Алина Бабенко Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Marketplace от Avito.tech

https://ods.ai/tracks/df26_mlavitotech ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 5 hours назад @ youtube.com
Валерий Карпузов | ИИ и МЫ. Нейросети сведут, влюбят и разведут. Что будет дальше?
Валерий Карпузов | ИИ и МЫ. Нейросети сведут, влюбят и разведут. Что будет дальше? Валерий Карпузов | ИИ и МЫ. Нейросети сведут, влюбят и разведут. Что будет дальше?

Спикер: Валерий Карпузов, AggregatorFX, Team Lead DS Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Funtech https://ods.ai/tracks/df26-ml-in-funtech ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 5 hours назад @ youtube.com
Мичил Егоров | Речевая платформа в production: от обучения моделей до работы под нагрузкой
Мичил Егоров | Речевая платформа в production: от обучения моделей до работы под нагрузкой Мичил Егоров | Речевая платформа в production: от обучения моделей до работы под нагрузкой

Спикер: Мичил Егоров, X5 Tech, руководитель команды речевые технологии и антифрод Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data и ML в Retail от X5.tech

https://ods.ai/tracks/df26-ml-in-retail

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 5 hours назад @ youtube.com
Георгий Смирнов | ARGUS: большой рекомендательный трансформер в системе с сотнями тысяч RPS
Георгий Смирнов | ARGUS: большой рекомендательный трансформер в системе с сотнями тысяч RPS Георгий Смирнов | ARGUS: большой рекомендательный трансформер в системе с сотнями тысяч RPS

Спикер: Георгий Смирнов, Поисковые сервисы и ИИ Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Practical ML ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 5 hours назад @ youtube.com
Иван Юрашку | Новая policy еще не работала: как сравнить ее со старой до A/B
Иван Юрашку | Новая policy еще не работала: как сравнить ее со старой до A/B Иван Юрашку | Новая policy еще не работала: как сравнить ее со старой до A/B

Спикер: Иван Юрашку, Сбер, руководитель направления анализа данных Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Reliable ML ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 5 hours назад @ youtube.com
Анна Алипова | Люди против плохих данных: как создать команду, которая научит LLM говорить правильно
Анна Алипова | Люди против плохих данных: как создать команду, которая научит LLM говорить правильно Анна Алипова | Люди против плохих данных: как создать команду, которая научит LLM говорить правильно

Спикер: Анна Алипова, MWS AI, Руководитель отдела AI-тренеров Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career https://ods.ai/tracks/df26-opencareer ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 4 hours назад @ youtube.com
Эдуард Колдунов | Мультимодальный суп: Превращение кучи рекламных баннеров в осмысленные кластеры
Эдуард Колдунов | Мультимодальный суп: Превращение кучи рекламных баннеров в осмысленные кластеры Эдуард Колдунов | Мультимодальный суп: Превращение кучи рекламных баннеров в осмысленные кластеры

Спикер: Эдуард Колдунов, Digital Budget, R&D Lead Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Advertising https://ods.ai/tracks/df26-ml-in-advertising

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 4 hours назад @ youtube.com
Алена Феногенова | Опасности “вайбкодинга” тестов: как не надо делать бенчмарки
Алена Феногенова | Опасности “вайбкодинга” тестов: как не надо делать бенчмарки Алена Феногенова | Опасности “вайбкодинга” тестов: как не надо делать бенчмарки

Спикер: Алена Феногенова, Исполнительный директор, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 4 hours назад @ youtube.com
Константин Крестников | Harness: новый подход к созданию AI-агентов
Константин Крестников | Harness: новый подход к созданию AI-агентов Константин Крестников | Harness: новый подход к созданию AI-агентов

Спикер: Константин Крестников, управляющий директор, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

5 days, 4 hours назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 3 weeks, 5 days назад
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

3 weeks, 5 days назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

1 month, 4 weeks назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

2 months, 3 weeks назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

3 months, 2 weeks назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

4 months назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

4 months, 2 weeks назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

4 months, 4 weeks назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

5 months, 2 weeks назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

5 months, 3 weeks назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

6 months, 2 weeks назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

6 months, 3 weeks назад @ lexfridman.com
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths #487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths

Irving Finkel is a scholar of ancient languages and a longtime curator at the British Museum, renowned for his expertise in Mesopotamian history and cuneiform writing.

He specializes in reading and interpreting cuneiform inscriptions, including tablets from Sumerian, Akkadian, Babylonian, and Assyrian contexts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep487-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://shopify.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/Chevron: Reliable energy for data centers.

7 months, 2 weeks назад @ lexfridman.com
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life #486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life

Michael Levin is a biologist at Tufts University working on novel ways to understand and control complex pattern formation in biological systems.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep486-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

(2:42:41) – Mind uploading(3:01:22) – Alien intelligence(3:16:17) – Advice for young people(3:22:46) – Questions for AGI

7 months, 4 weeks назад @ lexfridman.com
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy #485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy

David Kirtley is a nuclear fusion engineer and CEO of Helion Energy, a company working on building the world's first commercial fusion power plant by 2028.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep485-sc

See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript:

https://lexfridman.com/david-kirtley-transcript CONTACT LEX:

Feedback - give feedback to Lex: https://lexfridman.com/survey

AMA - submit questions, videos or call-in: https://lexfridman.com/ama

Hiring - join our team: https://lexfridman.com/hiring

Other - other ways to get in touch: https://lexfridman.com/contact EPISODE LINKS:

David's X: htt…

8 months, 1 week назад @ lexfridman.com
#484 – Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming
#484 – Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming #484 – Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming

Dan Houser is co-founder of Rockstar Games and is a legendary creative mind behind Grand Theft Auto (GTA) and Red Dead Redemption series of video games.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep484-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://box.com/aiUPLIFT Desk: Standing desks and office ergonomics.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(01:29) – Sponsors, Comments, and Reflections(11:32) – Greatest films of all time(23:45) – Making video games(26:36) – GTA 3(29:55) – Open world video games(32:42) – Character creation(36:09) – Superintelligent AI in A Bette…

8 months, 4 weeks назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 3 months, 1 week назад
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

3 months, 1 week назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

3 months, 2 weeks назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

4 months назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

4 months, 3 weeks назад @ microsoft.com
Ideas: Community building, machine learning, and the future of AI
Ideas: Community building, machine learning, and the future of AI Ideas: Community building, machine learning, and the future of AI

This week, machine learning researchers around the world will be attending the annual Conference on Neural Information Processing Systems, or NeurIPS.

In this series, we’ll explore the technologies that are shaping our future and the big ideas that propel them forward.

So around that time when I started my PhD at Penn, I was working in machine learning theory and algorithmic economics.

How had you experienced a lack of community or network of women in machine learning before the founding of WiML?

So particularly when working on topics related to fairness, I’ve ended up focusing a bunch on stuff to do with marginalized groups as part of my responsible AI work.

7 months, 4 weeks назад @ microsoft.com
Ideas: More AI-resilient biosecurity with the Paraphrase Project
Ideas: More AI-resilient biosecurity with the Paraphrase Project Ideas: More AI-resilient biosecurity with the Paraphrase Project

Today, I’m excited to talk about the Paraphrase Project, an effort I co-led exploring how advances in AI tools for protein design might impact biosecurity.

These “patches,” akin to those in cybersecurity, have now been shared with organizations globally to strengthen biosecurity screening.

The project highlights that the same AI tools capable of incredible good can also be misused, requiring us to be vigilant, thoughtful, and creative so we continue to get the most benefit out of AI tools while working to ensure that we avoid costly misuses.

So things like, how similar is this to that template, wild-type protein structure that we used as our conditioning information?

But I feel like broadly…

9 months, 3 weeks назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 6 часов назад
Social Choice for Fair Recommendations
Social Choice for Fair Recommendations Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social choice theory can help balance the needs of users, creators, and society. The conversation explores the future of recommendation algorithms and why fairness is a far more complex challenge than it first appears.

6 часов назад @ dataskeptic.com
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

3 weeks, 4 days назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

1 month назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

1 month, 1 week назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

2 months, 3 weeks назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

3 months назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

4 months назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 4 months, 2 weeks назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

5 months назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

5 months, 1 week назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

5 months, 3 weeks назад @ dataskeptic.com
Fairness in PCA-Based Recommenders
Fairness in PCA-Based Recommenders Fairness in PCA-Based Recommenders

In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users equally. David introduces the concept of "power niche users" - highly active users with specialized in…

6 months назад @ dataskeptic.com
Video Recommendations in Industry
Video Recommendations in Industry Video Recommendations in Industry

In this episode, Kyle Polich sits down with Cory Zechmann, a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares insights on why TikTok's algorithm works so well (clean data and massive interaction volume), the crucial …

7 months назад @ dataskeptic.com
Eye Tracking in Recommender Systems
Eye Tracking in Recommender Systems Eye Tracking in Recommender Systems

In this episode, Santiago de Leon takes us deep into the world of eye tracking and its revolutionary applications in recommender systems. As a researcher at the Kempelin Institute and Brno University, Santiago explains the mechanics of eye tracking technology—how it captures gaze data and processes it into fixations and saccades to reveal user browsing patterns. He introduces the groundbreaking RecGaze dataset, the first eye tracking dataset specifically designed for recommender systems research, which opens new possibilities for understanding how users interact with carousel interfaces like Netflix. Through collaboration between psychologists and AI researchers, Santiago's work demonstrate…

7 months, 1 week назад @ dataskeptic.com
Cracking the Cold Start Problem
Cracking the Cold Start Problem Cracking the Cold Start Problem

In this episode of Data Skeptic, we dive deep into the technical foundations of building modern recommender systems. Unlike traditional machine learning classification problems where you can simply apply XGBoost to tabular data, recommender systems require sophisticated hybrid approaches that combine multiple techniques. Our guest, Boya Xu, an assistant professor of marketing at Virginia Tech, walks us through a cutting-edge method that integrates three key components: collaborative filtering for dimensionality reduction, embeddings to represent users and items in latent space, and bandit learning to balance exploration and exploitation when deploying new recommendations. Boya shares insigh…

7 months, 3 weeks назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 3 days, 10 hours назад
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier 1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier

What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building. Additional mate…

3 days, 10 hours назад @ podtrac.com
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams 1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams

Dr. Catherine Williams, Chief Data Officer at the nonprofit Candid, was solving black-hole equations with pen and paper before she ever wrote a line of code. She earned a PhD in math researching general relativity and black holes, did postdocs at Stanford and Columbia and then became one of the very first data scientists, joining AppNexus back in 2012, around the same time “data scientist” became a job title at all. In this episode, she traces the field’s evolution from Bayesian models to BERT to today’s LLMs, and makes a compelling case that going deep on the underlying math matters more than ever, even now that AI can do the math for you. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

6 days, 10 hours назад @ podtrac.com
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents 1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents

In Episode #1010, Jon Krohn digs into “the advisor strategy”, a clever pattern that pairs a fast, cheap executor model with a frontier-class advisor it can consult mid-task, all inside a single API call. Every agent builder faces the same tension: frontier models plan best but cost too much to run on every turn, while small models fumble the decisions that matter. Anthropic’s advisor tool resolves it with roughly a one-line code change, and the benchmarks are startling: Sonnet with an Opus advisor scored higher than Sonnet alone while costing 11.9% less, and Haiku’s BrowseComp score more than doubled at 85% lower cost than Sonnet solo. Jon covers the newest Fable 5 numbers, the practical go…

1 week, 3 days назад @ podtrac.com
1009: How AI Is Quietly Saving Lives, with Steve Mock
1009: How AI Is Quietly Saving Lives, with Steve Mock 1009: How AI Is Quietly Saving Lives, with Steve Mock

In Episode #1009, Steve Mock (investor at Blumberg Capital, five-time entrepreneur and creator of aisavedme.org), joins Jon Krohn to explore the quiet layer of everyday AI adoption that rarely gets documented. After his 84-year-old father asked a deceptively simple question, “How does one use AI?”, Steve built a place for people to share how AI is actually helping them. The stories that came in surprised him: they’re rarely about the technology and almost always about human outcomes, caregiving, communication, learning, confidence and connection. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1009⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Intere…

1 week, 6 days назад @ podtrac.com
1008: The AI-Native Startup Playbook
1008: The AI-Native Startup Playbook 1008: The AI-Native Startup Playbook

In Episode #1008, Jon Krohn digs into Anthropic's 35-page Founder's Playbook and pulls out the practical guidance for each of its four startup stages: Idea, MVP, Launch and Scale. AI has erased the three bottlenecks that historically gated company-building — capital, headcount and technical skill — turning the founder from individual contributor into an "orchestrator of agents." Along the way, Jon covers the trap of mistaking building for validating, using AI as a structured devil's advocate against your own idea, the compounding danger of "agentic technical debt," two litmus tests for real product-market fit, and the three-layer moat that keeps a well-funded incumbent from copying you. His…

2 weeks, 3 days назад @ podtrac.com
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd 1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd

Benjamin Todd, co-founder and President of 80,000 Hours and author of the new Penguin Random House book 80,000 Hours: How to Have a Fulfilling Career That Does Good, joins Jon Krohn for a major update on career strategy in the AI era, his first appearance since before ChatGPT existed. Ben explains why “follow your passion” is backwards and why rare, valuable skills used to help others are what actually generate lasting fulfillment, the ABZ framework for planning under deep uncertainty, why the only durable move is to keep shifting onto whatever bottleneck AI can’t yet clear, and how a human-level digital worker becomes superhuman almost immediately. He and Jon also map the risk landscape, p…

2 weeks, 6 days назад @ podtrac.com
1006: In Case You Missed It in June 2026
1006: In Case You Missed It in June 2026 1006: In Case You Missed It in June 2026

In this month's episode of ICYMI, hear from Chip Huyen, Andrey Kurenkov, Frank Basso and Gilbert Eijkelenboom, discussing why moats are shifting toward physical systems and accumulated product intuition, how Astrocade built vibe coding before the term existed, what it's really like inside a deafeningly loud AI data center, why only 15% of people are technically self-aware and whether AGI requires anything like consciousness. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1006⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this episode you will learn: (00:00) The Cost of Bu…

3 weeks, 3 days назад @ podtrac.com
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom 1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom

Gilbert Eijkelenboom, bestselling author of People Skills for Analytical Thinkers and founder of the training firm MindSpeaking joins Jon Krohn to make the case that communication is a core data skill, not an optional extra. Gilbert shares the “And, But, Therefore” framework for turning dense analysis into a story stakeholders act on, the research suggesting only around 15% of people are genuinely self-aware (and how journaling, meditation, and exercise help close that gap), how childhood experiences install behavioral “algorithms” we carry into the workplace and why behavior change precedes attitude change, so doing small, uncomfortable things for 30 days can rewire how you see yourself. A…

3 weeks, 6 days назад @ podtrac.com
1004: Recursive Self-Improvement
1004: Recursive Self-Improvement 1004: Recursive Self-Improvement

Could an AI get good enough at AI research to build its own, more capable successor and kick off a compounding loop? That’s recursive self-improvement (RSI) and it surged into the conversation after Anthropic revealed that, as of May 2026, Claude wrote more than 80% of the code merged into its production codebase. In this Five-Minute Friday, Jon Krohn separates today’s AI-assisted coding from true RSI, walks through the accelerating evidence - METR’s shrinking task “time horizon,” Google DeepMind’s AlphaEvolve, Andrej Karpathy’s overnight training-tuner, weighs Jack Clark’s 60% bet that AI builds its own successor by 2028 against the compute, data and “marketing” skeptics. As ever, Jon land…

1 month назад @ podtrac.com
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso 1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso

Frank Basso, VP of Infrastructure at Lightning AI, joins Jon Krohn for a rare ground-level tour of the one layer of the AI stack the show had never covered in over a thousand episodes: the physical data center. Frank explains how Lightning AI provisions its 35,000-plus GPUs through hyperscale co-location, why everything new is liquid-to-chip cooled, how GPUs talk to each other over ultra-fast east-west networks, and what it’s actually like to stand inside a 110-decibel AI data hall. He also debunks the most persistent myths about data-center water and electricity use, and makes the case for fuel cells, nuclear power, and 800-volt DC distribution as the path forward. Additional materials: ⁠⁠…

1 month назад @ podtrac.com
1002: Fable 5: The Full Story from Capabilities to Drama
1002: Fable 5: The Full Story from Capabilities to Drama 1002: Fable 5: The Full Story from Capabilities to Drama

Anthropic’s Claude Fable 5 was the most capable AI model ever released to the public and it lasted just three days before the US government forced it offline. Jon Krohn unpacks both halves of the story: what makes Fable 5 special, and why it was pulled. Fable 5 and its locked-down sibling Mythos 5 are the same model separated only by safeguards, in a new “Mythos-class” tier above Opus. Jon covers its state-of-the-art benchmarks, premium $10/$50-per-million-token pricing, conservative safety classifiers, and the federal export-control directive, reportedly sparked by an Amazon-flagged “jailbreak” that took it down. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/100…

1 month, 1 week назад @ podtrac.com
1001: How AI Erased My Career Moat, an Episode #1001 Special: Jon Krohn interviewed by Kirill Eremenko
1001: How AI Erased My Career Moat, an Episode #1001 Special: Jon Krohn interviewed by Kirill Eremenko 1001: How AI Erased My Career Moat, an Episode #1001 Special: Jon Krohn interviewed by Kirill Eremenko

For this episode #1001 special, the tables are turned: SuperDataScience founder Kirill Eremenko takes the host’s chair and Jon Krohn is the guest. They trace Jon Krohn’s path from an Oxford neuroscience PhD to a New York hedge fund to founding the AI consulting firm Y Carrot, why he regrets leaving academia and how tools like Claude Code erased his hard-won technical moat and why that makes skilled engineers more valuable than ever. Along the way: whether AI is a bubble, Jevons paradox and the data-center boom, the RICE framework for choosing AI projects, the single biggest reason AI projects fail and how a well-built AI agent could give anyone “Christopher Nolan–like” focus. Additional mat…

1 month, 1 week назад @ podtrac.com
1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests
1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests 1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests

For this landmark 1,000th episode and the show’s 10-year anniversary, host Jon Krohn is joined by SuperDataScience founder Kirill Eremenko, who hosted the podcast for its first 400-plus episodes before handing over the reins. In a first for the show, the episode was recorded live with the audience invited to join on air, alongside surprise appearances from the team, longtime guests, and even Jon’s family. Together, Jon Krohn and Kirill look back on a decade of the podcast and field listener questions on AI’s biggest opportunities, the build-versus-buy dilemma, how to break into the field today, and how to stay grounded amid the relentless pace of AI. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

1 month, 2 weeks назад @ podtrac.com
999: What's Left to Build When Software Is Free, with Chip Huyen
999: What's Left to Build When Software Is Free, with Chip Huyen 999: What's Left to Build When Software Is Free, with Chip Huyen

Chip Huyen joins host Jon Krohn for this milestone episode 999 to talk about her record-breaking book "AI Engineering" the most-read title on the O'Reilly platform last year and how the AI landscape has shifted since her last appearance. Chip breaks down what separates AI engineering from machine learning engineering, makes the case for a "start simple" workflow, gets candid about the real costs of running LLMs in production, and shares why she's now fascinated by physical AI, robotics, and world models and why the durable problems worth solving are increasingly human ones. Jon Krohn guides the conversation from the practical content of the book through to where the field is heading next. A…

1 month, 2 weeks назад @ podtrac.com
998: In Case You Missed It in May 2026
998: In Case You Missed It in May 2026 998: In Case You Missed It in May 2026

In this month’s episode of ICYMI, Jon Krohn explores how AI agents are simultaneously creating new risks and unlocking powerful new ways of working with data. Hear from Anneka Gupta, Cal Al-Dhubaib, Trevor Manz, Jazmia Henry, Jeremy Mumford, and Jacob Miller, discussing why the old cybersecurity playbook breaks down in the age of Claude Mythos, how the notebook became an AI agent’s working memory, what it really takes to build a foundation model from scratch, and why failing slowly is the most expensive mistake an AI team can make. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/998⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email natali…

1 month, 3 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 1 week, 3 days назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 week, 3 days назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 week, 3 days назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 month назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

2 months, 1 week назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

2 months, 1 week назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 1 week назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 1 week назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

2 months, 1 week назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 1 week назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

3 months, 1 week назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

3 months, 1 week назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

3 months, 1 week назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

4 months, 3 weeks назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

4 months, 3 weeks назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

5 months, 3 weeks назад @ datascienceathome.com