Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 19 часов назад
Tried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]
Tried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

19 часов назад @ reddit.com
Stereo2Spatial: Convert Stereo Music Tracks to Spatialized Binaural Mixes [P]
Stereo2Spatial: Convert Stereo Music Tracks to Spatialized Binaural Mixes [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

22 часа назад @ reddit.com
short-paper at ACL/EMNLP/EACL [R]
short-paper at ACL/EMNLP/EACL [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 2 hours назад @ reddit.com
BMVC rebuttals update [D]
BMVC rebuttals update [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 3 hours назад @ reddit.com
Prism accidentally leaked [D]
Prism accidentally leaked [D] Prism accidentally leaked [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 3 hours назад @ reddit.com
TACL journal doubts [D]
TACL journal doubts [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 10 hours назад @ reddit.com
ACL ARR 2026 - can't seem to find the review issue report button anywhere? [D]
ACL ARR 2026 - can't seem to find the review issue report button anywhere? [D] ACL ARR 2026 - can't seem to find the review issue report button anywhere? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 13 hours назад @ reddit.com
EU AI Act OpenRAG: 933 legally structured chunks and BGE-M3 embeddings in one SQLite file [P]
EU AI Act OpenRAG: 933 legally structured chunks and BGE-M3 embeddings in one SQLite file [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 13 hours назад @ reddit.com
New Fable5/Opus4.8 harness called "Schema" claims 99% on ARC-3 [R]
New Fable5/Opus4.8 harness called "Schema" claims 99% on ARC-3 [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days назад @ reddit.com
whats the best and complete way to keep up with ai/ml news? [D]
whats the best and complete way to keep up with ai/ml news? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 1 hour назад @ reddit.com
Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]
Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R] Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 2 hours назад @ reddit.com
CfP | RTCA @ NeurIPS 2026 [R]
CfP | RTCA @ NeurIPS 2026 [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 4 hours назад @ reddit.com
Are Current AI Memory Architectures Optimizing for the Wrong Abstraction? [D]
Are Current AI Memory Architectures Optimizing for the Wrong Abstraction? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 5 hours назад @ reddit.com
ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level [P]
ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 8 hours назад @ reddit.com
PnP-CoSMo: A Multi-Contrast MRI Reconstruction Framework based on Content/Style Modeling [R]
PnP-CoSMo: A Multi-Contrast MRI Reconstruction Framework based on Content/Style Modeling [R] PnP-CoSMo: A Multi-Contrast MRI Reconstruction Framework based on Content/Style Modeling [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 8 hours назад @ reddit.com
Towards Data Science
последний пост 4 часа назад
Many Companies Use AI. Few Know How to Build an AI-Native Enterprise Data Platform.
Many Companies Use AI. Few Know How to Build an AI-Native Enterprise Data Platform. Many Companies Use AI. Few Know How to Build an AI-Native Enterprise Data Platform.

The key difference between an AI agent and a chatbot is that an AI agent can take actions instead of simply generating responses.

In the world of data, these AI agents are usually called data agents.

Data agents are designed to act as AI data analysts.

In my view, organizations should include at least the 3 key AI components in their data workflow – Data Agent, AI QA Agent and AI Governance & Observability.

Unlike traditional IT governance or data governance, AI governance and observability usually focus on the following areas:Prompt VersioningPrompt versioning means treating prompts like any other software artifact.

4 часа назад @ towardsdatascience.com
Loop Engineering with Adaptive PDF Parsing: Start Cheap, Pay for a Heavier Parser Only When the Page Needs It
Loop Engineering with Adaptive PDF Parsing: Start Cheap, Pay for a Heavier Parser Only When the Page Needs It Loop Engineering with Adaptive PDF Parsing: Start Cheap, Pay for a Heavier Parser Only When the Page Needs It

This article is the first of two parts on adaptive parsing, in Part III of Enterprise Document Intelligence, a series that builds an enterprise RAG system from four bricks: document parsing, question parsing, retrieval, and generation.

Both produce the same line_df shape, so the rest of the pipeline (question parsing, retrieval, generation) does not know which parser ran.

Keep most pages on PyMuPDF; pay for a deeper parser only on the pages a question needs.

The pipeline’s four bricks (document parsing, question parsing, retrieval, generation, introduced in Articles 5 to 8) each own one to three of the nine evaluation checks.

Question parsing: routing by intentQuestion parsing turns the que…

6 часов назад @ towardsdatascience.com
How to Improve Customer Retention in FinTech
How to Improve Customer Retention in FinTech How to Improve Customer Retention in FinTech

User retention in fintech depends on many factors: service quality, product functionality, loyalty mechanics, communication, and other parts of the user experience.

The experiment showed that the direction was right: the increased cashback offer helped improve user retention, and we saw statistically significant gains in retention.

Along with genuinely at-risk users, the pre-churn model also selected users who were less responsive to retention mechanics.

We split the segment into two groups: one without the uplift model and one with the uplift model.

This group showed the real churn level without intervention and served as a reference point for calibrating the pre-churn model.

8 часов назад @ towardsdatascience.com
How to Work Effectively with GPT-5.6
How to Work Effectively with GPT-5.6 How to Work Effectively with GPT-5.6

I’ve also tried the model for actual implementations, and I believe GPT-5.6 is able to work for longer to complete tasks and is basically more thorough in the way it does its work.

How to effectively apply GPT-5.6 to solve problemsUse cases for GPT-5.6Now I’ll get into how to effectively use GPT-5.6 to solve problems.

And then when I started using GPT-5.6, it performed worse because I didn’t remember to give it access to all of these tools.

OpenAI has basically all the same connectors that Claude Code has, and there’s no reason you shouldn’t give access to Codex if you already give access to Claude Code.

ConclusionIn this article, I covered my opinion on the latest OpenAI model, GPT-5.6, an…

1 day, 5 hours назад @ towardsdatascience.com
Using Classical ML to Empower AI Agents
Using Classical ML to Empower AI Agents Using Classical ML to Empower AI Agents

However, it turns out that agentic AI needs classical ML much more than we probably thought.

Why Classical MLHowever, I want to remind you that classical ML models can also be really valuable tools for your agent.

A well trained classical ML model is going to be vastly more accurate and trustworthy.

A well trained classical ML model is going to be vastly more accurate and trustworthy.

With a classical ML model you can identify the decisions made to get to your inference, and validate these against your subject matter expertise.

1 day, 6 hours назад @ towardsdatascience.com
Context Engineering Isn’t Enough — A Loop Engineering Experiment With No LLM Inside the Loop
Context Engineering Isn’t Enough — A Loop Engineering Experiment With No LLM Inside the Loop Context Engineering Isn’t Enough — A Loop Engineering Experiment With No LLM Inside the Loop

In AI engineering circles, people call this loop engineering.

What “Loop Engineering” Means, and Where the Term Comes FromIf you spend time on AI engineering Twitter, Substack, or LinkedIn, you have probably seen the phrase loop engineering.

Other practitioner write-ups extend this into a fuller stack, prompt engineering to context engineering to harness engineering to loop engineering, with each layer wrapping the one before it.

Why I Built This Without Calling a Single LLMEvery explainer I read about loop engineering assumes an LLM sits inside the loop, making the decisions at each turn.

https://dev.to/truongpx396/the-agentic-loop-a-practical-field-guide-mnc Adnan Masood, “Loop Engineerin…

1 day, 8 hours назад @ towardsdatascience.com
Analog AI Is Back, But Can It Survive Its Own Noise?
Analog AI Is Back, But Can It Survive Its Own Noise? Analog AI Is Back, But Can It Survive Its Own Noise?

Source: Gartner, June 2026This is the backdrop for a quiet resurgence of an old idea: analog computing.

And this article is about both halves of that story: why the pitch is truly compelling, and why what killed analog computing the first time hasn’t actually gone away.

Analog computing has always had a noise problemThis is exactly why analog computing lost to digital in the first place, decades before AI made it relevant again.

This is the real shape of “analog AI has a noise problem”, not a gentle accuracy tax, but a threshold the network was never prepared to cross.

IBM remains mostly in research mode, recently publishing on mapping mixture-of-experts LLM layers onto 3D analog memory arc…

1 day, 9 hours назад @ towardsdatascience.com
One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited
One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited

(A production RAG pipeline for PDFs: relational parsing, TOC retrieval, typed answers) we upgraded each of the four bricks: document parsing, question parsing, retrieval, and generation, and wired them into one clean, linear pipeline.

The first part, Upgrading a baseline RAG, brick by brick (link to come), upgraded each brick on its own.

}, "out": { "raw_keywords": ["RAG", "retrieval", "generation"], "corrected_keywords": ["RAG", "retrieval", "generation"], "expert_keywords_added": [], "final_keywords": ["RAG", "retrieval", "generation"], "expected_answer_shape": "single" } }Brick 3, retrieval.

What works, what breaksDocument parsingQuestion parsingRetrievalGenerationMake RAG generation ret…

1 day, 11 hours назад @ towardsdatascience.com
Prepare These 5 Assets Before Your AI Agents Take On More Work
Prepare These 5 Assets Before Your AI Agents Take On More Work Prepare These 5 Assets Before Your AI Agents Take On More Work

In my last article, Redesign Work Before You Add More AI Agents, I argued that companies should start by redesigning the workflow before rolling out more AI tools and agents.

The more AI workflow enablement projects I work on, the more convinced I become that the missing piece is preparing the workflow itself.

In this article, I want to show what you can do before handing recurring work to an AI assistant or agent.

I will reuse it in ongoing AI tasks so I do not need to explain the same context each time.

The Acceptance Test AssetYou need to know what failure looks like before AI output goes to a customer, a production system, or the public.

2 days, 5 hours назад @ towardsdatascience.com
Context Engineering for RAG Question Parsing: From a Raw Question to Typed Fields That Steer Retrieval and Generation
Context Engineering for RAG Question Parsing: From a Raw Question to Typed Fields That Steer Retrieval and Generation Context Engineering for RAG Question Parsing: From a Raw Question to Typed Fields That Steer Retrieval and Generation

Enterprise Document Intelligence builds enterprise RAG on four bricks (document parsing, question parsing, retrieval, generation).

From a user string to a typed row with five columns, then to two briefs each downstream brick can act on.

The context-engineering equivalent is the opposite: write out a typed row with named fields, so every downstream call gets the same shape and can address fields by name.

2.2 Compress: the retrieval briefThe retrieval brick has no business knowing the answer shape, the suggested model tier, or the clarification the user was asked.

Five fields RAG should extract from any question (Article 6B).

2 days, 6 hours назад @ towardsdatascience.com
How to Get the Most Out of Claude Fable 5
How to Get the Most Out of Claude Fable 5 How to Get the Most Out of Claude Fable 5

However, it’s now been returned to the Claude subscription, and anyone with the Claude subscription can access Claude Fable 5.

Why use Claude Fable 5The main reason you should be using Claude Fable 5 is simply that it is the most powerful coding model out there at the moment.

I would argue that in many cases GPT-5.5 and definitely GPT-5.6 are now better than Claude Opus 4.8, but they’re not better than Claude Fable.

However, you should notice that I did not specifically mention implementation of code ’cause I do believe that Claude Fable is not that much superior on code implementations compared to Claude Opus 4.8, for example, which is what I’ll cover in the next sectionHow to maximize Cla…

2 days, 8 hours назад @ towardsdatascience.com
Why Your Betas Explode: The Hidden Geometry of Multicollinearity
Why Your Betas Explode: The Hidden Geometry of Multicollinearity Why Your Betas Explode: The Hidden Geometry of Multicollinearity

The slide showed two beta coefficients side by side: Linear TV at +2.4, Digital TV at +1.8.

When we expand the quadratic form, we get:L ( β ) = y T y − y T X β − ( X β ) T y + β Τ X T X β L ( β ) = y T y − 2 β Τ Χ Τ y + β T X T X β L(β) = y^Ty -y^TXβ – (Xβ)^Ty + β^ΤX^TXβ \\ \\\\ L(β) = y^Ty -2β^ΤΧ^Τy + β^TX^TXβOptimization: Setting the Derivative to ZeroTo find the minimum of this loss function, we take the partial derivative with respect to the vector β.

The features are genuinely distinct; Linear TV spend and Digital TV spend really are two different media investments, on two different platforms, measured separately.

In statistical language, we say the two features are near-collinear, or …

2 days, 9 hours назад @ towardsdatascience.com
Don’t Let Claude Grade Its Own Homework
Don’t Let Claude Grade Its Own Homework Don’t Let Claude Grade Its Own Homework

Agent throughput broke the human review loopHallucinations by themselves are a manageable problem, since a very careful review or functional assessment would catch them.

That makes the human reviewer the bottleneck of the whole delivery pipeline, and the bottleneck is always the next thing you should automate.

Number(m[1]) : null; await github.rest.repos.createCommitStatus({ ...opts, sha: context.payload.pull_request.head.sha, context: 'Codex review', state: unresolved > 0 ?

The prompt contract is where this setup differs, and it is the core of our review agent:name: Codex PR Review description: Run openai/codex-action with the project's PR review prompt.

The trailing line gets parsed by t…

3 days, 5 hours назад @ towardsdatascience.com
Building Trustworthy Production RAG Systems Through Continuous Evaluation
Building Trustworthy Production RAG Systems Through Continuous Evaluation Building Trustworthy Production RAG Systems Through Continuous Evaluation

The steps are written so you can follow them for your own RAG application, not just read them as theory.

This is called a golden dataset, and skipping it is the most common mistake teams make while working on a RAG system.

A good golden dataset entry has three parts: the question, a correct answer written by someone who knows the domain, and which document contains the answer.

Question: {question} Ground Truth: {ground_truth} Generated Answer: {answer} Source Document Date: {doc_date} 1. numeric_accuracy - are all numbers and facts correct, not just plausible?

A golden dataset with twenty good questions and a RAGAS score you check by hand is already ahead of most RAG systems in production t…

3 days, 6 hours назад @ towardsdatascience.com
How I Mastered Data Structures and Algorithms for ML (In 6 Weeks)
How I Mastered Data Structures and Algorithms for ML (In 6 Weeks) How I Mastered Data Structures and Algorithms for ML (In 6 Weeks)

The majority of coding interviews in the data science and machine learning space are either on LeetCode or HackerRank, typically asking a data structures and algorithms question or something closely related.

Stop Learning Data Structures & AlgorithmsThe first step may seem counterintuitive, but it’s to actually stop trying to learn data structures and algorithms the traditional way.

For those of you unfamiliar with data structures and algorithms, or DSA for short, let me give you a quick definition of these terms:Data structures — Organising and storing data so it can be accessed and modified efficiently.

I even wrote a series of articles about DSA on my blog at the time, whilst I was takin…

3 days, 8 hours назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 2 days, 10 hours назад
The Sequence Opinion #896: Spark, Compute, and the Two Metas
The Sequence Opinion #896: Spark, Compute, and the Two Metas The Sequence Opinion #896: Spark, Compute, and the Two Metas

The occasion was the launch of Muse Spark 1.1, the second model out of Meta Superintelligence Labs and the first Meta model ever to ship with a price tag.

Two days earlier Meta shipped Muse Image, its first image generation model from the new lab.

Chips, datacenters, cloud, models, API, apps, devices.

At the layer where models meet users, the app and agent layer, Meta might be the favorite.

At the layer where models get made, the evidence is thin and the structural arguments cut against it.

2 days, 10 hours назад @ thesequence.substack.com
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break

OpenAI’s audit of SWE-Bench Pro shows why a precise score can still be a poor measure - and why coding agents may become essential tools for auditing the benchmarks that grade them.

A frontier coding score can look wonderfully precise: 80.3 percent, one decimal place, clean enough to rank models and anchor product claims.

That is the uncomfortable conclusion of OpenAI’s audit of SWE-Bench Pro.

OpenAI estimates that roughly 30 percent of the public benchmark is broken.

OpenAI has withdrawn its earlier recommendation that the field adopt SWE-Bench Pro.

3 days, 10 hours назад @ thesequence.substack.com
The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era
The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era

Looking back at the 2015 distillation paper, what’s striking isn’t the temperature trick or the dark-knowledge framing --- it’s the world the paper quietly assumed.

There was a fixed input distribution.

There was a teacher that produced a probability vector over a closed set of classes.

There was a student trained to match that vector.

Then language models arrived, and one by one, every assumption broke.

5 days, 6 hours назад @ thesequence.substack.com
The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack

In the opinion section, we are going to debate Meta’s opportunities and tremendous challenges to catch up with the AI frontier labs.

This week’s GPT-5.6, GPT-Live, ChatGPT Work, Grok 4.5, and Muse Spark 1.1 reveal a shift.

Meta’s Muse Spark 1.1 makes the race more crowded and cheaper.

AI Lab: LMMS-Lab, NTU MMLab, and MicrosoftSummary: This paper introduces SkillOpt-Lite, a minimal viable pipeline for autonomous agent skill optimization that replaces complex algorithmic architectures with a file-system-based trajectory exploration approach.

Muse Spark 1.1Meta introduced Muse Spark 1.1, the second version of its multimodal reasoning model.

6 days, 10 hours назад @ thesequence.substack.com
The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough
The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough

Quick note: For over two years now, we’ve been running The Sequence without sponsors.

Literally every week we receive tens of inquires to sponsor The Sequence.

The question is deceptively simple: what makes a domain a good domain for AI?

Not good in the sense of commercially interesting, but good in the sense that if you point a modern training pipeline at it, capability actually compounds.

Verifiability

1 week, 2 days назад @ thesequence.substack.com
The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI
The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI

Quick note: For over two years now, we’ve been running The Sequence without sponsors.

Literally every week we receive tens of inquires to sponsor The Sequence.

There’s a running joke in machine learning that the field’s most valuable model isn’t a transformer at all — it’s gradient-boosted trees fit on a CSV.

You hand it the whole problem — training rows, test rows, all of it — as one giant prompt, and it answers.

Understanding TabFM really requires understanding that lineage, so let’s start there.

1 week, 3 days назад @ thesequence.substack.com
The Sequence Knowledge #890: A Brief History of Model Distillation
The Sequence Knowledge #890: A Brief History of Model Distillation The Sequence Knowledge #890: A Brief History of Model Distillation

The story most people tell about knowledge distillation starts in 2015, with Geoffrey Hinton, Oriol Vinyals, and Jeff Dean introducing a clever softmax temperature trick and a phrase — “dark knowledge” — that immediately lodged itself in the field’s vocabulary.

The real history is quieter, more pragmatic, and worth recovering, because the conceptual moves the field made between 2006 and 2015 still define how we think about distillation today.

The vocabulary changed.

The diagrams changed.

The underlying question — what exactly is being transferred from a teacher to a student?

1 week, 4 days назад @ thesequence.substack.com
The Sequence Radar #889: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land Grab
The Sequence Radar #889: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land Grab The Sequence Radar #889: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land Grab

The opinion section, discusses that domains are a good fit for rapid progress of AI models vs which ones are challenging.

Subscribe and don’t miss out:📝 Editorial: Fable 5's Comeback, ZCode's Debut, Claude Science, and the $3.5B Deployment Land GrabThis week’s developments in AI provide a clear direction where the space is going: more capable models and the imperative of capable delivery capabilities.

Start at the model layer, where we just watched the first frontier model get hot-patched by a government.

One layer up, Anthropic launched what I’d call claude --science .

Claude ScienceAnthropic announced Claude Science, a workbench for scientists.

1 week, 6 days назад @ thesequence.substack.com
The Sequence Opinion #888: Everything You Need to Know About the AI in Space Race
The Sequence Opinion #888: Everything You Need to Know About the AI in Space Race The Sequence Opinion #888: Everything You Need to Know About the AI in Space Race

The core thesis of this essay is simple to state: space is becoming a new frontier for AI — and one of the most competitive ones.

When the scarce thing was ideas, the frontier was architectures; when it was data, the frontier was the open web; when it was FLOPs, the frontier was the fab.

Today the scarce thing is energy — grid capacity, cooling water, land, permits — and orbit is the one place in reach where energy is effectively unmetered and no zoning board has jurisdiction.

This essay discusses the core thesis of AI in space: value proposition, key players, architecture differences and much more.

The core thesis: compute is now an energy problem, and space is an energy solution

2 weeks, 2 days назад @ thesequence.substack.com
The Sequence AI of the Week #887: Meta's Autodata: When Models Learn to Make Their Own Lessons
The Sequence AI of the Week #887: Meta's Autodata: When Models Learn to Make Their Own Lessons The Sequence AI of the Week #887: Meta's Autodata: When Models Learn to Make Their Own Lessons

Today, we are covering an amazing paper published by Meta last week: https://arxiv.org/abs/2606.25996There is a quiet shift happening in AI training.

For years, the center of gravity was the model: more parameters, more GPUs, better architectures, longer context windows, better optimizers.

Meta’s new Autodata work flips that perspective.

Not “ask a strong model to generate a million examples and hope the distribution is useful.” Instead, Autodata treats data generation like a miniature research loop.

An AI agent creates examples, tests them, studies the failures, updates its recipe, and tries again.

2 weeks, 3 days назад @ thesequence.substack.com
The Sequence Knowledge #886: Demystifying Model Distillation
The Sequence Knowledge #886: Demystifying Model Distillation The Sequence Knowledge #886: Demystifying Model Distillation

The simplest way to understand knowledge distillation is to imagine a very expensive teacher and a very cheap student.

The teacher is a large model: smart, slow, high-capacity, expensive to run.

The student is smaller: faster, cheaper, easier to deploy, but usually less capable if trained in the standard way.

Distillation asks a very practical question:Can the student learn not only from the original dataset, but from the teacher’s behavior?

In other words, instead of training the small model directly on reality, we train it on reality as interpreted by the big model.

2 weeks, 4 days назад @ thesequence.substack.com
The Sequence Radar #885: Last Week in AI: Models, Games, and the Future of Evaluation
The Sequence Radar #885: Last Week in AI: Models, Games, and the Future of Evaluation The Sequence Radar #885: Last Week in AI: Models, Games, and the Future of Evaluation

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Models, Games, and the Future of EvaluationThis week in AI had the strange feeling of a stack trace resolving itself.

Frontier AI releases are starting to look less like software updates and more like controlled deployment of critical infrastructure.

Then came General Intuition’s new raise, which feels like the cleanest signal yet that the next data frontier is not text, or even video, but action.

AI Lab: Qwen TeamSummary: This research introduces Qwen-AgentWorld, a foundational language world model designed to simulate seven diverse agentic environments through long chain-of-thought reasoning.

🤖 AI Tech ReleasesGPT 5.6 SolOpenAI un…

2 weeks, 6 days назад @ thesequence.substack.com
The Sequence Opinion #884: Self-Driving Labs: The Laboratory That Chooses Its Next Experiment
The Sequence Opinion #884: Self-Driving Labs: The Laboratory That Chooses Its Next Experiment The Sequence Opinion #884: Self-Driving Labs: The Laboratory That Chooses Its Next Experiment

The scientist decides what to test, transfers samples between instruments, inspects the results, updates their mental model and chooses the next experiment.

A self-driving lab moves part of that loop into software.

It makes something, measures it, updates a model and chooses the next move.

A self-driving lab can run the first few hundred experiments, notice that most of the remaining design space looks unpromising and redirect itself toward better candidates.

The simplest mental model is a loop:design → make → test → learn → design again

3 weeks, 1 day назад @ thesequence.substack.com
The Sequence AI of the Week #883: Qwen is Getting Into Robotics
The Sequence AI of the Week #883: Qwen is Getting Into Robotics The Sequence AI of the Week #883: Qwen is Getting Into Robotics

For about three years now, the Qwen family has lived inside a rectangle.

It reads your code, looks at your screenshots, answers your questions, and the whole time it has been doing this behind glass.

The Tongyi Lab team put it plainly: seeing is not acting.

Let me explain why I think this is the right shape, and where I’d keep my skepticism.

The actual bottleneck is not intelligence, it’s tokenization

3 weeks, 2 days назад @ thesequence.substack.com
The Sequence Knowledge #882: A New Series About Distillation
The Sequence Knowledge #882: A New Series About Distillation The Sequence Knowledge #882: A New Series About Distillation

I am super excited about this series that deep dives into distillation techniques.

Bigger models.

The frontier model became one of the strangest artifacts in the history of computing: a single neural network that looks less like a program and more like a compressed civilization of patterns.

A coding agent does not always need a frontier model for every token.

It may need a smaller draft model, a specialized debugging model, or a distilled planner trained on expert trajectories.

3 weeks, 3 days назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 3 months, 2 weeks назад
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

3 months, 2 weeks назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

4 months, 4 weeks назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

4 months, 4 weeks назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing & LLMs — новый сезон Курс Natural Language Processing & LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

5 months, 2 weeks назад @ habr.com
Machine Learning Mastery
последний пост 1 week, 1 day назад
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach

The common pitfalls that show up once agent memory is implemented, and how to fix them.

Agent memory strategy deserves the same deliberate design as orchestration.

Why Is Choosing an AI Agent Memory Strategy Important?

A customer support agent, for example, might keep the current ticket in working memory, a customer’s subscription tier in semantic memory, past complaints in episodic memory, and a learned refund-handling routine in procedural memory.

As discussed, working memory, semantic memory, episodic memory, and procedural memory serve different purposes and require different storage and retrieval strategies.

1 week, 1 day назад @ machinelearningmastery.com
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls

openai import OpenAI as LlamaOpenAI from llama_index .

# Prerequisites: pip install openai python-dotenv # How to run: python raw_api_agent.py import os import json from dotenv import load_dotenv from openai import OpenAI load_dotenv ( ) client = OpenAI ( api_key = os .

Prerequisites:pip install openai langchain langchain-openai llama-index \ llama-index-llms-openai llama-index-embeddings-openai python-dotenv 1 2 pip install openai langchain langchain - openai llama - index \ llama - index - llms - openai llama - index - embeddings - openai python - dotenvHow to run: Save as three_ways.py and run python three_ways.py# three_ways.py # The same document Q&A task implemented three ways: # Raw …

1 week, 2 days назад @ machinelearningmastery.com
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering

Topics we will cover include:What tools and subagents are, and the key differences between them.

This article explains what tools and subagents are, where each fits, and how to make the choice every time.

When an agent calls a tool, the result lands back in the same context the agent is actively reasoning in — prior reasoning, tool result, and everything else together.

A database record, search result, or API response can often be consumed immediately.

If the answer is independent reasoning, context isolation, specialized capabilities, or parallel execution, a subagent is likely justified.

1 week, 4 days назад @ machinelearningmastery.com
The Complete Guide to Tool Selection in AI Agents
The Complete Guide to Tool Selection in AI Agents The Complete Guide to Tool Selection in AI Agents

tools = tools self .

tools = tools self .

threshold : return { "status" : "resolved" , "tool" : tool [ "name" ] , "confidence" : score , "attempts" : 1 } # Reformulate by stripping filler words.

tools = tools self .

full_catalog_tokens = sum ( estimate_tokens ( d ) for d in descs ) def _retrieve ( self , query : str , top_k : int ) -> list [ dict ] : query_vec = self .

1 week, 5 days назад @ machinelearningmastery.com
Context vs. Memory Engineering in Agentic AI Systems
Context vs. Memory Engineering in Agentic AI Systems Context vs. Memory Engineering in Agentic AI Systems

Share Post ShareIn this article, you will learn how context engineering and memory engineering solve different problems in agentic AI systems, and how the two disciplines meet at the point where retrieved memory enters the context window.

Most of the time, the problem lies in two areas that get built together, conflated, or skipped: context engineering and memory engineering.

Memory Engineering: Designing Persistent AI Memory SystemsOnce an inference call completes, memory engineering determines what deserves to persist and under what conditions it gets used again.

trust_level >= 0.5 )AI Agent Memory Design Guide – Working, Long-Term, and Procedural Memory with Forgetting and Staleness Mana…

2 weeks, 2 days назад @ machinelearningmastery.com
Context Window Management for Long-Running Agents: Strategies and Tradeoffs
Context Window Management for Long-Running Agents: Strategies and Tradeoffs Context Window Management for Long-Running Agents: Strategies and Tradeoffs

Share Post ShareIn this article, you will learn five practical strategies for managing context windows in long-running AI agent applications, along with the key tradeoffs each approach introduces.

Five distinct context management strategies: sliding windows, recursive summarization, structured state management, ephemeral context via RAG, and dynamic context routing.

Accordingly, shifting from “LLMs as prompt-response engines” to “(agent-endowed) LLMs as long-running background processes” turns context windows into a major AI engineering bottleneck.

For all these reasons, managing context windows in the long run requires specific strategies like sliding windows, tiered memory, and dynamic su…

2 weeks, 4 days назад @ machinelearningmastery.com
Model Context Protocol Explained in 3 Levels of Difficulty
Model Context Protocol Explained in 3 Levels of Difficulty Model Context Protocol Explained in 3 Levels of Difficulty

Share Post ShareIn this article, you will learn how the Model Context Protocol (MCP) standardizes the way AI applications connect to external tools and data sources, broken down across three levels of depth.

How the host, client, and server work together, and what happens when a model’s request flows through an MCP server.

The transport options, security risks, and deployment choices that matter once an MCP server is running in production.

Accessing information outside that context requires external tools.

Because both sides follow the same protocol, an MCP server can be used by any compatible MCP client without requiring a custom integration for that specific client.

2 weeks, 5 days назад @ machinelearningmastery.com
The AI Agent Tech Stack Explained
The AI Agent Tech Stack Explained The AI Agent Tech Stack Explained

According to Atlan’s research on AI agent memory, 95% of enterprise generative AI pilots delivered zero measurable ROI in 2025, with failure attributed to context readiness rather than model quality.

isoformat ( ) , "user" : user_input , "agent" : agent _ response } ) def load_recent_episodes ( n : int = 5 ) -> str : "" "Retrieve the last N episodes as a formatted string for injection into context."

Prerequisites:pip install langchain langchain-openai langchain-chroma langchain-text-splitters chromadb python-dotenv 1 pip install langchain langchain - openai langchain - chroma langchain - text - splitters chromadb python - dotenvHow to run: Save as rag_pipeline.py, add OPENAI_API_KEY to your…

3 weeks, 1 day назад @ machinelearningmastery.com
Agentic Workflow vs. Autonomous Agent: What’s the Difference?
Agentic Workflow vs. Autonomous Agent: What’s the Difference? Agentic Workflow vs. Autonomous Agent: What’s the Difference?

lower ( ) : return "billing" return "general" def extract ( raw_input : str ) -> str : "" "Step 1 -- always runs, always leads to step 2.

def handle_billing(text: str) -> str: return f"[BILLING TEAM] Routed: {text[:50]}" def handle_technical(text: str) -> str: return f"[TECH SUPPORT] Routed: {text[:50]}" def handle_general(text: str) -> str: return f"[GENERAL QUEUE] Routed: {text[:50]}" # The branch map IS the entire decision space.

def handle_billing ( text : str ) -> str : return f "[BILLING TEAM] Routed: {text[:50]}" def handle_technical ( text : str ) -> str : return f "[TECH SUPPORT] Routed: {text[:50]}" def handle_general ( text : str ) -> str : return f "[GENERAL QUEUE] Routed: {text…

3 weeks, 2 days назад @ machinelearningmastery.com
Context Windows Are Not Memory: What AI Agent Developers Need to Understand
Context Windows Are Not Memory: What AI Agent Developers Need to Understand Context Windows Are Not Memory: What AI Agent Developers Need to Understand

Topics we will cover include:Why a context window behaves like a stateless scratchpad rather than persistent memory.

Deeming a huge context window as “memory” is, in architectural terms, similar to buying a 25-foot-wide office desk because you are reluctant to acquire a filing cabinet.

Context WindowA context window in an AI model, particularly agent-based ones with underlying language models, is like a desk surface or a stateless scratchpad.

When passing an agent a conversation history spanning over 200K tokens (large context window), it isn’t remembering what happened at a previous step in time.

Generate a compact summary for the active prompt summary = summarizer_model.generate(raw_trans…

3 weeks, 3 days назад @ machinelearningmastery.com
Clustering Unstructured Text with LLM Embeddings and HDBSCAN
Clustering Unstructured Text with LLM Embeddings and HDBSCAN Clustering Unstructured Text with LLM Embeddings and HDBSCAN

Share Post ShareIn this article, you will learn how to build a text clustering pipeline by combining large language model embeddings with HDBSCAN, a density-based clustering algorithm, to automatically discover topics in unlabeled text data.

target } ) df = df [ df [ 'text' ] .

) print ( "Sample document:" ) print ( df [ 'text' ] .

cluster import HDBSCAN # Initializing HDBSCAN # min_cluster_size=8: we specified that each cluster must have at least 8 documents clusterer = HDBSCAN ( min_cluster_size = 8 , min_samples = 3 , store_centers = 'centroid' ) df [ 'cluster' ] = clusterer .

shape [ 1 ] ) ] ) reduced_df [ 'cluster' ] = df [ 'cluster' ] # Getting all unique pairwise combinations of the …

3 weeks, 4 days назад @ machinelearningmastery.com
Building Browser-Using AI Agents in Python
Building Browser-Using AI Agents in Python Building Browser-Using AI Agents in Python

# Prerequisites: pip install playwright && playwright install chromium # How to run: python scrape_books.py import asyncio import json from playwright .

"" page = await get_page ( ) await page .

"" page = await get_page ( ) await page .

"" page = await get_page ( ) await page .

"" page = await get_page ( ) try : await page .

3 weeks, 5 days назад @ machinelearningmastery.com
The Roadmap to Mastering AI Agent Evaluation
The Roadmap to Mastering AI Agent Evaluation The Roadmap to Mastering AI Agent Evaluation

Step 1: Understanding Why Agent Evaluation Is ImportantThe instinct when an agent fails is to treat it as a prompting problem: the system prompt needs to be clearer.

Step 2: Defining What Agent Evaluation Success Looks LikeEvaluation is only as good as its success criteria.

Step 4: Grading Agent Reasoning and Output Quality with Model-Based JudgesSome agent evaluation dimensions resist deterministic checking — output quality, tone, faithfulness to retrieved context, appropriate empathy.

Step 5: Matching Agent Evaluation Strategy to Agent TypeGrading strategies apply broadly, but agent type determines which graders carry the most weight and which failure modes to prioritize.

Summary of Key A…

1 month назад @ machinelearningmastery.com
Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM
Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM

How to build, run, and evaluate a zero-shot sentiment classification pipeline using scikit-learn-compatible syntax.

from sklearn.pipeline import Pipeline from skllm.models.gpt.classification.zero_shot import ZeroShotGPTClassifier # Define the end-to-end pipeline sentiment_pipeline = Pipeline([ ("cleaner", text_cleaner), # Updated to use Groq's active Llama 3.1 8B model ("llm_classifier", ZeroShotGPTClassifier(model="custom_url::llama-3.1-8b-instant")) ]) # Fit the pipeline # Note: For Zero-Shot classification, fit() doesn't train the LLM.

pipeline import Pipeline from skllm .

. . Actual : negative | Predicted : negative Review : This entry is certainly interesting for series fans ( like mys…

1 month назад @ machinelearningmastery.com
AI Agent Tool Design: What Works and What Doesn’t
AI Agent Tool Design: What Works and What Doesn’t AI Agent Tool Design: What Works and What Doesn’t

What Works in AI Agent Tool Design1.

The difference becomes clearer when comparing a multi-action tool against dedicated single-purpose tools:# Avoid: action-based multi-behavior tool @tool def manage_customer( action: str, customer_id: str | None = None, data: dict | None = None ): """ action: create | get | update | delete | suspend """ ... # Prefer: single-responsibility tools @tool def create_customer(data: CustomerInput) -> Customer: """Create a new customer record."""

... @tool def suspend_customer(customer_id: str, reason: str) -> SuspensionResult: """Suspend a customer account."""

. . # Prefer: single-responsibility tools @ tool def create_customer ( data : CustomerInput ) -> Custom…

1 month назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 1 day, 13 hours назад
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

1 day, 13 hours назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

2 months назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

4 months, 1 week назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

5 months, 2 weeks назад @ alexirpan.com
Authentic Imperfection
Authentic Imperfection Authentic Imperfection

* * *I’ve been thinking about the anger surrounding generative AI.

To keep things fair, he took the best human images and best AI images, meaning human art from famous artists, and AI art from prompters skilled at removing obvious tells of image generation.

When people complain about AI slop, I see it as a complaint against the deluge of default style AI images.

We’ve seen this happen in all forms: AI text, AI music, older forms of computer generated content like CGI.

As much as we celebrate imperfection, digital imperfection is a step too far.

8 months назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 4 months, 3 weeks назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

4 months, 3 weeks назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

5 months, 2 weeks назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 2 days, 12 hours назад
Our approach to bioresilience
Our approach to bioresilience Our approach to bioresilience

Today, Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience.

Inside our bioresilience programWe believe society must harness AI’s advancing capabilities to address infectious diseases and prepare for future outbreaks.

With this in mind, we are making our AI models and agents available to trusted partners to support progress across three key areas: prevention, detection and response.

Working in collaboration with governments and global health authorities to advance a diverse range of diagnostic and therapeutic strategies enables the Isomorphic Labs Drug Design Engine’s real-world impact for bioresilience.

To read more about our work and our call for new par…

2 days, 12 hours назад @ deepmind.google
Empowering India’s next generation of innovators with ATL Saathi
Empowering India’s next generation of innovators with ATL Saathi Empowering India’s next generation of innovators with ATL Saathi

A new contribution to Indian Education with Atal Innovation MissionWe believe behind every good student is a great teacher.

That’s why for over 20 years, Google has been dedicated to supporting the education ecosystem by introducing technology into teaching and learning through a teacher-led approach.

With foundational platforms like Google for Education and Google Classroom, we build products tailored to the needs of schools, keeping the teacher in the lead.

To further support the empowerment of educators, our new Google Educator AI Series ensures teachers are equipped with both the tools and the digital skills required for today's classrooms.

We see Gemini as a great tool to enable our pa…

5 days, 8 hours назад @ deepmind.google
Google DeepMind and A24 announce first-of-its-kind research partnership
Google DeepMind and A24 announce first-of-its-kind research partnership Google DeepMind and A24 announce first-of-its-kind research partnership

Today, Google DeepMind and A24 are announcing a first-of-its-kind partnership focused on research.

The collaboration pairs a world-leading research lab with the industry’s most filmmaker-forward studio to help artists develop new workflows and techniques.

This partnership creates a deep research and development collaboration between A24 and Google DeepMind spanning multiple projects over time.

This hands-on collaboration provides Google DeepMind with invaluable feedback and guidance from leading artists.

As A24 and Google DeepMind’s researchers work side-by-side to test, iterate and build, this partnership aims to expand what is possible in the future of entertainment.

2 weeks, 1 day назад @ blog.google
Start building with Nano Banana 2 Lite and Gemini Omni Flash
Start building with Nano Banana 2 Lite and Gemini Omni Flash Start building with Nano Banana 2 Lite and Gemini Omni Flash

Uploading audio references and scene extension is not yet supported in the Gemini API for this model.

Video references up to 3 seconds in duration are accepted by the API schema but are not correctly processed by the model at this time.

Gemini Omni is available in public preview starting today in Google AI Studio and the Gemini API.

Use Nano Banana 2 Lite as a high-speed image generation model, then pass that image as a reference to Gemini Omni Flash to animate it into a high-quality video.

To help you get started we created a few demo apps you can remix that let you experience how you can pair both Nano Banana 2 Lite and Gemini Omni Flash into one workflow.

2 weeks, 4 days назад @ blog.google
Introducing computer use in Gemini 3.5 Flash
Introducing computer use in Gemini 3.5 Flash Introducing computer use in Gemini 3.5 Flash

Making computer use safe in 3.5 FlashTo mitigate some of the prompt injection risks for agents operating in live environments, we use targeted adversarial training for computer use in Gemini 3.5 Flash.

We’re also releasing two optional enterprise safeguard systems that enable enterprises to:Require explicit user confirmation for sensitive or irreversible actions.

Automatically stop tasks if an indirect prompt injection is identified.

Taking a “defense-in-depth” approach, we encourage developers to combine these features with secure sandboxing, human-in-the-loop verification and strict access controls.

We are already seeing customers drive value with computer use.

3 weeks, 3 days назад @ blog.google
Unlocking UK house-building with AI-accelerated planning
Unlocking UK house-building with AI-accelerated planning Unlocking UK house-building with AI-accelerated planning

New UK government AI planning prototype built with Gemini aims to halve the time it takes to process homeowner applicationsAround the world, Governments are exploring how AI can deliver better public services, faster.

The UK is working to build 1.5 million new homes by 2029, but local planning authorities are often slowed down by dense paperwork and administrative backlogs.

To help get Britain building, we’re partnering with the UK government to help radically shorten the time it takes to process householder planning applications.

Following early trials in Barnet, Camden and Dorset, the government plans for the new AI planning tool to be made available to all councils nationally from 2027.

1 month назад @ deepmind.google
Securing the future of AI agents
Securing the future of AI agents Securing the future of AI agents

How we’re securing internal systems against increasingly capable and imperfectly aligned AIAI agents are transforming our relationship with technology.

In the U.S alone, AI agents could create $2.9 trillion in economic value by 2030.

That’s why we developed our AI Control Roadmap: a framework for building and managing the advanced AI we deploy within Google.

Similarly, our AI control system grants AI agents permissions based on their verified behavior, allowing us to build trust through controlled, incremental access.

In our AI Control Roadmap, we map security protocols to measurable milestones in AI capabilities on two critical fronts:

1 month назад @ deepmind.google
DiffusionGemma: 4x faster text generation
DiffusionGemma: 4x faster text generation DiffusionGemma: 4x faster text generation

While the AI research community has explored diffusion-based text generation for years, applying it to large models has remained a challenge.

DiffusionGemma changes this by shifting how models use hardware.

But when run locally for a single user, this word-by-word process leaves your dedicated GPU or TPU underutilized — it spends most of its time simply waiting for the next "keystroke."

By giving the computer's processor a larger chunk of work at once, DiffusionGemma utilizes your hardware to its full potential.

It upgrades your model inference from a single, sequential typewriter to a massive printing press that stamps the entire block of text simultaneously.

1 month, 1 week назад @ blog.google
Investing in multi-agent AI safety research
Investing in multi-agent AI safety research Investing in multi-agent AI safety research

Scaling AI Safety Research for a Multi-Agent WorldFor the past decade, we’ve focused on making individual AI models more capable, helpful and safe.

The funding call focuses on the study of how large-scale multi-agent AI systems behave as a group, and how we can provide frameworks to understand and mitigate against potential risks.

Scaling the frontier of multi-agent safety researchAlthough foundational frameworks for multi-agent safety exist, the rapid evolution of these systems requires an immediate, large-scale expansion of research.

A collaborative call to actionNo single lab can solve multi-agent safety alone.

Building realistic, reproducible environments to evaluate, compare and accele…

1 month, 1 week назад @ deepmind.google
Fluid, natural voice translation with Gemini 3.5 Live Translate
Fluid, natural voice translation with Gemini 3.5 Live Translate Fluid, natural voice translation with Gemini 3.5 Live Translate

Today, we’re taking our next step with the release of Gemini 3.5 Live Translate, our latest audio model for live speech-to-speech translation.

The model automatically detects 70+ languages and generates smooth, natural-sounding translated speech that preserves the speakers' intonation, pacing and pitch.

Unlike turn by turn systems that wait for the speaker to finish speaking before responding, 3.5 Live Translate generates speech continuously, balancing the trade-off between waiting for context to improve quality and translating immediately to stay in sync with the speaker.

Gemini 3.5 Live Translate is rolling out starting today across Google products:For developers in public preview via the…

1 month, 1 week назад @ blog.google
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Introducing Gemma 4 12B: a unified, encoder-free multimodal model Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Today, we are introducing Gemma 4 12B, our latest model designed to bring agentic multimodal intelligence directly to laptops.

Here’s an overview of what makes Gemma 4 12B unique:Novel unified architecture: No multimodal encoders.

Advanced reasoning: Benchmark performance nearing our 26B model, unlocking powerful multi-step reasoning and agentic workflows.

Benchmark performance nearing our 26B model, unlocking powerful multi-step reasoning and agentic workflows.

Small enough to run locally on consumer laptops with 16GB of RAM, it unlocks powerful multimodal and agentic experiences right on your machine.

1 month, 1 week назад @ blog.google
Powering the future of robotics in Europe
Powering the future of robotics in Europe Powering the future of robotics in Europe

That’s why we’re launching the Google DeepMind Accelerator: Robotics, a three-month program for early-stage robotics startups across Europe.

They’ll have access to our AI stack, technical expertise and Gemini robotics models.

Extend Robotics ( United Kingdom ): Provides teleoperation software and data pipelines that help train and fine-tune foundation models for real-world robotics applications.

Generative Bionics ( Italy ): Amplifies human potential by developing humanoid robots based on physical AI, developed in Europe but built to scale globally.

): Amplifies human potential by developing humanoid robots based on physical AI, developed in Europe but built to scale globally.

1 month, 1 week назад @ blog.google
Measuring the impact of learning with AI in Sierra Leone and beyond
Measuring the impact of learning with AI in Sierra Leone and beyond Measuring the impact of learning with AI in Sierra Leone and beyond

The results from this pre-registered trial suggest that AI can be a powerful pedagogical partner — not by replacing teachers, but by augmenting their reach.

Students using Guided Learning saw a gain of +0.258 standard deviations in their math scores compared to the control group.

In practical terms, this represents roughly 1.2 to 1.7 years of typical learning progress achieved within the eight-week trial.

To further understand the impact of Guided Learning on student learning, we are conducting a series of additional pre-registered RCTs globally.

Additionally, our support of the Global AI for Learning Alliance (GAILA) will accelerate these commitments and others through collective action.

1 month, 1 week назад @ deepmind.google
We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks
We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks

The Asia-Pacific region is a global engine for economic growth, but it's also highly vulnerable to climate change.

While green technologies are gaining momentum, a recent report shows they aren’t scaling fast enough to keep up with the region’s rising environmental risks.

Selected organizations will receive expert mentorship, tailored support and help integrating frontier AI and science AI models from Google AI experts into their projects or products.

If you're working on climate solutions, we want to help you scale your work.

The program kicks off with an in-person bootcamp in Singapore, and you can learn more and register your interest today.

1 month, 4 weeks назад @ blog.google
Fast-tracking genetic leads to reverse cellular aging
Fast-tracking genetic leads to reverse cellular aging Fast-tracking genetic leads to reverse cellular aging

Biologists Omar Abudayyeh and Jonathan Gootenberg are using Co-Scientist to help them blast through both.

Their lab runs huge genetic screens that flip thousands of genes on or off then reads how cells respond to these changes.

Co-Scientist is helping on two fronts.

Second, Co-Scientist speeds up the follow-through.

Having Co-Scientist analyse their screening data alongside the literature, that work is slashed to just a few days.

2 months назад @ deepmind.google
Google
последний пост 1 day, 5 hours назад
13 hands-on demos to build on Gemini Enterprise Agent Platform
13 hands-on demos to build on Gemini Enterprise Agent Platform 13 hands-on demos to build on Gemini Enterprise Agent Platform

Earlier this year, we introduced Gemini Enterprise Agent Platform, where you can build, scale, govern, and optimize agents.

Today, we’re sharing 13 demos that walk you through what Agent Platform can do.

The ambient expense agent codelab is the most complete "Agent Platform in action" demo in the set.

Deploy a stateful data science agent to Agent Runtime (formerly known as Agent Engine).

You deploy a multi-tool ADK agent on Agent Runtime that calls MCP servers on Cloud Run through Agent Gateway.

1 day, 5 hours назад @ cloud.google.com
Google is a Leader and positioned furthest in Vision and highest in Execution in the 2026 Gartner® Magic Quadrant™ for Conversational AI Platforms
Google is a Leader and positioned furthest in Vision and highest in Execution in the 2026 Gartner® Magic Quadrant™ for Conversational AI Platforms Google is a Leader and positioned furthest in Vision and highest in Execution in the 2026 Gartner® Magic Quadrant™ for Conversational AI Platforms

Building the next generation of customer experiences with Gemini Enterprise for Customer ExperienceEnterprise customer experiences are entering a new era.

Today, Gemini Enterprise for Customer Experience brings these capabilities together to give your customers a frictionless experience.

Built for production AIAt the center of Gemini Enterprise for Customer Experience is CX Agent Studio, Google’s platform for building intelligent customer experience agents.

This is why Gemini Enterprise for Customer Experience and CX Agent Studio run natively on Google Cloud’s complete, first-party AI stack.

Gartner research publications consist of the opinions of Gartner's research organization and should …

2 days, 2 hours назад @ cloud.google.com
Cloud CISO Perspectives: How AI leverages deep context as the defender’s advantage
Cloud CISO Perspectives: How AI leverages deep context as the defender’s advantage Cloud CISO Perspectives: How AI leverages deep context as the defender’s advantage

This ensures autonomy under human supervision, empowering engineering and security teams to eliminate backlogs and secure the software development lifecycle without sacrificing speed.

The key to countering this is enforcing Zero Trust for AI, and directing teams toward approved architectures with proper governance.

That means securing AI infrastructure requires building from the ground up, and not bolting on.

Fight AI with AI.

Learn more about how to secure your software lifecycle with Google AI Threat Defense.

2 days, 5 hours назад @ cloud.google.com
Three lessons in accelerating foundation model upgrades
Three lessons in accelerating foundation model upgrades Three lessons in accelerating foundation model upgrades

For most engineering teams, upgrading to a new model checkpoint means months of manual toil to verify performance.

And the industry is moving at breakneck pace – since 2023, we’ve announced six major model evolutions, bringing us to Gemini 3.5 today.

As part of that, our team built an agentic workflow that completes model upgrades in hours instead of months.

Their goal was to migrate to the latest out-of-the-box foundation model, guided purely by prompt engineering.

Build an agentic loop: You can use the Agent Development Kit within Gemini Enterprise Agent Platform to create your agent.

2 days, 5 hours назад @ cloud.google.com
IDC: Why the right networking approach is foundational to agentic AI
IDC: Why the right networking approach is foundational to agentic AI IDC: Why the right networking approach is foundational to agentic AI

Editor’s note: Today we hear from IDC on the results of its 2026 AI in Networking Special Report Survey exploring the enterprises' concerns about networking infrastructure to support the rise of agentic AI in their organizations.

Agentic AI specifically heightens these concerns by introducing more distributed and dynamic interactions across applications, services, APIs, tools, and data sources.

The right platform for agentic AI should be open, flexible, and able to evolve.

Businesses must meet their AI objectives while carefully managing dynamic agentic AI systems.

As agentic AI systems continue to evolve, the demands they place are unlikely to be addressed through best-of-breed point solut…

3 days, 5 hours назад @ cloud.google.com
Why AI apps fail in production (And how Google solved it)
Why AI apps fail in production (And how Google solved it) Why AI apps fail in production (And how Google solved it)

We are living in the golden age of the weekend AI side project.

Thanks to agentic engineering and LLMs, the time to go from a blank IDE to a functional local application has dropped from quarters to hours.

The data is sobering: only 5% of AI prototypes make it to production; the other 95% fall into the validation abyss.

For developers, watching people on social media ship lightning-fast AI deployments while you’re stuck in endless validation loops is maddening.

But as AI engineering leader Addy Osmani points out in our premiere of Emergent, unconstrained agentic orchestration inside an enterprise introduces an unpredictable blast radius.

3 days, 5 hours назад @ cloud.google.com
Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software
Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software

Gemini Enterprise: A unified system for the agentic eraA great foundation model is only as valuable as an organization's ability to safely put it to work.

Developers can build agents using Gemini 3.5 Flash on the Gemini Enterprise Agent Platform, or use it in your projects in Google AI Studio and Antigravity.

Get startedDownload the IDC MarketScape: Worldwide Foundation Model Software 2026 Vendor Assessment excerpt to learn why organizations are choosing Google Cloud.

Explore Gemini Enterprise today, or speak to your Google Cloud account representative to schedule a hands-on technical workshop.

IDC MarketScape: Worldwide Foundation Model Software 2026 Vendor Assessment, Doc #US54427726, Jul…

4 days, 3 hours назад @ cloud.google.com
Claude at scale on Google Cloud: Frontier AI, built for enterprise production
Claude at scale on Google Cloud: Frontier AI, built for enterprise production Claude at scale on Google Cloud: Frontier AI, built for enterprise production

Running frontier AI in production is demanding — accelerators to manage, latency to hold steady across continents, regulated data to keep in-region, and long-context requests to serve reliably.

Claude on Google Cloud is built for exactly this.

In our case, Claude brings the reasoning, and Google Cloud brings the managed infrastructure, global reach, and compliance posture that enterprises already run on.

Managed infrastructure that frees engineering timeClaude on Google Cloud runs on fully managed infrastructure, so enterprise teams ship features instead of building clusters.

Invoking Claude is operationally identical to invoking any other Google Cloud service: the same IAM policies, the sa…

4 days, 5 hours назад @ cloud.google.com
Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs
Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs

Furthermore, k8s-aibom treats the Kubernetes cluster state as a pure functional input: Identical cluster inputs produce byte-identical ML-BOM documents.

Commercial AI security platforms extend the picture with cloud-native posture management, but typically through external scanning shaped around vendor-specific data models.

Standard monitoring tools indicate that a container is running, but can’t prove whether an AI model was explicitly configured by a platform engineer or dynamically pulled by an autonomous script at runtime.

Getting startedIt’s rare that a technical solution like k8s-aibom can help mitigate the multi-faceted problem of shadow AI, impacting CISOs, governance, risk, and com…

5 days, 5 hours назад @ cloud.google.com
Frontier and Center: Who evaluates the evaluations?
Frontier and Center: Who evaluates the evaluations? Frontier and Center: Who evaluates the evaluations?

Yet this is exactly how we evaluate AI agents.

Along the way, the added fidelity exposed some deeper issues with the quality of emergent evaluation cases themselves.

Difficulty, measuredWhen it comes to retrieval, evaluation cases are often stratified into tiers of difficulty.

What we need is a rigorous approach that can modulate the difficulty of evaluation cases.

Therefore, we can adjust the difficulty of evaluation cases by adding or removing terms with varying levels of informative power.

1 week, 1 day назад @ cloud.google.com
Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud
Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud

Visit the blog to read more how BASF used AlphaEvolve to improve their existing planning and forecasting models by over 80%.

Visit the blog to read more about how Jetbrains used AlphaEvolve to improve their IDE performance by over 15-20%.

Kinaxis: Improving optimization and forecasting systems"Kinaxis researchers have used AlphaEvolve to materially improve both the speed and quality of highly mature forecasting and optimization algorithms.

Visit the blog to read more about how Klarna used AlphaEvolve to double Training Speed and improve performance for their foundational models.

Visit the blog to read more about how Schröedinger used AlphaEvolve to quadruple the speed of molecular discovery.

1 week, 2 days назад @ cloud.google.com
A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace
A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace

Here’s an overview of these architectural elements:Customer project: Where users discover agents via the dedicated Agent Marketplace category within Google Cloud Marketplace and interact with these agents through the Gemini Enterprise app.

Step 2: Review the organizational requirements to sell on MarketplaceJoin the Google Cloud Partner Network : If you're new to offering your solutions on Marketplace, join the Google Cloud Partner Network.

Nominate your agent for Google Cloud Marketplace by contacting your Google Cloud representative.

A2A Agent Card: Create an Agent Card, a JSON file declaring capabilities (skills), authentication methods, and service endpoints.

A2A agent cardTo list your …

1 week, 4 days назад @ cloud.google.com
Report: 83% of organizations need to upgrade their infrastructure to support agentic AI
Report: 83% of organizations need to upgrade their infrastructure to support agentic AI Report: 83% of organizations need to upgrade their infrastructure to support agentic AI

In this blog, we lay out the core insights from our research on how leading organizations are rethinking their infrastructure to build resilient, fluid foundations.

To fix this, organizations need fluid compute — the ability to dynamically match the right silicon to the right task while minimizing operational overheads.

For orchestration: General-purpose compute powered by CPUs is emerging as a critical component for driving AI control plane operations.

But as agentic AI scales, organizations are facing a new challenge: agent sprawl.

Agent Gateway gives you the visibility you need to see exactly how agents are sharing data.

1 week, 4 days назад @ cloud.google.com
20 questions for the Agentic Enterprise (and how Agent Platform can help)
20 questions for the Agentic Enterprise (and how Agent Platform can help) 20 questions for the Agentic Enterprise (and how Agent Platform can help)

That’s why we built Gemini Enterprise Agent Platform.

When integrated with Agent Platform, Model Armor intercepts prompts before they reach Gemini models, and intercepts responses before your application receives them.

This is where Agent Platform Threat Detection (part of Security Command Center) comes in.

Example: Build with Agents CLI hereGet started todayBy tackling these 20 questions early, you can build agents that actually do real work for your business — without keeping your security and operations teams up at night.

Get started with Gemini Enterprise Agent Platform here.

1 week, 4 days назад @ cloud.google.com
Drive proactive security, prioritize risks with Google Threat Intelligence and Wiz ASM
Drive proactive security, prioritize risks with Google Threat Intelligence and Wiz ASM Drive proactive security, prioritize risks with Google Threat Intelligence and Wiz ASM

To help you be more proactive by matching your real-world exposures with real-time adversary activity, we’ve begun integration efforts between Google Threat Intelligence and Wiz Attack Surface Management (ASM).

By connecting exposure and validated exploitable risks directly to real-time threat intelligence, we can help you detect and prioritize external-facing exploitable issues and uncover logic-driven vulnerabilities with AI scanning at the speed needed for today’s defenses.

This allows you to shift to a strategy that prioritizes actions based on the real-world threats that pose the greatest risks to your organization.

Combining these two perspectives on threats can help you move from rea…

1 week, 4 days назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 5 days, 5 hours назад
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

5 days, 5 hours назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

1 week, 2 days назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

1 week, 3 days назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

2 weeks, 4 days назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

2 weeks, 5 days назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

3 weeks, 2 days назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

3 weeks, 2 days назад @ microsoft.com
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis

At a glance Talos is an open-source tool for automated, iterative reanalysis of genomic data in rare disease.

Deployed across a prospective cohort of almost 5,000 undiagnosed patients, Talos delivered 241 new diagnoses (5.1% additional yield).

On monthly iterative cycles, analysts only needed to review one new variant per 200 patients, demonstrating that frequent, systematic reanalysis can be run sustainably.

Why genome reanalysis mattersGenomic testing has transformed the diagnosis of rare disease, but even with this advancement, more than half of patients remain undiagnosed after their first test.

Looking aheadTalos reframes genomic reanalysis from a rare, labor-intensive event into a con…

3 weeks, 3 days назад @ microsoft.com
Ire identifies another LOTUSLITE specimen
Ire identifies another LOTUSLITE specimen Ire identifies another LOTUSLITE specimen

At a glance Project Ire identifies a LOTUSLITE variant that shares TTPs (tools, tactics, procedures) with the public family but none of its indicators of compromise (IOC).

On Ire’s calibrationOne noteworthy observation in Ire’s report (opens in new tab) is worth highlighting first.

The Ire report does not surface a matching entry-point name, but it identifies that the behavioral shape is the same.

Ire never named LOTUSLITE in its report or chain of evidence.

Ire described the behavior precisely enough to make the mapping straightforward of this sample to LOTUSLITE.

1 month назад @ microsoft.com
Data Formulator 0.7: AI-powered data analytics for enterprise data
Data Formulator 0.7: AI-powered data analytics for enterprise data Data Formulator 0.7: AI-powered data analytics for enterprise data

At a glance Data Formulator 0.7 is an open-source AI-powered system for enterprise data analytics that combines data connectivity, agent-guided exploration, and visualization refinement in a shared workspace.

Enterprise teams increasingly rely on AI systems for analytics, but enterprise data workflows are often fragmented across storage systems and tools.

Listen now Opens in a new tabConnecting enterprise data with Data ConnectorsData Formulator helps teams bring enterprise data into an AI-ready workspace without needing to rebuild the same connections for every source of data.

Data Connectors provide persistent connections between enterprise data sources and Data Formulator, allowing analy…

1 month, 3 weeks назад @ microsoft.com
Extending Human Intelligence Through AI
Extending Human Intelligence Through AI Extending Human Intelligence Through AI

At a glance Modern AI systems are powerful not because they replicate human intelligence, but because they presuppose it, by extending structures already present in human cognition and language.

Understanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems.

Rather than asking whether AI systems are becoming intelligent in the human sense, these approaches ask a more basic question: What if AI systems work because they rely on structures that are rooted in human cognition?

In our recent paper, The Origins of Artificial Intelligence in Natural Intelligence, we argue that modern AI systems are best understood nei…

1 month, 3 weeks назад @ microsoft.com
MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models

Built as the next generation of Magentic-UI, it combines a redesigned app with a harness optimized for small models.

MagenticBrain and Fara1.5 are small models designed for orchestration and computer-use tasks, respectively.

Together, these releases explore how far agentic performance can be pushed with smaller models, codesigned tools, and an optimized execution harness.

Today, Microsoft Research AI Frontiers releases MagenticLite (opens in new tab), an experimental agentic application designed for small models.

The result is an agent that runs efficiently, keeps data on the user’s machine, and supports a broad range of agentic tasks.

1 month, 4 weeks назад @ microsoft.com
Vega: Zero-knowledge proofs for digital identity in the age of AI
Vega: Zero-knowledge proofs for digital identity in the age of AI Vega: Zero-knowledge proofs for digital identity in the age of AI

Vega puts these building blocks together into a single proof system.

The hashing problem, and how folding solves itA credential proof must do two expensive things: hash the credential bytes with SHA-256 and verify the issuer’s digital signature.

Making it zero-knowledge, cheaplyA proof system needs to be zero-knowledge: the verifier should learn nothing beyond the claim being proved.

Device bindingA zero-knowledge credential proof is only useful if it is tied to the person holding the credential.

The proof system powering Vega is already available as the open-source spartan2 (opens in new tab) project on GitHub.

1 month, 4 weeks назад @ microsoft.com
Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability
Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability

Our recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated workflows.

Using a controlled evaluation methodology, we examine how well information is preserved across these extended workflows.

We use chained transformation-and-inversion tasks that evaluate whether semantic content is preserved accurately across extended delegated workflows.

Azure AI Foundry Labs Get a glimpse of potential future directions for AI, with these experimental technologies from Microsoft Research.

At the same time, the findings should not be interpreted as evidence that AI systems lack practical value in real-world work today.

2 months назад @ microsoft.com
mimalloc: A new, high-performance, scalable memory allocator for the modern era
mimalloc: A new, high-performance, scalable memory allocator for the modern era mimalloc: A new, high-performance, scalable memory allocator for the modern era

mimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free.

The mimalloc memory allocator was initially designed in 2020 as a fast allocator for the state-of-the-art Lean (opens in new tab) and Koka (opens in new tab) programming languages developed at RiSE, both of which use novel compiler-guided reference counting (see Perceus).

ja .LBB0_generic leaq 7 ( %rsi ), %rax ; round to sizeof(void*) andq $-8 , %rax movq 232 ( %rdi , %rax ), %rcx ; rcx = heap->small_pages[index] movq 8 ( %rcx ), %rax ; block = rax = page->free testq %rax , %rax ; block == NULL?

Thus, mimalloc has three free lists per (64 KiB) mimalloc page, and effectively that …

2 months назад @ microsoft.com
MIT AI MIT AI
последний пост 1 day, 4 hours назад
Following the questions where they lead
Following the questions where they lead Following the questions where they lead

Ever since she was a child playing on her family’s farmland in Wisconsin, Bailey Flanigan was guided by her own selective, yet wide-ranging, curiosity.

“I found myself unmotivated to take all the AP [advanced placement] classes for the sake of it.

So Flanigan moved toward public health, where she researched microfluidic devices for HIV detection that could be used in low-resource settings.

After graduating from UW-Madison, Flanigan worked as a predoctoral research assistant in economics at Princeton.

“I feel so lucky to be studying these questions from within both political science and EECS, because I have the freedom to explore both the political and technical substance of tools for more d…

1 day, 4 hours назад @ news.mit.edu
A better way to turn 2D designs into 3D models for rapid prototyping
A better way to turn 2D designs into 3D models for rapid prototyping A better way to turn 2D designs into 3D models for rapid prototyping

The system generates new data based on the model’s abilities as it attempts to convert a 2D image into a CAD program.

“Nearly every physical product around us, from airplanes to appliances, begins its life as a CAD model.

For guesses that are nearly correct, GIFT adjusts them to become successful solutions.

The CAD models generated by VLMs using GIFT were better aligned with the shapes of ground-truth models.

In the future, the researchers want to expand GIFT so the framework can teach models to generate CAD programs that improve the performance and manufacturability of 3D models.

2 days, 17 hours назад @ news.mit.edu
3 Questions: Neural transparency and the future of AI design
3 Questions: Neural transparency and the future of AI design 3 Questions: Neural transparency and the future of AI design

Q: Your paper introduces “neural transparency,” a way to let everyday users peek inside an AI’s neural networks before their chatbot ever says a word.

“Neural transparency” means giving people something like a brain scan for AI.

Our study suggests that people have a blind spot when designing personalized AI.

In previous research, we documented cases of psychological harm associated with interactions with AI chatbots.

AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step.

3 days, 1 hour назад @ news.mit.edu
Helping AI models to meet the real world
Helping AI models to meet the real world Helping AI models to meet the real world

“In a sense, with a small amount of resource, you have to do a lot of heavy lifting,” he says.

“My interest was: How does one design such graphical models for generic, tabular data?” he says.

And each of the products that you manufacture has lots of small pieces that come from different parts of the world.

Shah adds that Celonis has specialized in digitizing and automating operations for more than 1,400 large companies around the world.

“A narrower focus comes with sharper technology,” he says, “but it’s broad enough that it’s very valuable.”Shah adds, “The recent buzzword that’s become pertinent in the modern AI popular press is a ‘world model.’ In a sense, this is trying to build the ente…

4 days, 1 hour назад @ news.mit.edu
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering

“The JARVIS challenge showed that AI can substantially accelerate safety-critical hardware engineering, but engineering judgment remains the decisive differentiator.

Manufacturing — not engineering design or analysis — remained the fundamental rate-limiting step,” says Professor Zolti Spakovszky, director of the MIT Gas Turbine Laboratory.

In weekly progress reviews, they would critically evaluate the student progress and assess how the students were using AI.

The 811 team had been resistant to using AI throughout the competition, trusting instead to their fundamentals and teamwork.

From the start of the JARVIS Challenge, younger students used Parley more frequently and cleverly, while the …

4 days, 3 hours назад @ news.mit.edu
How MIT students are helping to prevent cyberattacks
How MIT students are helping to prevent cyberattacks How MIT students are helping to prevent cyberattacks

To counter such threats, Lecturer Jungwoo Chun and Ford Professor of Urban and Environmental Planning Lawrence Susskind launched the MIT Cybersecurity Clinic in 2019.

Much like a legal or medical clinic, the course doubles as hands-on training for students and a pro-bono service to at-risk communities.

After completing instructional modules and passing a certification exam, students are assigned in teams to a client.

The Cybersecurity Clinic aims to round out the knowledge of students from every discipline.

In either case, Susskind and Chun check in periodically with clients for at least two years following each engagement.

5 days, 2 hours назад @ news.mit.edu
AI agents create virtual playgrounds to help robots get crucial training data
AI agents create virtual playgrounds to help robots get crucial training data AI agents create virtual playgrounds to help robots get crucial training data

It turns out that AI agents, or semi-autonomous programs that “think” and complete well-defined tasks, could help produce the lifelike virtual settings that robots need.

The most telling test: they dropped a pretrained robot policy — an AI controller trained largely on real-world data, which had never seen a SceneSmith scene — into the generated environments.

The team also teleoperated robots through the virtual spaces, guiding them to open cabinets, put away bottles, and navigate between rooms.

Behind the scenesThe agents that SceneSmith uses each have a well-defined role in the generative process, fleshing out scenes in stages.

It can take multiple hours to produce a single scene because …

5 days, 2 hours назад @ news.mit.edu
New method aims to keep kids safe from illegal AI-generated content
New method aims to keep kids safe from illegal AI-generated content New method aims to keep kids safe from illegal AI-generated content

When tested, the auditing procedure identified model variations that had been specialized to generate CSAM with 100 percent accuracy.

Auditing adaptationsRecent techniques have made it easier for users to specialize a generative AI model for their task through a process known as fine-tuning.

This has led to a wave of new generative AI model variants for a variety of purposes, like producing watercolor images that mimic an artistic movement.

They tested their method on variations of three types of models, comparing the results to ground-truth data from LoRA adaptors known for generating CSAM, other harmful images, and safe content.

Their method was 100 percent accurate in identifying models …

5 days, 17 hours назад @ news.mit.edu
Tiny robot boats build floating structures
Tiny robot boats build floating structures Tiny robot boats build floating structures

A team of MIT researchers sees it as a dynamic, Lego-like construction site.

Each robot, about the size of a dinner plate at 21 centimeters square, is a self-contained vessel with its own thrusters, sensors, and magnetic latches.

In that final mode, called collective transport, a planner charts a trajectory for the whole structure and each robot computes its own contribution.

“Our boats become more stable by joining together, like the ant raft, if you have waves or currents,” Hagemann says.

The team thanks MIT Sea Grant and Professor Michael Triantafyllou for providing the test tank.

1 week, 2 days назад @ news.mit.edu
How novice coders can develop AI programs for military applications
How novice coders can develop AI programs for military applications How novice coders can develop AI programs for military applications

We both wanted to understand better where and how AI could be used by nontechnical users in the military."

During the project, Lynch completed several professional development courses in AI and familiarized himself with both military and nonmilitary uses of the technology.

For the basis for his code generation, he used the paid models of three AI chatbots: Anthropic's Claude, OpenAI's ChatGPT, and Google's Gemini.

For example, he often encountered difficulties with the AI chatbots lacking hierarchical focus and modifying unrelated code sections.

Although AI can generate significant amounts of functional code, code review remains a bottleneck in this space.

1 week, 4 days назад @ news.mit.edu
Jesse Thaler named director of the Laboratory for Nuclear Science
Jesse Thaler named director of the Laboratory for Nuclear Science Jesse Thaler named director of the Laboratory for Nuclear Science

Professor Jesse Thaler has been named director of the MIT Laboratory for Nuclear Science (LNS), effective Aug. 1.

Thaler is a theoretical particle physicist who combines techniques from quantum field theory and machine learning to address outstanding questions in fundamental physics.

Mike Williams, professor of physics, will succeed Thaler as IAIFI director.

Established in 1946 to support nuclear and particle physics, LNS now encompasses research spanning cosmology, gravity, field theory, and quantum information science.

As head of LNS, Thaler will also oversee his home center of CTP-LI, which last year received a donation from the Leinweber Foundation to establish a network of theoretical …

1 week, 4 days назад @ news.mit.edu
Toward a future that preserves benefits of neurotechnology for all
Toward a future that preserves benefits of neurotechnology for all Toward a future that preserves benefits of neurotechnology for all

Sava’s concept was inspired by an internship at IBM, where she worked on a project with the PACE Center in London.

As advanced medical technology gets closer to hitting consumer markets, the need for guardrails on protected usage should increase.

What might begin as a neural implant to aid in communication could become a device used to police one’s innermost thoughts.

From its inception, the competition has consistently attracted undergraduate and graduate students from across a wide range of disciplines.

The judges also named four honorable mentions, each of whom received a $500 cash prize.

1 week, 5 days назад @ news.mit.edu
MIT in the media: Innovating and educating for the next 250 years of America
MIT in the media: Innovating and educating for the next 250 years of America MIT in the media: Innovating and educating for the next 250 years of America

Inspired by MIT’s motto, “mens et manus” (mind and hand), she shared: “We really want students to be able to use physical AI.

The economic impact of MIT on this country is equivalent to the 14th largest GDP in the world.

She further highlighted MIT for America, an initiative expanding access to calculus, a required course for institutions such as MIT, in under-resourced high schools nationwide.

“What we [ASU] learn from MIT is, where’s the edge of technology,” said Crow.

Kornbluth expressed her hope for MIT to continue its longstanding tradition of research and education in service of the nation’s next 250 years.

2 weeks, 3 days назад @ news.mit.edu
Q&A: What is agentic AI today, and what do we want it to be?
Q&A: What is agentic AI today, and what do we want it to be? Q&A: What is agentic AI today, and what do we want it to be?

A November 2025 report by MIT Sloan School of Management and Boston Consulting Group found that 35 percent of surveyed businesses had already deployed AI agents, while another 44 percent planned to implement agentic AI soon.

Q: What is agentic AI and how is it different from generative AI models like ChatGPT and Claude?

A: Agentic AI is AI that takes actions in the world.

Q: What are some promising applications of agentic AI?

Q: What does the future hold for agentic AI?

2 weeks, 4 days назад @ news.mit.edu
Inaugural Music Technology Research Showcase celebrates work of new graduate program’s initial students
Inaugural Music Technology Research Showcase celebrates work of new graduate program’s initial students Inaugural Music Technology Research Showcase celebrates work of new graduate program’s initial students

The MIT Music Technology and Computation (MTC) Graduate Program — launched in fall 2024 as a collaboration between the Music and Theater Arts Section in the School of Humanities, Arts, and Social Sciences (SHASS), and the School of Engineering (SoE) — presented its inaugural MIT Music Technology Research Showcase on May 13.

Each scholar presented inspiring exemplars of artful engineering that reflected the broader and burgeoning music technology scene at MIT.

“The MIT Music Technology and Computation Graduate Program taught me so much about the possibilities at the intersection of STEM and the arts," she says.

What does it mean to build music technology in this context?

All considered, the …

2 weeks, 5 days назад @ news.mit.edu
Berkeley AI
последний пост 1 week, 4 days назад
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

1 week, 4 days назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

2 weeks, 3 days назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

2 months, 1 week назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

2 months, 4 weeks назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

4 months, 1 week назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

6 months, 1 week назад @ bair.berkeley.edu
RL without TD learning
RL without TD learning RL without TD learning

RL without TD learningIn this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer.

We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning.

There are two classes of algorithms in RL: on-policy RL and off-policy RL.

We compared TRL with $n$-step TD learning with different values of $n$, from $1$ (pure TD) to $\infty$ (pure MC).

I still think one of the most important problems in RL (and even in machine learning) is to find a scalable off-policy RL algorithm.

8 months, 2 weeks назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 1 day, 2 hours назад
Transform your sales organization with Amazon Quick: your new agentic AI teammate
Transform your sales organization with Amazon Quick: your new agentic AI teammate Transform your sales organization with Amazon Quick: your new agentic AI teammate

Amazon Quick is an AI assistant for work that turns questions into answers, answers into actions, and actions into outcomes.

Companies like 3M, AWS Global Sales, and Amazon are already using Quick to unlock their sellers’ full potential.

See how AWS Global Sales transforms insights delivery, AWS sales teams stay current on new launches, and Amazon’s journey deploying Quick Suite across thousands of users.

Start by creating a dashboard using Amazon Quick Sight, the business intelligence service within Amazon Quick that’s used to build interactive visualizations from your connected data sources.

Explore the Amazon Quick for Sales page to learn more, or watch Quick in action in the Sales demo …

1 day, 2 hours назад @ aws.amazon.com
Introducing Mobile Layout for Amazon Quick dashboards
Introducing Mobile Layout for Amazon Quick dashboards Introducing Mobile Layout for Amazon Quick dashboards

Mobile Layout for Amazon Quick Free Form dashboards solves this by automatically rendering dashboards as a single-column, touch-optimized view that fills the device screen.

Mobile Layout for Free Form dashboards is now available in each supported AWS Region.

Mobile layout explainedMobile Layout is a rendering mode for Free Form dashboards.

In the Quick mobile app, readers can use a built-in view switcher to toggle between Mobile view (continuous scroll) and Desktop view (the original rendering).

Authors of Free Form dashboards benefit automatically because their published dashboards already render in Mobile Layout on smaller screens.

1 day, 4 hours назад @ aws.amazon.com
How Smartsheet built a remote MCP server on AWS
How Smartsheet built a remote MCP server on AWS How Smartsheet built a remote MCP server on AWS

The architecturally critical AWS services in the data path are:The detailed architecture flow is as follows:AI clients to API gateway layer to MCP Server: Requests pass through an API gateway layer (AWS WAF, AWS Shield, AWS Application Load Balancer, and OAuth validation) before reaching the MCP server on AWS Fargate.

MCP Server to Domain Services: The MCP server calls Smartsheet’s domain services through their APIs for transactional operations.

MCP Server to Intelligence Layer: The MCP server queries the Intelligence Layer built on Amazon Neptune and Databricks for cross-project agentic insights.

To handle and validate this pattern, Smartsheet built the MCP server to run on AWS Fargate for…

1 day, 5 hours назад @ aws.amazon.com
Build enterprise search for agents with Amazon Bedrock Managed Knowledge Base
Build enterprise search for agents with Amazon Bedrock Managed Knowledge Base Build enterprise search for agents with Amazon Bedrock Managed Knowledge Base

Syngenta Group uses Bedrock Managed Knowledge Bases to enable employees to create knowledge bases on demand, syncing data from SharePoint and Confluence for internal knowledge search and agentic RAG applications.

Rather than building separate pipelines for each format, Managed Knowledge Base provides fully managed parsing that automatically selects the right strategy per content type.

Where Managed Knowledge Base fitsAWS offers multiple options for building the knowledge retrieval that grounds agents, generative AI applications, and RAG pipelines.

The multimodal document parser, managed embedding model, managed re-ranker, and managed orchestration LLM are all included at no extra cost.

Conc…

2 days назад @ aws.amazon.com
Introducing Grok on Amazon Bedrock
Introducing Grok on Amazon Bedrock Introducing Grok on Amazon Bedrock

Why Grok 4.3 is a great fit for agentic and reasoning workloadsAccording to xAI, Grok 4.3 is built for enterprise work where accuracy matters.

How you access Grok 4.3 on Amazon BedrockGrok 4.3 runs on Mantle, and accessing it differs from models that use the Amazon Bedrock Runtime API.

The model ID is xai.grok-4.3 :from openai import OpenAI client = OpenAI( api_key="", base_url="https://bedrock-mantle.us-west-2.api.aws/openai/v1", ) response = client.chat.completions.create( model="xai.grok-4.3", messages=[ {"role": "user", "content": "In one sentence, what is Amazon Bedrock?"}

But if you generated a long-term Amazon Bedrock API key for exploration, delete it from the Amazon Bedrock console…

2 days, 2 hours назад @ aws.amazon.com
Building a restaurant telephony AI host with Amazon Bedrock AgentCore and Amazon Nova 2 Sonic
Building a restaurant telephony AI host with Amazon Bedrock AgentCore and Amazon Nova 2 Sonic Building a restaurant telephony AI host with Amazon Bedrock AgentCore and Amazon Nova 2 Sonic

The system uses Amazon Bedrock AgentCore to host and run the agent and Amazon Nova 2 Sonic for real-time speech, connected to a restaurant backend through the Model Context Protocol (MCP).

Amazon Elastic Container Registry (Amazon ECR), AWS CodeBuild, and Amazon Simple Storage Service (Amazon S3) build and store the agent container image.

AgentCore Runtime lists and calls the available tools from AgentCore Gateway using the MCP protocol.

Deploy in a Region where Amazon Nova 2 Sonic, Amazon Chime SDK PSTN Audio, and AgentCore Runtime are all available.

Processing voice with Amazon Bedrock AgentCore and Amazon Nova 2 SonicThe agent runs on AgentCore Runtime.

2 days, 5 hours назад @ aws.amazon.com
Built Technologies builds an AI-powered document intelligence solution on AWS to power agents across real estate finance
Built Technologies builds an AI-powered document intelligence solution on AWS to power agents across real estate finance Built Technologies builds an AI-powered document intelligence solution on AWS to power agents across real estate finance

Built Technologies, a real estate finance software provider, processes over $500B in real estate projects.

The company deployed an AI-powered document processing engine on Amazon Bedrock and the AWS Intelligent Document Processing (IDP) Accelerator.

Why real estate finance needs AI-powered document intelligenceReal estate finance is document-heavy, fragmented, and highly contextual.

By using the AWS Intelligent Document Processing Accelerator and Amazon Bedrock, Built developed a reusable document intelligence solution that demonstrates how real estate finance companies can modernize document processing workflows.

To learn more about using Amazon Bedrock for document processing, see Documen…

3 days, 3 hours назад @ aws.amazon.com
Agentic vision: Building visual intelligence with Amazon Bedrock and MCP servers
Agentic vision: Building visual intelligence with Amazon Bedrock and MCP servers Agentic vision: Building visual intelligence with Amazon Bedrock and MCP servers

We are converging the three key technologies: Computer Vision, Strands Agents, and the Model Context Protocol (MCP).

Computer vision MCP serversOur implementation is composed of two servers namely the CV server and the OpenSearch server.

The Inline Agent serves as the orchestrator, efficiently coordinating with the MCP Server to access a suite of computer vision tools.

ConclusionBy integrating the powerful AI capabilities of Amazon Bedrock with standardized MCP protocols, we’ve demonstrated how modern computer vision applications can be both sophisticated and accessible.

We’d love to hear how you’re using the Computer Vision MCP servers.

3 days, 3 hours назад @ aws.amazon.com
Monitor Amazon SageMaker Pipelines cross-account with custom Amazon CloudWatch dashboards
Monitor Amazon SageMaker Pipelines cross-account with custom Amazon CloudWatch dashboards Monitor Amazon SageMaker Pipelines cross-account with custom Amazon CloudWatch dashboards

Amazon SageMaker Studio provides monitoring for SageMaker Pipelines within a single account and Region.

Organizations can use services like Amazon CloudWatch, AWS Lambda, Amazon DynamoDB, and Amazon EventBridge to build dashboards that track SageMaker Pipelines executions across multiple AWS environments, tailored to their observability requirements.

In this post, we present a solution designed to centralize the monitoring of SageMaker Pipelines across AWS accounts and Regions using Amazon CloudWatch custom dashboards.

Combined, the two stacks collect, process, and display aggregated information to users through the following workflow:When a step of a SageMaker Pipeline changes status, Amaz…

3 days, 3 hours назад @ aws.amazon.com
Multi-agent social intelligence with Strands Agents and Amazon Bedrock
Multi-agent social intelligence with Strands Agents and Amazon Bedrock Multi-agent social intelligence with Strands Agents and Amazon Bedrock

Multi-agent systems built with Strands Agents and Amazon Bedrock AgentCore can automate this social intelligence at scale.

With multi-agent orchestration, you assign each source to a specialist agent, then fuse results through a dedicated analysis agent that spots cross-source patterns.

This post shows how Thrad.ai deployed a multi-agent system with Strands Agents and Amazon Bedrock AgentCore that automates the pipeline from prospect discovery through personalized email generation.

After both agents finish, the Analysis Agent scores each prospect-trend pair from 0 to 100 using Claude Sonnet 4.6 on Amazon Bedrock.

ConclusionYou now have a blueprint for building multi-agent social intelligenc…

4 days, 2 hours назад @ aws.amazon.com
Accelerating software delivery with agentic QA automation using Amazon Nova Act – Part 2
Accelerating software delivery with agentic QA automation using Amazon Nova Act – Part 2 Accelerating software delivery with agentic QA automation using Amazon Nova Act – Part 2

In a previous post, we introduced QA Studio, a reference solution for agentic QA automation built with Amazon Nova Act.

Test suites for organized regression testingWith QA Studio, you can group individual use cases, each validating a specific user journey, into collections called test suites that run together.

The suite execution page provides an aggregate view: how many use cases succeeded, failed, or are still running.

The QA Studio CLI ( qa-studio ) provides this interface.

Installing and authenticatingThe QA Studio CLI is part of the project’s GitHub repository.

4 days, 4 hours назад @ aws.amazon.com
Scaling UX testing with Amazon Nova Act: A new approach to user flow analysis
Scaling UX testing with Amazon Nova Act: A new approach to user flow analysis Scaling UX testing with Amazon Nova Act: A new approach to user flow analysis

Amazon Nova Act offers a different approach to these challenges.

This makes Nova Act a powerful tool for automated UX testing because it mimics human reasoning when navigating interfaces.

Claude then generates detailed test instructions at multiple levels of granularity, creating the step-by-step interaction paths for Amazon Nova Act to test.

Orchestration layer – The orchestration layer manages test execution at scale:Amazon DynamoDB stores generated test flows with metadata and execution parameters.

Amazon Nova Act agents execute user flows in parallel browser sessions.

4 days, 4 hours назад @ aws.amazon.com
Scaling medical content review at Flo Health with Amazon Bedrock – Part 2
Scaling medical content review at Flo Health with Amazon Bedrock – Part 2 Scaling medical content review at Flo Health with Amazon Bedrock – Part 2

While general-purpose AI tools have shown impressive capabilities, they present significant risks for medical content review.

AWS proof of concept as a starting pointThe PoC’s success presented us with an exciting opportunity to transform how we approach medical content review.

AI content generation systemAfter successfully onboarding Amazon Bedrock as the foundation for our AI content review workflow, we wanted to reuse the same service to accelerate content creation.

At Flo Health, this approach has allowed us to maintain high content standards while scaling our capacity.

ConclusionIn this post, we showed how Flo Health’s engineering team transformed the MACROS proof of concept into a pro…

4 days, 5 hours назад @ aws.amazon.com
ScienceSoft’s HIPAA-compliant AI voice scheduler built on AWS
ScienceSoft’s HIPAA-compliant AI voice scheduler built on AWS ScienceSoft’s HIPAA-compliant AI voice scheduler built on AWS

Healthcare organizations need efficient scheduling solutions, and ScienceSoft’s AI voice assistant, powered by Amazon Nova Sonic and Amazon Bedrock Guardrails, shows how responsible AI can deliver that.

Voice AI is emerging as a transformative technology in healthcare settings, and AWS Partner ScienceSoft is at the forefront of developing responsible AI applications for the industry.

The responsible AI solutionScienceSoft’s AI voice scheduler addresses these challenges by combining the conversational capabilities of Amazon Nova Sonic with the responsible AI framework of Amazon Bedrock Guardrails.

Figure 1 — ScienceSoft’s HIPAA-compliant AI voice scheduler architecture on AWSThe technical fo…

4 days, 5 hours назад @ aws.amazon.com
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock

Today, GPT-5.6 Sol, Terra, and Luna from OpenAI are generally available on Amazon Bedrock, bringing the smartest family of models from OpenAI yet to Amazon Bedrock’s next-generation inference engine built for high-performance, security and reliability.

More ways to put GPT-5.6 on Amazon Bedrock to workAlongside GPT‑5.6, OpenAI launched ChatGPT Work, an agent in ChatGPT for larger, multi-step tasks.

Run the smartest family of models from OpenAI on Amazon Bedrock and get frontier intelligence with the security and scale of AWS.

Get started with Sol, Terra, and Luna in the Amazon Bedrock Console or programmatically through the Responses API.

To learn more, see the Amazon Bedrock documentation …

5 days назад @ aws.amazon.com
NVIDIA
последний пост 1 day, 6 hours назад
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI

Unlike a generative model responding to a prompt, an agentic model must plan, use different tools and recover from problems it encounters mid-run.

Agentic AI introduces a new compute pattern for post-training, making it the central workload of the agentic era and the primary driver of intelligence per dollar.

That means that every improvement to cost per token flows directly into intelligence per dollar.

Agentic Post-Training DemystifiedPost-training is where intelligence is built.

NVIDIA NeMo open libraries, such as NeMo Gym for training environments and NeMo RL for distributed post-training, turn post-training from bespoke research code into repeatable infrastructure.

1 day, 6 hours назад @ blogs.nvidia.com
Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW
Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW Sharpen the Sword, Skip the Downloads — ‘Onimusha: Way of the Sword’ Is Coming to GeForce NOW

Onimusha: Way of the Sword is coming to GeForce NOW at launch, with the playable demo available this week.

Plus, GeForce NOW officially launches in India, moving from beta to public availability — meaning gamers can sign up without a waitlist.

Onimusha: Way of the Sword will arrive on the cloud at launch on Thursday, Sept. 3.

GeForce NOW now supports UPI payments, giving gamers across India a fast, secure and convenient way to purchase memberships and day passes.

— the off-the-rails, train‑wrecking action game — is now available to stream on GeForce NOW, letting members embrace pure chaos from almost any device.

2 days, 8 hours назад @ blogs.nvidia.com
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI

To meet that need, NVIDIA today introduced the T3000 and T2000, new modules based on the NVIDIA Thor architecture that enable mass-market robotics and edge AI applications at scale.

Going Wide on Edge AI With T2000The Jetson T2000 brings Thor architecture to a broader range of edge AI systems.

With the introduction of the new NVIDIA Jetson modules, NVIDIA now offers a scalable edge AI platform spanning performance from 70 TOPS to 2,000 teraflops, enabling developers to address virtually any edge AI workload.

These skills support the entire Jetson portfolio, including Jetson Thor and Jetson Orin, enabling developers to run more capable workloads on lower-memory configurations.

Delivering Cos…

2 days, 22 hours назад @ blogs.nvidia.com
NVIDIA and Japan Bring Full-Stack AI and Robotics to Every Industry
NVIDIA and Japan Bring Full-Stack AI and Robotics to Every Industry NVIDIA and Japan Bring Full-Stack AI and Robotics to Every Industry

Home to leading manufacturers, robotics pioneers and infrastructure builders, Japan is one of the world’s centers of AI — building across the full stack with NVIDIA technologies.

NVIDIA and SEGA Celebrate 30 Years of Innovation, Bringing ‘VIRTUA FIGHTER CROSSROADS’ and Other Legendary SEGA Games to NVIDIA RTX Spark 🔗NVIDIA and SEGA are celebrating more than three decades of collaboration by bringing VIRTUA FIGHTER CROSSROADS and future SEGA titles to NVIDIA RTX Spark — a new superchip for slim Windows laptops and compact desktop PCs.

SEGA will support RTX Spark, giving gamers new ways to experience SEGA’s iconic franchises, including the upcoming VIRTUA FIGHTER CROSSROADS.

The expanding NVI…

3 days, 10 hours назад @ blogs.nvidia.com
Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize
Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize

Open models like NVIDIA Nemotron are built for customization — helping enterprises and nations build AI that’s controllable, trustworthy and tailored to their needs.

From Using AI to Owning IntelligenceSpecialized AI, such as autonomous agents and applications, are built with customized open models.

The most effective agentic AI applications are systems of models where open models work alongside leading frontier models, each fulfilling the job it does best.

Customization Enterprises Can TrustOpen models give enterprises something closed models cannot: full control to customize, inspect and improve AI against business needs.

Learn more about NVIDIA Nemotron open models and try them at build.…

4 days, 4 hours назад @ blogs.nvidia.com
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency

Because of this, performance per watt — a metric that can’t be gamed, only earned through real-world results — is the foundation for AI factories.

Across the newest generation of leading open models, NVIDIA GB300 NVL72 delivers up to 25x performance per watt compared with the NVIDIA Hopper generation.

Moreover, software keeps improving performance over time: On DeepSeek V4, performance per watt improved by up to 5x in a single month.

That’s why leading AI labs such as Anthropic and OpenAI use NVIDIA Blackwell NVL72 systems to run inference.

Fireworks AI deploys GLM 5.2 on the NVIDIA Blackwell platform, enabling production deployments for customers including Cursor and Factory AI.

4 days, 6 hours назад @ blogs.nvidia.com
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo Synthetic Data Generation for Financial AI Research with NVIDIA NeMo

The workflow combines NVIDIA NeMo Data Designer for structured generation, NVIDIA NeMo Curator for semantic deduplication, and NVIDIA Nemotron models for high-throughput headline synthesis.

GPUs 0–3 dedicated to vLLM inference (4-way tensor parallelism, 448 concurrent requests); GPUs 4–7 ran NeMo Curator semantic deduplication with Ray.

", prompt=f"""Generate a realistic financial news headline \ for the category: {{{{ category }}}} {format_examples_for_prompt(examples_by_category)} Generate a single headline for the "{{{{ category }}}}" category.

Run the recipe — Reproduce or extend this pipeline using NeMo Data Designer for structured synthetic generation and NeMo Curator for scalable sem…

1 week, 2 days назад @ developer.nvidia.com
GeForce NOW Turns Up the Heat With New GeForce RTX 5080-Powered Toronto Server
GeForce NOW Turns Up the Heat With New GeForce RTX 5080-Powered Toronto Server GeForce NOW Turns Up the Heat With New GeForce RTX 5080-Powered Toronto Server

This GFN Thursday brings more games, more power and more ways to play on GeForce NOW.

The cloud gaming service is expanding with a new GeForce RTX 5080-powered server in Toronto, bringing dedicated high performance in the cloud closer to members across the region.

It leads the way for GeForce NOW bringing native touch control to the game, coming soon.

A new GeForce RTX 5080-powered GeForce NOW server is coming to Toronto, expanding service in the region and bringing dedicated cloud gaming performance closer to local members.

Ultimate members can stream across PCs, Macs, handhelds, mobile devices, TVs and more with GeForce RTX 5080-class power in the cloud.

1 week, 2 days назад @ blogs.nvidia.com
Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72
Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72 Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72

GPU-accelerated Presto brings low latency to your analytical workloads, keeping you and your agents unblocked and iterating as fast as possible.

GPU-accelerated Presto uses NVIDIA cuDF algorithms for peak performance and NVIDIA NVLink for the fastest GPU-to-GPU communication.

Presto GPU running with one B200 GPU showed 2.5x faster runtime compared to an eight-node Presto CPU cluster, and Presto GPU running with eight B200 GPUs showed 8.2x faster runtimes compared to an eight-node Presto CPU cluster.

Presto GPU running with three B200 GPUs showed 3.6x faster runtimes compared to a 10-node Presto CPU cluster, and Presto GPU running with eight B200 GPUs showed 7.8x faster runtime than a 10-nod…

1 week, 3 days назад @ developer.nvidia.com
NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness
NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

Harness Engineering, Not Fine-TuningLangChain’s team ran Nemotron 3 Ultra against its public Deep Agents benchmark suite, then analyzed the deep agent’s execution traces to find exactly where it lost points.

NemoClaw for LangChain Deep Agents and the tuned Nemotron 3 Ultra model profile are available now.

Developers can pull the tuned Deep Agents harness directly from LangChain, or use the NemoClaw for LangChain Deep Agents blueprint as a starting point for building specialized agents from scratch.

Learn more about NVIDIA NemoClaw for LangChain Deep Agents and NVIDIA Nemotron.

Stay up to date on agentic AI, NVIDIA Nemotron and more by subscribing to NVIDIA news, joining the community, and f…

1 week, 3 days назад @ blogs.nvidia.com
AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters
AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters

AI factories need a CPU with max single-threaded performance to maximize AI factory revenue and agent performance.

How Max Single-Threaded CPUs at Scale Are Built to Run the Agentic LoopAn AI agent doesn’t stop running after a single request.

At the core of Vera is Olympus, NVIDIA’s custom CPU core, which delivers 50% higher instructions per cycle than NVIDIA Grace.

Rigel is NVIDIA’s next-generation Arm v9.2 CPU core, delivering higher per-core performance than Olympus while keeping the same silicon footprint.

Learn more about the NVIDIA Vera CPU.

1 week, 4 days назад @ blogs.nvidia.com
NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community
NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community

Open source AI has shown how quickly developers can innovate when models, data and tools are shared.

NVIDIA and Hugging Face are collaborating to bring the NVIDIA Isaac GR00T 1.7 open, reasoning vision language action (VLA) model for humanoid robots and the NVIDIA Isaac Teleop framework to LeRobot — Hugging Face’s open source library for robotics — with NVIDIA Cosmos 3, a frontier model for physical AI, planned soon.

“Open source is how a field turns advanced research into something people can study, adapt and build on,” said Thomas Wolf, cofounder and chief science officer at Hugging Face.

NVIDIA’s continued partnership with Hugging Face connects NVIDIA’s 3 million robotics developers with…

1 week, 4 days назад @ blogs.nvidia.com
How Open Models Are Driving AI Research
How Open Models Are Driving AI Research How Open Models Are Driving AI Research

This year’s accepted papers reveal a clear direction: open frontier models and open AI infrastructure have become foundational to how modern AI science gets done.

Approximately 2,000 accepted papers cite NVIDIA GPUs, and 145 cite NVIDIA Nemotron — a family of open models, including open datasets — as the foundation for new research.

Hundreds more draw on NVIDIA Cosmos, NVIDIA Isaac GR00T, BioNeMo and other NVIDIA open model families, spanning physical AI, robotics, autonomous vehicles and biomedical research.

AI for life sciences was fueled by NVIDIA BioNeMo open models and research contributions that help researchers understand protein function, molecular behavior and genetic code.

Sakana …

1 week, 5 days назад @ blogs.nvidia.com
How Nations Are Deploying AI for Strategic Priorities
How Nations Are Deploying AI for Strategic Priorities How Nations Are Deploying AI for Strategic Priorities

Why AI Capabilities MatterThe urgency for countries to build and deploy AI capabilities has grown with the rise of generative and agentic AI, which is reshaping markets, inspiring new industries and transforming existing ones — from gaming to healthcare.

Ingredients of a National AI StrategyThere are five ingredients of a national AI strategy:AI Imperative: Domestic AI capabilities are critical to economic growth, national security, cultural preservation and innovation — with responsible, trustworthy AI aligned to local policies as well as national goals.

AI-Ready Workforce: A wide spectrum of local AI skills and talent, plus basic AI literacy across the population.

National AI Strategies U…

1 week, 5 days назад @ blogs.nvidia.com
Joyride Through July With 12 Games Coming to GeForce NOW
Joyride Through July With 12 Games Coming to GeForce NOW Joyride Through July With 12 Games Coming to GeForce NOW

Summer is heating up — and GeForce NOW is taking players along for the ride.

Start the month with Monopoly: Star Wars Heroes vs. Villains, bringing a galaxy far, far away to the iconic board-game franchise, alongside 12 new games joining the cloud this month.

Plus, don’t let the sun set on the biggest GeForce NOW savings of the year.

GeForce NOW makes it easy to take the battle between the light and dark sides across nearly any device.

Plus, check out this spreadsheet, made by a community member, featuring discounted games streaming on GeForce NOW and build out a bigger library at the best bargains during the Steam Summer Sale.

2 weeks, 2 days назад @ blogs.nvidia.com
Facebook
последний пост 3 days, 4 hours назад
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

3 days, 4 hours назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

5 days, 5 hours назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

2 weeks, 4 days назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

3 weeks, 1 day назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

1 month, 3 weeks назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

2 months назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

2 months, 4 weeks назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

3 months назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

3 months, 1 week назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

3 months, 2 weeks назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

3 months, 2 weeks назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

3 months, 2 weeks назад @ engineering.fb.com
Friend Bubbles: Enhancing Social Discovery on Facebook Reels
Friend Bubbles: Enhancing Social Discovery on Facebook Reels Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend bubbles in Facebook Reels highlight Reels your friends have liked or reacted to, helping you discover new content and making it easier to connect over shared interests.

Friend bubbles enhance the social experience on Facebook Reels by helping you discover content your friends enjoy, creating a shared viewing experience and sparking new conversations.

Along with additional optimizations in the underlying method, this approach enabled us to ship friend bubbles while preserving core Reels performance.

Friend bubbles work because the signal is high value: It adds meaningful social context that helps people decide what’s worth watching.

Engagement also scales consistently with the number …

4 months назад @ engineering.fb.com
Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation
Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation

Meta’s Ranking Engineer Agent (REA) autonomously executes key steps across the end-to-end machine learning (ML) lifecycle for ads ranking models.

Powering these interactions are highly sophisticated, complex and massively distributed machine learning (ML) models that continuously evolve to serve both advertisers and people who use the platforms.

Optimizing these ML models has traditionally been time-consuming.

To address this, Meta built the Ranking Engineer Agent, an autonomous AI agent designed to drive the end-to-end ML lifecycle and iteratively evolve Meta’s ads ranking models at scale.

ML training jobs run for hours or days, far beyond what any session-bound assistant can manage.

4 months назад @ engineering.fb.com
Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps
Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps

Nowhere is this more apparent than in mobile security, where a single class of vulnerability can be replicated across hundreds of call sites scattered throughout a sprawling, multi-app codebase serving billions of users.

Meta’s Product Security team has developed a two-pronged strategy to address this:Designing secure-by-default frameworks that wrap potentially unsafe Android OS APIs and make the secure path the easiest path for developers, andLeveraging generative AI to automate the migration of existing code to those frameworks at scale.

The result is a system that can propose, validate, and submit security patches across millions of lines of code with minimal friction for the engineers w…

4 months, 1 week назад @ engineering.fb.com
Uber Engineering
последний пост None
neptune.ai neptune.ai
последний пост 7 months, 2 weeks назад
We are joining OpenAI
We are joining OpenAI We are joining OpenAI

Piotr Niedźwiedź, CEO/CTO and founder of neptune.aiI’m excited to share that we’ve entered into a definitive agreement to be acquired by OpenAI, subject to closing conditions.

We are thrilled to join the OpenAI team and help their AI researchers build better models faster.

Neptune is a metrics dashboard company.”We’ve worked closely with OpenAI to create the metrics dashboard that helps teams building foundation models.

Our future with OpenAINeptune will join OpenAI and continue to support AI researchers with tools to monitor, debug, and evaluate frontier models.

We are looking forward to working with top AI researchers and supporting OpenAI’s mission of ensuring that AGI benefits all of hu…

7 months, 2 weeks назад @ neptune.ai
Synthetic Data for LLM Training
Synthetic Data for LLM Training Synthetic Data for LLM Training

For instance, financial data is highly sensitive and protected by very strict regulations, and synthetic data mimics the real data distribution without revealing customer information.

Read more about how leading foundation model teams curate their training data and other topics in the State of Foundation Model Training Report 2025.

Choosing the right synthetic data generation technique depends on the type of data and its complexity.

Synthetic tabular data generation is a promising direction to overcome these challenges by learning the distribution of the tabular data.

Post-processingAs the distribution of tabular data is highly complex, it makes the synthetic tabular data generation very ch…

8 months, 1 week назад @ neptune.ai
What are LLM Embeddings: All you Need to Know
What are LLM Embeddings: All you Need to Know What are LLM Embeddings: All you Need to Know

TL;DR LLM embeddings are the numerical, vector representations of text that Large Language Models (LLMs) use to process information.

Unlike their predecessor word embeddings, LLM embeddings are context-aware and dynamically change to capture semantic and syntactic relationships based on the surrounding text.

What are the applications of LLM embeddings?

Word EmbeddingsSparse Word Embeddings One-Hot Vectors 1970s TF-IDF1980s Co-Occurrence MatrixStatic Word Embeddings Word2Vec 2013 GloVe 2014Contextualized word embeddings ELMo 2018 GPT-1 2018 BERT 2018 LLAMA 2023 DeepSeek-V1 2023 GPT-4 2023Static word embeddingsStatic word embeddings, such as word2vec in 2013, marked a significant development.…

8 months, 2 weeks назад @ neptune.ai
Detecting and Fixing ‘Dead Neurons’ in Foundation Models
Detecting and Fixing ‘Dead Neurons’ in Foundation Models Detecting and Fixing ‘Dead Neurons’ in Foundation Models

TL;DR Dead neurons silently waste compute and reduce effective model capacity in foundation models.

Dead neurons’ impactRecent studies into dead neurons in the context of foundation models show interesting, albeit worrying, results.

These large reported fractions of dead neurons in foundation models are a concern from a computational perspective.

Before we move on to discuss how to detect and fix dead neurons, let’s touch upon an important distinction between dead neurons and vanishing gradients.

Further reading How to Monitor, Diagnose, and Solve Gradient Issues in Foundation Models Read moreVisualizing activation distributionsIs your foundation model suffering from dead neurons?

8 months, 3 weeks назад @ neptune.ai
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training

In the first part of this series, we covered the fundamentals of instruction fine-tuning (IFT).

def calculate_irs(instruction, output, reference_model): evaluation_prompt = f""" Instruction: {instruction} Model Output: {output} Rate how well the output follows the instruction on these criteria: 1.

| SourceHINT addresses a computational inefficiency in standard instruction fine-tuning: repeatedly reprocessing the same task instruction with every input example.

Read more about foundation model training infrastructure and other topics in Neptune’s 2025 State of Foundation Model Training Report.

First, during initial instruction fine-tuning across multiple diverse tasks, the model learns genera…

8 months, 4 weeks назад @ neptune.ai
How to Optimize LLM Inference
How to Optimize LLM Inference How to Optimize LLM Inference

Large Language Model (LLM) inference at scale is challenging as it involves transferring massive amounts of model parameters and data and performing computations on large tensors.

In the following, we’ll use the Llama model family architecture as a specific example to understand the LLM workload at inference.

For a far more detailed analysis of the LLM workload at inference, see the chapter All About Transformer Inference in the book How to Scale Your Model, published by Google DeepMind.

See also How to Run LLMs Locally Read moreA quick primer on hardware for LLM inferenceA typical LLM inference cluster consists of several nodes, each with a multi-core CPU and multiple accelerator devices, …

9 months, 1 week назад @ neptune.ai
A Researcher’s Guide to LLM Grounding
A Researcher’s Guide to LLM Grounding A Researcher’s Guide to LLM Grounding

In this article, we’ll explore the fundamental concepts of LLM grounding as well as strategies for optimally grounding models.

What is LLM grounding?

LLM grounding is analogous.

If relevant knowledge cannot be inferred from the data, then LLM grounding cannot yield more relevant responses.

When grounding LLMs using RAG, consider retaining only a few of the top hits (i.e., top-k) for your retrieval queries.

9 months, 3 weeks назад @ neptune.ai
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 4 months, 1 week назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

4 months, 1 week назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

6 months, 3 weeks назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

6 months, 3 weeks назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

6 months, 3 weeks назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

7 months назад @ youtube.com
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff) [Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)

https://arxiv.org/abs/2510.17558 Abstract:

We propose an extension of the decoder Transformer that conditions its generative process on random latent variables which are learned without supervision thanks to a variational procedure. Experimental evaluations show that allowing such a conditioning translates into substantial improvements on downstream tasks. Author: François Fleuret Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the con…

8 months, 2 weeks назад @ youtube.com
[Video Response] What Cloudflare's code mode misses about MCP and tool calling
[Video Response] What Cloudflare's code mode misses about MCP and tool calling [Video Response] What Cloudflare's code mode misses about MCP and tool calling

Theo's Video: https://www.youtube.com/watch?v=bAYZjVAodoo

Cloudflare article: https://blog.cloudflare.com/code-mode/ Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8…

9 months назад @ youtube.com
[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)
[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant) [Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)

Paper: https://arxiv.org/abs/2508.21038 Abstract:

Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-following, coding, and more. These new benchmarks push embeddings to work for any query and any notion of relevance that could be given. While prior works have pointed out theoretical limitations of vector embeddings, there is a common assumption that these difficulties are exclusively due to unrealistic queries, and those that are not can be overcome with better training data and larger models. In this work, we demonstrate that we may encounter these theoretical limitations in realist…

9 months, 1 week назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 2 days, 9 hours назад
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

2 days, 9 hours назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

1 month назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

1 month назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

1 month, 1 week назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

1 month, 1 week назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

1 month, 3 weeks назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

3 months назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

3 months, 3 weeks назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

3 months, 3 weeks назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

3 months, 4 weeks назад @ youtube.com
Bacteria Grid Puzzle Solution
Bacteria Grid Puzzle Solution Bacteria Grid Puzzle Solution

Part of a monthly series of puzzlers, in collaboration with MoMath and Peter Winkler

3 months, 4 weeks назад @ youtube.com
The most underappreciated formula | Exploring high-dimensional spheres
The most underappreciated formula | Exploring high-dimensional spheres The most underappreciated formula | Exploring high-dimensional spheres

On the volumes of higher-dimensional spheres

Explore the 3b1b virtual career fair: See https://3b1b.co/talent

Become a supporter for early views of new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Thanks to UC Santa Cruz for letting me film there, and special thanks to Pedro Morales-Almazan for arranging everything. My video on Numberphile with a fun application of this problem: https://youtu.be/6_yU9eJ0NxA Timestamps:

0:00 - Introduction

1:01 - Random puzzle

6:16 - Outside the box

14:35 - Setting up the volume grid

21:14 - Why 4πr^2

25:21 - Archimedes in higher dimensions

36:17 - The general formul…

4 months, 3 weeks назад @ youtube.com
The lattice bacteria puzzle
The lattice bacteria puzzle The lattice bacteria puzzle

Part of a series of monthly puzzles, done in collaboration with MoMath.

https://momath.org/mindbenders

5 months назад @ youtube.com
Solution to the ladybug clock puzzle
Solution to the ladybug clock puzzle Solution to the ladybug clock puzzle

Solution to last month's probability puzzle.

5 months назад @ youtube.com
The Hairy Ball Theorem
The Hairy Ball Theorem The Hairy Ball Theorem

Unexpected applications and a beautiful proof.

Looking for a new career? Check out https://3b1b.co/talent

Supporters get early access to new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Credits:

Senia Sheydvasser: Co-writing and sphere deformation animations

Paul Dancstep: Those lovely fluffy sphere animations Vince Rubinetti: Music Timestamps:

0:00 - To comb a hairy ball

1:24 - Applications

8:46 - The puzzle of one null point

12:12 - The proof outline

16:41 - Defining orientation

21:44 - Why inside-out is impossible

25:59 - 3b1b Talent

27:44 - Final food for thought ------------------ These animati…

5 months, 2 weeks назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 2 days, 5 hours назад
AI Helped Them Code Faster… But At A Cost
AI Helped Them Code Faster… But At A Cost AI Helped Them Code Faster… But At A Cost

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/AI-assistance-coding-skills 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

2 days, 5 hours назад @ youtube.com
The Hidden World Inside An AI
The Hidden World Inside An AI The Hidden World Inside An AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://transformer-circuits.pub/2025/linebreaks/index.html Paper for reindeer vision change - https://royalsocietypublishing.org/rspb/article/280/1773/20132451/50765/Shifting-mirrors-adaptive-changes-in-retinal 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli …

3 days, 7 hours назад @ youtube.com
New AI Just Reinvented Minecraft Worlds
New AI Just Reinvented Minecraft Worlds New AI Just Reinvented Minecraft Worlds

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://xandergos.github.io/terrain-diffusion/

https://modrinth.com/mod/terrain-diffusion

https://github.com/xandergos/terrain-diffusion Source video for some parts of the footage: https://www.youtube.com/watch?v=irE4tcDtUIg 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fi…

6 days, 5 hours назад @ youtube.com
DeepSeek's New AI Speed Hack Is Amazing
DeepSeek's New AI Speed Hack Is Amazing DeepSeek's New AI Speed Hack Is Amazing

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The DeepSeek paper is available here:

https://arxiv.org/abs/2607.05147v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 week, 4 days назад @ youtube.com
Game Physics Just Got 170 Times Faster
Game Physics Just Got 170 Times Faster Game Physics Just Got 170 Times Faster

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://arxiv.org/abs/2506.06494 Sources:

https://www.youtube.com/shorts/Tx7167DXr8U

https://www.youtube.com/watch?v=55F9dY2Y1zc 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

2 weeks, 1 day назад @ youtube.com
This New AI Model Changes Everything
This New AI Model Changes Everything This New AI Model Changes Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers GLM 5.2: https://z.ai/blog/glm-5.2 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

2 weeks, 3 days назад @ youtube.com
DeepSeek Just Solved AI's Billion Dollar Problem
DeepSeek Just Solved AI's Billion Dollar Problem DeepSeek Just Solved AI's Billion Dollar Problem

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2602.21548 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi #deepseek

3 weeks, 5 days назад @ youtube.com
This is OpenClaw On Steroids
This is OpenClaw On Steroids This is OpenClaw On Steroids

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://recursivemas.github.io/

https://github.com/RecursiveMAS/RecursiveMAS 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://felicia.hu

4 weeks, 1 day назад @ youtube.com
Claude AI Knows More Than It Tells You
Claude AI Knows More Than It Tells You Claude AI Knows More Than It Tells You

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/natural-language-autoencoders

https://transformer-circuits.pub/2026/nla/index.html 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month назад @ youtube.com
NVIDIA's New Free AI - A Gift To All of Us
NVIDIA's New Free AI - A Gift To All of Us NVIDIA's New Free AI - A Gift To All of Us

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Nemotron 3 Ultra paper is available here:

https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/ Free Rendering course and source code:

https://users.cg.tuwien.ac.at/zsolnai/gfx/rendering-course/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://f…

1 month назад @ youtube.com
AI Agents as "Games Masters"? 🎮🔥
AI Agents as "Games Masters"? 🎮🔥 AI Agents as "Games Masters"? 🎮🔥

Check the pinned comment for the link to the full interview. Could AI agents eventually become the "Games Master" driving your gaming storylines? We explore the concept of AI assisting players or creating dynamic, non-scripted narratives. Discover how AI is currently being tested inside immersive game environments to change how we play. 🧠 Hashtags: #aiingames #gaming #ai #gamedev #futuretech

1 month, 1 week назад @ youtube.com
DeepMind’s New AI Found A Strange New Way To Think
DeepMind’s New AI Found A Strange New Way To Think DeepMind’s New AI Found A Strange New Way To Think

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://github.com/google-deepmind/alphaproof-nexus-results

https://arxiv.org/html/2605.22763v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 1 week назад @ youtube.com
Meet the AI "Co-Scientist" Changing Everything 🤖🧪 #ai
Meet the AI "Co-Scientist" Changing Everything 🤖🧪 #ai Meet the AI "Co-Scientist" Changing Everything 🤖🧪 #ai 1 month, 2 weeks назад @ youtube.com
Claude Opus 4.8: Lying Machine No More
Claude Opus 4.8: Lying Machine No More Claude Opus 4.8: Lying Machine No More

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers Anthropic's Opus 4.8: https://www.anthropic.com/news/claude-opus-4-8 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 2 weeks назад @ youtube.com
A Second Nobel Prize for AlphaFold? 🧬🏆 #alphafold #deepmind #nobelprize #science #ai
A Second Nobel Prize for AlphaFold? 🧬🏆 #alphafold #deepmind #nobelprize #science #ai A Second Nobel Prize for AlphaFold? 🧬🏆 #alphafold #deepmind #nobelprize #science #ai

Check the pinned comment for the link to the full interview. We're discussing whether a "second order Nobel" prize is on the horizon for AI-driven science. With over 3 million researchers already using AlphaFold, the real-world impact is already historic. Hear what the experts think about what comes next for scientific discovery! 🔬

1 month, 2 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 2 weeks, 2 days назад
Омни-модели будущего 🚀
Омни-модели будущего 🚀 Омни-модели будущего 🚀

Что они будут уметь — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 weeks, 2 days назад @ youtube.com
Качество модели взлетело... без мультимодального RL?
Качество модели взлетело... без мультимодального RL? Качество модели взлетело... без мультимодального RL?

О росте мультимодального качества рассказал Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

2 weeks, 5 days назад @ youtube.com
Работа с данными — это скучно?
Работа с данными — это скучно? Работа с данными — это скучно?

А почему — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

3 weeks, 2 days назад @ youtube.com
Как приготовить SFT 🍲
Как приготовить SFT 🍲 Как приготовить SFT 🍲

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

3 weeks, 4 days назад @ youtube.com
Почему мультимодальные модели — это база 🤖
Почему мультимодальные модели — это база 🤖 Почему мультимодальные модели — это база 🤖

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

4 weeks назад @ youtube.com
Омни-модель: что это за зверь такой
Омни-модель: что это за зверь такой Омни-модель: что это за зверь такой

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month назад @ youtube.com
Borealis — как обучить аудио-LLM по цене MacBook
Borealis — как обучить аудио-LLM по цене MacBook Borealis — как обучить аудио-LLM по цене MacBook

На конференции Data Fest 2026 в Белграде независимый исследователь Александр Николич рассказал практическую историю создания аудиоязыковой модели Borealis с бюджетом, сопоставимым со стоимостью MacBook. Больше контента для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
Better LLM pre-training in NVFP4
Better LLM pre-training in NVFP4 Better LLM pre-training in NVFP4

At Data Fest 2026 in Belgrade, Andrei Panferov from the Institute of Science and Technology Austria introduced Quartet II, a novel method for NVFP4 pre-training that recovers SOTA accuracy. He outlined the core challenges of low-precision LLM training and presented CUDA kernels tuned for Blackwell GPUs, ready for integration into real training pipelines. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
Как безопасно выкатывать новые версии продуктовых AI-агентов
Как безопасно выкатывать новые версии продуктовых AI-агентов Как безопасно выкатывать новые версии продуктовых AI-агентов

На Data Fest 2026 в Белграде Дмитрий Коршунов, Team Lead ML в Ecom, показал, как безопасно обновлять продуктовых AI-агентов с помощью системы автометрик. На примере агента Яндекс AI для турецкого рынка он объяснил, как фиксировать регрессии до прода, сравнивать версии и принимать решение о релизе, когда простой «Hello, Agent» уже позади. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents

At Data Fest 2026 in Belgrade, Karina Romanova, Senior LLM Research Engineer, presented HGRPO — a hierarchical modification of GRPO for multi-turn dialogue agents. Applied to a booking agent in Yandex Alice, the method improved truthfulness by 8.0 percentage points and reduced dialogue length by 10.7%. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей

На Data Fest 2026 в Белграде Вячеслав Костров, ML-инженер в Яндексе, рассказал, как uplift-модели решают бизнес-задачи Лавки: от персональных скидок до показа продуктовых подборок. Он разобрал постановку uplift-задачи, подбор метрик и построение политик, а также практические приёмы с лагранжианом и uplift-деревьями для баланса ограничений. Всё это — на примере реальных внедрений и с разбором типичных ошибок. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #Mu…

1 month, 1 week назад @ youtube.com
Hacks and Defenses in Automatic Kernel Generation
Hacks and Defenses in Automatic Kernel Generation Hacks and Defenses in Automatic Kernel Generation

На Data Fest 2026 в Белграде Егор Коновалов, ML-инженер, разобрал хаки, которые находят LLM-агенты, когда генерируют GPU/TPU-код: от тривиального обхода numerical tolerance до изощрённых атак на timing-измерения и эксплуатации дыр в test harness. А ещё Егор показал, какие методы защиты реально работают, а какие создают ложное чувство безопасности. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
Поиск по архивам: как мы переходим к осознанному распознаванию текста
Поиск по архивам: как мы переходим к осознанному распознаванию текста Поиск по архивам: как мы переходим к осознанному распознаванию текста

На Data Fest 2026 в Белграде Дарья Виноградова, лид команды компьютерного зрения, представила два важных майлстоуна архивного поиска: новую архитектуру распознавания текста и выделение смысловых структур. Эти изменения делают поиск человечнее — теперь можно искать не слова среди текста, а человека среди людей. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
Real-time video generation: where we are and what comes next
Real-time video generation: where we are and what comes next Real-time video generation: where we are and what comes next

At Data Fest 2026 in Belgrade, Andrey Filatov from KREA AI broke down the current state of real-time video generation: which architectures dominate, how they differ, and what challenges arise from compute limits and memory bottlenecks. He also covered production solutions like distillation and caching, and shared his outlook for the next 2–3 years: what will soon become possible and which bottlenecks the industry still overlooks. More content for developers: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #Reinforceme…

1 month, 1 week назад @ youtube.com
AI-генерация учебного контента и проверка открытых ответов студентов
AI-генерация учебного контента и проверка открытых ответов студентов AI-генерация учебного контента и проверка открытых ответов студентов

Доклад из секции ML & Education конференции Data Fest 2026 в гостях у Яндекса «AI-генерация учебного контента и проверка открытых ответов студентов». Спикер — Денис Королёв, доцент, МИЭМ НИУ ВШЭ. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 1 week назад @ youtube.com
ML Trainings ML Trainings
последний пост 1 day, 6 hours назад
Анастасия Никулина | Кластерный подход к прогнозу CTR для миллионов рекламных товаров
Анастасия Никулина | Кластерный подход к прогнозу CTR для миллионов рекламных товаров Анастасия Никулина | Кластерный подход к прогнозу CTR для миллионов рекламных товаров

Спикер: Анастасия Никулина, Wildberries Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Advertising https://ods.ai/tracks/df26-ml-in-advertising

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 6 hours назад @ youtube.com
Олеся Булгакова | Реранкер в ритейле: как превратить рекламу в рекомендацию
Олеся Булгакова | Реранкер в ритейле: как превратить рекламу в рекомендацию Олеся Булгакова | Реранкер в ритейле: как превратить рекламу в рекомендацию

Спикер: Олеся Булгакова, X5 Tech инженер нейронных сетей Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Advertising https://ods.ai/tracks/df26-ml-in-advertising

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 6 hours назад @ youtube.com
Игорь Жданов и Елена Ковылянская | Применение методов ML в системе расчёта контрагентского риска
Игорь Жданов и Елена Ковылянская | Применение методов ML в системе расчёта контрагентского риска Игорь Жданов и Елена Ковылянская | Применение методов ML в системе расчёта контрагентского риска

Спикеры: Игорь Жданов и Елена Ковылянская, Управляющий директор, лидер команды Количественные Финансы Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data Fusion https://ods.ai/tracks/df26-data-fusion ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 6 hours назад @ youtube.com
Василий Сизов | RAG для комплексных ситуационных вопросов
Василий Сизов | RAG для комплексных ситуационных вопросов Василий Сизов | RAG для комплексных ситуационных вопросов

Спикер: Василий Сизов, Лидер кластера "CRM и клиентский опыт", СМБ Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data Fusion https://ods.ai/tracks/df26-data-fusion ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

1 day, 6 hours назад @ youtube.com
Иван Смольянинов | Real-time аналитика на ClickHouse
Иван Смольянинов | Real-time аналитика на ClickHouse Иван Смольянинов | Real-time аналитика на ClickHouse

Спикер: Иван Смольянинов, X5 Tech, руководитель команды интеграции данных Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data и ML в Retail от X5.tech

https://ods.ai/tracks/df26-ml-in-retail ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 8 hours назад @ youtube.com
Мария Шабалкова | ABsalute: автоматизация A/B
Мария Шабалкова | ABsalute: автоматизация A/B Мария Шабалкова | ABsalute: автоматизация A/B

Спикер: Мария Шабалкова, X5 Tech, начальник отдела развития платформы проверки бизнес-гипотез Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data и ML в Retail от X5.tech

https://ods.ai/tracks/df26-ml-in-retail ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 8 hours назад @ youtube.com
Дмитрий Колодезев | Как AI меняет рынок IT труда - ремейк доклада 2025 года, на свежих данных
Дмитрий Колодезев | Как AI меняет рынок IT труда - ремейк доклада 2025 года, на свежих данных Дмитрий Колодезев | Как AI меняет рынок IT труда - ремейк доклада 2025 года, на свежих данных

Спикер: Дмитрий Колодезев, Promsoft, директор Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career https://ods.ai/tracks/df26-opencareer ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 8 hours назад @ youtube.com
Антон Воронов | AI bar raising - как выжить в эпоху AI-зации
Антон Воронов | AI bar raising - как выжить в эпоху AI-зации Антон Воронов | AI bar raising - как выжить в эпоху AI-зации

Спикер: Антон Воронов, Avito Подработка, Technical Unit Lead Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career https://ods.ai/tracks/df26-opencareer ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 8 hours назад @ youtube.com
Никита Калганов | Единая веб-платформа для инженеров данных
Никита Калганов | Единая веб-платформа для инженеров данных Никита Калганов | Единая веб-платформа для инженеров данных

Спикер: Никита Калганов, X5 Tech, ведущий системный инженер данных Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data и ML в Retail от X5.tech

https://ods.ai/tracks/df26-ml-in-retail ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 6 hours назад @ youtube.com
Сергей Маслов | Модель авторизации в Lakehouse без даунтайма
Сергей Маслов  | Модель авторизации в Lakehouse без даунтайма Сергей Маслов | Модель авторизации в Lakehouse без даунтайма

Спикер: Сергей Маслов, X5 Tech, старший корпоративный архитектор Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data и ML в Retail от X5.tech

https://ods.ai/tracks/df26-ml-in-retail

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 6 hours назад @ youtube.com
Максим Павлов | Inside X5 Tech: что мы поняли, внедряя LLM и AI-сервисы в большой компании
Максим Павлов | Inside X5 Tech: что мы поняли, внедряя LLM и AI-сервисы в большой компании Максим Павлов | Inside X5 Tech: что мы поняли, внедряя LLM и AI-сервисы в большой компании

Спикер: Максим Павлов, X5 Tech, директор по развитию продуктов искусственного интеллекта Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data Strategy ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 6 hours назад @ youtube.com
Фронтирные модели и неожиданные примеры в разработке
Фронтирные модели и неожиданные примеры в разработке Фронтирные модели и неожиданные примеры в разработке 3 days, 14 hours назад @ youtube.com
Робот-водитель и его страховка
Робот-водитель и его страховка Робот-водитель и его страховка 3 days, 14 hours назад @ youtube.com
Наука и экономика: как глобализация меняет науку
Наука и экономика: как глобализация меняет науку Наука и экономика: как глобализация меняет науку 3 days, 14 hours назад @ youtube.com
Маленькие модели и их будущее в следующем году
Маленькие модели и их будущее в следующем году Маленькие модели и их будущее в следующем году 3 days, 14 hours назад @ youtube.com
Primer Primer
последний пост 6 months, 2 weeks назад
Taking AI Doom Seriously For 62 Minutes
Taking AI Doom Seriously For 62 Minutes Taking AI Doom Seriously For 62 Minutes

Patreon: https://www.patreon.com/primerlearning

80,000 Hours: 80000hours.org/primer https://www.desmos.com/calculator/a5pfjtr4tr Other connections:

Discord: https://discord.gg/NbruaNW

Twitch: https://www.twitch.tv/justin_helps

Store: https://store.dftba.com/collections/primer Reddit: https://www.reddit.com/r/primerlearning/

Bsky: https://bsky.app/profile/justinhelps.bsky.social

Twitter: https://twitter.com/primerlearning Links to other resources:

https://yoshuabengio.org/2024/07/09/reasoning-through-arguments-against-taking-ai-safety-seriously/

https://www.youtube.com/c/robertmilesai

https://www.youtube.com/@Siliconversations

https://www.youtube.com/@Go-Meta

https://www.youtube.com/@Dwarkes…

6 months, 2 weeks назад @ youtube.com
Simulating a single brain cell
Simulating a single brain cell Simulating a single brain cell

Patreon:

https://www.patreon.com/primerlearning Helpful resources if you want to learn more about neural networks

https://www.youtube.com/@AndrejKarpathy

https://course.fast.ai/

https://www.youtube.com/@WelchLabsVideo

https://www.youtube.com/@3blue1brown Early papers. These probably aren't helpful for understanding the concepts in this video, but if you're interested in history.

The Perceptron – A perceiving and recognizing automaton: https://bpb-us-e2.wpmucdn.com/websites.umass.edu/dist/a/27637/files/2016/03/rosenblatt-1957.pdf

The Perceptron: A probabilistic model for information storage and organization in the brain: https://www.ling.upenn.edu/courses/cogs501/Rosenblatt1958.pdf A Logical…

9 months, 3 weeks назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 2 weeks, 4 days назад
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

2 weeks, 4 days назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

1 month, 2 weeks назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

2 months, 1 week назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

3 months, 1 week назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

3 months, 3 weeks назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

4 months, 1 week назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

4 months, 2 weeks назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

5 months назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

5 months, 2 weeks назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

6 months назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

6 months, 2 weeks назад @ lexfridman.com
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths #487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths

Irving Finkel is a scholar of ancient languages and a longtime curator at the British Museum, renowned for his expertise in Mesopotamian history and cuneiform writing.

He specializes in reading and interpreting cuneiform inscriptions, including tablets from Sumerian, Akkadian, Babylonian, and Assyrian contexts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep487-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://shopify.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/Chevron: Reliable energy for data centers.

7 months, 1 week назад @ lexfridman.com
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life #486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life

Michael Levin is a biologist at Tufts University working on novel ways to understand and control complex pattern formation in biological systems.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep486-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

(2:42:41) – Mind uploading(3:01:22) – Alien intelligence(3:16:17) – Advice for young people(3:22:46) – Questions for AGI

7 months, 2 weeks назад @ lexfridman.com
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy #485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy

David Kirtley is a nuclear fusion engineer and CEO of Helion Energy, a company working on building the world's first commercial fusion power plant by 2028.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep485-sc

See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript:

https://lexfridman.com/david-kirtley-transcript CONTACT LEX:

Feedback - give feedback to Lex: https://lexfridman.com/survey

AMA - submit questions, videos or call-in: https://lexfridman.com/ama

Hiring - join our team: https://lexfridman.com/hiring

Other - other ways to get in touch: https://lexfridman.com/contact EPISODE LINKS:

David's X: htt…

8 months назад @ lexfridman.com
#484 – Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming
#484 – Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming #484 – Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming

Dan Houser is co-founder of Rockstar Games and is a legendary creative mind behind Grand Theft Auto (GTA) and Red Dead Redemption series of video games.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep484-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://box.com/aiUPLIFT Desk: Standing desks and office ergonomics.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(01:29) – Sponsors, Comments, and Reflections(11:32) – Greatest films of all time(23:45) – Making video games(26:36) – GTA 3(29:55) – Open world video games(32:42) – Character creation(36:09) – Superintelligent AI in A Bette…

8 months, 2 weeks назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 2 months, 4 weeks назад
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

2 months, 4 weeks назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

3 months, 1 week назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

3 months, 3 weeks назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

4 months, 2 weeks назад @ microsoft.com
Ideas: Community building, machine learning, and the future of AI
Ideas: Community building, machine learning, and the future of AI Ideas: Community building, machine learning, and the future of AI

This week, machine learning researchers around the world will be attending the annual Conference on Neural Information Processing Systems, or NeurIPS.

In this series, we’ll explore the technologies that are shaping our future and the big ideas that propel them forward.

So around that time when I started my PhD at Penn, I was working in machine learning theory and algorithmic economics.

How had you experienced a lack of community or network of women in machine learning before the founding of WiML?

So particularly when working on topics related to fairness, I’ve ended up focusing a bunch on stuff to do with marginalized groups as part of my responsible AI work.

7 months, 2 weeks назад @ microsoft.com
Ideas: More AI-resilient biosecurity with the Paraphrase Project
Ideas: More AI-resilient biosecurity with the Paraphrase Project Ideas: More AI-resilient biosecurity with the Paraphrase Project

Today, I’m excited to talk about the Paraphrase Project, an effort I co-led exploring how advances in AI tools for protein design might impact biosecurity.

These “patches,” akin to those in cybersecurity, have now been shared with organizations globally to strengthen biosecurity screening.

The project highlights that the same AI tools capable of incredible good can also be misused, requiring us to be vigilant, thoughtful, and creative so we continue to get the most benefit out of AI tools while working to ensure that we avoid costly misuses.

So things like, how similar is this to that template, wild-type protein structure that we used as our conditioning information?

But I feel like broadly…

9 months, 2 weeks назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 2 weeks, 2 days назад
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

2 weeks, 2 days назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

3 weeks, 4 days назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

1 month назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

2 months, 2 weeks назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

2 months, 3 weeks назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

3 months, 3 weeks назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 4 months, 1 week назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

4 months, 3 weeks назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

5 months назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

5 months, 2 weeks назад @ dataskeptic.com
Fairness in PCA-Based Recommenders
Fairness in PCA-Based Recommenders Fairness in PCA-Based Recommenders

In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users equally. David introduces the concept of "power niche users" - highly active users with specialized in…

5 months, 3 weeks назад @ dataskeptic.com
Video Recommendations in Industry
Video Recommendations in Industry Video Recommendations in Industry

In this episode, Kyle Polich sits down with Cory Zechmann, a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares insights on why TikTok's algorithm works so well (clean data and massive interaction volume), the crucial …

6 months, 3 weeks назад @ dataskeptic.com
Eye Tracking in Recommender Systems
Eye Tracking in Recommender Systems Eye Tracking in Recommender Systems

In this episode, Santiago de Leon takes us deep into the world of eye tracking and its revolutionary applications in recommender systems. As a researcher at the Kempelin Institute and Brno University, Santiago explains the mechanics of eye tracking technology—how it captures gaze data and processes it into fixations and saccades to reveal user browsing patterns. He introduces the groundbreaking RecGaze dataset, the first eye tracking dataset specifically designed for recommender systems research, which opens new possibilities for understanding how users interact with carousel interfaces like Netflix. Through collaboration between psychologists and AI researchers, Santiago's work demonstrate…

7 months назад @ dataskeptic.com
Cracking the Cold Start Problem
Cracking the Cold Start Problem Cracking the Cold Start Problem

In this episode of Data Skeptic, we dive deep into the technical foundations of building modern recommender systems. Unlike traditional machine learning classification problems where you can simply apply XGBoost to tabular data, recommender systems require sophisticated hybrid approaches that combine multiple techniques. Our guest, Boya Xu, an assistant professor of marketing at Virginia Tech, walks us through a cutting-edge method that integrates three key components: collaborative filtering for dimensionality reduction, embeddings to represent users and items in latent space, and bandit learning to balance exploration and exploitation when deploying new recommendations. Boya shares insigh…

7 months, 1 week назад @ dataskeptic.com
Designing Recommender Systems for Digital Humanities
Designing Recommender Systems for Digital Humanities Designing Recommender Systems for Digital Humanities

In this episode of Data Skeptic, we explore the fascinating intersection of recommender systems and digital humanities with guest Florian Atzenhofer-Baumgartner, a PhD student at Graz University of Technology. Florian is working on Monasterium.net, Europe's largest online collection of historical charters, containing millions of medieval and early modern documents from across the continent. The conversation delves into why traditional recommender systems fall short in the digital humanities space, where users range from expert historians and genealogists to art historians and linguists, each with unique research needs and information-seeking behaviors. Florian explains the technical challen…

7 months, 3 weeks назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 1 day, 10 hours назад
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents 1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents

In Episode #1010, Jon Krohn digs into “the advisor strategy”, a clever pattern that pairs a fast, cheap executor model with a frontier-class advisor it can consult mid-task, all inside a single API call. Every agent builder faces the same tension: frontier models plan best but cost too much to run on every turn, while small models fumble the decisions that matter. Anthropic’s advisor tool resolves it with roughly a one-line code change, and the benchmarks are startling: Sonnet with an Opus advisor scored higher than Sonnet alone while costing 11.9% less, and Haiku’s BrowseComp score more than doubled at 85% lower cost than Sonnet solo. Jon covers the newest Fable 5 numbers, the practical go…

1 day, 10 hours назад @ podtrac.com
1009: How AI Is Quietly Saving Lives, with Steve Mock
1009: How AI Is Quietly Saving Lives, with Steve Mock 1009: How AI Is Quietly Saving Lives, with Steve Mock

In Episode #1009, Steve Mock (investor at Blumberg Capital, five-time entrepreneur and creator of aisavedme.org), joins Jon Krohn to explore the quiet layer of everyday AI adoption that rarely gets documented. After his 84-year-old father asked a deceptively simple question, “How does one use AI?”, Steve built a place for people to share how AI is actually helping them. The stories that came in surprised him: they’re rarely about the technology and almost always about human outcomes, caregiving, communication, learning, confidence and connection. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1009⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Intere…

4 days, 10 hours назад @ podtrac.com
1008: The AI-Native Startup Playbook
1008: The AI-Native Startup Playbook 1008: The AI-Native Startup Playbook

In Episode #1008, Jon Krohn digs into Anthropic's 35-page Founder's Playbook and pulls out the practical guidance for each of its four startup stages: Idea, MVP, Launch and Scale. AI has erased the three bottlenecks that historically gated company-building — capital, headcount and technical skill — turning the founder from individual contributor into an "orchestrator of agents." Along the way, Jon covers the trap of mistaking building for validating, using AI as a structured devil's advocate against your own idea, the compounding danger of "agentic technical debt," two litmus tests for real product-market fit, and the three-layer moat that keeps a well-funded incumbent from copying you. His…

1 week, 1 day назад @ podtrac.com
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd 1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd

Benjamin Todd, co-founder and President of 80,000 Hours and author of the new Penguin Random House book 80,000 Hours: How to Have a Fulfilling Career That Does Good, joins Jon Krohn for a major update on career strategy in the AI era, his first appearance since before ChatGPT existed. Ben explains why “follow your passion” is backwards and why rare, valuable skills used to help others are what actually generate lasting fulfillment, the ABZ framework for planning under deep uncertainty, why the only durable move is to keep shifting onto whatever bottleneck AI can’t yet clear, and how a human-level digital worker becomes superhuman almost immediately. He and Jon also map the risk landscape, p…

1 week, 4 days назад @ podtrac.com
1006: In Case You Missed It in June 2026
1006: In Case You Missed It in June 2026 1006: In Case You Missed It in June 2026

In this month's episode of ICYMI, hear from Chip Huyen, Andrey Kurenkov, Frank Basso and Gilbert Eijkelenboom, discussing why moats are shifting toward physical systems and accumulated product intuition, how Astrocade built vibe coding before the term existed, what it's really like inside a deafeningly loud AI data center, why only 15% of people are technically self-aware and whether AGI requires anything like consciousness. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1006⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this episode you will learn: (00:00) The Cost of Bu…

2 weeks, 1 day назад @ podtrac.com
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom 1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom

Gilbert Eijkelenboom, bestselling author of People Skills for Analytical Thinkers and founder of the training firm MindSpeaking joins Jon Krohn to make the case that communication is a core data skill, not an optional extra. Gilbert shares the “And, But, Therefore” framework for turning dense analysis into a story stakeholders act on, the research suggesting only around 15% of people are genuinely self-aware (and how journaling, meditation, and exercise help close that gap), how childhood experiences install behavioral “algorithms” we carry into the workplace and why behavior change precedes attitude change, so doing small, uncomfortable things for 30 days can rewire how you see yourself. A…

2 weeks, 4 days назад @ podtrac.com
1004: Recursive Self-Improvement
1004: Recursive Self-Improvement 1004: Recursive Self-Improvement

Could an AI get good enough at AI research to build its own, more capable successor and kick off a compounding loop? That’s recursive self-improvement (RSI) and it surged into the conversation after Anthropic revealed that, as of May 2026, Claude wrote more than 80% of the code merged into its production codebase. In this Five-Minute Friday, Jon Krohn separates today’s AI-assisted coding from true RSI, walks through the accelerating evidence - METR’s shrinking task “time horizon,” Google DeepMind’s AlphaEvolve, Andrej Karpathy’s overnight training-tuner, weighs Jack Clark’s 60% bet that AI builds its own successor by 2028 against the compute, data and “marketing” skeptics. As ever, Jon land…

3 weeks, 1 day назад @ podtrac.com
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso 1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso

Frank Basso, VP of Infrastructure at Lightning AI, joins Jon Krohn for a rare ground-level tour of the one layer of the AI stack the show had never covered in over a thousand episodes: the physical data center. Frank explains how Lightning AI provisions its 35,000-plus GPUs through hyperscale co-location, why everything new is liquid-to-chip cooled, how GPUs talk to each other over ultra-fast east-west networks, and what it’s actually like to stand inside a 110-decibel AI data hall. He also debunks the most persistent myths about data-center water and electricity use, and makes the case for fuel cells, nuclear power, and 800-volt DC distribution as the path forward. Additional materials: ⁠⁠…

3 weeks, 4 days назад @ podtrac.com
1002: Fable 5: The Full Story from Capabilities to Drama
1002: Fable 5: The Full Story from Capabilities to Drama 1002: Fable 5: The Full Story from Capabilities to Drama

Anthropic’s Claude Fable 5 was the most capable AI model ever released to the public and it lasted just three days before the US government forced it offline. Jon Krohn unpacks both halves of the story: what makes Fable 5 special, and why it was pulled. Fable 5 and its locked-down sibling Mythos 5 are the same model separated only by safeguards, in a new “Mythos-class” tier above Opus. Jon covers its state-of-the-art benchmarks, premium $10/$50-per-million-token pricing, conservative safety classifiers, and the federal export-control directive, reportedly sparked by an Amazon-flagged “jailbreak” that took it down. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/100…

4 weeks, 1 day назад @ podtrac.com
1001: How AI Erased My Career Moat, an Episode #1001 Special: Jon Krohn interviewed by Kirill Eremenko
1001: How AI Erased My Career Moat, an Episode #1001 Special: Jon Krohn interviewed by Kirill Eremenko 1001: How AI Erased My Career Moat, an Episode #1001 Special: Jon Krohn interviewed by Kirill Eremenko

For this episode #1001 special, the tables are turned: SuperDataScience founder Kirill Eremenko takes the host’s chair and Jon Krohn is the guest. They trace Jon Krohn’s path from an Oxford neuroscience PhD to a New York hedge fund to founding the AI consulting firm Y Carrot, why he regrets leaving academia and how tools like Claude Code erased his hard-won technical moat and why that makes skilled engineers more valuable than ever. Along the way: whether AI is a bubble, Jevons paradox and the data-center boom, the RICE framework for choosing AI projects, the single biggest reason AI projects fail and how a well-built AI agent could give anyone “Christopher Nolan–like” focus. Additional mat…

1 month назад @ podtrac.com
1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests
1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests 1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests

For this landmark 1,000th episode and the show’s 10-year anniversary, host Jon Krohn is joined by SuperDataScience founder Kirill Eremenko, who hosted the podcast for its first 400-plus episodes before handing over the reins. In a first for the show, the episode was recorded live with the audience invited to join on air, alongside surprise appearances from the team, longtime guests, and even Jon’s family. Together, Jon Krohn and Kirill look back on a decade of the podcast and field listener questions on AI’s biggest opportunities, the build-versus-buy dilemma, how to break into the field today, and how to stay grounded amid the relentless pace of AI. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

1 month назад @ podtrac.com
999: What's Left to Build When Software Is Free, with Chip Huyen
999: What's Left to Build When Software Is Free, with Chip Huyen 999: What's Left to Build When Software Is Free, with Chip Huyen

Chip Huyen joins host Jon Krohn for this milestone episode 999 to talk about her record-breaking book "AI Engineering" the most-read title on the O'Reilly platform last year and how the AI landscape has shifted since her last appearance. Chip breaks down what separates AI engineering from machine learning engineering, makes the case for a "start simple" workflow, gets candid about the real costs of running LLMs in production, and shares why she's now fascinated by physical AI, robotics, and world models and why the durable problems worth solving are increasingly human ones. Jon Krohn guides the conversation from the practical content of the book through to where the field is heading next. A…

1 month, 1 week назад @ podtrac.com
998: In Case You Missed It in May 2026
998: In Case You Missed It in May 2026 998: In Case You Missed It in May 2026

In this month’s episode of ICYMI, Jon Krohn explores how AI agents are simultaneously creating new risks and unlocking powerful new ways of working with data. Hear from Anneka Gupta, Cal Al-Dhubaib, Trevor Manz, Jazmia Henry, Jeremy Mumford, and Jacob Miller, discussing why the old cybersecurity playbook breaks down in the age of Claude Mythos, how the notebook became an AI agent’s working memory, what it really takes to build a foundation model from scratch, and why failing slowly is the most expensive mistake an AI team can make. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/998⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email natali…

1 month, 1 week назад @ podtrac.com
997: How This AI Startup Hit 20M Users (No Moat)
997: How This AI Startup Hit 20M Users (No Moat) 997: How This AI Startup Hit 20M Users (No Moat)

Dr. Andrey Kurenkov returns to the show to talk about Astrocade's astronomical growth from pre-alpha to over 20 million engaged users, what it actually takes to build a vibe-coding platform that scales, and how the broader AI landscape has shifted since his last appearance. Andrey shares behind-the-scenes lessons from building B2C user-generated content products, why the real moat is community rather than tech, and his current thinking on humanoid robotics, AGI, and the AI risks people actually overlook. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/997⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode…

1 month, 2 weeks назад @ podtrac.com
996: TrueFoundry’s Nikunj Bajaj on How to Get $100M Returns on AI Agent Deployments
996: TrueFoundry’s Nikunj Bajaj on How to Get $100M Returns on AI Agent Deployments 996: TrueFoundry’s Nikunj Bajaj on How to Get $100M Returns on AI Agent Deployments

TrueFoundry co-founder and CEO Nikunj Bajaj speaks to Jon Krohn about how enterprises like Nvidia and Siemens are realizing returns of over $100 million from single agent deployments, the AI gateway architecture that makes it possible to connect, observe, and govern agents at scale, and why the familiar advice to “start small” is the wrong way to roll out AI agents inside a large organization. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/996 Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.⁠⁠⁠ In this episode you will learn: (01:21) What TrueFoundry does and why agents in production nee…

1 month, 2 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 1 day, 13 hours назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 day, 13 hours назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 day, 13 hours назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 weeks, 4 days назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

2 months назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

2 months назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

2 months назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 4 weeks назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

2 months, 4 weeks назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

2 months, 4 weeks назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

2 months, 4 weeks назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

4 months, 2 weeks назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

4 months, 2 weeks назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

5 months, 2 weeks назад @ datascienceathome.com