HomeVideos

LM Studio Bionic: Open-Source Agent Harness (First Impression)

Now Playing

LM Studio Bionic: Open-Source Agent Harness (First Impression)

Transcript

281 segments

0:00

In this quick video, we're going to go

0:01

over LM Studio Bionic, a new tool

0:03

released by LM Studio, which is

0:05

essentially a genetic harness for

0:07

open-source models, for open weights

0:08

models. And in my opinion, this is more

0:10

important than ever with everything

0:11

going on in AI. I do believe open-source

0:13

and open weights is something that

0:14

everybody should become familiar with.

0:16

LM Studio is one of my favorite ways to

0:17

run open-source models on my computer.

0:19

I've done a few videos on them already.

0:21

But LM Studio is pretty much just a

0:22

model runner, a way to run LLMs. And

0:25

today, it's not really about LLMs

0:26

anymore, it's about running LLMs with an

0:28

agents. So, ChatGPT has now ChatGPT work

0:31

and Codex, and that is essentially what

0:33

LM Studio is doing. But, the promise of

0:35

LM Studio is you could run it with local

0:37

models, open weight models, where you're

0:38

technically not paying a third-party

0:40

provider. Everything is technically

0:42

staying local on your computer, it's

0:43

private, and you have all these models

0:45

you can choose from that are optimized

0:46

to work on your computer. Now, I'm just

0:48

going to give you a first look. I've

0:49

been running this for last day since it

0:50

came out. I'm running it with Gemma 4

0:52

12B and GPT-OSS. I'm going to tell you

0:54

what I've discovered, the pros and cons.

0:56

But, first, let's just look at the

0:57

announcement really quickly. Bionic is

0:58

an AI agent for getting real work done

1:00

with open models, including coding,

1:01

research, and complex work with

1:03

documents and files. You can use local

1:05

models or switch to open-source models

1:06

in the cloud for heavier tasks, all

1:08

while staying in control of your privacy

1:10

and AI spend. For all LM Studio Bionic

1:12

users, we commit to zero data retention

1:14

and never training on your data. A

1:16

Bionic agent excels at coding and

1:18

document work, voice input with

1:20

state-of-the-art local voice

1:21

transcription, flexible model execution.

1:24

You can run it either locally, LM Link,

1:26

so LM Studio running on a different

1:27

computer, or use a largest frontier

1:30

open-source models through LM Studio

1:31

secure cloud. So, essentially, this is a

1:33

soft launch for their own cloud model

1:36

services, kind of like what Ollama has

1:38

been doing. They're also releasing this

1:40

offline voice transcription, kind of

1:41

like Whisper Flow or Aqua Voice, and is

1:43

running off of Vox Row by Mistral. And

1:45

you can apparently run it in almost any

1:47

app, so this may be a side feature that

1:49

might replace one of the subscriptions

1:50

you already have. It essentially has two

1:52

modes, just like ChatGPT and Claude,

1:55

work and code. For coding, you obviously

1:57

want it to have shell access, just like

1:59

Claude Code, Code X, etc. And here's the

2:01

thing, Bionic works with powerful

2:03

open-source models like GLM 5.2, Kimiko

2:05

2.7 code. I'm sure Kimiko 3 is going to

2:08

come soon. And by the way, we get a

2:09

trust me, bro, these Chinese models are

2:12

running on US infrastructure, but

2:14

they're not really going to too much

2:15

details, they're just saying trust me.

2:16

And for document work, slides and

2:18

sheets, it's more in a sandbox, just

2:20

like Claude Co-work. You can have it

2:21

organized local folders, edit files,

2:24

summarize materials, all running locally

2:26

on your computer with those open models.

2:27

They support checkpoints and rollbacks.

2:30

Here you can see the local models, and

2:31

what's cool about this is if you're

2:32

already running LM Studio on your

2:33

computer and you download a bunch of

2:35

local models, they obviously work now

2:36

within Bionic. The one thing is you have

2:38

to turn off your LM Studio server, which

2:40

is nice because you don't have to have

2:41

all these things running. That's how I

2:43

would connect other things to LM Studio

2:45

before. And then here they talk about

2:46

their cloud inference with zero data

2:47

retention with GLM 5.2, Kimiko 2, and

2:51

they point out here when using cloud

2:52

models, your requests are processed

2:54

transiently and are not retained after

2:56

the request completes. So, the

2:57

difference between this and previous LM

2:59

Studio is LM Studio was able to run LLMs

3:02

and use tools, but not make edits, not

3:04

really take a big actions. But Bionic is

3:06

able to do that. You create multiple

3:08

projects, we'll create a new one now.

3:09

We'll call this for video, and you can

3:11

see down here it could either be code or

3:13

work. So, just like Claude Co-work or

3:15

Claude Code. Now, by default, it starts

3:17

in work, and if you do that, it won't

3:19

have shell access. So, in my case, I

3:20

like to use code, and you can choose the

3:22

models. So, here we have my three models

3:24

I have installed at the moment. One

3:25

little bug I realized is that even

3:26

though we just created this new project,

3:28

you have to go and select it. At that

3:30

point, it will ask you what directory

3:31

you want to work in. So, we're just

3:32

going to say Bionic test, and we'll say

3:34

what tools you have access to. But this

3:36

release makes it kind of clear of what's

3:37

really going on in AI right now. The

3:39

main workflows are now work or code.

3:41

Frontier models are getting too

3:42

expensive, and people want to run and

3:44

probably will need to run open-source in

3:46

the near future. But again, for LM

3:48

Studio, I think the main release here is

3:49

their cloud inference, models that they

3:51

are providing with a subscription that

3:53

are hosted in the US. By the way, I

3:55

think that Ollama's cloud models aren't

3:57

officially guaranteed to run in the US,

3:59

but they mainly are. Whereas here

4:01

they're saying inference is being

4:02

provided on US servers. Again, trust me,

4:05

bro. So anyways, going back, we see here

4:07

it tells me the tools it has access to.

4:09

Now, the biggest caveat here or what to

4:10

take away from this is that open source

4:12

or open weights models that you're

4:13

running locally on your computer will

4:14

not compare with cloud code or codex or

4:17

anti-gravity or any of these frontier

4:19

models. These are trillion parameter

4:21

models that are running on huge

4:23

infrastructure that you just won't have

4:24

on your computer. That means you have to

4:26

adjust your approach. The approach in

4:27

how much context you want to use, the

4:29

approach in how much you're going to

4:30

trust your agent. Because even last week

4:32

there were plenty of reports of GPT 5.6

4:35

deleting files off of people's computer.

4:37

And that's not a dig at OpenAI. My point

4:39

is these agents still make mistakes. And

4:41

if we're talking about a small model, a

4:43

12 billion parameter model, or a 9

4:45

billion parameter model, or even a huge

4:47

model, they can still screw up. And when

4:49

you're giving these models access to

4:50

your computer, you have to understand

4:52

this is a smaller model. It has less

4:53

instructions, has less reasoning

4:55

capabilities, and it still has access to

4:57

your full file system. So keep all of

4:59

that in mind. That being said, if we go

5:01

into settings, we see here the ability

5:03

to log in. This is where you make your

5:04

LM Studio account and then access their

5:06

cloud models. I haven't done that yet.

5:08

You can turn on web search, but you need

5:09

to be logged in for that. In

5:10

integrations, they have connected apps,

5:12

MCB servers. So they have some out of

5:13

the box. And what I want you to focus on

5:15

here is unlike with LM Studio, which

5:18

ports over your models, it doesn't port

5:19

over your MCB servers. So you're going

5:21

to have to reinstall each MCB server. It

5:23

will support both local and remote MCB

5:26

servers. But what I want you to notice

5:27

here is one thing that's missing, and

5:30

it's crucial, and I'm sure it's coming,

5:31

is skills. They don't have support for

5:33

skills right now. And in coding agents

5:35

specifically, skills are so important

5:37

because they're reusable prompts. And

5:38

I've done plenty of videos on skills.

5:40

But the fact that skills are missing

5:42

here means that the CLI tools I'm using

5:44

are less powerful or less capable

5:46

because the Bionic agent can't read the

5:48

skills that go with the CLI tools. Now,

5:50

let's say you don't have LM Studio

5:51

installed and you want to just start

5:53

with Bionic and see all the different

5:54

models that are available, you could go

5:56

down to settings explore and choose

5:57

whatever model you want. And then the

5:59

most important setting that you should

6:00

probably change is model defaults. What

6:03

is the context window you want your

6:04

models to run? Now, this depends on your

6:06

computer. You have to calculate and you

6:08

can have Cloud Code or Codex or even one

6:11

of these models calculate what the ideal

6:13

context window you should define for

6:15

whatever model you're running. So, now

6:17

I'm just going to have it use Bright

6:18

Data CLI to tell me about the current

6:19

situation with Fable 5 and if it's going

6:21

to subscription plans. So, I'm going to

6:23

send it off and now we're going to watch

6:24

it use tools, specifically the Bright

6:26

Data CLI to go get it. But again,

6:28

because it's missing the Bright Data

6:29

skills, because there's no way to load

6:31

skills into Bionic, it doesn't use the

6:33

CLI as best possible. So, it's using the

6:35

CLI to understand what arguments it's

6:37

able to run. I also like, by the way,

6:39

how we could actually see the thinking.

6:40

The reasoning has been hidden or

6:42

truncated on Cloud and also on ChatGPT

6:45

on Codex. And here on Bionic, we're able

6:47

to actually see the thinking tokens.

6:49

Probably still summary, but still pretty

6:50

good. And there we go, July 18th,

6:52

emerging reports suggest that it's being

6:54

reintegrated into subscription plans

6:56

with a usage cap. So, overall, I think

6:57

this is really cool. I like LM Studio, I

6:59

like to run local models with it. Local

7:01

models are limited, they're smaller, it

7:03

really depends on your computer and even

7:04

if you have the most powerful computer

7:06

on the market, it's probably not going

7:07

to compare to anything you're using with

7:09

Fable or GPT 5.6 Soul. So, you have to

7:12

manage your expectations, but not only

7:14

manage those, manage how you expect to

7:16

work and get work done. It's all

7:17

possible. It's a workflow shift and I

7:19

suggest at least playing with this

7:20

understanding, getting to know open

7:22

source, getting to know open weights.

7:24

Now, what LM Studio Bionic is doing here

7:26

isn't that new. I've been actually

7:27

running open source models via Ollama

7:29

and Cloud Code for a few months now.

7:31

Bionic is essentially an agentic

7:32

harness. Cloud Code is an agentic

7:34

harness and Cloud Code is still my

7:36

favorite. It is very built out, it is

7:38

feature rich, but that being said, it is

7:40

not so efficient with small models

7:42

because it's built for large models. So,

7:44

the benefit here is if Bionic is built

7:47

with running smaller models, the smaller

7:50

context windows in mind, I believe it is

7:52

better positioned, obviously, for

7:54

running open weight models in the

7:56

future. That's my first take on LM

7:57

Studio Bionic. I've been running it for

7:59

last day since it came out on very small

8:01

coding and work tasks. It is capable. It

8:04

is missing some features, but I do like

8:05

LM Studio. I like what they've done and

8:07

I'm sure it will only get better. So, if

8:09

you're out of failure usage or out of

8:10

5.6 usage or you specifically run open

8:13

weights models, I suggest giving Bionic

8:15

a shot. For any questions or feedback,

8:17

drop in the comments below. If you

8:18

haven't done so already, subscribe to

8:19

the channel. It really helps me grow.

8:21

Thank you guys for watching and have a

8:22

great weekend.

Interactive Summary

LM Studio Bionic is a new agentic harness tool from LM Studio designed for running open-source and open-weight models on local computers. It allows users to execute complex tasks like coding and document management while prioritizing privacy and control. Unlike previous versions of LM Studio, Bionic acts as an AI agent capable of taking actions on a user's file system, offering both local execution and a cloud-based option for heavier tasks. While the tool shows great potential for those looking to avoid third-party provider dependencies, it currently lacks certain features like 'skills' and requires users to manage their expectations regarding the reasoning capabilities of smaller local models compared to massive frontier models.

Suggested questions

4 ready-made prompts