HomeVideos

Vibe Coding With GLM 5.2

Now Playing

Vibe Coding With GLM 5.2

Transcript

356 segments

0:00

In this video, I'm gonna be vibe coding with the newly released GLM

0:03

5.2, which released this morning from ZAI as a result of the US government

0:10

banning access to Fable 5 and Mythos 5

0:14

By this point, I'm sure that most of you know that Anthropic has suspended

0:18

access to Fable 5 and Mythos 5, and this does matter to the release of

0:22

GLM 5.2 because ZAI basically saw this as a marketing opportunity.

0:27

If you look at their post, they literally say the future of AI is open and it

0:31

belongs to the people, which is crazy.

0:34

They weren't planning on releasing this model, but they saw this

0:36

as an opportunity, and it's only available to GLM coding plan users.

0:40

So it's not available in their chatbot.

0:43

It's not available via API.

0:45

So for the purpose of this video, I have purchased a GLM

0:48

coding plan for $65 a month.

0:50

Note that GLM coding plans are now like 100% more expensive than they

0:54

were just a couple months ago.

0:56

Like literally the last time I bought one of these subscriptions, it was like $35.

1:00

So the price has gone up significantly because the price of

1:03

GLM models have also gone up a lot.

1:05

But I'm excited to test the new model

1:07

I currently have GLM 5.2 being run through BridgeBench.

1:10

That's what this open code instance is doing.

1:12

But I also wanna test GLM 5.2 on that horror house game

1:17

that we did with Fable 5.

1:19

If you guys remember from my previous video, I tested Fable 5 building out a

1:23

complete horror house game, and I think it would be very interesting to test

1:27

GLM 5.2 with the same exact prompt.

1:30

So I'm gonna drop this in the BridgeMind Discord community

1:33

if you guys wanna check out the prompt that is being used for this.

1:35

But I am gonna drop it in here just so that we can test out the

1:40

capabilities of GLM 5.2 on the same exact test that we gave to Fable 5

1:45

One thing I wanna mention real quick just so you guys can see it in real

1:48

time, is that the rate limiting on the GLM coding plan is absurd.

1:53

I literally deal with these rate limits every time that I use a ZAI

1:57

coding plan because they do not have the resources to be able to

2:02

serve the model from their server.

2:04

So if you guys look at this, literally I'm getting rate limited

2:06

here, I'm getting rate limited here.

2:08

And yes, I'm running a few subagents, but running five to nine subagents should

2:13

not cause me to get rate limited the way that I'm getting rate limited right now.

2:18

So if you're thinking about getting a GLM coding plan, just have that in mind that

2:22

running subagents and actually being able to work in a normal Vibe coding workflow

2:26

where you have multiple agents and tasks being worked on at the same time is

2:30

gonna be much harder for you than normal.

2:32

And it looks like those rate limits have just subsided now.

2:35

But still, that is not a great thing to have when you're

2:38

spending $65 a month on a plan

2:40

I'm also going to drop in a stealth game test.

2:43

So this is from a member of the Bridge Mine community.

2:46

You guys can look at the prompt here, but this is going to be a 3D

2:49

stealth action game in Three.js.

2:51

So I'm gonna submit that prompt here, and we'll see what happens there

2:55

GLM 5.2 does seem a lot faster than GLM 5.1.

2:59

If you guys look at this here, we're actually getting a

3:01

substantial amount of output in a relatively short amount of time.

3:05

It seems like those rate limits have now subsided.

3:07

We've got all of our tasks running here.

3:09

We've got our games being worked on here and here.

3:11

And the output, I'm actually pretty happy about it.

3:14

It's much better than the release of GLM 5.1.

3:17

If you guys remember that, I was getting five hundred and twenty-nine errors.

3:20

I was getting four zero one errors.

3:21

I was getting four hundred errors.

3:22

So it seems like ZAI has figured out how to serve the models a little bit better

3:26

I did just drop in another prompt that is going to create a Minecraft clone.

3:31

So this agent here is currently working on a Minecraft clone, and guys, you

3:35

can see it's been working only for a few minutes, and we are getting

3:39

quite a lot of output from GLM 5.2.

3:43

So we're gonna see what it does on Bridge Bench for the speed test,

3:46

but to me, it seems like it's definitely faster than GLM 5.1.

3:50

Another thing that I also wanna test this on is creating a Remotion video.

3:54

So let's drop in the Bridge Agent desktop app and then drop in the Remotion

3:58

directory, and then I'm just gonna give it a quick prompt here with Bridge Voice.

4:01

I want you to do a complete in-depth review of the Bridge Agent desktop

4:06

app and our Remotion directory, and then I want you to create a

4:09

new marketing video for the release of the Bridge Agent desktop app.

4:14

Note that this needs to be completely independent from other

4:17

Remotion videos that you see.

4:19

Do not reference any other Remotion videos.

4:21

I want you to create a Remotion video that is 30 seconds long and is

4:25

a digital twin of the actual Bridge Agent desktop app and is fast-paced

4:30

and easy to follow and uses actual demos of the Bridge Agent desktop app.

4:34

All right, guys, here is the moment of truth.

4:36

So we have the horror game completed.

4:39

We have the Shadow Protocol game completed.

4:41

We have our Minecraft game, which is about to be done.

4:44

We have our Remotion video, and Bridge Bench is on its last step.

4:48

So let's pull up Shadow Protocol here and just check out what we have.

4:52

So this says high-value target, so we're targeting Dr. Malachi Voss.

4:57

And okay, it looks like we're just gonna literally be in,

5:00

a building and be fighting.

5:02

Oof.

5:02

Okay.

5:03

All right, first issue that we're already seeing is that the

5:05

lighting, you guys can see that I can't actually see anything, okay?

5:08

So this is an attention to detail that we experience with a lot of

5:11

these Chinese models, where it is

5:13

it does not get the attention to detail that is needed, okay?

5:16

So what I'm gonna do is I'm just gonna stop this, okay?

5:19

So that was not good, okay?

5:21

And this is an attention to detail that some of these Chinese models really do

5:25

struggle with, is that with Fable 5, what we saw before it was banned, is

5:30

that pretty much anything that you gave it, it would be a perfect one shot.

5:34

perfect lighting, perfect game.

5:36

And GLM 5.2, not there

5:41

So let's click to enter, and right away, definitely not as good of

5:45

graphics as Fable 5, but if you do look around, it does look like this works.

5:50

Is there a Post-It note or something in here that we can use?

5:52

Let's read the note.

5:52

"Three locks, three keys somewhere in these rooms.

5:55

The lo- the front door is your only way out.

5:57

Keep the light on.

5:58

Do not let the silence fool you.

6:00

If it starts to hunt, run.

6:02

The door will open then." Okay, so let's pull up here, and let's

6:05

see if we can find these keys.

6:07

I'm just gonna try and speed run this.

6:08

Let's open this door.

6:09

So yeah, look at this key.

6:10

The animations and the UI and the graphics is nowhere near

6:15

Opus or Fable-level capabilities.

6:16

here is the second key, and let's just take a look at the note.

6:19

" Don't take them all at once.

6:20

It wakes when the third is moved.

6:22

I only needed one to see the door.

6:24

God, give me for the… forgive me for the rest." Okay, so apparently this thing is

6:28

going to wake up and chase us or something like that, or do a jump scare at, what,

6:32

I guess when we find the third key.

6:33

So let's try and see if we can find this third key.

6:35

Oh, whoa.

6:35

That, like, sent chills up my spine, that, sound effect.

6:38

That was cool.

6:39

Okay, so I just picked up the third key.

6:40

Is anything gonna happen now?

6:41

Let's see.

6:42

Looking around.

6:42

I don't see anything, guys.

6:44

There's just this creepy poster.

6:46

I don't know.

6:46

What do you guys think?

6:47

Drop a comment down below based off what you guys see.

6:49

So I have all three keys.

6:51

I don't see anything chasing me.

6:52

Is nothing gonna chase me?

6:53

It doesn't seem like anything's gonna chase me.

6:55

Just kinda run around and see if we can trigger anything, but it seems

6:59

like it's not actually working, guys.

7:00

So we're seeing the same issue.

7:02

Nothing is actually happening.

7:03

when we used this with Fable 5, as soon as we picked up that third

7:06

key… Oh, and look at this, guys.

7:08

I can't even get out.

7:09

Even though I have all three keys, I can't escape.

7:12

It says… Look at this.

7:13

"Objective: find three keys, escape through the front door." I

7:16

have all three keys, and I can't escape through the front door.

7:20

And guys, I'm gonna say it again, this is what we're seeing from the Chinese models

7:24

Let's check out this Remotion video that was created by GLM 5.2

7:27

for the Bridge Agent desktop app.

7:29

So first things first, this is not good.

7:32

right off the bat, not loving what I'm seeing.

7:34

It thinks, it searches, it ships.

7:35

It does have pretty much a digital twin of the UI.

7:39

here's the task board, but this is nothing that we haven't seen before

7:42

from other models, and that is… I like that a little bit right there,

7:46

but, like, really a purple gradient.

7:48

This is-- look at the text here, "Built by agents, built for builders.

7:52

Get agent free." I don't know.

7:54

I-- You guys can be the judge of that one.

7:55

I don't think that's good, but you guys can let me know what you think in the

7:58

comment section down below, but I'm gonna not give it a pass on that one.

8:01

Just again, lack of attention to detail

8:04

Here's the Minecraft game that we created.

8:06

So it's called Voxel Sandbox.

8:07

Let's click Play.

8:08

And guys, I can't even move.

8:11

I can't even move around to… I can't do anything.

8:13

Okay.

8:14

Okay.

8:15

All right.

8:15

So this doesn't work

8:16

This is one of the most frustrating parts about Chinese models, is that I don't

8:21

trust them in production code bases.

8:24

If you can't get a Minecraft game where I can at least navigate the map

8:29

and, like, walk around, then how am I supposed to trust you in a production

8:33

code base where actual customers and actual subscribers are on the line?

8:38

And for me, that's what earns you the Bridgemind stamp of approval,

8:41

is can I use you and trust you in my real vibe coding workflow?

8:45

And from what I've seen there, that doesn't look like I can trust GLM 5.2.

8:49

But let's look at BridgeBench.

8:50

We do have the results.

8:52

Here is the Flappy Bird game test.

8:53

So this is a game.

8:55

You guys can see the UI.

8:56

I don't know how I feel about the bird's beak, but that's a thing in itself, but

8:59

it is a functional Flappy Bird game.

9:01

It looks fine.

9:02

I'm not gonna play that too much.

9:03

Here is the infamous lava lamp.

9:05

So this is probably the fattest lava lamp.

9:07

I don't know if this is just a bar being balanced up on the top here, but

9:10

that's not gonna get a pass from me.

9:12

Here's the thunderstorm over the city.

9:14

Let's see the lightning effects.

9:15

That looks pretty good.

9:16

Thunderstorm over city is all right, and let's see another lightning

9:19

bolt just for reference sake.

9:20

Okay, that looks pretty good.

9:21

coral reef underwater, that looks pretty good there.

9:25

Everything looks good But remember, at the same time,

9:28

these are HTML file tests, okay?

9:30

This is completely different than oh, this is something I'm gonna use

9:32

in a real-world production workflow.

9:34

This is like just to give us a visual representation of 'Hey, how good is the

9:37

model?' This game here is pretty good.

9:39

I have to say, I like this.

9:40

GLM 5.1, I'm pretty sure that it does rank pretty high on like Design Arena,

9:44

and that little bug there where it doesn't destroy him if it gets behind him.

9:47

But that's fine.

9:48

that looks good.

9:49

But let's check out the leaderboards and see where GLM 5.2 is ranking.

9:53

So let's check out.

9:54

Whoa, look at this.

9:55

It ranks number six for hallucination.

9:57

Do you guys see this here?

9:58

So it ranks sixth for hallucination, so it looks pretty good on hallucination.

10:02

For refactoring, it ranks fourth, and that's the model's ability to be able

10:05

to rewrite code without breaking it.

10:08

BS, whoa!

10:09

We just got our first one hundred percent score on the BS bench.

10:13

It just got a hundred percent.

10:15

So this is the benchmark that basically it pu- it will push back against nonsensical

10:20

premises or confidently invent answers.

10:23

And look at this, GLM 5.2, we just got our first perfect score

10:27

on the BS benchmark from GLM 5.2.

10:30

Where was GLM 5.1 on this?

10:31

GLM 5.1 was way down here at twenty-six, so that is actually a huge

10:36

leap in terms of the BS benchmark.

10:38

And then, whoa, over here on reasoning too, it does very well on reasoning.

10:42

Number one in reasoning and even beat out Claude Fable Five.

10:46

So this is actually one of the hardest benchmarks on BridgeBench.

10:49

It's a thirty-task benchmark for grounded multi-reasoning over code, logs, config

10:53

specs, and conflicting artifacts.

10:55

This is the hardest one.

10:57

The scores on here are very close together, so the fact that it

11:00

scores one point three percent ahead of Claude Fable Five says a lot.

11:06

But the problem is, is that with the initial test that we just did, like

11:10

the Minecraft test and the horror game test, it just seemed to leave

11:14

out some very critical details.

11:16

And then look at this, guys.

11:17

This is what I noticed when it was running.

11:19

The speeds of GLM 5.2 are actually off the charts compared to GLM 5.1.

11:26

Look at this.

11:26

GLM 5.1 was all the way down here.

11:29

It was number twenty-one.

11:31

It ran at eighty-three point nine tokens per second, and all of a sudden,

11:34

with GLM 5.2, they've been able to speed up the model significantly.

11:38

It ran almost, yeah, over three times faster than GLM 5.1.

11:42

It runs at two hundred and ninety-six tokens per second.

11:45

So that actually does change things, okay?

11:48

And I'll tell you why.

11:50

It's performing very well in reasoning.

11:52

It's performing very well on hallucination.

11:55

It's performing very well on BS, and it's performing well on speed.

11:58

It's also performing well on debugging.

12:01

It scored an 86.6.

12:03

And then big thing here is also cost.

12:06

And, the cost, the cost is very low.

12:08

If we go to OpenRouter, we'll have to see what it does cost.

12:12

I wasn't-- 'cause I'm in the Jill coding plan, and the way I typically will do

12:15

this is through OpenRouter, so I didn't get a cost here for this cost benchmark.

12:18

But when you do look at it, guys, I'm assuming that this is gonna

12:21

be the same cost as Jill M 5.1.

12:22

We'll see, okay?

12:24

But it looks like it's gonna be the same cost as Jill M 5.1, and if it

12:27

is, this is gonna be something that you could potentially be using in

12:30

like Hermes agents because of the speed, because of the reasoning

12:34

capabilities, and because of the BS.

12:37

And that's one thing that I noticed.

12:38

It's hey, I typically will use GPT 5.5 in Hermes agents and Bridge agent

12:43

because it's just so intelligent, and I can use it with my existing Codex CLI.

12:49

But the thing is, is that this model is cheap enough, and it's fast

12:53

enough, and it's smart enough that this could potentially be a really

12:57

good model for Hermes agents, okay?

13:00

Check this out, guys.

13:01

While I was making this video, I reached 71% used on my five-hour quota

13:06

and 14% used on my weekly quota, and I only used 37 million tokens, which you

13:12

may say, "Oh, that's a ton," but it's actually not that much for the model.

13:17

It's a super cheap model, and if I kept going, I would've already

13:22

reached my five-hour quota.

13:23

If I was in a real vibe coding workflow, I would've obliterated this quota, okay?

13:28

Because I would've been running way more tasks.

13:30

In the grand scheme of things, I only ran this through… Hey, I ran it

13:33

through BridgeBench, yes, but I also ran it through one, two, three, four

13:37

tasks that weren't that difficult.

13:39

But at the same time, I also am using GLM 5.2, which is a super cheap

13:44

model, so why am I hitting usage limits on a model that is this cheap?

13:48

hey, if I'm using Opus 4.8 in Claude Code and I get usage limits after,

13:51

two hours or whatever using it.

13:52

Hey, I'm using Opus 4.8, rest in peace Fable 5, hopefully we get it back.

13:56

But the usage limits here are a little bit too much for me.

13:59

So with that being said, moment of truth here, I am not going to be putting

14:03

the Bridge Mine stamp of approval on this model for the following reasons.

14:08

First of all, you're not gonna get a ton of usage out of this.

14:10

I'm paying… What is it?

14:11

My plan is the $65 a month plan, and I'm using a model that is supposedly

14:17

supposed to be only, a dollar per million on the input, and I already

14:20

got 71% of the way through my five-hour limit and 14% of the way through my

14:24

weekly limit just from, a one-hour test.

14:27

That doesn't make any sense to me.

14:29

I'm not gonna be able to use that in my real vibe coding workflow.

14:31

Also, if you guys look back at the games that we created, the

14:35

Minecraft game I couldn't even move.

14:37

The horror game didn't actually function where it had, a monster that came after

14:42

you, or you could even, once you had the three keys, to be able to get out.

14:46

And then with this, whatever this was, it was like the first-person

14:49

shooter game, it was just completely dark, and it didn't look good.

14:52

And then with the Remotion video, it just had no, touch to it.

14:55

It had no creativity.

14:56

So I don't see anything about GLM 5.2 that's calling to me

15:00

Other than the fact that it is very fast, much faster than GLM 5.1, and

15:06

it does do well on BS benchmarks and reasoning benchmarks, which makes it

15:10

suitable for Hermes agents potentially.

15:13

But really the reason that I'm not gonna be using this in a vibe coding workflow

15:16

is attention to detail and reliability is more important now than ever.

15:22

And I think that we saw this with Fable 5.

15:24

With Fable 5, you would basically be able to give it a prompt, and it would

15:27

have so much attention to detail.

15:29

It would be able to one shot pretty much anything that you asked it.

15:33

And it's terrible because we basically got this incredible thing

15:37

that has now been taken away from us, and it is very frustrating.

15:39

Maybe you guys are getting that from this video.

15:41

But I, I'm not seeing that with GLM 5.2.

15:43

GLM 5.2, it's not meant to compete with Fable 5, but that's the core issue that

15:48

I have with Chinese models is that if you make something, if you make a horror game

15:52

and you make it so that one of the biggest pieces of functionality, the ability

15:56

for me to escape the house and for there to be any like ghost that is actually

16:01

going to appear, that's a huge piece of functionality that just is nonexistent.

16:06

It makes it so that I would have to prompt it maybe two, three, maybe

16:09

even four times to be able to get what I wanted out of the model.

16:12

Whereas with Frontier models, Opus 4.8 obviously is far more capable.

16:16

But hopefully we get Fable 5 back because with that model, you could

16:20

pretty much do anything with that model.

16:21

So we'll see what happens today and what Anthropic gives us today,

16:25

what updates they give to us.

16:26

But GLM 5.2, you guys aren't gonna be seeing me work on this

16:29

in my vibe coding workflow.

16:30

But definitely an interesting model and an interesting drop from ZAI just

16:35

as a competitor to say, "Hey, you're gonna block the American models? guess

16:39

what? The future of AI is open, and it belongs to the people." So shout out

16:43

to ZAI for open sourcing their models.

16:45

I think that is a phenomenal way of doing this and great work on your marketing,

16:49

but not so great on your model because I personally will not be using it.

16:53

So with that being said, if you guys haven't liked, subscribed,

16:56

or joined the BridgeMind Discord community, make sure you do so.

16:58

And with that being said, I will see you guys in the future.

Interactive Summary

The video provides a review and hands-on test of the newly released GLM 5.2 model from ZAI, which was launched as an open-source alternative following the suspension of access to Anthropic's Fable 5 and Mythos 5. The speaker tests the model on several code-generation projects, including a 3D stealth game, a horror house game, a Minecraft clone, and a Remotion video. Although GLM 5.2 demonstrates impressive performance on the BridgeBench benchmarks—specifically achieving a perfect score on the BS bench, top reasoning capabilities, and three times the speed of GLM 5.1—it struggles significantly with attention to detail and game mechanics in practice. Due to broken gameplay loops, strict rate limiting, and restrictive quota limits on the $65/month coding plan, the speaker ultimately decides not to endorse GLM 5.2 for real-world vibe coding workflows.

Suggested questions

4 ready-made prompts