HomeVideos

Claude Voice Uses Your MCP Servers Now (I Tested It vs ChatGPT)

Now Playing

Claude Voice Uses Your MCP Servers Now (I Tested It vs ChatGPT)

Transcript

261 segments

0:00

Hey, can you tell me about the latest or

0:03

the upcoming call on Monday that I have

0:05

from cal.com?

0:06

>> Sure, let me pull that up from cal.com

0:07

for you.

0:08

>> I'm really excited about this. I've been

0:09

waiting for this for a long time. We

0:11

just got two huge improvements to voice

0:13

AI from both OpenAI in ChatGPT and Codex

0:16

and Anthropic in Claude. They both work

0:18

differently and are better at different

0:19

things. I'll explain everything. But

0:21

this is something I've personally been

0:22

very excited about for a long time

0:23

because until now the voice features in

0:25

ChatGPT or in Claude can't really do

0:27

anything. They can't really use tools.

0:28

You can talk to it, you can do a little

0:29

web search, but it can't use Connectors,

0:31

it can't control things. So I've been

0:32

building my own voice AI. I'll show you

0:34

that later, but I don't think it's even

0:35

relevant anymore because the new version

0:37

both OpenAI's version and Anthropic's

0:39

versions can now use tools depending on

0:41

the platform. And that means we can talk

0:43

to our AI and not just dictate. We can

0:45

talk, it can use tools, take actions,

0:47

and talk back to us. We don't even have

0:49

to hit the keyboard anymore. And that's

0:50

a huge deal. So first of all, OpenAI

0:52

made a huge deal about theirs. They made

0:53

a release video essentially bringing

0:55

voice back to ChatGPT desktop. But now

0:57

instead of ChatGPT desktop, whether it's

0:59

regular chat, work, or Codex, you can

1:02

have a voice agent that sits with you

1:04

and can control and take actions. This

1:05

will be very powerful moving forward

1:07

where it can actually read back its

1:08

responses and even control your

1:09

computer. One thing I want to point out

1:11

about this, to get it to work, you'll

1:12

see this chat icon, different from the

1:14

dictation, it will only show up in a

1:16

blank chat with no text entered. So hit

1:18

this, we see the famous animation, we

1:20

say, "Can you play Runaway by Kanye West

1:22

on Spotify?"

1:23

>> Yeah.

1:23

>> So right now it's controlling my

1:24

computer and still listening to me.

1:26

Let's see. So we're able to get it to

1:28

control our computer and play Spotify.

1:30

Now that's a very basic example, but

1:32

essentially what's going on here is

1:34

ChatGPT has computer use access and it

1:36

can run its voice agent as an

1:38

orchestrator to direct different agents

1:40

to do things on your computer. So it's

1:41

coding or it's work or it's anything

1:42

else. And I know a lot of us have been

1:44

using dictation tools like Aqua Voice or

1:46

Whisper Flow or Super Whisper or one of

1:48

the alternatives and that's really cool.

1:50

But being able to talk directly and hear

1:52

it talk back and not go back and forth

1:53

is really powerful. You can just

1:54

imagine. So ChatGPT and got this

1:56

advanced live voice on desktop, but

1:59

Anthropic did it a bit differently. But,

2:01

I just want to point out here, OpenAI

2:02

has the best voice model. They have this

2:04

real-time, which you can interrupt, it

2:05

will still process, and doesn't cut it

2:07

off completely. It's a great tool.

2:09

Anthropic is essentially the old type of

2:10

voice, which is speech-to-text and

2:12

text-to-speech. So, you talk, it

2:14

transcribes it, it sends it to the

2:15

model, it responds, and puts it back in

2:17

a voice for you. So, it's a bit slower.

2:19

But, Anthropic's release is nicer

2:21

because now it can use connectors, it

2:23

can use your MCPs in voice mode. They're

2:25

saying that the headline is now you can

2:26

choose different models. So, before

2:28

Claude Voice was using Haiku, their

2:29

smallest model. Now, you can choose if

2:31

it's Opus, Sonnet, or Haiku. Fable is

2:33

not an option, that's probably a good

2:34

thing cuz if you're like me, I'm already

2:36

at 100% of Fable usage. So, Claude Voice

2:38

is only available in chat. You have this

2:41

voice icon similar to ChatGPT's. Let me

2:42

just show you. First, we're going to use

2:44

ChatGPT Voice. I'm going to do it all on

2:45

my phone. And the reason we're going to

2:46

do that is I want to show you how it has

2:47

or doesn't have access to tools. So, I'm

2:49

going to open up ChatGPT Voice. Do you

2:51

have access to my Gmail?

2:52

>> Hang on. And not directly, no. If your

2:55

Gmail is connected through a supported

2:57

connector in this chat, then I can help.

2:59

Otherwise, I can't see your inbox.

3:00

>> It's connected to the connector.

3:02

>> All right. In that case, I can use it if

3:04

it's exposed to me here.

3:05

>> It's exposed to you here.

3:06

>> Checking. Okay, if it's exposed, the

3:08

quickest test is to ask something

3:10

specific like, "What's my latest

3:12

>> What is my latest email?"

3:14

>> Checking now. I can't see your Gmail

3:16

from here. The connector isn't actually

3:19

available to me in this chat.

3:20

>> Trust me, Gmail's connected here. I'm

3:22

even going to send it in the same chat.

3:23

What's my latest email? It's going to

3:25

use the connector. If it's not too

3:26

white, you'll be able to see it here.

3:28

Essentially, it just shows me my last

3:29

email that I sent to myself for this

3:31

video. It is a free 10-minute discovery

3:33

call between me and someone special.

3:35

Now, let's go to Claude. And you're

3:36

going to see why I'm so excited about

3:37

this. One of the MCPs or connectors I

3:38

built was a cal.com MCP. So, I could

3:41

have whatever agent I'm using do

3:42

research or change or send replies to

3:44

people that book calls with me. So, now

3:46

we're going to open up Claude on my

3:47

phone, and we're going to press the

3:48

voice. Hey, can you tell me about the

3:51

latest or the upcoming call on Monday

3:53

that I have from cal.com.

3:55

>> Sure, let me pull that up from cal.com

3:57

for you. It's just listed as a free

3:58

10-minute discovery call. So, there's no

4:00

specific topic attached to it. The only

4:01

details I have are that it's with Steve

4:03

Rogers at his email steve@avengers.com.

4:05

Want me to dig into anything else like

4:07

check your email for context on him?

4:08

>> Can you do research and tell me who he

4:09

is with Bright Data?

4:10

>> Sure, let me see what I can dig up on

4:12

him through Bright Data. Well, here's

4:14

the thing. That name and email are

4:15

basically a placeholder. Steve

4:17

Rogers@avengers.com

4:19

is Captain America. So, either someone

4:21

booked with a fake name or it's a test

4:22

booking. Uh there's no real person I can

4:24

dig up here.

4:25

>> So, I'm editing this video as I ride my

4:26

indoor bicycle. I realized I didn't do a

4:28

one-to-one test, right? So, we're just

4:30

going to open up Claude again on my

4:32

phone and I'm just going to ask it the

4:33

same question I asked ChatGPT.

4:36

Can you look at my email and tell me the

4:37

last email I got?

4:39

>> Sure, let me take a look. Your most

4:40

recent email came in today from cal.com.

4:42

It's a booking confirmation for a free

4:45

10-minute discovery call between you and

4:46

Steve Rogers scheduled for Monday, July

4:48

27th at 10:45 a.m.

4:50

>> So, apples to apples, it was able to use

4:52

the same Gmail connector that was

4:53

connected to Claude that ChatGPT wasn't

4:55

able to use even though it was able to

4:57

use it in its text interface. So, the

4:59

point is both Claude's new voice mode

5:01

and ChatGPT's voice mode have the

5:03

ability to take actions. They can

5:04

actually do stuff for us. The voice

5:06

technology for ChatGPT is a lot better,

5:08

but it's only running on your desktop,

5:10

only running on the computer for now.

5:11

But, Anthropic finally added connectors

5:13

or MCP use to their voice mode so that

5:15

it can actually take actions from your

5:17

voice. So, I don't have to sit at the

5:18

computer. I can have it use my

5:19

connectors either on the computer or on

5:21

my phone when I'm on the go to take

5:22

actions, to do research, to tell me

5:24

what's going on. And that is huge.

5:25

That's something I've been waiting for.

5:26

And I'm sure you can imagine all the use

5:28

cases where you want to direct your AI

5:29

to do something without having to type

5:31

on your phone or type on your computer

5:32

and just do it by voice. And now that is

5:34

possible. And I really believe voice is

5:37

the interface of the future for better

5:38

or for worse. I know that this will be a

5:40

huge productivity unlock and I'm so

5:41

happy that both Frontier Labs released

5:43

this on the same day. It It not be a new

5:45

model or a usage reset, but it's pretty

5:47

cool. Now, I just want to show you my

5:49

voice model. It's pretty funny. I'm

5:50

probably going to take it offline cuz I

5:52

won't need it anymore. Essentially, this

5:54

is connected to Better Claude. This is

5:55

my own personal version of Open Claude

5:58

that is running in Claude code on my Mac

5:59

mini. I'll link to that video above. I

6:01

use my Better Claude every day via

6:02

remote control, but what I was really

6:04

missing was the ability to talk to it,

6:06

similar to what ChatGPT or Claude didn't

6:08

have until today. I wanted to be able to

6:10

take actions or look at my databases,

6:12

look at my schedule, look at my emails.

6:14

So, I built it myself. And I modeled it

6:16

after ChatGPT's voice where it could

6:17

spawn multiple threads. So, I have an

6:19

orchestrator running on a small model,

6:21

and then it's able to send tasks to

6:22

Better Claude. So, let's just see. So, I

6:24

have it as a call, right? It's like a

6:26

phone call. What's up, man?

6:27

>> Not much, just standing by. What's on

6:29

your mind?

6:29

>> Can you tell me about my latest email?

6:31

>> Let me grab that for you. That's a

6:32

Netlify bot notification about a PR.

6:35

Looks like feedback on an editor

6:36

feature. Want me to pull the full thread

6:38

on that one? I can have Better Claude

6:39

check that for you.

6:40

>> So, yeah, basically, I built what I

6:42

could build, similar to what OpenAI was

6:44

offering. It's speech-to-text to speech

6:46

running on my Mac mini where I could

6:48

call into it and have it run tasks, have

6:50

it do actions. It's obviously slower and

6:52

not as efficient as anything that OpenAI

6:54

or Anthropic can do. And the creepy

6:56

thing is it was using my voice that I

6:57

created in my 11 Labs video like a year

6:59

ago. So, essentially, I was literally

7:01

talking to myself. Super weird. I don't

7:03

think I need that anymore. So, yeah, I'm

7:05

really into this about this. I like

7:06

doing research and getting updates while

7:07

I'm walking or on bike rides or driving.

7:10

This is huge. The voice interfaces from

7:12

OpenAI and Anthropic now have the

7:14

ability to take

7:16

So, I hope you found this video helpful

7:18

or insightful. If you have any questions

7:20

or feedback, drop them in the comments

7:21

below. Thank you guys for watching and

7:22

have a great weekend.

Interactive Summary

This video covers the significant advancements in voice AI technology from OpenAI's ChatGPT and Anthropic's Claude. The speaker highlights how these platforms have evolved from simple dictation tools to agents capable of using tools and connectors to perform actions directly. While ChatGPT's voice model is praised for its real-time, fluid interaction capabilities on desktop, Anthropic's Claude voice mode is highlighted for its ability to utilize connectors and MCPs on both desktop and mobile, enabling users to perform research or interact with their emails and schedules on the go. The speaker discusses these developments as a major productivity unlock and concludes by showcasing a personal, custom-built voice AI project that, while functional, is likely now redundant due to these official updates.

Suggested questions

2 ready-made prompts