Hudson Labs CEO Kris Bennatti: Financial AI Still Gets the Numbers Wrong
1422 segments
For the foreseeable future, if you want
[music] to make sure that your numbers
are definitely correct, you do need some
pre-processing and some built-in checks.
>> How have you thought about that
transition of the instructions from the
investor to the order?
>> You know, prompting is becoming becoming
a a lot less important. And I believe
we're the only people on the market that
that are willing to do a no
hallucination guarantee, [music]
which does give you a bit of a sense of
how much risk there is in the average
financial AI software.
>> [music]
>> All right. Welcome to the next episode
of Invest with AI, where we explore the
intersection of investing and
fundamental research. I'm really excited
to have Chris Brinati here with us
today. Chris was the co-instructor for
the first iteration of the Fundamental
Edge AI accelerator.
Um, has been, you know, someone who's
really taught me a lot about deploying
AI in investment research, and is the
founder and CEO of Hudson Labs. So,
thanks so much for being here, Chris,
and maybe start just telling us a little
bit about your background, how you got
into AI and finance.
>> Awesome. Thanks so much for having me.
As always, a pleasure.
Uh, yeah, I'm I'm Chris Brinati. I've
been around for a minute now.
Uh, I got my start
I
way back when as a data scientist at a
corporate governance advisory firm,
uh, where I focused on sort of pre-LLM
securities filing processing,
and was
building my career as AI was making
really big moves, and LLMs specifically
in the research community. And this
would be would have been back in say
2017
2018,
we started seeing auto suggest on our
phones. Uh the research community was
getting super excited about LLMs. And
yet at my workplace, I was being hailed
as, you know, like the ultimate
innovator running uh essentially all
word-based counts on Canadian corporate
disclosure. Uh
and saw that and saw, you know, sort of
recognized the fact that
often times, at least historically in
finance,
uh we're in corporate governance and in
the corporate fields we're many years
behind consumer technology uh and we
were not adopting LLMs nearly as quickly
as,
you know, iPhone,
Gmail, all of these these cool
applications, quit my job, co-founded
Hudson Labs with uh the LLM guy here in
Canada, who is Suhas Pai. If you haven't
purchased his book, Designing LLM
Applications, I highly recommend doing
it. It's published with O'Reilly. It's
now been translated into over 10
languages, including Polish. Uh and uh
Suhas does a lot of really cool uh
training and onboarding around AI
agentic uh
modeling right now. So, do give him a
follow and uh attend some of his
seminars.
Uh yeah, and then we we uh launched what
was then Bedrock AI, now Hudson Labs.
And we focus on AI for institutional
finance. Our
big contributions to AI research have to
do with materiality ranking uh and
precision. So, being able to
maintain accuracy over very large
contexts.
I'd be able to pull numbers,
pull guidance,
pull sentiment in a way that's
consistent, etc. We also
uh
started our business with a forensic
risk score
and
for for many years now have been popular
among the short community, among plenty
of class action lawyers, and D&O
insurers, as well as big buy-side firms.
>> Great. Great. Maybe um
I I I always found
um I always get like super excited on
things and then I call you and you're
like, "Well, let's talk about the
precision
of uh of these elements." And precision
is so critical
for finance use cases. You know, you can
sort of when so many AI things create
these impressive-looking
demos,
but if the numbers aren't at a level of
precision, they're just not
institutionally useful. How do you think
about just the broad challenge of
precision with large language models,
which are non-deterministic
by nature? And And what are a few of the
key vectors that you've thought about
for years about making these tools more
precise?
>> Yeah, it's a really tricky problem and
it's been a persistent problem. It's
definitely getting better
from a from a generalist perspective,
but
for the foreseeable future, if you want
to make sure that your numbers are
definitely correct,
you do need some pre-processing and some
built-in checks. It's pretty much
impossible to just take a generalist
model, apply it to finance, and be like,
"Okay, great. These This is all going to
be right."
Um for reference,
we did a test on all of the
state-of-the-art models
a few months ago now.
Uh, so
it's been before
before 5.5, whatever the the cloud and
open AI models were a few months ago,
and whenever you asked for a metric for
more than four to eight periods, you'd
get an incorrect number 30% of the time.
Uh, so when you think about
hallucination right now, there's
mostly a problem or the biggest problem
when you start
getting beyond comfortable context
limits.
Uh, and and that's one of the big
problems we focus on.
Uh, it will continue to get better, but
as you've seen so far, it's getting
better much much much more slowly than
things we've seen like reasoning, uh,
mathematical elements, all of those
things. We did, as of yesterday,
formally launch our no hallucination
guarantee. Uh, so we're pretty feeling
incredibly confident about our ability
to not pull incorrect numbers, uh,
restatement adjusted error-free numbers
across a five-year period. Uh, that's
all done using AI architecture rather
than just thinking about the last
generative step. Um, and, you know, it's
it's taken us a while to, um,
be able to do that five-year period
restatement adjusted perfectly, uh, but
we're we're feeling really excited about
that, and I believe we're the only
people on the market that that are
willing to do a no hallucination
guarantee, uh, which does give you a bit
of a sense of of
how much risk there is in the average
financial AI software.
>> What did What does that entail to
uh, gain the conviction to offer a
guarantee like that? Like you said, pre-
pre-processing and built-in checks. Like
what is that? How hard is that problem?
What are the big work streams to sort of
take the general models and convert them
to that hallucination-free layer.
>> Yeah, uh
if you want to think about
hallucination,
you I think it's chapter 11 of Suhas's
textbook.
He talks about the the place where you
end up with hallucination is when
the model doesn't know the answer and it
doesn't know it doesn't know the answer.
That is the highest risk
time when you're not feeding
the model the information it needs
and it doesn't know that.
So, as you can imagine,
when you start
moving beyond these comfortable context
limits, that's when you start seeing
hallucination in a generalist cell.
How do we overcome that? We talked about
preprocessing.
So, once you get towards more than say
four eight quarters,
you need to start using search and
retrieval
to get the right information.
And one of the reasons so many tools,
I'm not going to name names, but have
issues in that long time horizon
is because search and retrieval is
really hard.
If you're just looking for revenue and
you're searching for it again across 20
press releases,
you're going to be pulling in a lot of
noise.
And if you think about how complicated
giving getting a revenue right,
depending on what you're asking for, can
be,
you know, maybe you're looking for a sub
segment
9 month
and instead you're going to be getting
numbers
in your retrieved element that are
consolidated and three-month
Uh
so those are the long horizon complex
queries. In order to fix that problem,
you do need some pre-processing where
your embeddings need to have information
about what is this, what time period is
it related to,
uh and that allows the model to be able
to select the information it needs and
actually use it. And when we talk about
guardrails or checks, it's about making
sure
you have specific pathways
uh that exist so that the model knows
that you're asking for a KPI poll and it
should be able to expect this
information and if it's not getting it,
it just gives you an error.
>> Can I ask a uh
lazy uh layperson question?
How um if you think of I definitely
appreciate the context window challenge,
how does this like how does something
like Notebook LM
solve this where presumably it's it as a
user that doesn't understand the
technical in and out of it, it feels
like the context window is much larger.
Just Like I can give it 500 files.
It seems to be more accurate in the uh
nature of it. It's not as intelligent.
It's like pure retrieval.
Um but can you How does that compare?
Like like what's happening under the
hood in a Notebook LM versus a
generalist model versus a Hudson's uh
labs um
approach?
>> So one of the reasons that Notebook LM
feels better is because it has a more
intelligent search.
Yeah. So it's going to be it's not
finance specific. Obviously, it's not
numeric. It's not built for pulling
KPIs. You are going to still even with
Notebook LLM, you are going to Notebook
LM.
Uh you are going to run into that um
sort of eight quarter plus
failure mode. I don't know if you've
tried it.
Uh
so
for the same reasons around search.
But it is using a more advanced search
mechanism. So it's not about it's not it
doesn't have a a better context window,
it just has better search.
>> Got it.
>> Yeah.
>> And so you Hudson
you being finance experts, you can
basically customize that search process
way better than any than than a
generalist approach.
>> So if you want to think about Hudson
Labs versus Notebook LM versus just
having a Claude instance,
what's happening
the back end is a big part of the
difference. So whenever we take in an
SEC filing, a torrent script, etc., it's
we use vector databases. So we're
converting
a sentence, a word, a paragraph to an
embedding.
Are you uh familiar with or a vector?
Yeah.
>> Mhm.
>> So
there's lots of different ways that you
can do that.
>> Yeah.
>> And you should be doing that. But if
you're using Claude,
your data isn't in a vector database,
which is one of the way reasons that it
can be so expensive.
Uh
and if you're using Notebook LM, it's
uh
it's essentially search rather than a
vector database.
>> Got it.
>> But if you have a vector database or you
have a more sophisticated back end, you
can
do a lot of really cool things in
pre-processing that facilitate both
accuracy, relevance, and completeness on
the front end.
>> Got it.
>> So in addition to to be able to if you
have these things as as embeddings, you
can
give metadata to the model
in addition to
containing better information about what
that sentence is.
And you can do cool things with your
embeddings. For instance, the way that
we store embeddings, we have meta
declarations not just about is this
uh a KPI or not, is it forward-looking
historical, but also information about
sentiment.
Which is why I believe we're the only
platform where you can do
sentiment-based screening. So, ask like,
who are the most stressed out CEOs?
Those types of queries. Cuz if you're
using that like LM and it's just using
search, how does it search for stress?
There's no keyword that you can use
there.
Uh but if you had embeddings that have
information about sentiment, you can do
really cool searches.
>> This was a This was a huge
problem for the institutional finance
use cases of chatbots in default 25.
And in all of my experimentation going
into ChatGPT or Claude, almost anything
quantitative was not helpful due to this
sort of accuracy, like 60 70%
benchmark. Part of the way the ecosystem
is aimed to solve that is through MCP
and connectors into an agentic
workspace. And so now I can connect
into, rather than web searching for a
number,
I can connect into the MCP of a Dilupa
or a FactSet, etc. through my agentic
workspace. Is that How do you think
about sort of what you're doing and sort
of you you've aimed to solve this
problem in the way you just explained?
How would you evaluate the way that the
industry has tried to solve this problem
with that MCP and connector
uh movement into agents?
>> So, I know the last time we talked about
this, I had complained about the Claude
MCP introducing hallucinations into
Hudson Labs results, which had been very
frustrating.
Uh
and I would say I've been very, very
impressed with
how MCPs have evolved since. They're a
lot less brittle than they used to be.
Uh I you know, I think I think it's a
good approach. I think
um
having connectors that deal with the
vectorization, the pre-processing,
uh doing all of this back-end work for
you and pulling that into your own
universe is
is a good idea. Uh
there's
definitely some
trade-offs if you're using MCP. If
you're using Hudson Labs via MCP,
it's going to be a lot less efficient.
There's
sort of a
one of the ways that
Cloud has become less brittle is by
using more calls.
Uh so it so it is a much less efficient
way of using Hudson Labs, but it is much
better than it I don't know the last
time we talked about this, Brett, but um
but it it's
it's a lot better than it was before. I
mean, there's still
some situations where it becomes harder
to control. Uh we we don't offer uh no
hallucination guarantee if you are using
Hudson Labs through MCP because
it's it's not deterministic, right? It's
it's using another model on top of ours
and sometimes it can introduce
things that we didn't expect or you
didn't expect, but at this point we feel
that the benefits outweigh the risks,
which wasn't true a year ago.
>> Interesting.
>> Yeah.
>> And that's principally like the the
evolution has principally been been
what? Just the the the model's getting
better, the engineering around it? Like,
how has that brittleness problem been
uh uh been uh mitigated?
>> Model's getting better,
more guidance around how to set up an
MCP. I would say a lot of issues right
now with
usage of, say,
any connector,
they're not
They're mostly solvable through
better API API documentation, better MCP
documentation, facilitating more
endpoints. Uh so, it's it's
because there's a little bit more more
documentation, more flexibility, it
becomes easier to guide
these MPCs MCP servers to to make better
use of your own
own data.
>> Mhm.
And I hear, and I don't fully understand
this, I hear some people say, like, oh,
XYZ vendor has a their API's great or
their MCP's great. Their MCP's not so
good. Like, what is the skill in in sort
of like, for a neophyte, like, what's
the process of taking your data into an
MCP?
And where is the skill in that? Like,
how how do some vendors sort of do that
well? How do some vendors fail in that
in your in your assessment?
>> Like many things,
I think it's a function of
effort
and attention to detail
because it does take effort to break
down
your products and make it
agentically usable.
You're translating
essentially what is was originally a
user interface
into something that makes the most sense
to an LLM. And what makes sense to an
LLM doesn't always make sense to a human
being. I think, you know, because we
have Suhas, we're particularly good at
making
our our tools LLM readable. That is
another thing we talked a little bit
about,
you know, what's different between
Hudson Labs, what's going on in the back
end. And one of the things that we do do
differently is when you ask a question
as a human being,
we we have a translation step
and make your prompt more LLM friendly,
particularly
when we're talking about market search,
where
those search and retrieval steps and how
you you ask that question really, really
matters.
Uh and that applies to
to MCP as well.
>> Yeah. How have you thought about
you know, how have you thought about the
evolution of prompts to to skills? I
find myself really not prompting
much anymore. I have a few sort of
conversion prompts, but I find my
day-to-day investment process has really
centered around skills into an agentic
workspace.
How have you thought about that transit
that transition of the instructions from
the investor to the to the portal?
>> Yeah, the goal as a provider is
provide the thing the customer wants
with as little instruction as possible.
So, we think about that a lot. You know,
if you just type in SSS,
we will guess that you want same-store
sales and bring that back to you. And
that's not something that we were able
to do or doing
2 years ago. So,
you you're right, you know, prompting is
becoming
becoming a a a lot less important.
>> Can I Can I ask [clears throat] about on
the MCP uh
>> Mhm.
>> kind of uh creation and evolution.
I'm going to put my like power user open
claw hat on, so it's not maybe directly
relevant to financial services and a
but there's a big push even in that
community from MCPs to commit to CLIs.
So, like even more like And again, it's
going to be beyond my my technical
understanding, but CLIs being more
efficient and even more
friendly for
um agents to communicate with.
I I just my question for you is like is
that
is that true? Is that um
how do you see that playing out? I as an
example, like in my own workspace I've
been using the Google CLI instead of the
Google connector. It's like so much more
efficient and so much faster. I've been
using a Notion CLI instead of the MCP
server. Um and the results are much
better.
Um and so I'm wondering
uh if that's something that you are all
are thinking about or how you kind of
think of that evolution or the
relationship between MCP and CLI.
>> It's not something that I've spent a lot
of time thinking about. I'm sure our
technical team has.
And the one thing I would also say is
you'd be
surprised by
you know, we're all on Twitter a lot. We
see sort of
the the the the coolest things that
everyone's doing
with AI. A lot of our customers
don't even know what an MCP is.
So,
it's going to be
hopefully a minute before we prioritize
doing a CLI
given
it it it's tough. I think it's there's a
there's a huge range of where consumers
are are at right now, and you've got to
cater to to all of them, but
we're
we're okay with where we're at right
now.
Yeah.
>> How how have you thought about the in
sort of business model perspective about
the around the agent path and sort of my
experience having demoed many dozen of
the the the sort of chatbot finance
chatbots last year.
I sort of ultimately found like three or
four that were really good that I liked
and Hudson Labs was one of those.
And
that was a good news. The bad news is it
sort of remained this cumbersome
exercise to log in to portrait and
Hudson Labs and Alpha Sense and
>> Yeah.
>> sort of trying to remember where to go
to ask what question wasn't a
a great
user experience
contrasted to now an agentic workspace
with my skills embedded. I don't have to
sort of copy and paste prompts and my
connectors.
It's a much more delightful seamless
experience.
Um
And so how But how do you think about
sort of shifting what what Have you seen
that as well amongst your user base? And
how do you think about sort of shifting
the Hudson Labs
business model to align with that agent
path?
>> Yeah.
My perspective is in 2026 integrate or
die. You have to be
the future analyst is coding they're
building their own tools.
I'm relatively technical. I would never
buy a product that doesn't have an API.
That's going to become everyone going
forward.
So you have to offer integrations. You
have to offer a product that works for
people who are building their own
dashboards, building their own skills.
And we
we're doing that. We're we're including
all of that uh functionality and making
sure that we have connectors that work
for
a variety of
of use cases.
Um
The other way we're thinking about it is
we have this bifurcation in the market
right now, where there's a lot of users
like you, Brett,
who are
really excited about their skills, being
able to pull things from other all these
different places. They're using
perplexity computer, doing all of the
cool stuff. They love the MCP
experience. And we have these big
enterprise contracts with people who can
only use co-pilot.
Uh they've
never tried any of this before, and they
are looking to us
to become
the the place where where they
collect information and bring bring
information into us and and our thinking
about us
as
that place where you go to collect
everything and almost be the server
itself.
>> Mhm.
>> So, we're prioritizing integrating with
other people right now, um but also
thinking about how do we make sure
uh that our users who don't have access
to these more powerful tools can take
advantage of Hudson Labs in a more
extensible way, so that we also have
connectors coming in, not just going
out.
>> Got you. So, almost sort of a
bifurcation where in some like I think
everyone's searching for a single pane
of glass
right now. So, in some instances, you
would be that single pane of glass, and
you would have a sort of ingress MCP
connectors in that. Is that Is that Is
that right? And then you'd also sort of
have egress MCP into another single pane
of glass?
>> Yeah. And not just MCP, API, everything.
All all of the integrations, yeah.
>> How How do you sort of think about the
the durable advantage of of Hudson Labs
in that sort of, you know, single pane
of glass elsewhere? Like what are the
key We can talk about Friends of
Glassdoor, obviously, which is sort of a
differentiated uh data. What What do you
sort of, like I'm trying to sort of
figure out what the institutional grade
stack of connectors
looks like, how many that is,
uh where where do I get my all data
transcripts, etc. Like where do you
think What do you see as the TAM or the
market share of Hudson Labs in that
institutional grade fundamental agent
stack?
>> We're gaining popularity in model
building
and hope to corner that market with the
no hallucination guarantee, being able
to pull all of the the the information
you actually need. Obviously, there's
lots of tools like Delupa that do where
out of the box modeling, but for people
who are doing it themselves,
um but generally, where we think about
our notes versus others. So, we've got
the no hallucination elements, which is
a
a precision element, which is something
institutional users really care about,
whether that's single company or
multi-company. Uh but we do expect
eventually generalist models in the next
5 years to catch up from a hallucination
perspective, and we don't think that,
you know, 10 years from now, we can
still be competing on a a precision
element, particularly from a single
company perspective.
Uh so, really, where we provide a lot of
value is in situations where you need to
be doing
really good search and retrieval. Right
now, that does apply to single company
workflows, cuz you can't get good
results over, you know, a a longer time
period and pull those numbers accurately
if you want the more statement adjusted.
If you don't want it to be full of of of
junk.
Uh but
where we offer a ton of value
now and in the future is being able to
retrieve information across
10,000, 20,000 companies.
Uh I think we have
52 million data points in our our vector
vector database at this point. And you
can do things with AlphaSense that you
literally cannot do with any other
platform like pull out uh cool
sentiment-based elements, be able to
find the things you're looking for
quickly using materiality ranking, using
our uh
our declarations around metadata, etc.
So,
uh
To some extent, mostly right now big
firms
don't have sophisticated search and
retrieval. They're using the out of the
box brute force force methodology, and
you mentioned the fact that they don't
care about spend.
And it's okay that these big firms are
spending like $30,000
uh every couple weeks trying to create
information without a back end.
But, the problem is that it's not just
about cost when we're thinking about
search and retrieval and and and using a
a brute force mechanism across a large
data set.
It's also about accuracy. Um you know,
if you try to use a brute force
methodology to
you know, I'm an accountant, look up
accounting policy changes using Claude
right now, you end up missing so much
information because it's hard to find.
Uh so, there's this this big demand for
how do we do not just single company
deep dives, but how do we do
AI-based
screening, dataset creation,
high-precision dataset creation
uh go the extra step
uh in these multi-company workflows.
>> Is it fair to say that the brute So, the
brute force cuz presumably like with
every new model release, the brute force
powers gets get better? Is it If I'm
hearing you correctly, brute force can
give you like
like it might improve your breadth like
your breadth capabilities, but you still
have to manage the accuracy. Like brute
force is not going to be as helpful to
maintain that accuracy component. Is
that
>> Brute force is
So, what's a topic you want to learn
about right now? Okay.
>> What's a topic?
Um
let's see.
I'm trying to learn about
uh data lakes and data warehouses, like
the distillation of information across
different structures.
>> There is a lot of information out there
on the web about data lakes. There's a
lot of information in SEC filings.
Let's say data lakes
there's
15
articles about
not data lakes, but other types of
databases that are competitors and don't
have the word data lake in them.
If you want to use the brute force
approach right now, you're only going to
be able to effectively search for things
that are easily findable.
Got it. Right?
>> Yeah.
>> So, if you're thinking about getting a
complete of the land using AI search,
you need a way to to discover
information that is related but not
easily searchable.
And that's true
for a lot of AI style search.
One reason I often point to sentiment
is because it becomes really clear in
that that example. If you go, "Okay,
find
frustrated managers" or find managers
that are show evidence of deflection.
That's not something
that you can type in and search for.
And if you think of the AI model, how is
it going to find it if if you can't
search for it either?
And you need some way of making that
information findable so that when you
pull that into the context window,
there's the model has something to work
with.
Yeah.
>> Talk to us a little bit about um
you you're an accountant the forensic
accounting score, which is um part of
the
one of the delightful capabilities of AI
is to take uh
a previously highly cumbersome process
like doing a deep forensic accounting
analysis and sort of distill that into a
a score which has the quantitative
elements but also the unstructured data
elements.
Uh so this was sort of something that
caught my attention
very early on and I've been a big fan of
the the Hudson risk score approach. Can
you sort of just crack that open for us
and and walk through uh what you've been
doing there?
>> Yeah.
So
just to to take a quick step back,
people have been trying to predict
securities fraud since the beginning of
time.
And historically
uh when we used to try to do fraud
protection, we would use
generally things like days sales
outstanding, how quickly is revenue
growing, accruals,
these financial metrics.
Have you ever tried doing that, Brett?
>> Oh, yeah.
>> Yeah. Yeah. And the problem is that it
does
find fraudulent companies, but it also
finds a ton of just normal companies.
It's a very noisy way of predicting
fraud. It's a the reason why when you
read, you know, a Hindenburg report or
when you're pitching a short to your
boss, you're not going to say days sales
outstanding increased by X percentage.
You're saying, "Hey, look, the CFO quit.
They're selling the product to their
mom.
There's millions of dollars of
off-balance sheet debt.
The the market thinks
is is really excited about this
contract, but the contract's actually
with a related party and the market
doesn't know that. So, those are the
things that that create a good short
thesis. So,
uh we take that information, integrity
of the management team, turnover at the
top, you know, the number of times
they're changing their segment
disclosure,
um
aggressive accounting policies,
off-balance sheet risk, related party
transactions,
uh
dependence,
uh
governance.
Fun fact, uh dual class governance is
highly predictive of both growth and
fraud. Uh all of these elements and
using LLMs, we're able to take all of
these previously very qualitative
elements of risk that, you know,
maybe you'd be able to look at and go,
"Hey, that's sketchy." But, we can
convert that into a mathematical
representation of
of risk. How important is this weird
family-related party relationship? And
then use those quantifications to
predict fraud. So, uh our model looks at
all of these different types of
qualitative risks and gives says, "Given
that they restated once in the last 3
years and the CFO
uh
quit
and they changed their useful life
estimate, what's the likelihood they'll
be subject to an SEC enforcement action
or uh litigation related to the
securities fraud?"
>> How have you um
And a part of part of the reason I love
that is sort of like heretofore, like
tracking segment, you know, disclosure
changes over time at scale,
there was just like no way to
there's no way to do that. Like all
these little signals into a into a
collective mosaic. Like yes, you get
this sort of footprint of like
this feels like, you know, there's some
obfuscation here,
right? And there was no way to scale
that scale that before before.
Um
How do you think about, like the actual
like algorithm of like putting that into
a score? Like how do you how how have
you thought about weighting those
various components?
>> We let machine learning do the
weighting. So,
we have
a forensic model
that
classifies or or represents the
underlying risk.
And then we take those and create
variables,
quantitative variables,
and then we run
a standard ML process where we have
a big database of companies that were
demonstrably fraudulent either because
they had an SEC investigation
or a settled class action lawsuit
related to fraud and we let the machine
machine learning model figure out the
weights.
Which are often not the weights that
a human short seller would apply which
is is somewhat interesting.
>> Interesting. How have you tried How How
long has that model been live for and
how do you track like efficacy? Is there
any way to sort of distill that down
into a batting average or hit rate or
alpha, etc.?
>> Yeah, so
it's been a while since we've done a
backtest versus
share price. Um
but I believe there's about a 14-point
differential
uh in performance. So uh underperforming
by about 15% for for high-risk companies
back back when we did it, but uh we've
been doing this since 2019.
Uh very popular among actual teams at
D&O insurers. So we
do a lot of tracking based on
SCA securities class actions which is
the most important the fraud indicator
for them. Uh so high-risk companies are
three times more likely to be subject to
an SCA of any type of any kind. And
every company that has a score of 70 or
higher has about a one in three chance
of SEC enforcement. So it's it's very
high precision.
Uh yeah.
And of course the other two out of three
also likely frauds just not subject to
to SEC enforcement.
>> Yeah. Yeah.
>> Yeah.
>> Conversation we had one time is is um
just how sort of natively good good LLMs
have been at identifying risk and I
think the sort of conclusion was there's
a lot of evidence in the training corpus
to sort of pattern recognize like am I
am I remembering that conversation
correctly like why why why have LLMs
been sort of why why is this been
sort of a a strong native skill of of AI
the sort of risk dimension dimension?
>> We've had mixed results actually with
out-of-the-box LLMs.
>> Okay.
>> So
some things that have worked well are
you find headwinds those types of
requests. Where we've seen issues
is with more complex elements.
So
I've actually been somewhat disappointed
with
out-of-the-box LLMs and being able to
understand say a
bank's risk
or off-balance sheet risk in a way
that's a bit more nuanced.
So
yeah. I think there's still a couple
gaps in understanding that personally I
would have expected to be fixed by now.
Yeah. I I yeah I think
the reason for that though is that there
is just not a lot of good forensic
information on the web to learn from.
>> Okay.
>> Where there is a lot of good information
about general risks like operational
risks etc. But there's not a ton of good
forensic accounting stuff out there.
Up until two years ago when we until we
wrote a blog post on it if you googled
what is a critical audit matter it would
tell you the wrong thing.
>> Yeah. There's sort of a weird circ too
like when I'll search for things about
buy-side and AI like it'll bring my own
blog post back to me so you probably
>> [laughter]
>> Yeah. Yeah, no.
>> no, I want someone smarter than me
telling [laughter] me
Yeah, I don't want to read my own.
>> Yeah.
>> I don't want to read my own blogs, so
but a lot of Hudson Labs stuff shows up
there, too, so that's funny sort of like
getting your own documents fed back to
you.
Um
>> Honestly, the reason we created a blog
post about critical audit matters is
because it was
it was bothering me because I kept
reading short reports that would have
referenced the Google result for
critical audit matter, which had clearly
been written by AI.
Uh
and there's all of these short reports
going out being like, "Oh, you know,
this company has a critical audit matter
related to revenue. Ergo, fraud." And
I'm going, "You know, every single
software company has a critical audit
matter related to revenue. This is
we need to we need to clear this up."
So, yeah. Uh
I I do think there is still
still a few gaps in understanding.
Much, much better than it was
a year ago, but in some of the more
nuanced disclosure
there's
I'm still finding
some reasoning error.
>> Okay, makes sense.
>> Yeah.
>> What um super super super interesting
um
and I think sort of like a very
low-calorie way for investors to just
make better decisions, like having some
sort of like systema- system Like if I
sort of think about like the agent path
in the agent stack, like you know, some
sort of like forensic accounting risk
element in that agent stack is is
critical. Like if I if I can do this in
a seamless way before I make a trading
decision, having sort of a risk
checklist, forensic risk score
Like hey, this you know, there's a one
in three chance of this thing being a
fraud, like I'd like to know that before
I make the trading decision instead of
after I make the trading decision. Like
odds of restatement risk, etc., which
those are sort of very painful aches to
sort of wake up to
having the systematizing your decision
dashboard just seems like an absolute no
no brainer to me.
>> Agreed.
And one of the things that's cool, we
talked a little bit about cost and
efficiency
and
it's
if you
try to run Hudson Labs risk scores
out of the box even for us
uh
it costs hundreds of thousands of
dollars to run the Hudson Labs forensic
risk score over a 10-year period.
Um
which is a sort of somewhat interesting
elements of AI and one of the reasons
that we're now finding it harder to
support quant funds, which I think is a
pretty
just because you know, whenever you're
updating something in that stack, if you
have to re run 10 years of data uh
it's it becomes harder to to offer the
state-of-the-art functionality to to
purely quant funds who need that
backtesting, which I think is a pretty
interesting development in AI where
we've all been thinking about like oh
aren't LLMs amazing for quant funds? Uh
but we end up running into these cost
constraints where we're really only able
to offer state-of-the-art functionality
to
people who don't need 10 years of data.
>> Yeah.
How's it work? I mean
how's it work? I mean so I haven't tried
this yet, but you know, this is a
conversation is inspiring me to try
this. So I'm a healthcare investor.
Um can I create like a risk skill where
it's like a three-page risk checklist on
the name pulling in
all of my underlying research on those
companies, but also pulling in the you
know, the Hudson Labs forensic risk
score pulling from the 52 million you
know, data point vector database via
MCP.
Like is that something where I could
sort of use press that button and get a
get sort of a custom made risk risk
report on any of my companies?
>> We should we should chat. Let's chat
about trying to get you set up.
>> Cool.
>> All right.
>> Sounds great.
>> Can I ask about on this cost question?
The
there's so much happening, right? So you
have Fable that is 10x more expensive
than Opus. You have all these a lot of
my clients have become heavy co-work
users and guess what? That's not a cheap
harness.
Turns out.
Then you have Chinese models, open
weight models. How are you seeing this
kind of
cost like and and also people care about
like people stopped caring about people
didn't care about cost three months ago.
And now a lot of people maybe some of
the mega funds excluded
don't care.
What's your take on the shifting token
tokenomics landscape as it relates to
investment decision-making?
>> Yeah, it's
it's going to be interesting.
One thought I have is
let's all max out our subscriptions in
the short term because
they're not going to stick around.
Um
one thing that we
tell our customers, which is a a funny
thing to be telling customers, is Hudson
Lab subscriptions are profitable for
Hudson Labs. So you can, you know, trust
that we're not going to
randomly hike your prices.
Um
Yeah, it's it's it's it's it's going to
be interesting and it and it does come
back to this
we're creating these really
extraordinary outputs in part by using
very inefficient
very inefficient ways of of using AI,
very inefficient search, very
inefficient like using
a million calls instead of one
in order to get something you want. And
I don't think the fact
that it's so expensive and we're not
seeing any of the cost
matters. It just means that when this
starts to
collapse on us and our our prices start
to go up, there's going to be a lot more
a lot of people are going to start
thinking a lot more about AI
infrastructure, AI architecture, and
what's actually happening on the back
end instead of just thinking about that
top level LLM, the top level model. It's
going to be that's going to be the next
push is we're going to move beyond the
generation step and really think about
how are we storing data, where are we
getting data from?
There's so many ways to make make all of
this more efficient, which is one of the
reasons why Hudson Labs can run data
sets using long running agents and it,
you know, it only costs us 100 bucks.
But if you just went to an Opus 4.8
API, you're going to accidentally spend
13 grand. You know, it it it all of
these problems
are are solvable.
They just require they require more.
>> I wonder too, right? Cuz
these you mentioned like it's at the app
layer, like this this cost decision's
going to have to be made at the
individual contributor layer, the app
layer, the model layer, the model
selection layer. Like do you think one I
mean you you a horse in the race, but do
you think like one of the levers matters
most on like
optimizing against cost or optimizing
for cost?
>> There's a reason open source tools like
open code are having such a big moment
right now.
People who are thinking about cost don't
want to get locked into one model
provider. It doesn't make sense.
Um open code is a uh they're good
friends of ours and and and they're
getting huge because developers are
frustrated with oh if you you're using
quad code, only getting to use Anthropic
models, Anthropic probably won't be king
forever. So, yeah, if if if if you're
thinking about cost, it's
a good idea to to make sure that you can
you can switch providers, you're not
locking into to one model provider in
perpetuity. We know a lot of bigger
banks, bigger enterprise customers that
have built their own routers rather than
relying on somebody else's MCP, so they
do have more control.
Uh when we think about AI and where
we're going, it's yeah, it it's all
about being able to control your own
outcomes and make your own choices.
>> We've covered a lot of uh ground, Chris.
Any other, you know, core debates that
you're having with clients or peers or
friends or
um sort of big questions that you think
still need to be resolved that we should
talk about?
>> One thing that's somewhat interesting
is uh
have you tried using any of the big
models for extracting soft guidance?
>> No. Tell me about it.
>> Or do it, yeah. So, we built a guidance
model 2 years ago now
that we thought would only have a moat
for about 9 months.
The reason we bought built the guidance
model is because LLMs are brilliant, but
when they think about tense, they focus
when they think about whether or not
something's future or past, they focus
on the tense of the verb.
So, we talked about this before brought
where, you know, American Eagle
frequently says CapEx is now expected to
be.
And if you run state-of-the-art models
and say get me all management guidance,
it's missing a bunch of guidance.
Um
and that appears to still be true. Which
I find absolutely fascinating,
especially because so many of these big
model providers haven't been investing
in finance-specific training.
Now that I've said it on this podcast,
it'll probably be
fixed tomorrow, but [laughter]
>> No, I'll I'll I'll I'll tell the nerds
to the invest what they have to invest.
>> invest in it. Yeah, yeah.
>> Although it's probably being scraped
into many many agents, I'm sure
probably. Um
that's interesting. Um
>> Yeah.
>> Yeah, fast fascinating. Okay, great. Um
Yeah, what why do you think like I hear
this a lot like the Opus model of agents
have more been more specifically finance
trained than the ChatGPT? Like do you
have any sense of like why that's been
the case?
>> Why
>> Has that been your experience?
>> are more finance?
>> Yeah, like is it just the intention of
the training process? Like
>> I think a lot of the big firms are as
now specifically focusing more on
finance because A, they realized it was
a major gap before and B, realizing it
that this is a consumer base actually
willing to pay.
>> Yeah.
>> Yeah.
>> Yeah, I think it was a little bit of
galvanizing moment when Anthropic did
their a finance day and showed the
finance was the second largest vertical
behind tech.
Um
So, and finance
>> Complexity as well. It's Yeah,
obviously.
>> Finance AI is definitely having a
moment. We even have a We even have a
specific finance and investing podcasts
out now. Invest with AI, available on
all podcast platforms. Wherever you get
your podcasts, like and subscribe as the
kids say. So, uh
thank you so much for for being with us,
Chris. This has been really uh really
helpful and uh
part of what we're trying to sort of uh
do in this podcast is uh so sort of like
a flashlight on a dark evening, just see
3 ft in front of us.
I don't really know. I don't think any
of us know what what what lies around
the bend, but um we hope to continue to
have you on to
sort of see a little bit further down
the road and under stand at least
understand where we're at in the moment.
So, thanks for Thanks for everything
you've you've shared and what's the best
way for people to get in contact with
you if they want to learn more about
Hudson Labs?
>> hudson-labs.com
Come and give us a view. And I'm Chris
Pennati, you can find me on LinkedIn, X
everywhere.
>> Great. Well, thanks so much for the time
today.
>> Thank you.
>> I'll bother you soon with uh more
questions, I'm sure.
>> Can't wait.
>> Mhm.
Ask follow-up questions or revisit key timestamps.
The video features an in-depth conversation with Chris Brinati, founder and CEO of Hudson Labs, regarding the integration of AI into institutional finance. They discuss the challenges of precision and hallucinations in large language models (LLMs) when handling financial data, the importance of robust pre-processing, vector databases, and specialized search mechanisms for long-horizon financial queries. Brinati introduces Hudson Labs' 'no hallucination guarantee' and explores the evolution of agentic workspaces, the role of Model Context Protocol (MCP), and the transition from prompt-based interactions to agentic skills in the investment research process.
Videos recently processed by our community