Shan Lilja Headshot
Episode 97
Guest: Shan Lilja

Designing Support Journeys with Multimodal AI w/ Shan Lilja, Mavenoid

Overview

According to Service Council’s Voice of the Field Service Engineer survey, 60% of frontline workers report that the products they service have grown significantly more complex over the past five years. As assets grow more complex, so does the support required to install, maintain and repair them. While traditional text-based AI excels at addressing common and routine issues, support teams often face challenges when problems are nuanced or involve multiple data types, such as images, video, voice, or structured data. In these cases, incomplete context can lead to extended downtime, additional delays, and increased frustration for both technicians and customers.

Join us as Service Council’s Chief Research Officer, Gerardo Pelayo, sits down with Shan Lilja, President and Founder of Mavenoid. Together, they’ll discuss how service leaders can design more effective escalation paths that capture the full context of a customer’s issue, enabling faster and more accurate resolutions while preserving the essential role of human expertise.

Shan is the co-founder and CEO of Mavenoid. Prior to Mavenoid, he co-founded the internet startup Gruppi, acquired by 24 Media Network, and was an early employee at Palantir.

Topics: Intelligent Service, Technology

About the Host

00:12
Gerardo Pelayo, Ph.D.
Hello. Good morning everybody. Happy New Year. Thank you for joining the In Service Podcast. I’m Gerardo Pelayo, the Chief Research Officer at Service Council and really glad that you’re here today to kick off 2026 with a theme, a topic that I think is really interesting, really innovative and it’s going to give us a very good half an hour to talk about it. And that topic is going to be designing Support Journeys with Multimodal AI. And I’m pleased to have Shan Lilia, he’s the president and founder of Mavenoid, join me in this conversation. We’ll get to know you in a little bit, Shan, but thank you for being here. Good, good. And so I’ll tell you a little bit of why I think this topic is important.

00:56
Gerardo Pelayo, Ph.D.
And it brought me back to an article that was published in HBR Harvard Business Review back in 2024 where over 5,000 consumers were interviewed about their hyper personalized experiences. And even though 80% of them wanted to have their support experiences, their service experiences to be hyper personalized to their liking, the downside is that 2/3 of them said that they had a personalized experience that had been negative either because it was disruptive, it felt invasive, it was inaccurate. And the dilemma here is we want to be met where we are. We want the information to be produced in the way in which we learn, in the way in which we consume it the easiest. But the attempts to do so in the past have not always been great. So today we’re going to learn about where the values in that.

01:55
Gerardo Pelayo, Ph.D.
How do we think about this problem at scale and what are some of the lessons that we can learn so that we do this in an effective way to improve the service experience overall. So before I get into the conversation, a little bit of housekeeping, a reminder that today’s podcast is recorded. It will be available across any streaming platform that you use, including our social media platforms, our website, our LinkedIn web page, and for that LinkedIn audience, we see you, we invite you to comment as we get into the conversation. If you agree, if you disagree, if you have any questions or follow ups that you’d like to add and infuse the conversation with. We welcome that. So without further ado, Shan, pleasure to have you here again and I’ve had the chance to interact with you before.

02:48
Gerardo Pelayo, Ph.D.
We’ve done a great webinar as well. But for the audience that might be less familiar, can you share a little bit of introduction about yourself and about Mavenoid.

02:58
Shan Lilja
Well, yeah, so I started Mavenoid back in 2018 and before that I actually started another Internet startup. So I’ve kind of been doing companies my whole life. The one job I had where I started my career was at Palantir, which is a Silicon Valley based company. And just always been very interested in the nature of intelligence consciousness, AI, things that are very close to what we’re doing at Mavenoid, which is to automate customer support. And at Mavenoid, the thing that’s a little bit different is that we don’t just kind of build AI agents for customer service in general, but we actually focus on physical products and devices. So our customers are big brands that everyone will kind of know. Companies that produce coffee machines and fridges and standing desks and all kinds of physical things you see around in your home.

04:08
Shan Lilja
And it turns out that those kind of companies have very unique support needs. So we’re trying to solve kind of support automation for hardware in particular.

04:21
Gerardo Pelayo, Ph.D.
Good, good. Very, very interesting. And so within that and part of the title, it’s this multimodal AI is one of those things that we assume that we know what it is. But I’ve learned that it’s important to have clarity. And I would like to start the conversation by level setting. And who better than you to explain what do we mean when we say multimodal AI?

04:48
Shan Lilja
So multimodal basically just means that you have multiple modalities. So you don’t just have text, but you also have visual input output, you have audio input output, and really any kind of type of data. So it could also be sensor data from physical products, for example. But the way to think about it, the most important things are really visual and audio information in addition to text. So you can kind of see and speak and not just kind of scribble walls of text, but you can actually take a photo with your smartphone, for example, when you’re describing a problem and that photo becomes part of the support experience.

05:28
Gerardo Pelayo, Ph.D.
And that’s really relevant. I remember even when I was a child, I heard and I learned that we all internalize information in different ways, right? Sometimes some of us like to listen to it, some of us need to see it on a whiteboard or like some videos, as you mentioned. And the other thing that to me is very interesting is that it’s both inputs and outputs. And I think that’s a very important distinction because we’ve seen some forms of AI that have progressively been getting better at capturing information that is trapped in different modes from the structured and unstructured text, but then we would get into tables and diagrams and even videos. But most of what I’ve seen actually consumes all of this and then produces text based outputs. So the ability to do this in that wider spectrum, I think it’s critical.

06:33
Gerardo Pelayo, Ph.D.
And so my next question around that is, you talked about the type of assets. Can you get a little bit deeper into the type of examples that actually benefit from this multimodal approach to providing support for the customers?

06:48
Shan Lilja
Yeah. So I mean, in our case, some typical examples would be troubleshooting, for example. Helps a lot to see what the user sees. If you’re troubleshooting a product, if you excuse me, I just have an incoming call here, let me cancel it. Yes. So troubleshooting and then, you know, if you onboard a user to a new product, it helps to show them. It helps to, you know, for example, being able to speak to them if they have their hands busy while they are trying things. If you have a warranty claim, for example, it can be a much easier experience both for the user but also better for the company if you can take photographs to verify something that has broken, for example. So kind of reduces fraud.

07:41
Shan Lilja
It’s basically anytime you would have an advantage by having a human on the scene together with the customer. This basically lets the AI be on the scene with the customer. So it’s rare when it comes to physical products that there are not use cases where it benefits from this. But even yeah, so like in general, anytime where it helps to see or use audio or any other type of modality, it’s helpful.

08:14
Gerardo Pelayo, Ph.D.
Good. And the subtext that I get from that, and correct me if I’m wrong, but it’s not necessarily the Pareto that we usually go for, which is the standard. The 70, 80% of you know what to do. It’s kind of straightforward to consume the information. But it’s the long tail of complex scenarios where text based sometimes is not enough and it would lead to a lack of clarity or to having to revisit the information multiple times to do it without confidence and therefore try to resort to other means and require that person to person interaction. Is that, would you agree with that?

09:00
Shan Lilja
Yeah, that’s a very. So I gave you some examples, but I guess the way I would kind of characterize the main benefit there is you can often really make the customer experience much more natural. You can reduce avoidable friction like you can reduce effort a lot for the customer by, you know, letting them just show you rather than having to describe something or, you know, While they are showing you something with an image or with a video that they can also speak at the same time. That’s really much more natural for somebody who has a problem with something and trying to communicate to support what they need help with. So you know, the really big win from a kind of customer input perspective is reducing the friction for them to express themselves.

09:52
Shan Lilja
And it’s often even like my grandma or there are a lot of older people who actually struggle just with writing walls of text to chatbots. It seems intuitive to a lot of people, but a huge percentage of users actually struggle with expressing themselves in text to a bottle. I think that’s something that’s kind of true in general. And then the other thing is it can also help with actually improving the resolution rate and the accuracy in the support experience because you can ground things much more, you can have much more information density basically if you can take a picture of how a situation looks like. It also constrains the AI from hallucinating in some situations, for example, is harder to bullshit with an image than with text only if you’re just speaking.

10:47
Shan Lilja
If you mess with an image, the user will immediately see that the image doesn’t match reality. And that’s kind of much harder way to kind of give the user plausible sounding but kind of inaccurate instructions correct.

11:04
Gerardo Pelayo, Ph.D.
To get these different modalities to cross check the information and to prevent many of the errors. I, I like that. And, and I. So the part about reducing friction really lands with me. But, but I, I want to play devil’s advocate for a little bit and talk about some of the roadblocks that we see also in particular as we connect multimodality to AI. Because operationalizing AI is still one of the things that most service leaders are struggling with. In fact, for our state of remote support and self service resolution, for example, 45% of service leaders said that inability to operationalize AI was the main reason for the challenges that they faced. This was the number one ranked reason by them.

11:53
Gerardo Pelayo, Ph.D.
And part of the explanation I, I believe is, or I related that part of the story is that even the technicians that are consuming the information currently are struggling with that. And the key element there is trust. So my question to you is, have you come across, have you found that multimodality influences trust in any way, positive or negative? And if you can elaborate a little bit on that.

12:31
Shan Lilja
Yeah, so to my previous point, because it’s more difficult to kind of bullshit the end user with. For example, if you’re producing Images as output like you’re describing. Hey, do like it says in this image, please. That actually can make the user trust the system much more because they can immediately like they know when to trust AI, so to speak, because if they see an image and that image matches the product they have in front of them, they can trust that this is a real action that I can take with this product. For example, I think if you are text only, it definitely becomes harder for the user to know when to trust AI and when not to compared to if you use visual instructions, for example.

13:21
Shan Lilja
I think also the reduced complexity in the experience itself because it’s more natural to communicate with voice and visuals and showing rather than just telling and so on. I think that also increases trust. Like you feel like this is like, it’s almost like you’re facetiming with an expert when you go through the support experience rather than kind of just being forced to just write text. So yeah, I think, I suppose that’s kind of trust from a user perspective. If you look at it from. Again, you’re trying to play devil’s advocate, right? So let me also try to play a little bit devil’s advocate. I guess one thing that we encounter people talking about is what will happen when you can film the environment of the end user. What about nudity? All kinds of things can show up that you don’t see in text.

14:23
Shan Lilja
You know, you capture things in the background that you shouldn’t be capturing and things like that. I think those are still things that kind of need to be solved. It’s, you know, my gut kind of instinct when it comes to, you know, safety or security issues like that. Is that like long term? I think it’s not the right solution to handicap these systems and use less information to. In order to not capture bad information. I think there will be new norms and to some extent perhaps technical solutions as well that will kind of mitigate these things. But right now, to be fair, it is something we get asked a lot like this seems to be one of the first questions are especially the nudity questions. What happens if things are inappropriate and they start to show up and things like that.

15:18
Shan Lilja
Long term I’m not afraid of that, but short term, I think that it’s a hump that needs to be resolved for sure.

15:27
Gerardo Pelayo, Ph.D.
Yeah, yeah, but, but what you were describing also, I, I do believe it outweighs the resistance that you brought up, which it’s great that you did. But part of the message that I got was the customer gets more Confidence, even about themselves, about am I following the instructions correctly? Because it’s not just am I reading and interpreting in the right way, but because I’m seeing the video, because I’m listening to it, I feel better that I’m doing the steps right and therefore that this should work. If you see that accuracy in the steps to begin with, and as we get into that, I’m listening to it as well. One of the benefits that I’ve heard is also the.

16:18
Gerardo Pelayo, Ph.D.
The impact that it can have on safety, for example, which is an objective that we’ve seen service leaders pay an increasing level of attention towards, along with mental health. So has that come across in your conversations? How would you say multimodality influences safety? The scenario that I just to make it a little bit more specific that I was thinking of is if you need both hands to actually work on the piece of hardware, or you’re balancing other things while you’re doing it instead of holding your phone and having to read through it, if you can listen to it and be doing it live without distracting your eyes from the place that you’re working on, I would expect that it makes it less likely that you have an accident or things like that. Is that part of the storyline, part of the argument?

17:24
Shan Lilja
Sure, yeah. In general, it’s easier to break things up piecemeal. It becomes much more natural to kind of, you know, wait until you get visual confirmation that a state has changed, for example. So it’s like, you know, hey, do this, and then you look at the state has changed. You ask the user to do something else. I think it’s, I mean, to be fair. Right. I think it’s probably too early to say. I mean, if we’re talking about physical bodily harm. Right, like safety in that sense, probably the most important thing is how accurate are your instructions. What you don’t want to do is to give the user some wrong advice that causes them to electrocute themselves. You want to make sure that there’s instructions that are grounded in real knowledge and that the user interprets, if it’s instructions, that they interpret instructions correctly.

18:24
Shan Lilja
So I think from that perspective, having a richer experience, having a more multimodal experience probably improves safety. But I think safety is also about a much broader set of issues. It’s a challenge in any AI support interaction. So I guess speculatively, I guess multimodal would help that a little bit. But I think it’s. I think it’s too early to say for sure.

18:58
Gerardo Pelayo, Ph.D.
Yeah. Yeah. I think specific scenarios, like when that Visual would be clear on which cable to cut or what, which cable to stay away from, things like that. Sure, yeah, Having an image, having a video would be helpful, but so I spoke a little bit about that customer experience and how that changes when multimodal becomes accessible based on the interactions that you’ve had. Because you spoke also about how automation is the goal for this type of support experiences. What’s necessary, what are some of the best practices for this customer interactions to be clear, to be engaging, to be effective. So that automation is the chosen path of interaction for support.

19:59
Shan Lilja
Yeah. So I mean, I think if we’re talking about the clarity of the interaction, how effective it is and so on, I think that the number one objective of customer support, automation in general, like AI support in general, is to save user effort. So if you’re saving user effort, if you’re avoiding needless friction, if you don’t kind of add to the hassle that the user already experiences by having to call support and kind of being annoyed, that’s like the number one way to make that experience feel effective, feel clear. And our kind of deep belief, deep conviction about what we’re doing is that multimodal support is actually a surprisingly good way to reduce perceived effort for the end user.

20:55
Shan Lilja
Like it’s a surprisingly good way to reduce friction because it’s so natural for human beings to just talk, to just show things, to show the scene rather than describe the scene with text, for example, and in some cases actually get sensor data from the product automatically and be able to tell all kinds of things about a lawnmower based on the battery that’s left, error codes and so on, but in general, just more sensors involved. It’s a really good way to increase clarity because if you made an experiment and one set of users could only describe things with words, they could only scribble text, and one set of users could snap photos, film the scene and so on. The second group of users, you would find that it’s much more clear for them, it’s faster for them to do things, everything feels much more smooth.

21:49
Shan Lilja
So I think it’s one of the best vectors for reducing friction in customer experience is to make the experience more multimodal.

22:00
Gerardo Pelayo, Ph.D.
Yeah. And as you were talking, one thought that occurred to me is people have their scart issues with automated support, right? Including, or perhaps most notoriously, with the automated calling systems. And you have to press a 3 and then a 7 and then a 6, and by the time you get to it, you forgot your question probably. And I think multimodal has an opportunity to, to invite people to self select, given that it’s not limited to audio only or to text only, but as you said, it’s really meeting them where they are in the way that they act most naturally. And, and that’s what makes it different from other self support mechanisms that have been attempted in the past that it does feel more natural.

22:57
Gerardo Pelayo, Ph.D.
And within that self selection, I believe there are going to be situations still where humans need to be part of the support experience. Part of the answer is when does a human need to lead the interaction? When do you start with a multimodal AI approach? Even if you go for the ladder, if it needs to be escalated to a person, how does that happen? I’ve heard the resistance of both from a customer’s perspective and from the person providing the support. Are we going to have to start all over again? So can you speak a little bit of how would that human AI interaction work once a problem gets escalated to a human?

23:55
Shan Lilja
Okay, so you’re saying in situations where the AI can’t solve the problem, how would multimodal help in that situation?

24:04
Gerardo Pelayo, Ph.D.
Yes. So does the tier 2 support individual? Do they have to start all over? Is there anything that they can leverage from the previous conversation? How does that happen?

24:16
Shan Lilja
Yeah, the short answer is that multimodal provides much richer context. That’s very useful for the human agent. If they have to pick up a session that has escalated, for example, they can look at the scene. There might be photos, there might be video, there might be sensor data. They might know the error codes that has been in a product. They might know other information about the product. Like I mentioned, battery power, whatever. They know, of course, what has been said from an audio perspective, they can get that summarized very quickly and they see any text input that the user has given as well. If they have entered their name or any form date or whatever, it might have been more appropriate to ask the user to spell things out. They will see all of that in one place before they pick up the session.

25:16
Shan Lilja
So I think that what it does is it reduces frustration a lot for the user because they don’t have to explain things. Again, that’s very obvious. They are seeing something that’s annoying in front of them and they have described it already, they have snapped a photo and so on. So it really saves a surprising amount of frustration and time for the, you know, for the end user and for the agent. You know, you would think that this is kind of more marginal than it actually is like when you see it a few times, you really, you get hit by the fact that, wow, this actually there’s a lot of relevant context here that, you know, you’re not capturing just with pure text. I said short answer. It was a longer answer than I was, than I intended.

25:58
Shan Lilja
But I guess, you know, maybe somewhat more interestingly, the way that we look at support is that there are repetitive issues that come over and over to a customer support organization. It’s the same type of issue. And some of those repetitive issues are very quick and easy, very routine. They can probably be handled with traditional ways of doing automated support, even IVRs or decision trees, whatever. But then there are kind of more complex, much more time consuming requests, right? Like it could be something that takes half an hour to go through with the customer. And typically those things that are much more time consuming, even though they’re repetitive, they tend to be always escalated to human agents.

26:48
Shan Lilja
Like it’s kind of taken for granted that they’re just certain if it’s a complicated troubleshooting request, for example, you know, you can’t really troubleshoot with AI for like 25, 30 minutes, then you tend to escalate. So the other kind of big advantage with multimodal AI and just better AI support in general is that you can handle much more time consuming, much more complex requests that are still repetitive. There’s still the same things coming over and over. People are asking, how do I use this product? Or there’s a certain error that takes a long time to kind of figure out how to fix. And so it’s not just that if you escalate, you save a lot of time, but it’s also that you can solve much harder requests that are precisely the ones that eat up a lot of time for human agents.

27:40
Gerardo Pelayo, Ph.D.
Good. That was a very thorough response. Thank you for that clarity. And before we run out of time, I really want to get into the value conversation because you painted a really good picture about what’s possible. What are the right use cases, how does your interact with the rest of the support experience? One of the things that we observed in our 2025 State of AI and Service Technology, and we’re launching the 2026 version in under a month. But was that most of the near term impact that service leaders were going for was around the customer experience, the frontline experience. But when it came to the impact that they wanted to see, in order to confidently scale those AI innovations, 72% selected at least one of the financial metrics, either higher Revenue, lower costs, higher margins. What is your advice?

28:45
Gerardo Pelayo, Ph.D.
What is your perspective on how should organizations measure the success from multimodal AI both in the near and the long term?

28:59
Shan Lilja
Yeah, that’s a good question. So I guess the interesting way to answer this question rather than just say, look, there’s a lot of ways that I have opinions about how success should be measured in general in AI support. But I think the interesting thing about multimodal AI support in particular is that it’s basically proxies for how it reduces friction in the end user experience. So my view on customer support is that the essence of it is to reduce avoidable friction, reduce avoidable effort, make it more smooth for the customer. That’s more important than saving cost, more important than saving time alone. It’s really like take the hassle out, that doesn’t need to be there. That’s the number one objective. And multimodal is really great at reducing hassle. It’s really great at reducing friction.

29:55
Shan Lilja
In the experience, for all the reasons that I mentioned earlier in our conversation, you show rather than tell. An image is worth more than a thousand words. It’s very natural, organic, intuitive for people to interact in this way. So I think the measure that I would look at that would be really interesting is to be much more ambitious about measuring how friction is reduced. So for example, look at not just resolution rate, but abandonment rate. So people, if it’s more natural to engage, maybe they won’t abandon the AI conversation as much as they’re doing currently. Or look at actually the time saved, how much shorter are the interactions even though you’re solving the same problems. And I think measuring the reduced friction is super hard, but it’s actually, I think the most important thing to measure.

30:45
Shan Lilja
Like that’s true whether multimodal or not multimodal or whatever. But I think that’s, I’d say this often to customer support leaders, but if there’s one thing that I would do differently than what they’re doing, it’s a bit of a proactive way to put it. But that’s basically to try to measure the saved user effort, even though it’s hard.

31:10
Gerardo Pelayo, Ph.D.
Yeah, no, I, I agree. And it influences the satisfaction of the frontline as well, which we talked about. But as talent retention becomes or stays a big challenge across service leaders, this definitely helps. And customer retention and growth, it was the number one focus area for service leaders last year. I expect it will be high up in the rankings this year as well. We’ll get the results in a couple of weeks. So. So yes, that reduction of friction and the focus on measuring it and understanding and doing something about it really resonates. Sean, you’ve illuminated so much, but I, I would be failing if we didn’t have our first podcast of the year and I didn’t ask you about the outlook. So what is it that you’re most excited about as we start the year?

32:13
Shan Lilja
I actually think that 2026 will be the year in my domain, which is AI support automation, where multimodal becomes a thing because obviously already voice has become a big thing. Everyone is Talking about how LLMs can be used to automate phone calls. But I think there are a lot of extremely good multimodal models out there and people actually are used to talking with ChatGPT with multimodal or they use Gemini for multimodal conversations. They take photographs of their car when the car breaks down and actually it’s starting to become an end user experience, including when the end users are customer support leaders. What happens to be a customer Support leader uses ChatGPT, whatever. So I think multimodal in support is not going to be 2027. I think the big breakthrough year for multimodal support will be this year.

33:09
Gerardo Pelayo, Ph.D.
I love that optimistic view. For your sake and for the betterment of service, I hope that you’re right and you’ve helped significantly in having us understand when and how to go across it. So finally, just very quickly, I always like to wrap up with a little bit about yourself. Is there any passion project outside of work, something that you’re looking forward to personally that you’d like to share with the community a little bit about Shannon himself.

33:47
Shan Lilja
I have to answer that. Yesterday I learned that I’ll become an uncle. I can’t answer anything other than that. That’s my passion project for 2026.

33:59
Gerardo Pelayo, Ph.D.
Absolutely. That’s a great answer. That’s very hard to compete with. Shan, I want to thank you so much for being here today again for the community. Thank you for listening in. Today’s podcast reminder was recorded, will be available across our platforms and we have a very exciting year ahead of us. Podcast and all sorts of digital and in person events for you to join. So looking forward to that. Have a great year everybody and see you soon.

Related Content

Lead, Share, Partner With Us

The Service Council is made up of global leaders who are actively committed to the exploration of critical topics, insights and best practices. Please join our team of professionals and experts. We’re stronger and better together.