Look Broadly, Think Deeply, Turn AI into Value — Designing AI for Real-World Use

2026.09.07

Making AI usable in the real world requires thinking about more than model performance. Business decisions, system design, evaluation, and operations need to be considered as one whole. In this article, Hideaki Okamoto, a Forward Deployed Engineer (FDE) at PayPay, draws on experience in implementation, evaluation, and operations to discuss how to connect AI to value in the real world.

Hideaki Okamoto

PayPay Corporation, Product Division, AI Transformation Department.

After researching machine learning and computer vision at university and graduate school, he worked at SoftBank on AI product research and development as well as software implementation. He joined PayPay in 2023 and currently works as an FDE on AI applications and real-world operations design.

Introduction

When I worked on a real-time voice AI implementation and tested it over a phone connection, generating an answer was not enough. When I looked at the timing of the conversation and the stability of the responses, there were still issues to address before real-world use. Tracing the logs showed that the challenge was not the AI model’s ability to generate an answer, but where to divide a user’s speech, when to return generated audio, and how to handle interruptions during playback.

As development progressed, I was reminded that having AI generate a correct answer and turning that answer into an experience users can trust are two different things. There is a long distance between the two, involving business understanding, experience design, evaluation, and operations. Closing that distance one step at a time is what engineering AI for real-world use means to me.

I began working with AI through research on machine learning and computer vision at university and graduate school. Since then, I have worked on making AI usable within products and business operations. At PayPay, I am involved in a range of AI initiatives, working with people in the business to organize problems and carrying that work through design, implementation, and validation.

AI is not the goal; it is a means of delivering value. In this article, I draw on development experience to share what I have learned while working on AI applications at PayPay.

Let the Work Determine Where AI Belongs

When discussing AI, it is easy to start with technical and methodological questions such as model selection, RAG (a mechanism that retrieves relevant knowledge and uses it to generate answers), tool integration, and evaluation infrastructure. These are all important. Choosing an appropriate model and architecture, and designing a suitable development process, are essential to building AI that works in practice.

However, the first question should not be what we want AI to do. We first understand what the problems are in the workflow, where decisions are required, and where time and effort are being spent. Rather than starting with what AI can do, we work backward from the operational bottleneck to determine where AI can create value.

Should AI answer questions, gather and organize information, connect multiple systems and tools to move a workflow forward, or support human judgment? The answer depends on AI’s role within the entire workflow, not on the performance of AI in isolation.

When considering customer-facing operations, for example, we first need to understand how the work classifies inquiries and what decisions determine how each one is handled. Which inquiries can be handled with a standard response? Where is additional confirmation needed? Which topics require careful handling? Which cases are handled by a person or a specialist team? Only after organizing these operational decisions can we see where AI should be placed and what it should be responsible for.

This is not merely prompt design; it is also business process design. In some cases, improving a workflow or an existing system is better than introducing AI. If a problem can be solved with fixed rules, process improvements, a simple business application, or better visibility, choosing one of those options is also a decision made by working backward from the purpose.

What Voice AI Revealed About the Difference Between “Correct Answers” and “Usable Experiences”

With voice AI, generating a correct answer and creating a conversation that works as an experience are different things.

The value is not in AI speaking by itself. It is in helping users reach the information they need without getting lost, while enabling the service to provide stable interactions over time.

In a chat interface, users can look at the screen while they wait, even if there is some delay. In voice interactions, however, small differences in pauses and response length can significantly change the impression of a conversation. We need to design not only the answer itself, but the flow of the conversation: where to decide that a user has finished speaking, when to start responding, how to ask for clarification when speech was not recognized well, and how to handle overlapping speech.

Between the telephony platform and the real-time voice AI, audio and control messages travel in both directions, so the conversation state must be handled consistently. In one implementation I worked on, a Bridge handled that connection and state management. It was implemented with Python’s websockets library and ran on ECS. The Bridge relayed audio and control messages while managing speech boundaries, the start of AI processing, response-audio transmission, and playback state. In CloudWatch Logs, each process could be traced independently.

During a phone call, the system must continue receiving connection-maintenance messages and the next audio input even while AI is returning response audio. We therefore separated audio transmission from the receiving process and ran it asynchronously in the Bridge.

However, when we tried the system over an actual phone call, the next conversation turn sometimes started before the previous response had finished playing, causing responses to accumulate while waiting for playback. By tracing the logs in chronological order, we found that the next utterance was being processed as a normal conversation turn during the brief interval between the Bridge starting to send audio and receiving the playback-start notification from the telephony platform.

We therefore managed the period from the start of response-audio transmission through playback completion as one continuous “AI responding” state. Utterances received during that period were held temporarily and processed in order after playback completed. This change prevented responses from accumulating while waiting for playback and allowed the conversation to proceed in the correct order. Keeping one process from blocking and advancing the conversation correctly are different problems. In real-time voice AI, we need to design not only asynchronous processing, but also which parts of the conversation state should be treated as one continuous sequence.

The figure below shows this state management and how speech received during playback is handled.

To evaluate the voice AI’s response as a whole, we separated and checked voice activity detection, end-of-turn detection, finalization of speech-recognition results, the start of AI processing, the first audio output from speech synthesis, and playback completion. In particular, we looked separately at the time until audio started playing and the time until the full response completed, so that perceived waiting time would not be confused with internal system processing time. We also checked the processing state and timing in the logs, listened to the actual audio, and confirmed that the conversation did not contain unnatural pauses or interruptions.

Finalizing the end of speech earlier can make the response faster, but it also increases the risk of treating a pause while the speaker is thinking as the end of the utterance and starting a response too soon. Conversely, waiting longer can reduce false detections, but the slower response can damage the rhythm of the conversation. We traced these trade-offs in the logs and made it possible to evaluate where the experience had broken down.

Generating everything in real time is not always the best option. For example, for an opening message whose content never changes, using audio generated in advance can create a faster and more stable experience than generating it with AI each time. On the other hand, when the content changes according to the user’s situation, AI can add value by responding based on the context. Separating the parts that should use AI from the parts that should follow a fixed process was another important decision in designing the conversation as a whole.

Another difficult challenge was deciding how to switch cases where an incorrect response could have a significant impact to a safer handling path. Which cases should receive the normal response, and which should switch to a fixed message, additional confirmation, or human handling? We first organized the operational categories and expected behavior, then defined the range of responses AI may provide and the conditions for switching to a safer handling path as guardrails. Here, guardrails do more than identify inputs that require caution; they define the scope of AI’s responsibility for responses and what happens when that scope is exceeded.

However, the same intent can be expressed in different ways because of paraphrases, variation in speech-to-text results, and surrounding context. Even when the system behaves as expected for one expression, it does not necessarily handle every input in the same category consistently.

If we adjust the prompt, context, or decision logic for each individual failure case, that case may improve while another paraphrase is missed or an ordinary input is detected too aggressively. We therefore avoid stopping at local tuning. We include paraphrases and speech-to-text variations from the same category, as well as ordinary inputs that are likely to cause false positives, in the evaluation set, and check both missed detections and over-detections.

We maintain the baseline evaluation set and metrics, while accumulating newly failed cases for regression checks. After a change, we use the same baseline, including the existing cases, to confirm whether the intended improvement was achieved and whether the quality of other cases was affected. If the expected quality is not met, we roll back to the previous state. Building this sequence into operations makes it possible to improve the guardrails continuously.

The design target is not limited to the AI model. It also extends to the applications, infrastructure, and operations that support it. The value of engineering AI for real-world use lies in looking broadly across the business and technology, quickly shaping and trying ideas, and learning from failures to make improvements.

Turning Development Speed into Sustained Value

AI development has at least two sides. One is using AI as a developer to accelerate research, design, implementation, and review. The other is developing products and systems that incorporate AI to solve problems in the business.

I use ChatGPT, Codex, and Claude Code for research, comparing design options, implementation, and review, choosing among them according to the nature of the work and the development phase.

This has made it faster to compare options, draft specifications, implement ideas, and identify evaluation perspectives. However, that speed does not automatically become value for the business. We need to verify the assumptions, constraints, and failure behavior of AI-generated designs and code, and connect the resulting speed to business understanding, design, evaluation, and operations. The developer remains responsible for the final judgment.

For continued use in the business, we also need to consider more than accuracy: speed, cost, stability, security, logs, evaluation, updates, rollback, operational burden, and other factors.

RAG and FAQ integration are not complete once the system can search and generate an answer. We need to treat the knowledge, answer controls, and evaluation as one system: what information may be referenced, how updates are reflected, how the basis for an answer can be checked, how incorrect answers are suppressed, how quality is evaluated after changes, and how the system can be rolled back when there is a problem.

These evaluation and operational practices are sometimes called MLOps or LLMOps. The key point is whether the product can continue to grow with changes in the business, the field, and the technology, and remain usable over time.

Start from the Problem, and the Technology Choices Change

This way of thinking is not limited to voice AI.

In another operational-support initiative, a large-scale RAG platform and a complex AI-agent workflow were initially among the options. As we examined the actual work, however, we concluded that a simpler workflow would fit the requirements better: use an LLM to interpret and organize inputs, then return results based on fixed business rules and defined knowledge.

For workflows that require careful judgment while checking multiple sources and pieces of evidence, AI should not make the final decision. Instead, it should organize the information needed for a human decision and show the evidence that needs to be checked. Separating information shown to the user from information used as internal decision support is important. The greater the safety and accountability requirements of a workflow, the more important it becomes to define the scope of AI responsibility, explainability, auditability, and operational guardrails.

On the other hand, when working backward from the purpose shows that a deeper technical approach is necessary, we combine multiple technologies. In an environment where connections to external services are restricted, we have built a RAG application that combines a local LLM, an embedding model, and a search index so that information retrieval and text generation can be completed within the environment. We designed the distribution and update process as part of the same effort. We have also used LangGraph to control an AI-agent platform in situations where multiple AI agents divide roles to advance a single workflow, including evaluation, auditing, and observability of the execution results.

The important thing is not to decide in advance whether to use complex technology or a simple system. First, ask whether AI is really necessary. Even when AI is appropriate, ask whether RAG or a complex AI agent is necessary. Determine this step by step from the business needs and constraints. Combine technologies deeply when necessary, and simplify the architecture when they are not. Both are forms of engineering that work backward from value in the real world.

The ideas developed through choosing the necessary architecture, evaluating it, and improving it should not remain limited to individual product improvements. Preserving knowledge about evaluation, auditing, updates, and operations as organizational knowledge, standards, or intellectual property, and making it reusable, is also part of designing for sustained value.

Moving Between the Business and Technology

In the initiatives I am involved in at PayPay, we work with people who understand the operation to organize business decisions, turn them into requirements, implement them, test them in conditions close to actual use, and return to the next improvement based on logs and evaluation results. Collaboration with people who bring expertise in operations, products, development, infrastructure, and security is essential.

Large-scale services have diverse user touchpoints and business operations. This creates many opportunities to use AI, but it also requires us to think about reliability, safety, and continued operations in addition to convenience. The appeal of working on AI at PayPay is being able to participate from problem definition through design, implementation, evaluation, and operational improvement, and to grow AI into something people actually use.

As a Forward Deployed Engineer (FDE), I value moving back and forth between the business and technology and connecting this entire process. Rather than translating the words of the business directly into technical requirements, I try to understand the purpose and decisions behind the work. Rather than only showing what is technically possible, I work with others to consider how it can be operated. This back-and-forth is how I help move AI from validation toward value.

Look Broadly, Think Deeply, Turn AI into Value Quickly

Look broadly. See not only technology, but also the business, product, data, operations, and the broader business context.

Think deeply. Consider the options, assumptions, risks, quality, evaluation, and scope of responsibility.

Turn ideas into value quickly. Run the cycle of research, prototyping, implementation, feedback, and improvement to bring the solution closer to something people can use in practice.

AI is certainly making this entire cycle faster. But building quickly is not the goal by itself. What matters is thinking quickly, trying quickly, learning quickly, and connecting the results to a product that can continue to grow as conditions change.

Using AI effectively is only part of the work. We also need to think about where to place AI, where not to use it, and how to bring it into the business. Through my work at PayPay, I want to keep developing that ability and turn AI from a technical experiment into value that reaches the people and operations it is meant to serve.

Open Positions

Career