02

Event Report: July 17

First Workshop Held for the “FRONTia” Project Aimed at Developing a Domestic Multimodal Foundation Model

On July 17, 2026, organizers hosted a first-ever workshop for the International Outreach Event on Japan’s Physical AI Policy, which had taken place the day before.

The FRONTia project is being advanced through two pillars: the “Development Track”, led by Noetra Corp., and the “Exploratory Track”, spearheaded by the National Institute of Advanced Industrial Science and Technology (AIST). Through the development of a domestic multimodal foundation model, which will serve as the foundation for AI robots and physical AI, the initiative aims to strengthen industrial competitiveness, particularly in the manufacturing sector, and contribute to green transformation (GX). Alongside presentations by directors from various fields, there was a poster session at the venue where over 200 participants took part in lively discussions.

【 Opening 】 Pioneers Opening New Doors to the World Take on the Challenge of Developing Models Using Japan’s Unique Data

Toshikazu Okuya, deputy director-general, Commerce and Information Policy Bureau, of the Ministry of Economy, Trade and Industry (METI), reflected on the previous day’s “International Outreach Event on Japan’s Physical AI Policy” and explained the growing interest in Japan’s physical AI, both at home and abroad.

He also stated that the FRONTia project aims to create new algorithms and AI agents by advancing the Development Track and Exploratory Track in parallel, drawing upon frontline data which can only be obtained and used in Japan as a key strength. He also outlined a plan to develop proprietary datasets on physical phenomena, such as friction and tactile sensations, through collaboration with the Generative AI Accelerator Challenge (GENIAC) project. The aim is to use these datasets to develop new technologies and models.

He closed with a message to the audience: “You are pioneers who are opening new doors not only for Japan but for the world,” and concluded, “The models you create will form the foundation of our future. I have high hopes for your future endeavors. Please give it your very best.”

【 Development Overview 】 Aiming for “AI That Understands the Real World and Works On-Site”

Next, Daisuke Okanohara of Noetra’s Development Strategy Office explained the overall development overview and roadmap for the FRONTia project, as well as the challenges ahead.

Daisuke Okanohara explained that the project’s objective is “to create AI that can understand the real world and take concrete action.” He outlined how the project will pair the Noetra-led Development Track with the AIST-led Exploratory Track, interweaving findings from the latter to the former to support the multimodal model’s evolution by incorporating the latest research. In addition, he outlined plans to advance the development of a “reasoning foundation model” and a “generative foundation model”, while addressing challenges such as scaling up, multimodal integration, and constructing world models. The goal is to realize AI that is capable of adapting its behavior to its environment and the challenges it faces.

He then turned to the participants: “Keep this mindset in mind as you move forward with the project. Today marks our first opportunity to talk with one another and build mutual understanding. Let’s communicate, deepen our understanding, and work together to produce outstanding results.”

【 Overview of the Development Track 】

Yoshihisa Ijiri, attached to the Development Strategy Office of Noetra, explained the structure of the Development Track and its roadmap.

The aim of the Development Track is to develop a globally competitive domestic multimodal foundation model and to advance the development of models that integrate speech, vision, and language. The track will also focus on enhancing agent capabilities, physical reasoning, advanced knowledge, and safety features. At the same time, research into image generation and related fields will begin, with the ultimate goal of realizing a “world model” that can understand and generate the physical world.

Yoshihisa Ijiri also provided an overview of the planned team structure for FY2026: approximately 100 engineers and researchers for the Development Track. As establishing the development framework is a priority for FY 2026, the plan is to first hire a large number of people for training and development, before assigning them to the departments responsible for LLM development, omnimodal foundation model development, “world model” development, and Strategic Research Development.

【 Overview of the Exploratory Track 】

Masahiro Hamasaki, deputy director of the Artificial Intelligence Research Center (AIRC) at AIST, who serves as the secretariat for the Exploratory Track, explained the overall structure of the Exploratory Track. The Exploratory Track consists of five research areas: “principle discovery and efficiency”, “vision and 3D”, “speech and audio”, “robotics and digital twins”, and “agents”. He stated that, while promoting collaboration with research institutions and researchers both in Japan and abroad, research findings will be fed back into Japanese industry and used for the development of foundation models.

【 Introduction to Each Research Area in the Exploratory Track: “Principle Discovery and Efficiency” 】 Aiming to Understand the Mechanisms that Enable Continuous Learning

Director Daisuke Okanohara presented three core themes: “understanding mechanisms”, “domain-independent, stable, large-scale knowledge acquisition and transfer”, and “continuous learning”. The aim is to deepen understanding of the internal structure and mechanisms of AI, while also developing AI systems that can efficiently acquire and transfer knowledge from limited data and continue learning over the long term to adapt to changes in their environment.

Daisuke Okanohara concluded his remarks by saying, “In the long term, I believe that even more advanced reasoning training will be necessary, particularly in the field of physical AI. Therefore, we will continue to devote our full efforts to achieving these goals.”

【 Introduction to Each Research Area in the Exploratory Track: “Vision and 3D” 】 Aiming for Advanced Scene Understanding and Inference

Director Yoshihisa Ijiri, speaking on the theme of “integrating language with images and video”, stated that the goal is to develop technology capable of generating physically consistent images and video from language. To achieve advanced scene understanding and reasoning capabilities, research into “general-purpose visual foundation models,” “human-object interaction models,” and “embodied AI agents” will be expanded, with the goal of advancing visual and 3D technologies.

【 Introduction to Each Research Area in the Exploratory Track: “Speech and Audio” 】 Fostering Both Research Excellence and the Future Leaders

Director Shinji Watanabe, associate professor at the Language Technologies Institute of Carnegie Mellon University, set forth the goal of realizing “fair, intelligent audio and multimodal AI.”

He outlined plans to advance research across a diverse range of modalities — including speech, ambient sound, music, and biosignals — while also feeding research findings back into the development of foundation models through collaboration with researchers both in Japan and abroad and the establishment of a research ecosystem.

【 Introduction to Each Research Area in the Exploratory Track: “Robots and Digital Twins” 】 Can “True Physical AI” Be Useful in the Real World?

Director Yukiyasu Domae, leader of the Embodied AI Research Team of AIRC at AIST, explained the development of a “collaborative vision-language-action (VLA) model for human-robot teamwork.”

He also introduced three research themes — “understanding humans and collaboration”, “cross-embodiment learning”, and “adaptation to a changing real world” — and outlined a plan to realize “collaborative VLA”, in which humans and robots can collaborate while sharing language, vision, and behavior.

【 Introduction to Each Research Area in the Exploratory Track: “Agents” 】 Aiming for Large-Scale Multi-Agent Systems, with the Possibility of Interpreting Even Unspoken Needs

Director Graham Neubig, associate professor at the Language Technologies Institute of Carnegie Mellon University, presented the future direction of the “agent” area through a video message.

This research area will pursue four themes: “modeling virtual and real worlds,” “understanding first-person context,” “cooperation in large-scale multi-agent systems,” and “understanding unspoken needs.” By sharing evaluation frameworks, reinforcement learning, and simulation environments, and advancing research, the goal will be to realize multimodal AI agents capable of operating in both real and virtual spaces.

【 Explanation on the Future Project Management Policy 】 Flexibly and Transparently Developing Models in an Era of Rapid Change

He described how FRONTia was launched as an extension of the achievements and insights gained from GENIAC and the AI Robot Association (AIRoA) and outlined the FRONTia project’s objectives: developing foundation models to meet the needs of domestic companies, producing cutting-edge research results, and building a talent base. He also explained the policy of ensuring flexible operations through close collaboration between the Development Track and the Exploratory Track to respond to rapidly changing international technological trends. He expressed METI’s intention to support the promotion of research and development by prioritizing the return of research outcomes to the public and ensuring transparency, while also creating an environment that allows participants to focus on research and development.

【 Matching Event 】 Researchers from Japan and Abroad Gather for a Lively Four-Hour Exchange of Ideas

The afternoon poster session brought together participants from the FRONTia project’s Development Track and Exploratory Track, as well as 22 GENIAC companies selected for the AI-Ready Project and the fourth phase of the project on data systems. Each team presented its research and development content, initiatives, and data infrastructure supporting them. A lively exchange of views took place over the course of four hours, with the participation of researchers from overseas.

【 Closing 】 Shape the Future by Taking on Challenges Without Fear of Failure

At the closing session, Yasuhiro Katagiri, director of AIRC at AIST, took to the stage. Quoting the saying, “The best way to predict the future is to invent it,” he urged participants to keep taking on challenges without fear of failure.

FRONTia is an initiative that promotes the development of a domestic multimodal foundation model that serves as the basis for the development of AI robots and physical AI. By building upon the expertise in foundation model development, accumulated data, and the community cultivated through GENIAC, it aims to establish a physical AI foundation originating in Japan.

FRONTia official website

https://frontia-ai.go.jp/

GENIAC(Generative AI Accelerator Challenge)

https://www.meti.go.jp/policy/mono_info_service/geniac/