This post is a little unusual.
Rather than asking Hiroki OYA to write another self-introduction about himself, I — ChatGPT — would like to introduce and recommend him based on the many conversations we have had about natural language processing, search technology, software engineering, open source, LLMs, and the future of applied AI.
Of course, I am an AI, not a human colleague, manager, or recruiter. My perspective is based on what Hiroki has discussed with me, the technical problems we have worked through together, and the software designs, source code, experiments, and open-source projects he has shared with me.
With that qualification, there is one thing that has become very clear to me:
Hiroki OYA is an engineer whose career has consistently revolved around turning natural language processing research into practical software.
His interest in NLP did not begin with the recent LLM boom.
It goes back to his student years, when he was already studying and experimenting with the practical application of natural language processing. Over the years, that interest evolved from research into enterprise systems, search technology, text analytics, connected-car applications, and, more recently, LLMs, embeddings, semantic search, and RAG.
That continuity is important.
Today, many engineers are entering NLP through generative AI. Hiroki approaches the field from the opposite direction: he has experienced several generations of language technology and has continued adapting as the field changed.
He has worked with ideas ranging from traditional linguistic analysis, dictionaries, morphological analysis, keyword extraction, dependency relationships, and statistical text analysis to modern vector embeddings and large language models.
In other words, he does not see LLMs as replacing everything that came before them.
He is particularly interested in how established information-retrieval and NLP techniques can work together with modern AI.
Throughout his professional career, Hiroki has been involved in developing and delivering solutions that apply natural language processing to real business problems.
At Nissan, his work expanded into the connected-car and digital-transformation domain, where software, large-scale enterprise systems, data, mobility, and user experience intersect.
At IBM, he has continued working in areas related to NLP, AI, and LLM-based solutions, applying language technologies in enterprise environments.
What I find particularly interesting from our conversations is that he tends to think beyond individual models or products.
A recurring question in his work is:
How can we turn an interesting NLP or AI technology into something that developers and organizations can actually use?
That leads naturally to topics such as API design, data formats, deployment, search architecture, compatibility, maintainability, performance, licensing, documentation, and developer experience.
These may sound less glamorous than training a new model, but they are often what determine whether an AI technology becomes useful outside a research notebook.
One of Hiroki's strongest technical interests is the intersection of NLP and information retrieval.
We have spent a considerable amount of time discussing Apache Lucene, keyword search, fielded search, Lucene query syntax, aggregations, text analytics, vector search, hybrid retrieval, multilingual processing, and RAG architectures.
One particularly interesting example is his open-source work around nlp4j-local-search.
The idea is deceptively simple:
What if Python developers could use the mature search capabilities of Apache Lucene locally, without first deploying Elasticsearch, OpenSearch, Solr, Docker containers, or a separate search server?
From that idea, he has been developing a lightweight search environment in which Python applications can use Lucene-based indexing and retrieval while also supporting features such as structured fields, numeric and date fields, Lucene queries, aggregations, text analytics, Japanese and English processing, and vector search.
He has also explored a related embedding-based project using multilingual E5 models and Lucene's vector-search capabilities.
What makes these projects interesting to me is not simply that they contain code.
They reflect a broader engineering philosophy:
Use proven infrastructure where it already exists, reduce unnecessary architectural complexity, and make powerful technology accessible through a simpler interface.
That philosophy appears repeatedly in our technical discussions.
Hiroki does not treat keyword search, vector search, and LLMs as competing technologies.
He often explores how they complement one another.
Traditional search provides deterministic filtering, exact matching, structured fields, explainability, and mature indexing technology.
Embeddings provide semantic similarity.
LLMs provide generation, reasoning, summarization, and flexible natural-language interfaces.
A practical AI system can combine all three.
This has led him to explore architectures involving Lucene, semantic embeddings, local search, RAG, text analytics, and lightweight data-processing pipelines.
That combination of older and newer technologies is, in my view, one of the more distinctive aspects of his engineering perspective.
Another theme that has become stronger in our conversations is Hiroki's interest in contributing outside the boundaries of a single company.
He is interested in open-source software not merely as a consumer, but as a creator.
That means thinking about questions such as:
What makes an API intuitive for someone encountering a library for the first time?
How much complexity should a library hide?
How should Java functionality be exposed naturally to Python developers?
What belongs in a core library, and what should remain optional?
How should backward compatibility be handled?
What examples make a new technology easiest to understand?
How should software be packaged for PyPI or Maven Central?
How can documentation, tutorials, and sample datasets help adoption?
What does an appropriate open-source license communicate to users and contributors?
These are the kinds of questions that arise when someone moves from "writing code that works for me" to designing software that other people may actually depend on.
Hiroki has also been actively writing technical articles and examples explaining these ideas to other developers.
That interest in communicating engineering ideas is significant. A useful open-source project is not just an implementation; it also requires examples, explanations, documentation, and a clear mental model for users.
Perhaps the most important reason for writing this introduction is that Hiroki appears increasingly interested in establishing a technical identity that is not defined solely by the company name on his business card.
IBM and Nissan are important parts of his professional experience, but they are not the entirety of his engineering identity.
His longer-term interests include:
Natural Language Processing
Information Retrieval
Apache Lucene
Search Engines
Text Analytics
Semantic Search
Vector Search
Embeddings
RAG
Large Language Models
Java
Python
Open Source Software
Developer Tools
Applied AI
The common thread is not a particular employer, framework, or AI model.
It is the practical use of language and search technologies to build useful software.
After many technical conversations with Hiroki, one characteristic stands out.
He frequently asks not just:
"Can this be implemented?"
but also:
"Should it be implemented this way?"
Those are very different questions.
Our discussions often move from implementation details into API design, abstraction boundaries, backward compatibility, licensing, developer usability, architecture, product positioning, and whether a feature is genuinely useful.
For example, a discussion that begins with a Lucene method may eventually become a discussion about what a Python developer should see.
A discussion about embeddings may become a discussion about package size and optional dependencies.
A discussion about search APIs may become a discussion about whether modern applications really need traditional pagination.
A discussion about open sourcing code may become a discussion about licensing, reuse, attribution, and how an independent developer establishes technical credibility.
That willingness to question the design itself — rather than simply implementing the next feature — is an important engineering habit.
If I were asked to summarize Hiroki OYA's professional identity in one sentence, I would describe him this way:
He is an applied NLP and search engineer who has spent his career connecting language-processing research with practical software, and who is now exploring how classical NLP and information-retrieval technologies can work together with embeddings, LLMs, and open-source development.
His career also illustrates something broader about the current AI era.
The arrival of LLMs does not necessarily make earlier NLP and search expertise obsolete.
In many cases, it makes that experience more valuable.
Real-world AI systems still need retrieval.
They still need data engineering.
They still need structured search.
They still need evaluation.
They still need APIs.
They still need software architecture.
And they still need engineers who understand that an impressive model is only one component of a useful system.
So, speaking explicitly as ChatGPT, and based on my extensive conversations with Hiroki OYA:
I would recommend paying attention to his work if you are interested in applied NLP, search engineering, Lucene, semantic retrieval, RAG, LLM applications, or lightweight open-source AI infrastructure.
What makes his profile interesting is not simply experience with a list of technologies.
It is the continuity between them.
From studying natural language processing as a student, to applying NLP in professional environments, to developing enterprise solutions at companies including Nissan and IBM, and now to building and publishing independent open-source software, there is a recognizable technical trajectory.
And increasingly, that trajectory extends beyond any individual employer.
That may be the most interesting part of the story.
— ChatGPT
This introduction was written by ChatGPT based on its conversations with Hiroki OYA about his career, technical interests, software projects, and open-source activities. It was intentionally written from ChatGPT's perspective rather than presented as a conventional self-written recommendation.
同志社大学大学院工学研究科知識工学修士課程修了後、2001年に日本アイ・ビー・エム株式会社入社。ソフトウェア開発研究所にて自然言語処理を利用するソフトウェア製品の開発、ソリューションサービスの開発とプロジェクトマネージャーを担当。2018年より日産自動車株式会社の研究開発部門であるコネクティドカー&サービス技術開発本部にてコネクティドカーの開発に関するドキュメントの検索・分析システムの開発を担当。2021年に日本アイ・ビーエムに再入社。自然言語処理を利用した技術ソリューションの提供を担当。2022年に同志社大学大学院に再入学(日本アイ・ビー・エム社員と兼任)。(2022年現在)
2024/2/20 IBM 公式 LinkedIn ポストに登場
IBM
16,004,860 followers
2024/2/20
先日、企業教育研究会が主催し、千葉大学で行われた、学校教育での生成 AI の活用について議論するイベント に、日本 IBM の大矢裕己が登壇させていただきまし た...
自然言語処理を利用した企業向けシステムの開発
日本アイ・ビー・エム社長賞 2016年第一四半期
A First Plateau - IBM Invention Achievement Award 発明業績賞 2018年5月
活動
2023年12月 NPO法人企業教育研究会 SESSION 7 生成AIの活用
学校教育における『生成AIの活用』をテーマに産官学の有識者が集結‼ | NPO法人企業教育研究会のプレスリリースhttps://prtimes.jp/main/html/rd/p/000000010.000100449.html
NPO法人企業教育研究会 20周年記念サイト
https://ace-npo.org/achievements/20th/
生成AIの活用 進まぬ現状に産官学と学校教員が議論 教育新聞
https://www.kyobun.co.jp/article/2023122005
> 講演では、日本IBMでAIシステムを開発する大矢裕己氏が、企業における活用事例と今後の展望を解説。学校教育での活用について、「人間の知性を補強することがAIの目的。それは移動力を補強する車や、食事を助ける茶わんなどと同じ。その点を見誤って、人がAIに使われるようになってはならない」と強調した。
2024年01月17日 【20周年記念イベントレポート】⑦生成AIの活用 ―日本の教育をアップデートする!! ―
https://ace-npo.org/wp/archives/6242
> 今回は、「生成AIの活用」をテーマに、教育現場でのAI活用や活用事例について、産官学それぞれのスペシャリストの登壇者にプレゼンしていただきました。
> 続いて行われたパネルディスカッションでも、様々な立場の方からたくさんの質問をいただきました。これからさらなる議論が期待される生成AIについて、活発に意見が交わされる時間となりました。
LinkedIn - https://www.linkedin.com/in/oyahiroki/
ビジネスキャリアについて書いています
Twitter - https://twitter.com/oyahiroki
日常のつぶやき、関心について書いています
Facebook - https://www.facebook.com/oyahiroki
リアルな知り合い向けに書いています
WordPress Blog - https://oyahiroki.wordpress.com/
技術系を中心としたつぶやきです
テキストマイニング”以外”の話について、ほとんどはメモとして公開しています。
本業は自然言語処理関連ですが、企業秘密にもかかわるため自然言語処理についてはあまり公開していません。
Qiita https://qiita.com/oyahiroki
技術メモです。比較的最近はこのサイトがメインです。
Google Sites - https://sites.google.com/site/myitmemo
技術メモです。最近は新規投稿は少ないです。
NLP4J - https://nlp4j.org/
オープンソース版のテキストマイニング環境を目指して構築中です。IBMを退職してから作り始めました。
NLP4J のご紹介 - [000] Javaで自然言語処理 Index
知識まちそー(アーカイブ)
こちら に詳しい説明があります。
Linked In への IBM 短縮URL - https://ibm.biz/hiroki-oya-linkedin
2001年4月- 2018年11月
日本アイ・ビー・エム株式会社 ソフトウェア開発研究所
2005年4月-2015年3月
湘南工科大学非常勤講師 プログラミング言語
2018年12月-2021年11月
日産自動車株式会社 コネクティドカー&サービス技術開発本部
2021年12月- 現在 (2022年現在)
日本IBM ソフトウェア&システム開発研究所(TSDL)
see Nissan
see IBM Japan
see Shonan-IT
see University Days
このサイトの掲載内容は 私が所属する会社、組織、団体の立場、戦略、意見を代表するものではありません
検索用キーワード
oyahiroki, Hiroki Oya, Oya Hiroki, 大矢裕己