LOTRO: Project Questa

A few years ago, there was a fan created project to provide voice narration of NPC dialogue in the MMORPG, The Lord of the Rings Online. This mainly involved capturing the text in the quest dialogue windows and then sending it to OCR software, before using a text-to-speech program to read it. LOTROVoice as it was known, broadly worked but the creator stopped developing it as it proved troublesome to maintain. However, I recently discovered Khazad Voice TTS which is a separate project that does pretty much the same. I went so far as to install this “immersive AI narrator for LOTRO” after seeing a YouTube video of it being used successfully. Sadly I ran into issues configuring Khazad Voice TTS, because I have two monitors and LOTRO runs on the secondary (by my choice) and not the primary. I couldn’t “lasso” the dialogue windows in LOTRO during the calibration set up, which effectively prevented me from completing the process.

A few years ago, there was a fan created project to provide voice narration of NPC dialogue in the MMORPG, The Lord of the Rings Online. This mainly involved capturing the text in the quest dialogue windows and then sending it to OCR software, before using a text-to-speech program to read it. LOTROVoice as it was known, broadly worked but the creator stopped developing it as it proved troublesome to maintain. However, I recently discovered Khazad Voice TTS which is a separate project that does pretty much the same. I went so far as to install this “immersive AI narrator for LOTRO” after seeing a YouTube video of it being used successfully. Sadly I ran into issues configuring Khazad Voice TTS, because I have two monitors and LOTRO runs on the secondary (by my choice) and not the primary. I couldn’t “lasso” the dialogue windows in LOTRO during the calibration set up, which effectively prevented me from completing the process.

Despite the set backs, I still like the idea of having the quest dialogue in LOTRO narrated. Hence I thought it would be “fun” (a highly subjective and situational term) to try and develop something myself. It is not exactly something that falls into my skills set as my background is in enterprise infrastructure from the Windows NT era (that dates me). However, there are a lot of resources available today that are helpful, so I decided to see if I could put together a text-to-speech narrator, similar to what others have produced but with some greater quality of life features. I also wanted to use commonly available open source resources and software that is either already available on an existing Windows 11 PC, or can be easily added. The overall design was to be constructed in a modular fashion, so if better facilities became available, it would be easy to upgrade. The budget for this project was effectively nothing apart from my time.

Let it suffice to say that I have spent the past few days experimenting and I have made some interesting progress and have a working build. I shall go into much more detail in a later post but for the present, I have used the following in this project. DXCam, OpenCV, Windows OCR and Microsoft Speech API (SAPI). A python script then uses these tools to successfully narrate quest dialogue. So far, it does this within roughly 0.7-0.8 seconds, which isn’t too bad. The default voice is not ideal but can be addressed later. The core concept is demonstrably viable but there remains the need for a lot of refinement. The stylised fonts used in the quest dialogue can be problematic to read for the OCR software. The image capturing software does not support HDR and the overly bright images further compound the OCR process. Quest dialogue windows that are large and require the player to scroll down, also present a problem. However, I think these can all be addressed and the process refined and made more accurate and polished.

It is also worth noting that the way in which this project has been designed, means that it could be used for other older games that have mainly text based interactions with the player. It doesn’t have to be exclusively used with LOTRO. Theoretically, it would simply require some minor tweaks so that the OpenCV tool could familiarise itself with a new standard format of dialogue windows. Perhaps at a later date, I can cobble together a fancy front end with calibration tools for such a task. However, that is a long way off at present. As it stands today, this tool broadly does what it is intended to do in a functional but unsophisticated and unpolished way. Furthermore, it runs completely separately from LOTRO, so it should not violate any TOS. As for the name of the project, it is Questa which is the Quenya word for speech or language. It seems fitting. Finally I would like to stress that even with a good TTS service with a pronunciation guide, it is still no substitute for quality voice acting.

NB I shall endeavour to post some YouTube videos showing Project Questa working, later today.

Read More