Sometimes I want to extract text from MS Word documents or PDFs. The former can be done with abiword and the like, and the latter with pdftotext, which comes with xpdf, for example. But when it comes to Ichitaro, there are not many tools available.
In the end, there seem to be only a product from the Data Conversion Laboratory or xdoc2txt. The latter is for Windows, so to run it on Linux, the former is the only option. But it costs several hundred thousand yen, so it would be fine for a company but impossible for an individual…
Related posts

ID&IT Is Today: I’ll Be Leading the Closing Session
ID&IT is being held today. This year, I will be leading the closing session. It is entitled “Digital Transformation of Business and Identity” Using…

Meaningful Consent, Part 2: Crowdsourced Evaluation of Terms of Service—The ToS-DR Project
In an earlier article, “Meaningful Consent and the Standard Information Sharing Label,” I explained that it is unrealistic to assume users read, understand, and consent to…

I Will Deliver the Closing Session at the ID Management Conference on October 9
At the ID Management Conference to be held on the upcoming October 9 (Friday), I will appear in the closing session (17:55-18:35). The topic is: “Digital…
