The uber-geeks of Google have announced their intention to release (re-release?) Tesseract OCR. Interesting, but nothing to get very excited about at this point - unless you are truly into the whole open-source/linux movement. There may be promise here one day - especially considering that everything Google touches tends to turn to gold. The Google team brings to mind a slew of Indiana Jones types scouring the digital-Amazon (the river, not the shop-at-home site) for lost relics of a bygone era.
Outside of the announcement by Google I had a hellava time finding anything on "Tesseract OCR". Plenty of neat stuff on
Tesseracts, but not much on the OCR namesake - and I have the
library resources of a major university to play with. I am actually suprised I did not find more flash based sites with Tesseracts - it would seem that math types would spend more time creating little tessies and releasing them into the wild - but they are all probably playing with
SodaRace.
I finally found
something on Linux.com that provides a fairly decent overview of the program (history, what it does, tech notes, etc.). This article puzzles me as it seems like these folks have been hiding in a server closet for years having never heard of some of the exceptional commercial OCR programs available in 2006 instead of harkening back to a program that was
"one of the top-tier performers at UNLV's OCR competition in 1995" NINTEEN-NINETY-FIVE!!! Are you sh**ting me?!?! And there are people out in the world excited about resurrecting this thing. What - I ask - is the point! I remember OCR technology in 1995 and it sucked compared to what we have today. Give me my Abbyy Fine Reader and a sweet scanner anyday over some long forgotten OpenSource OCR program.
The
SourceForge.net dowload site provides merely a one-line description:
"A commercial quality OCR engine originally developed at HP between 1985 and 1995. In 1995, this engine was among the top 3 evaluated by UNLV. It was open-sourced by HP and UNLV in 2005." Of course it was open-sourced by UNLV - they were probably cleaning out old storage and figured something
this good should not be tossed.
Hey, I know - let's give it to Google (he, he) they'll take anything! After all - it did win their in-house competition more than a decade ago. And, of course, we all know
how well things are going over at HP - they have bigger things to worry about then where UNLV tosses their trash. [off the record - UNLV will only acknowledge a long-term platonic relationship with Tesseract].
Thanks Google - but I'll pass. Please bring us more nifty stuff like the
Google Personalized HomePage. I have basically shunned all other options and am working on moving
several years worth of PDA dated into the trusted hands of Google. If you have not yet discovered the tabs in the GPHP you are truly missing out! I am also very much looking forward to where
Google Scholar is headed - even the reference librarian for my program is excited.
Labels: access_tech, E-text, open_source