Database Ideas, please
Author
Discussion

HiRich

Original Poster:

3,337 posts

292 months

Monday 18th July 2005
quotequote all
I've (rather foolishly) volunteered to produce a digitised archive of magazine articles. I know what I want to achieve, but would be grateful for your advice on how.

I will be scanning in (and potentially using OCR to covert to editable text) all references to a particular subject. The files (which may be jpegs, pdfs, text documents, etc - almost anything goes) will be scanned, trimmed, generally cleaned up, and properly filed by issue date. I then want to keyword each file (a semi-automatic system would be nice, but I'm resigned to doing it manually), add a description, and use this to produce a searchable database. Users can then search the database by a series of AND/OR selections to produce a shortlist of relevant entries. Then they can call up the master documents.

Now, I want to be able to issue the archive on CD. That presumably means I need to be able to incorporate a reader-only programme on the disc that is:
- Free
- Mac and PC compatible
- Ideally ready to run straight off the disc, rather than including an installer.
I keep thinking that it would be nice to export the database in an html format (or similar) so that users can access it with their browser.

I have a copy of Extensis Portfolio. Although a bit clunky, it can basically do this job I want. But does anyone have recommendations, suggestions or general comments to guide me in the right direction. Any thoughts would be appreciated

P.S Whilst not a computer numpty, I'm not a techy. So please keep it reasonably simple.

Plotloss

67,280 posts

300 months

Monday 18th July 2005
quotequote all
You could have a table using just about any RDBMS which contains your keywords etc and then storing the scanned image as a BLOB per record.

Its not going to be the fastest thing in the world (comparing it to straight text) but something like mySQL will run anywhere and you could install it off a CD then run then copy the table definition and data off the CD into the RDBMS

zaktoo

1,401 posts

270 months

Monday 18th July 2005
quotequote all
To do it really well would be a massive undertaking. I'm also not sure that your distribution model is the way to go. What about a central server, with restricted access, access via internet?

Those issues aside, I'd guess your best bet is archive images in a hierarchical file structure, dump OCRed text to table for indexing, have pointers to actual scanned images from the table. Slap on a web front-end, and Bob's your auntie.

The tricky bit is comparing OCRed results to actual text. You'll find some funnies, and IMHO without physically proofreading every OCRed word, you'd be making an archive that might be less useful than intended. Therefore use the text only for basic indexing, and point the reader to the actual scanned image for legibility. If readers would like, they can OCR it and proofreead it themselves, on a per-need basis (or take a dump (ahem!) of your OCRed text & proof it themselves)...

Just a few rambling thoughts... hope they help.

Ciao

Zak

ATG

23,803 posts

302 months

Monday 18th July 2005
quotequote all
Rougly how many images are we talking about and how many keywords per image?

editted to add: in particluar if there aren't an overwhelming number of key words, you could build a file containing a tree of keywords that you would search from a Java script running in a browser; i.e. your viewer app is a few DHTML pages stored on CD. That ought to run happily enough on either a PC or a Mac. Each keyword node could link to one or more images stored on the disc. If the tree was small enough, you could store it as XML and search it using a standard XML parser DOM object whotsit.

>> Edited by ATG on Monday 18th July 16:07

HiRich

Original Poster:

3,337 posts

292 months

Tuesday 19th July 2005
quotequote all
Thanks for your ideas (even if Plotloss' reads as "You want the what in the what-what?"

To give you an idea of the project size, it would of the scale of recodring every news item from a season of English Premiership football (including International matches) from one season. I reckon that would be about:
- 5,000 elements (match reports, bews stories, interviews, photos)
- 1,000 Keywords (player names, teams, venues, etc.)
- And that means from 1-30 keywords per article.

So I appreciate that it is potentially a huge task. However, I have completed a more graphics-based package very similar in Portfolio. There is no time limit (I'm treating it as a hobby) so I can process one magazine at a time, digitising, keywording and adding a description to each element, before signing it off and moving on to the next magazine.

Although my OCR software appears to work rather well, it's not totally essential. And since I know the basic information fairly well, I could speed read the OCR'd document and spot errors fairly quickly.

I appreciate the idea of a central server, but this is for a very small (poor) organisation. Usage will also be rather infrequent. And in time I will be adding whole extra 'chapters' (using my example, if the first database is the Sun, later chapters would be the Mirror, the Times/Sunday Times and so on). So I feel more comfortable working towards reducing a room full of paper archive to a set of CDs. Since the structure (nearly all the keywords) will be the same, I can always then combine them at a later date.