{"id":2892,"date":"2026-09-16T22:31:20","date_gmt":"2026-09-17T02:31:20","guid":{"rendered":"https:\/\/wordpress.forensicpath.us\/?p=2892"},"modified":"2026-09-16T23:17:16","modified_gmt":"2026-09-17T03:17:16","slug":"update-on-using-ai-for-cataloging-old-articles","status":"publish","type":"post","link":"https:\/\/wordpress.forensicpath.us\/index.php\/2026\/09\/16\/update-on-using-ai-for-cataloging-old-articles\/","title":{"rendered":"Update on using AI for cataloging old articles"},"content":{"rendered":"<p>So, a couple of you strongly suggested that I not use my subscriptions to Grok or ChatGPT as models for Hermes Agent, since that eats up so many tokens.\u00a0 You suggested I use ollama and the gemma4 models.\u00a0 I gotta tell you, it worked great.\u00a0 The key for me was that ChatGPT was profoundly helpful in teaching me to set it up.\u00a0 Basically, the workflow goes something like this:<\/p>\n<p>I do a first pass that goes recursively through the directory and uses local tools (e.g. pdftotext, tesseract) for text retrieval and OCR. It uses ollama with gemma4:E4b to do the initial classification.\u00a0 Those documents that aren&#8217;t classified with high confidence then go through a second pass that uses gemma4:12B to do better OCR and classification.<\/p>\n<p>And it does it all locally, so no token costs!\u00a0 It takes about 5-6 seconds per document on average for classification.\u00a0 I have about 15000 documents per year in my backups (of which there are really only about 4000-5000 unique documents).\u00a0 So, it&#8217;s about one night to go through one year of backups.\u00a0 \u00a0It&#8217;s working great.<\/p>\n<p>I have beome a big fan of ollama and local models.\u00a0 It all loads and runs on my NVIDIA GPU.<\/p>\n<p>The key for me, though, is that I would never have figured out how to write the schema and commands for ollama without being tutored by ChatGPT.\u00a0 It turned out that there were some issues with using Hermes agent on ollama, but it was just as easy to write a little python script to do it.\u00a0 One odd thing is that parallelization slowed things down.\u00a0 Doing this with one &#8220;worker&#8221; was faster than two.<\/p>\n<p>When this is done, I&#8217;ll move to images.\u00a0 ChatGPT says the ollama and gemma should be great for classifying images in my backups.\u00a0 We&#8217;ll see.\u00a0 \u00a0I would include the schemata and prompts, but it was vibe coded with ChatGPT and each pass runs around 3000 lines.<\/p>\n<p>In any case, if any of you are like me and have 20 years of backups sitting there. this turns out to be a great way to get a handle on all those old files.\u00a0 Thanks to those of you who told me to use ollama and a local model.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>So, a couple of you strongly suggested that I not use my subscriptions to Grok or ChatGPT as models for&hellip;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[99],"tags":[],"class_list":["post-2892","post","type-post","status-publish","format-standard","hentry","category-forensic-pathology"],"_links":{"self":[{"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/posts\/2892","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/comments?post=2892"}],"version-history":[{"count":2,"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/posts\/2892\/revisions"}],"predecessor-version":[{"id":2894,"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/posts\/2892\/revisions\/2894"}],"wp:attachment":[{"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/media?parent=2892"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/categories?post=2892"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wordpress.forensicpath.us\/index.php\/wp-json\/wp\/v2\/tags?post=2892"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}