Trust but Verify
The use and reliance on AI driven open-source intelligence gathering is exploding but the need for human verification has not gone away.
AI offers wonderful possibilities but human judgment and contextual understanding are still irreplaceable. Photo: This is Engineering/Pixabay.
By now it’s commonly known. Back in late 2021, while Russia was massing troops on the border with Ukraine, news organizations and other curious minds knew exactly what was happening months before government reporting was released. Anyone who knew the right Telegram channel or Twitter accounts to follow essentially had their own sources on the ground. Long a part of regular work for the authors of this article, the gathering of open-source intelligence or OSINT became mainstream. Since then, governments and private enterprises have rushed to take advantage of this new firehose of intelligence just as generative AI has burst on the scene to disrupt the sector. However, this new marriage faces a familiar problem: trust.
OSINT refers to the process of gathering, filtering, verifying and analyzing publicly available information. This can include social media posts, satellite imagery, news reports, breached databases that have been released publicly, forums, or even leaked government documents. In recent years, OSINT practitioners have built an ecosystem of tools and methods to handle an explosive growth in data. A quick search for “OSINT” on the open code repository GitHub results in over 13,000 tools that do everything from scraping social media profiles to finding accounts associated with a given email. Now, large language models (LLMs) such as Google’s Gemini and OpenAI’s ChatGPT and other AI systems are entering the fray, offering to speed up every stage of that pipeline: from triage to analysis.
The partnership seems natural. The same technology that can do deep research for a college report and write it just as well as an above-average student can also be used to comb through and summarize hundreds of online sources. AI tools can also extract names, phone numbers, and email addresses from hundreds or thousands of pages of text that serve as leads for additional research. Generative AI and other tools hold the promise of shrinking days or months of research and analytic work into minutes.
Unfortunately, there’s the nagging trust issue. Even before generative AI, the Internet was full of misinformation, conspiracy theories, and fake social media profiles. The wide availability of tools like Gemini and ChatGPT have made the problem much worse. “We are getting into a forest of artificial personas where we won’t be able to see the ‘living’ trees,” said a colleague of ours who works in OSINT for a major Canadian financial institution.
“When I present on emerging fraud trends, I often play a game with the participants: I show them a series of photos and voice clips and ask them to guess which ones are real,” but they’re “shocked to find out they’re wrong”, said Jess Ben Arosh, a Toronto-based OSINT analyst. Ironically, this flood of AI “slop” is another reason AI’s rapid adoption for OSINT collection and analysis is so vital. With so much data out there and the added need to be able to discern what’s real and what isn’t, human capabilities are quickly becoming too little to meet the challenge.
The trust issue extends to the use of the tools themselves, with OSINT practitioners commonly expressing worries about the reliability of the LLMs’ output and analysts’ potential over-reliance on LLM’s to critically think about results and conduct analysis with the right context. “I recently tested Gemini’s ‘deep research’ function by asking it to research me. It found my OSINT work as well as my old IMDB page from when I was acting, but it mistakenly concluded there were two different people with my name in Toronto, which is highly unlikely. That’s a good example of where human judgment and contextual understanding are still irreplaceable,” said Ben Arosh, who has extensive training in OSINT.
“OSINT is an art more than a science, we’re hunting information primarily generated by humans to answer requests for information (RFI) in a timely and relevant fashion,” said Chris St. Germain, host of the OSINT Output podcast. “We’re following bread crumbs to piece together unknown pieces of an unclear puzzle. Simply put, AI cannot do that, it can simulate and imitate, but it can’t create in the same way. Every RFI is different, which requires adaptive methodologies, disciplined research, creative thinking, and a little bit of luck.” “The deeper challenge is over-trusting AI,” said David Taxer, a cybersecurity professional and OSINT researcher based in the US. “Every result it generates must be verified, because over-reliance leads to flawed assumptions.”
These concerns are well-placed. LLMs often “produce overconfident, plausible falsehoods, which diminish their utility” for researchers. Such outputs are better known as ‘hallucinations’ that contain all kinds of incorrect information, although recent LLMs have significantly improved their accuracy. Still, LLMs are nowhere near good enough to be trusted with a full investigation without a human in the loop to sift through what is produced. Still, with OSINT proving to be an increasingly vital source of information and the flood of information (and misinformation) growing exponentially as LLMs flourish, the concerns need to be addressed and new technologies need to be adopted. For instance, video and image authentication software can help weed out media generated by AI, while coding assistants like Claude Code can allow analysts to quickly build custom tools to gather and parse data from different sources as the OSINT environment evolves.
In Canada, the tension is particularly acute. Intelligence agencies such as CSIS remain tightly bound by privacy, legal, and institutional constraints. According to a 2023 report conducted by researchers at the University of British Columbia, about CSIS, AI is rarely used to source OSINT. The same report highlighted new risks posed as AI unlocks mass OSINT analysis to non-state actors, highlighting the asymmetrical threat that arises as national organizations delay the use of technologies that non-state actors readily adopt.
Meanwhile, the Canadian federal government is attempting to catch up: its June 2025 “Guide on the Use of Generative AI” instructs departments to experiment with AI only where risks are manageable and mandates safeguards on accuracy, fairness, security and transparency. However, even this step in the right direction may overstate the risks LLMs pose and create an impediment to AI adoption. The guide lists “deploying a tool (for example, a chatbot) for use by the public” as a “higher-risk” use of AI. Using AI to assist national security OSINT investigations would likely be considered a few degrees riskier.
The discrepancy between AI’s promise for OSINT and institutional uptake is where the story lives. How will Canada and its allies adapt when civil society, private firms, and even activist groups are more agile in applying AI-driven OSINT? This is the tension – many of these groups are simply less risk averse, and in many cases, the cost of “getting it wrong” isn’t as high for them as it is for governments and law enforcement. What does “lag” look like when the pace of change is growing exponentially? And how do you build trust in systems whose internal logic may be inscrutable?
These are the hard questions, but it is critical to answer them. For better or for worse, more information is openly available than ever before and the amount continues to explode. Maintaining security, diplomatic advantage, and economic competitiveness are totally dependent on nations and companies being able to take advantage of the only tools that are capable of processing the deluge.
