Key Takeaways

  • The first cycle of online competitive intelligence — the raw data grab — is where federal criminal exposure most often begins, because intent is formed at the point of access, not at the point of use.
  • The Computer Fraud and Abuse Act (18 U.S.C. § 1030) criminalizes exceeding authorized access, and corporate investigators routinely cross that line by scraping password-protected portals or breaching terms-of-service barriers without realizing they are committing a federal felony.
  • Economic espionage prosecutions under 18 U.S.C. § 1832 do not require a foreign government nexus; the statute reaches any theft of trade secrets intended to benefit a competitor, and what looks like benign market research can easily meet the elements of a criminal information.
  • Preserving a defensible record of authorization, consent, and scope in Cycle 1 is the difference between a lawful competitive analysis and a federal indictment; after-the-fact cleanup cannot retroactively cure an access that was criminal the moment it occurred.

The First Mouse Click That Summons a Grand Jury: Why Cycle 1 Is Already a Crime Scene

In my 25 years as a federal prosecutor and now as a defense attorney, I have watched intelligent businesspeople destroy their careers in the very first phase of competitive intelligence gathering — the phase they invariably dismiss as harmless. Cycle 1 is the intake: the systematic harvesting of information from websites, social media platforms, subscription databases, and industry portals. The executive or investigator believes they are simply doing what everyone else does, collecting market intelligence to stay ahead. What they fail to understand is that federal criminal statutes do not wait to apply until valuable secrets are downloaded; they attach at the moment of unauthorized access. The Computer Fraud and Abuse Act, 18 U.S.C. § 1030(a)(2), makes it a felony to intentionally access a computer without authorization or to exceed authorized access and thereby obtain information. When a competitor’s employee uses a former colleague’s still-active login to pull pricing sheets from a client portal, that is not a gray area — that is a completed violation the instant the server processes the request. I have seen search warrants executed on corporate headquarters before the second cycle of analysis even began, because the FBI traced an IP address to an unauthorized database query that lasted less than ninety seconds.

The reason Cycle 1 is so perilous is that it takes place in an atmosphere of institutional self-deception. Companies hire third-party research firms and instruct them to be “aggressive,” while studiously avoiding written directives about methods. The researchers, incentivized to deliver the most granular intelligence, create shell accounts on competitor websites, scrape data behind login walls, or use anonymizing tools to circumvent IP-based access restrictions. Every single one of those actions potentially satisfies the “access” element of the CFAA, and the Department of Justice has become increasingly willing to charge organizations, not just rogue individuals. The 2022 update to the DOJ’s charging policy on CFAA violations specifically carved out space for prosecuting cases where access was obtained through deception or violation of a clear access barrier, meaning that even a terms-of-service pop-up can supply the notice required for criminal intent. I have sat across the table from CEOs who genuinely believed that because the information was ultimately publicly available somewhere, the path they took to reach it faster was irrelevant. That belief is legally catastrophic — the statute does not ask whether the information was secret, only whether the computer was protected and the access was unauthorized.

Another critical misunderstanding during Cycle 1 is the assumption that competitive intelligence is shielded by the First Amendment or by the defense of legitimate business purpose. The federal appellate courts have consistently rejected that proposition. The CFAA criminalizes conduct, not speech, and the conduct of clicking through a security barrier is no more protected than picking a lock on a filing cabinet. Additionally, when the intelligence-gathering targets a domestic company and the gathered materials include anything that derives independent economic value from not being generally known and is the subject of reasonable efforts to maintain secrecy, the investigator has just put their hands on what the Economic Espionage Act of 1996, 18 U.S.C. § 1832, defines as a trade secret. In Cycle 1, nobody knows yet which data points will ultimately prove valuable; they simply vacuum everything into a repository. That vacuum is a liability sinkhole, because a single downloaded customer list that was not intended for external eyes can support both a CFAA charge and a trade secret theft charge simultaneously. I have defended cases where the entire government theory of prosecution rested on the contents of a .zip file assembled during the first ten days of a competitive research project.

The fact patterns that federal agents find most compelling in Cycle 1 are those that involve affirmative circumvention: password guessing, credential stuffing, exploiting known vulnerabilities in a competitor’s customer-facing portal, or paying an insider for temporary access to a gated resource. Those are not edge cases; they are standard operating procedure in certain corners of the competitive intelligence industry, dressed up with euphemisms like “human-source intelligence cycles” or “grey-hat reconnaissance.” When I was prosecuting, I would look at the forensic timeline of access attempts and immediately seek to establish that the defendant encountered a clear warning or barrier and then took deliberate steps to bypass it. That moment — the deliberate bypass — is the strongest evidence of the mens rea required for a felony conviction. And because every server log retains a timestamp and IP address, that evidence does not fade. It sits in a data center waiting for a federal grand jury subpoena, often for years after the project has been archived and forgotten by its creators.

There is no such thing as a cheap cure for a contaminated Cycle 1. Once the government can prove that the initial collection was tainted by an unauthorized access, it will argue that everything flowing from that collection — every slide deck, every strategy memo, every revenue projection derived from the tainted data — is fruit of a criminal tree and admissible as intrinsic evidence of the offense. I have seen companies settle civil trade secret cases for substantial sums, only to discover that the civil compromise does not bind the U.S. Attorney’s Office, which can still indict on the same facts. The only defensible approach is to ensure that the very first click in a competitive intelligence campaign is within the bounds of clearly authorized access. That means written authorization agreements, documented public-access-only protocols, and a strict prohibition on any credential-sharing or barrier-circumvention. Retroactively scrubbing the record does not work because the server logs at the target company are not within the investigating company’s control to alter or delete.

When “Open Source” Is the Prosecution’s Favorite Phrase: The Mirage of Lawful OSINT Collection

Competitive intelligence professionals love the term OSINT — open-source intelligence — because it carries an air of military legitimacy and suggests that everything being done is legal by definition. In my experience as both a prosecutor and a defense attorney, the term is a mirage that has led more than a few well-meaning analysts straight into a federal criminal investigation. The problem is that the law does not recognize “open source” as a safe harbor; it recognizes the Computer Fraud and Abuse Act, the Stored Communications Act (18 U.S.C. § 2701), the Digital Millennium Copyright Act’s anti-circumvention provisions (17 U.S.C. § 1201), and a host of state computer crime statutes. If a website is publicly viewable, but the terms of service prohibit automated scraping, using a crawler to collect data may constitute unauthorized access under the CFAA, as the Ninth Circuit acknowledged in hiQ Labs, Inc. v. LinkedIn Corp., though the precise boundaries remain the subject of fervent litigation. The Supreme Court’s decision in Van Buren v. United States narrowed the CFAA’s “exceeds authorized access” clause somewhat, holding that an individual authorized to access a computer system does not exceed authorization merely by accessing information for an improper purpose, but the Court explicitly left untouched the question of access barriers — gates that must be opened through code-based or authentication-based means. When a competitive intelligence database is structured as a gated portal requiring login credentials, even if the credentials are freely registrable, the creation of multiple fake accounts to avoid rate limits or to mask the scope of collection plainly implicates the “without authorization” prong.

The DOJ’s current approach to online scraping and crawling focuses heavily on the technological measures the target has deployed to limit access. If a competitor has implemented a robots.txt exclusion, an IP-blocking mechanism, a CAPTCHA, or a rate-limiting algorithm, and your Cycle 1 operation deliberately circumvents those measures, you have given the government a roadmap to proving criminal intent. I have reviewed indictments where the central factual allegations were not that the defendant stole some exotic trade secret, but simply that the defendant programmed a scraping tool to spoof user agents and rotate IP addresses after their initial queries were blocked. That is a technological arms race that becomes a criminal conspiracy narrative in a courtroom. The National Security Division’s cyber unit, alongside the Computer Crime and Intellectual Property Section, has increasingly deployed electronic search warrants to seize the scraping infrastructure itself — servers, source code, proxy lists — and the agents executing those warrants have been trained to view every line of code designed to evade detection as an admission of consciousness of wrongdoing.

The second wave of exposure in an OSINT Cycle 1 is the Wire Fraud statute, 18 U.S.C. § 1343. When an intelligence gathering operation uses interstate wire communications — and every HTTP request crosses state or international lines — the question is whether the access scheme contains a fraudulent misrepresentation. Creating a false identity to register for a competitor’s webinar in order to download the attendee list and presentation materials after the event is a textbook wire fraud scheme if the target company would not have granted access to a known competitor. I have prosecuted cases where the false representation was nothing more than an email address domain that concealed the true employer of the registrant. The Department of Justice views the sanctity of online credentials and identity representations with the same gravity it once reserved for false statements to federal agencies. A single sign-up page checkbox certifying that the registrant is not a competitor, checked under a false identity, creates criminal exposure that may lie dormant until the commercial relationship between the two companies sours and a referral is made to the FBI.

The only reliable way to conduct Cycle 1 open-source intelligence without building a federal case file is to operate on a principle of “consent-first” collection. If a resource is behind any access control — password, registration, paywall, non-disclosure clickthrough — written consent from the target or a court order is required. If the resource is publicly viewable without any barrier, it should be accessed only by an identified employee of the company, using the company’s real IP addresses, without any evasion technology. Any instruction to a contractor or research firm to “do whatever it takes” that is not accompanied by a rigorous legal review of the specific access methods to be deployed is an invitation to disaster. In my practice, I advise clients that the cost of a pre-collection legal audit — reviewing the target sites, analyzing their terms and access controls, and crafting a protocol — is a fraction of the cost of responding to a single federal grand jury subpoena, which can run into seven figures and consume executive attention for years.

Proof of Life in the Dark: How Federal Agents Reconstruct Cycle 1 from Logs You Forgot Existed

One of the most sobering conversations I have with clients facing competitive intelligence investigations is explaining the full inventory of electronic evidence the government already possesses before the first interview is ever requested. Federal agents do not begin with the question, “What did you do?” They already know what you did. They have the server logs from the victim company, which show every incoming IP address, every user-agent string, every timestamp, every failed login attempt, and every successful data extraction. They have the subscriber records from the internet service providers, which map those IP addresses to physical locations and billing accounts. They have the email provider’s records if Gmail, Outlook, or a corporate Exchange server was used to register fake accounts. And they have the cloud storage logs from wherever the collected intelligence was aggregated. In Cycle 1, the collection itself leaves an indelible and often highly detailed trail. The Stored Communications Act, 18 U.S.C. § 2703, gives the government the authority to compel disclosure of the contents of electronic communications and stored files from service providers through a warrant, a 2703(d) order, or a subpoena, depending on the age and nature of the records, and those mechanisms are routinely deployed early in these investigations, long before a target becomes aware of the inquiry.

What makes the forensic reconstruction particularly devastating is the metadata layer that competitive intelligence analysts rarely consider. A simple web scraper built with Python and deployed from a corporate laptop leaves behind a cascade of artifacts: the local browser cache, the command-and-control server logs for the scraping orchestrator, Slack messages discussing the project, and VPN connection logs if the analyst attempted to mask their location. When I was on the prosecution side, the first thing I would do after receiving the victim company’s logs was to draft a preservation letter and a series of warrants directed at the suspect’s own corporate infrastructure. The most incriminating evidence is almost never found in the targeted competitor’s data; it is found in the suspect’s own files: internal strategy documents that say “Use a VPN and register as a student to get past the paywall,” time entries billing the client for “grey-hat acquisition phase,” and exported CSV files with naming conventions that betray their origin. A single email from a manager saying “We can’t use the standard portal because they’ll recognize our domain” is often all the mens rea a prosecutor needs to get an indictment.

The government’s technical capacity has evolved dramatically over the last decade. The FBI’s Computer Analysis and Response Teams can reassemble fragmented scraping scripts, recover deleted virtual machines, and trace cryptocurrency payments used to purchase access credentials on darknet marketplaces. In competitive intelligence cases, the government increasingly charges violations of the access device fraud statute, 18 U.S.C. § 1029, when lost or stolen credentials are used, because each use of a pilfered password counts as a separate count, creating stacking exposure that routinely drives guideline ranges above ten years. I have defended a case in which the only thing that elevated a straightforward CFAA misdemeanor into a fifty-count felony indictment was the agent’s discovery that the research firm had purchased a batch of compromised competitor employee credentials through a cybercrime forum, believing that the buffer of an intermediary would insulate them from criminal liability. It did not. The intermediary was an undercover FBI agent operating a honeypot marketplace, and the entire purchase was recorded in real-time audio and video.

Another dimension of the forensic case that surprises clients is the government’s ability to correlate multiple Cycle 1 projects across time to establish a pattern of knowing and intentional criminal conduct. If a company has engaged five different outside firms over three years, each of which accessed a competitor’s gated portal under false pretenses, a federal prosecutor will consolidate those episodes into a single conspiracy charge under 18 U.S.C. § 371, alleging an overarching scheme to defraud and to violate the CFAA. The 371 conspiracy charge is the darling of federal white-collar prosecutions because it is broad, flexible, and carries its own five-year statutory maximum. I have watched the government weave together a compelling narrative of a corporate culture of trade secret theft using nothing more than the preserved logs and internal communications from a series of discrete projects that, in the minds of their sponsors, had nothing to do with one another. The pattern itself becomes the evidence of willfulness, and the defense of “everyone does it” collapses under the weight of the company’s own documented repetition.

The Trade Secret Boomerang: How Cycle 1 Collection Can Turn the Collector into the Defendant Even When the Target Never Sues

It is a pervasive myth in the business community that criminal trade secret prosecutions require a complaining victim. In truth, the United States Attorney’s Office can and does initiate investigations under 18 U.S.C. § 1832 based on third-party referrals, press reports, or intelligence gathered in a completely unrelated investigation. In Cycle 1, the collector is often laser-focused on obtaining information about the competitor and never pauses to consider that the mere act of obtaining that information, if it meets the statutory definition of a trade secret,