The first claim made by this article that I question is the claim that the MD5 algorithm, which is a function mapping an infinite space onto a finite one, produ
| This article is rated Stub-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | ||||||||||||||
| ||||||||||||||
The first claim made by this article that I question is the claim that the MD5 algorithm, which is a function mapping an infinite space onto a finite one, produces a unique signature for a file. This is mathematically impossible on its face (a function cannot be injective if the cardinality of its domain exceeds the cardinality of its codomain), and furthermore the article on MD5 states that it is not even collision resistant (i.e., it is not difficult to compute two inputs that produce the same key).
The second claim, which is repeated twice, is that examiners can reach conclusions based upon matches from the HashKeeper database with "statistical certainty." This phrase is ambiguous and misleading. There is no such thing as 100% certainty in any statistical context; inferences can only be made with levels of certainty less than 100% (i.e., 99%, 95%, etc.).
These claims should be either amended or rewritten by an expert more familiar with the database and related mathematics.—Kbolino (talk) 17:02, 7 April 2010 (UTC)
Having only slight knowledge in this arena, I was perplexed by the terms 'good' and 'bad'. Does these terms mean 'valid' and 'corrupted' (as in the data itself) or 'ethical' and 'evil', respectively? LorenzoB (talk) 18:26, 15 September 2014 (UTC)
A "bad" file would be a file that is contraband - illegal to possess regardless of circumstances. Typically, that would be child pornography. Bad files allow an investigator to focus the examination of a computer hard drive.
A "good" file would be one the source of which is known. Normally, these would be files from operating systems and freshly installed program files from reputable vendors. There would be no reason for an investigator to review these files. By way of example, an fresh install of Windows 10 adds tens of thousands of files to a hard drive. Hash those files, store the hashes as known good files and when examining a seized hard drive, ignore those tens of thousands of files.
The files not identified as good should be examined because they could contain evidence of a crime. Which is the purpose of co nducting a forensic examination.
Hashing can also be used to identify, in a practical sense, identical files (practical because, as pointed out by Kbolino, there is a non-zero probability of a false positive) and can be used to focus copyright and piracy investigations. Hashkeeper was not intended for those kinds of investigations. Pndfam05 (talk) 17:17, 17 January 2017 (UTC)
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.