Summary from literature review
| Ref | Main idea | Contribution | Techniques | Data | Performance | Remarks |
|---|---|---|---|---|---|---|
| Sharif et al. (2008) | Eureka framework for enabling static Internet malware binaries analysis | Course-grained execution tracker applying a heuristic based and a binary n-gram statistical trigger to estimate when to stop the malicious process image | API resolution techniques IDA-Pro disassembler De-obfuscation | Corpus of 1,291 malware instances, 479 malicious executables from spam traps, 435 malicious executables - Honeynet | 97.7% Spam malware corpus unpacking 93.3% honey net malware corpus unpacking Unpacking of 90 binaries/hr | Automated classification of malware |
| Lengyel et al. (2014) | DRAKVUF- dynamic MA system | Improves stealth by enforcing scalability, fidelity, stealth and isolation conserving resources | Hardware virtualization extensions and the Xen hypervisor | 1,000 samples from shadow server | Memory saving of 62.4% | Automated classification of malware |
| Ucci et al. (2019) | A survey on MA through machine learning techniques | Novel concept of MA economics, malware anti-analysis techniques, etc. | A qualitative analysis | Processing one million malware per day | 86% accuracy using 3 as minimum n-grams size | Tuning strategies to balance metrics such as accuracy and cost in designing MA environment |
| Schultz et al. (2001) | A data mining framework for automatically detecting new malicious binaries | Method for detecting previously undetectable malicious executables | Naive Bayes, multimodal-naive Bayes, RIPPER standard statistical cross-validation | Data set of 4,266 programmes 3,265 malicious binaries and 1,001 clean programmes | Multi-naive Bayes yielded highest detection rate 97.76% | Extension of learning algorithms to make use of byte-sequences |
| Sethi et al. (2017) | A framework for detecting and classifying malware | Intelligent MA framework for dynamic and static analysis of malware samples based on similarity | J48, SMO and random forest Cuckoo sandbox for malware analysis | 220 Samples of malicious and benign files | 100%, 99% and 97% detection rate, and 100%, 91% and 66.67% classification, respectively | Data set of 220 samples needs to be expanded |
| Ref | Main idea | Contribution | Techniques | Data | Performance | Remarks |
|---|---|---|---|---|---|---|
| Eureka framework for enabling static Internet malware binaries analysis | Course-grained execution tracker applying a heuristic based and a binary n-gram statistical trigger to estimate when to stop the malicious process image | API resolution techniques | Corpus of 1,291 malware instances, 479 malicious executables from spam traps, 435 malicious executables - Honeynet | 97.7% Spam malware corpus unpacking | Automated classification of malware | |
| DRAKVUF- dynamic MA system | Improves stealth by enforcing scalability, fidelity, stealth and isolation conserving resources | Hardware virtualization extensions and the Xen hypervisor | 1,000 samples from shadow server | Memory saving of 62.4% | Automated classification of malware | |
| A survey on MA through machine learning techniques | Novel concept of MA economics, malware anti-analysis techniques, etc. | A qualitative analysis | Processing one million malware per day | 86% accuracy using 3 as minimum n-grams size | Tuning strategies to balance metrics such as accuracy and cost in designing MA environment | |
| A data mining framework for automatically detecting new malicious binaries | Method for detecting previously undetectable malicious executables | Naive Bayes, multimodal-naive Bayes, RIPPER standard statistical cross-validation | Data set of 4,266 programmes 3,265 malicious binaries and 1,001 clean programmes | Multi-naive Bayes yielded highest detection rate | Extension of learning algorithms to make use of byte-sequences | |
| A framework for detecting and classifying malware | Intelligent MA framework | J48, SMO and random forest | 220 Samples of malicious and benign files | 100%, 99% and 97% detection rate, and 100%, 91% and 66.67% classification, respectively | Data set of 220 samples needs to be expanded |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.