Menu ▾ ▴

Home

Evgeny Vasiliev

Welcome to your wiki!

This is the default page, edit it as you see fit. To add a new page simply reference it within brackets, e.g.: [SamplePage].

The wiki uses Markdown syntax.

Project Members:


Discussion

  • Evgeny Vasiliev

    Evgeny Vasiliev - 4 days ago

    HAV screenshots:

     
  • Evgeny Vasiliev

    Evgeny Vasiliev - 4 days ago

    Technical Overview and Manifesto

    Over 35 years in software development, I have witnessed multiple generations of antivirus solutions, the evolution of threat detection approaches, and the endless cycle of the arms race between malware authors and security vendors. Yet the core problem has remained fundamentally unchanged.
    A new malicious file emerges. It is detected, analyzed, classified, and added to signature databases or detection models. Then a modified version appears that bypasses existing rules — and the cycle continues.
    I have long been interested in a different question: what if we build detection not around "do we know this file?" but around "how characteristic is this file's behavior and structure of a malicious program?"
    This is how the experimental project HAV — Heuristic AntiVirus — came to be. It is not a commercial product, nor an attempt to replace existing solutions. It is a research experiment aimed at exploring how far one can go in malware detection using static and dynamic analysis, statistics, and heuristic models. The project repository is publicly available.
    Classical signature-based approaches work well against known threats: given a known sample or family, one can describe it with a signature or rule. This is fast, reliable, and relatively cheap. However, this approach has a fundamental limitation: a signature must first exist before it can be used. For a new sample, such information may not yet be available.
    Modern antivirus products, of course, are no longer limited to signatures — they employ reputation systems, sandboxes, behavioral analysis, machine learning, emulation, and many other techniques. But I was interested in a narrower question: what level of detection quality can be achieved if the primary focus is on structural and behavioral analysis, with signatures playing only a supporting role? This is a testable engineering hypothesis.
    Goals and Requirements
    At the initial stage, I formulated several key requirements for the experimental scanner:
    ● minimal dependence on constantly updated signatures;
    ● local processing without mandatory cloud infrastructure;
    ● compact database size;
    ● low resource consumption;
    ● ability to analyze unknown samples;
    ● minimization of false positives;
    ● simple architecture.
    The last point is particularly important to me.
    A security tool should not interfere with the user's work. If an antivirus blocks legitimate files, consumes significant system resources, or constantly demands attention — the user will eventually look for ways to disable it. This defeats the entire purpose of protection.
    How HAV Works
    Static Analysis
    The file is treated as a structured object, not merely as a sequence of bytes.
    Depending on the format, the following are analyzed:
    ● executable file structure;
    ● sections;
    ● data sizes and distributions;
    ● imported and exported functions;
    ● strings;
    ● code and data layout characteristics;
    ● packing and obfuscation indicators;
    ● entropy;
    ● byte distributions;
    ● watermarks, manifests, taggants, forms, resources, GUIDs, certificates;
    ● statistical properties.
    An unusual structure can also occur in legitimate software. The presence of a particular API does not by itself make a program malicious.
    Therefore, HAV does not use a "feature found → file malicious" logic. Instead, it employs a cumulative assessment. On average, about three features contribute to a decision. Years of experience have made it possible to identify the most characteristic features — to the point that in some cases even a cumulative assessment is redundant.
    Statistical Analysis
    Byte sequences are analyzed using distributions, n-grams, frequency characteristics, and other statistical properties.
    The idea is straightforward: when examining large collections of files, statistical regularities emerge across different classes of programs. The model's task is not to find a specific sequence matching a specific virus, but to determine how similar the analyzed object is to known classes of malicious software. The key is to isolate the most distinctive features.
    The result is a kind of "statistical fingerprint" of the file.
    Linguistic Analysis
    A separate area of focus is the analysis of strings and textual sequences. The following are of particular interest:
    ● API calls;
    ● command lines;
    ● file names;
    ● URLs;
    ● execution parameters;
    ● characteristic sequences;
    ● obfuscation indicators;
    ● character distributions.
    What matters is not only the specific string, but also its context. In the future, the emphasis on linguistic analysis will be strengthened — the potential of this direction is considerable.
    Why a Weighted Model Is Used
    One of the most common problems in heuristic detection is excessive aggressiveness. It is easy to build an antivirus that flags almost anything suspicious. But along with malware, it will also block a vast amount of legitimate software.
    HAV therefore uses a weighting system. For example:
    ● feature A → +2
    ● feature B → +5
    ● feature C → +1
    ● feature D → –3
    Features are not considered in isolation. Certain combinations significantly increase the likelihood of maliciousness, while the presence of other characteristics may decrease it.
    It is the combination of features that determines the final verdict. This allows HAV to abandon the "if X is present — delete" logic in favor of a more nuanced approach: "to what extent does the observed set of characteristics correspond to malicious behavior?"
    The scanner is built as a multi-layered heuristic model. Each layer analyzes the file from a specific perspective.
    Importantly, the exact number of layers and their coefficients constitute know-how and are not disclosed. This helps preserve the uniqueness of the approach and avoids giving attackers ready-made targets for adaptation.
    Dynamic Analysis
    Static analysis has an obvious limitation: code can be deliberately hidden. Packing, obfuscation, dynamic resolution of function addresses, and other techniques significantly complicate analysis.
    Therefore, HAV also employs dynamic analysis. Code is executed in a controlled environment using emulation (for executables) and interpretation (for scripts). This allows observation not only of what is contained in the file, but also of what it attempts to do. Programs that appear relatively benign during static analysis are of particular interest.
    System Call and Event Chain Anomalies
    Instead of searching for signatures, I analyze the file's behavior during execution — primarily at the level of system calls and their sequences. Many observed actions are assigned the status of named events and are accounted for within chains.
    Why this works:
    ● System calls are the language through which a program communicates with the OS. They cannot be imitated or concealed without breaking the program's logic.
    ● Even if a file is packed with ENIGMA, THEMIDA, or VMProtect, the call pattern remains stable.
    ● Legitimate software exhibits different call patterns than malware. HAV has learned to distinguish them.
    Supported Formats
    The emulator works not only with Windows executables:

    Format Method Specifics
    PE 32/64-bit System call emulation 64-bit code is emulated in a 32-bit environment via call adapters
    ELF (Linux) Static analysis ("feature skeleton") Separate config
    VB Native code emulation Separate config
    DOTNET IL code emulation Separate config
    This approach covers all relevant platforms using a unified decision-making logic and a shared API implementation base. ARM and Android remained outside the project scope due to time and resource constraints.
    Emulation Specifics: Compatibility and Complex Protectors
    Due to the large number of Windows versions and critical changes affecting software compatibility, I am forced to maintain a somewhat averaged version of Windows during emulation. This approach ensures maximum reproducibility of results, although it somewhat reduces detection efficacy.
    Complex Protectors: THEMIDA, MoleBox, and Others
    On files packed with THEMIDA, the emulator must reproduce multitasking inside the protector's virtual machine — all within a single thread. Processing such samples takes 3–5 minutes, which is acceptable given the complexity of the protection.
    Files packed with, for example, MoleBox require processing a massive number of system calls (up to 40 million). The emulator handles this in minutes, returning a fully unpacked file for further analysis.
    Most commercial solutions on such samples either hang or produce false positives.
    Signature Traps: For 20% of Complex Cases
    Heuristics and emulation cover 70–80% of all cases. For the remaining 20% that do not yield to these methods, a targeted mechanism is used: a universal signature is generated. Such signatures do not become obsolete in the foreseeable future — they cannot be bypassed without breaking the file's functionality.
    There are orders of magnitude fewer of them than in classical AVs, so the database remains compact (60 MB) and requires much less frequent updates (maintained by 1–2 analysts).
    Limitations
    Dynamic analysis should not be regarded as a universal solution. Any emulation has its limitations. Malware can:
    ● detect the execution environment;
    ● alter its behavior when emulation is detected;
    ● introduce timing delays;
    ● check for the presence of specific components not accounted for in the emulator;
    ● execute its malicious payload only when certain conditions are met.
    The scanner does not promise "100% protection" — that would be marketing. I openly acknowledge that:
    ● certain file types may require model refinement;
    ● the experimental status means the project is still evolving;
    ● HAV is ready for testing on real-world data and awaits feedback.
    Why I Prefer Not to Rely on a Massive Database
    A large signature database is a perfectly viable tool. But it comes at a cost. It must be:
    ● created;
    ● verified;
    ● updated;
    ● distributed;
    ● stored;
    ● tested for conflicts;
    ● kept in sync with the current threat landscape.
    HAV shifts the center of gravity from a database of known samples to models and statistical features. This allows it to remain compact and independent of constant updates, without compromising detection quality.
    In the current experimental version, the database occupies on the order of tens of megabytes, and all decisions are made locally.
    This does not mean that signatures are no longer needed. They remain a useful supplement where precise knowledge of a sample allows a definitive decision.
    But how small can this part of the system be? Repeated weekly tests on a highly representative selection of malicious files have not yet required a significant increase in the number of records. The internal test set includes virtually all collection files from MalwareBazaar, VirusShare, and VirusSign. Moreover, as heuristics improve, many records in the current database become redundant. If necessary, the database can be reduced to 40–50 MB.
    HAV can work with an arbitrary number of database files — their load order and the sequence of records within them are irrelevant.
    Important: HAV contains know-how that could become a target for industrial espionage, and whose understanding could ease the work of attackers. This includes, in particular, protection against ARC bombs, reliable detection of packed files with classification without unpacking, and others. Therefore, some technologies remain behind the curtain.
    False Positives
    At present, the results are encouraging. HAV is continuously tested on a large collection of clean software — totaling 15 TB. This includes hundreds of thousands of containers and tens of millions of files: virtually all versions of Windows and Microsoft software, about 30 popular Linux distributions, nearly the entire content of software sites such as SourceForge, APKPure, SoftPortal, ComSS, and the Nuget and NPM repositories (from the latter, only JS scripts — about 12 million). New versions of popular software are tracked when possible.
    On this entire mass of files, false positives did occur — this is a natural trade-off: the higher the detection rate, the more false positives. However, in percentage terms, they account for 0.01–0.1% (approximately 3,000–4,000 cases), and for scripts — 0.001–0.01%. All these cases are known, individually analyzed, and blocked by individual records. This figure is qualitatively below the industry average.
    Diagnostic Potential
    Since HAV also tracks anomalies at the level of system calls, it can be viewed as a diagnostic tool as well. Even a false positive is a signal: the program likely contains a bug in its code. A bug that compilers do not catch, that static analyzers such as PVS Studio or Clang Static Analyzer do not find, and that the "cloud AI" certainly does not notice.
    The HAV emulator observes how the code behaves at runtime. It detects anomalies that are not visible in the source code. It finds code fragments that "smell" suspicious. In modern software — especially in large OOP programs with thousands of inheritance hierarchies, abstractions, and dependencies — such anomalies are far more common than one might think: I have recorded about 100 cases of clearly incorrect code.
    I do not place great hopes on this side effect, but if we were to focus on it...
    What Has Been Achieved
    At the current stage, the project shows encouraging results on internal datasets. The detection rate is approximately 99.5% on the collections mentioned above, including fresh files up to the present day.
    However, a caveat is necessary here. This figure is not proof that HAV detects 99.5% of all existing malware.
    The result may depend on the composition of the test set, family distribution, sampling methodology, the time gap between training and testing, and many other factors. On Malware Bazaar, for instance, there are days dominated by APK files, days of scripts, days of ELF files, and so on.
    Unfortunately, I do not have full access to a collection as vast as VirusTotal, which also affects the objectivity of the assessment.
    Therefore, I consider the percentage itself far less interesting than the ability to build a reproducible, independent benchmark. This direction currently represents the greatest interest for the project.
    High sensitivity alone is not a sufficient quality criterion. At a minimum, two metrics must be considered together:
    ● TPR (True Positive Rate) — the proportion of malicious samples detected;
    ● FPR (False Positive Rate) — the proportion of clean files incorrectly classified.
    It is their balance that deserves primary attention
    Confusion Matrix
    A proper evaluation should look roughly like this:
    Actually Malware Actually Clean
    HAV: Malware TP (True Positive) FP (False Positive)
    HAV: Clean FN (False Negative) TN (True Negative)
    From this matrix, one can derive:
    ● detection rate;
    ● false positive rate;
    ● precision;
    ● recall;
    ● F1;
    ● ROC/PR curves;
    ● results per family.
    Such measurements are far more informative than the advertising-like (though confirmed on our test sets) figure of "99.5%". In the future, I would like to publish results in exactly this format.


    Is It Possible to Do Without Constant Updates?
    This is where it gets interesting.
    Theoretically, one would like to build a model that performs well without constant feeding of new signatures. In practice, this is a very difficult task.
    Malware evolves as well. If a defender has learned to detect a certain set of features, the attacker has an incentive to modify the program so that it retains its malicious functionality while changing the observable characteristics.
    Therefore, it would be wrong to claim that an antivirus will never need updates again.
    The goal of the experiment is more modest: to test how much the model's lifetime between updates can be extended while maintaining acceptable detection performance.
    ● If it turns out that updating it once a month is sufficient — that is interesting.
    ● If it turns out that weekly updates are required — that is also a useful result.
    ● If it becomes clear that the model degrades rapidly without constant updates — that is also a result.
    But my ultimate target is a one-year cycle — because I know what needs to be done and how to do it. The only thing in short supply is time.


    Why I Decided to Take This On Independently
    An antivirus is a very complex system. A commercial product includes drivers, services, cloud components, sandboxes, databases, telemetry systems, updates, and many other subsystems. All of this is acceptable for a commercial product.
    But sometimes it is useful to step back and ask: what is the minimal system capable of solving the problem? One does not have to build an industrial-scale combine right away. One can start with a small program and test a single specific hypothesis.
    Invitation to Collaborate
    I would be interested to discuss whether this project could find application in your research efforts or serve as a useful independent tool.
    I am particularly interested in people who would not simply write "works/doesn't work," but would try to break it, to refute my arguments. If you work in malware research, reverse engineering, machine learning, binary analysis, or simply enjoy subjecting programs to harsh tests — this is precisely the case where criticism is more valuable than praise.
    I am interested in:
    ● independent testing;
    ● new sample sets;
    ● false positive hunting;
    ● false negative hunting;
    ● adversarial testing;
    ● comparison with other approaches;
    ● algorithm analysis;
    ● suggestions for improving the model.
    Conclusion
    I do not think the signature-based approach is dead. HAV can also handle signatures. There is no problem in fully automatically generating a database of records from existing detections — the issue of conflicts has long been resolved. But its size would be around 500 MB. Although database size does not affect scan speed in any way, why then did I write all of this?
    I do not believe that machine learning will solve the malware problem. Nor do I think that one small project can replace all the systems that large teams have been building for decades.
    But I am certain of something else. Sometimes it is useful to try to keep things as simple as possible. Not to build yet another massive infrastructure.
    Instead, ask a simpler question: what do malicious programs have in common as a class of objects? If an answer to this question exists, can it be turned into an algorithm?
    HAV is one such attempt. And I am very pleased with its results. But for now, it is an experiment. And that is precisely why I am making it public. Not to declare victory over the antivirus industry. But to see how far one can go by starting from a different direction.
    After 35 years — it is finally time to try.


    Open Questions
    Of course, many ethical aspects and nuances remain unresolved and unsettled. They should be addressed in the next version. For example:
    ● how to handle non-functional programs whose execution depends on dates, file names, command-line parameters, OS version or specifics, or trial protection timing;
    ● what to do with clearly "broken" files, of which there are a considerable number in malware collections;
    ● whether adware and malware should be fundamentally distinguished? (In my view, advertising is worse than viruses.)


    Project Status
    The project's source materials are published on GitHub. The project is distributed free of charge. In the public version, file deletion is disabled (to avoid any risk), leaving only notification.
    https://github.com/evgenvasiliev62-code/hav
    https://github.com/evgenvasiliev62-code/hav/releases/download/hav/hav01.zip

     

Log in to post a comment.