Skip to content

Repository files navigation

dragonfly-client-rs

Modular compute nodes capable of scanning packages and sending results upstream to a control server, written in Rust.

Set up

This section goes over how to set up a client instance locally and via Docker.

Refer to the Environment variables section for information on what environment variables are necessary.

Local

Requirements

1. Set the appropriate environment variable pointing to the YARA installation

export YARA_LIBRARY_PATH='/path/to/yara/libs'

2. Build the binary with cargo

cargo build --release

3. Run the built binary

./target/release/dragonfly-client-rs

Docker

Requirements

1. Build and tag the image

docker build --tag vipyrsec/dragonfly-client-rs:latest .

2. Run the container

docker run --name dragonfly-client-rs vipyrsec/dragonfly-client-rs:latest

Docker Compose

Requirements

Run the service

docker compose up

How it works: Overview

The follow is a brief overview of how the client works. A more extensive writeup can be found towards the bottom of this page.

The client is comprised of a few discrete components, each running independently. These are the scanning threadpool, the loader thread, and the sender thread.

  • The Scanning Threadpool - Downloads and scans the releases.
  • The Loader Thread - This thread is responsible for requesting jobs from the API and submitting them to the threadpool.

Performance, efficiency, and optimization

The client aims to be highly configurable to suit a variety of host machines. The scanner processes one package and one distribution at a time. Compressed downloads, expanded archives, individual scan targets, archive entry counts, and distributions per package are bounded independently. The defaults target a 512 MiB background-worker container while preserving substantial headroom for compiled YARA rules, ZIP metadata, allocator overhead, and filesystem cache.

How it works: Detailed Breakdown

This section attempts to describe in detail how the client works under the hood, and how the various configuration parameters come into play.

The client can be broken down into a few discrete components: The scanner threads, the loader thread, the sender thread. We will first explore in detail the workings of each of these components in isolation and then how they all fit together.

The scanner thread(s) are what do most of the heavy lifting. They use bindings to the C YARA library, and most of this code can be found in scanner.rs. The way this program models PyPI data structure is as so: There are "packages" (or "releases") which is a name/version specifier combination. These "packages" are comprised of several "distributions" in the form of gzipped tarballs or wheels (which behave similarly to zip files, hence the use of the zip crate). Each distribution is comprised of a flat sequence of files (the hierarchical nature of the traditional file/folder system has been flatted for our use case). The main entry point interface to the scanner logic is via scan_all_distributions. This loops over the download URLs sequentially, stages each compressed distribution on temporary disk, validates its resource limits, and extracts it. Files are scanned individually from disk by YARA. Only the highest-scoring file and unique matched rules are retained for each distribution. After every distribution finishes, the client unions matched rule identifiers across the complete package and submits one package verdict to the API. Its score is the maximum independently calculated distribution score; evidence from mutually exclusive distributions is not added together. The inspector URL identifies the highest-scoring file in that distribution.

The client requests at most one job per configured worker, up to DRAGONFLY_BULK_SIZE, and scans those packages concurrently. Each package's distributions and files remain sequential. The default worker count follows the machine's available CPU parallelism; constrained deployments can set DRAGONFLY_THREADS=1 to guarantee sequential package processing. Empty and failed job requests are retried after DRAGONFLY_LOAD_DURATION seconds.

The client authenticates every Dragonfly API request with a Cloudflare Access service token. The source code of the YARA rules is compiled (very much like compiling regex) and stored in shared state. Then, the necessary threads are spawned. Once a threadpool task has finished scanning, it sends its results over the Dragonfly HTTP API.

Environment variables

Below are a list of environment variables that need to be configured, and what they do

Variable Default Description
DRAGONFLY_BASE_URL https://dragonfly.vipyrsec.com The base API URL for the mainframe server
DRAGONFLY_CF_ACCESS_CLIENT_ID Environment-specific Cloudflare Access service-token client ID
DRAGONFLY_CF_ACCESS_CLIENT_SECRET Environment-specific Cloudflare Access service-token client secret
DRAGONFLY_THREADS Available parallelism / 1 Concurrent package workers; set to 1 for sequential constrained deployments
DRAGONFLY_LOAD_DURATION 60 Seconds to wait between each API job request
DRAGONFLY_BULK_SIZE 20 Upper bound for job request, also capped by the worker count
DRAGONFLY_MAX_ARCHIVE_ENTRIES 4096 Maximum number of entries in one archive
DRAGONFLY_MAX_DISTRIBUTIONS 32 Maximum number of distributions in one package
DRAGONFLY_MAX_DOWNLOAD_SIZE 33554432 Maximum compressed distribution size in bytes
DRAGONFLY_MAX_EXPANDED_SIZE 67108864 Maximum total expanded distribution size in bytes
DRAGONFLY_MAX_SCAN_SIZE 16777216 Maximum individual file size passed to YARA in bytes

About

Rust client for Dragonfly.

Resources

Code of conduct

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages