分析网站访问日志,追踪 AI 爬虫的抓取行为
Combined Log Format。日志文件中包含 IP、时间、请求方法、路径、状态码、User-Agent。
/www/wwwlogs/你的域名.log
/var/log/nginx/access.log
支持 .log 和 .txt 文件
把下面的脚本保存到服务器,加个 cron 每天自动上传前一天的日志。
#!/bin/bash
# 站擎 AI 爬虫日志自动上传
# 使用前:
# 1. 到 /account 页面生成「API 上传令牌」
# 2. 到 /dashboard 找到站点 ID(点站点进详情页,URL 里的 ?id=xxx)
# 3. 替换下面的两个变量
# 4. 保存为 /root/zhanqing-upload.sh 并 chmod +x
# 5. crontab -e 加:0 5 * * * /root/zhanqing-upload.sh
SITE_ID="你的站点ID"
UPLOAD_TOKEN="你的上传令牌"
LOG_FILE="/www/wwwlogs/zq.kuailexue.com.cn.log"
# 取最近 3 万行日志
tail -n 30000 "$LOG_FILE" > /tmp/zq-crawler-tmp.log
# 用 Python 打包成 JSON(避免依赖 jq)
python3 << PYEOF
import json, urllib.request
with open('/tmp/zq-crawler-tmp.log', 'r', errors='ignore') as f:
log_text = f.read()
body = json.dumps({"siteId": "$SITE_ID", "logText": log_text}).encode('utf-8')
req = urllib.request.Request(
"https://zq.kuailexue.com.cn/api/ai/crawlers/upload",
data=body,
headers={
"Content-Type": "application/json",
"X-Upload-Token": "$UPLOAD_TOKEN",
},
)
try:
resp = urllib.request.urlopen(req, timeout=120)
print(resp.read().decode())
except Exception as e:
print("Error:", e)
PYEOF
rm -f /tmp/zq-crawler-tmp.log
echo "[$(date)] 日志上传完成"
提示:上传令牌在 我的账户 页面生成。