Generative AI, in Japanese, on your own machine.
AI engineer and technical consultant in Tokyo. I build local AI tools, train small Japanese language models, and turn research implementations into usable software for Windows, Linux and Apple Silicon.
日本語 · Hugging Face · Zenn · X
sokudan — Japanese decisions without text generation
A 314.6M-parameter ModernBERT-ja model that returns typed choice, score and bool answers with probabilities in one forward pass. Includes an HTTP server, a TypeScript SDK and a CLI; inference uses MLX on supported Apple Silicon Macs or PyTorch on CUDA / CPU.
v0.3.0 is available through five channels: PyPI · npm · Homebrew tap · Scoop bucket · GHCR.
The npm package combines a CLI wrapper and a typed HTTP SDK. Homebrew and Scoop install launchers; the container targets Linux amd64 CPU. CI covers installation, SDK/server integration and container model inference. Installation guide · Model weights
ComfyUI-JoyAI-Video-Edit — edit video with text instructions
An unofficial integration of JoyAI-Video-Edit, with a persistent WSL2 worker, model load/unload controls, cancellation and memory guards. v0.1.0 pre-release, tested on Windows 11 with an RTX 5090: 840×480 landscape output, two inference steps, up to 600 input frames.
Validation compares against a reference with the same low-memory placement patches; an unmodified upstream end-to-end run was resource-blocked. Production outputs are not bitwise repeatable. Setup and validation scope
I publish code, setup instructions and measured results together. Hardware, inputs, upstream differences and untested paths are documented in each project's README. ComfyUI integrations are community projects; upstream code and weights retain their own licenses.
1LM → 2LM → 3LM explores character tokenization, subwords and modern decoder architectures on a single machine. Base models are trained from scratch; GAL variants study the trade-off between persona fine-tuning and general ability.
| Series | Focus | Repositories |
|---|---|---|
| 1LM · 11.5M | Character-level mini-GPT | MLX · Blackwell |
| 2LM · 13.81M | SentencePiece, dialogue data and fixed evaluation | MLX · Blackwell |
| 3LM · 35.66M | RoPE, RMSNorm, SwiGLU and overnight pretraining | MLX · Weights |
| GAL variants | Persona data, fine-tuning and regression measurement | 2LM MLX · 2LM Blackwell · 3LM MLX · Weights |
| Project | What it provides |
|---|---|
| Qwen3-TTS-JP | Windows-native TTS, voice design and cloning; multilingual Web UI and Whisper transcription. |
| Qwen3-TTS-Mac-GeneLab | Apple Silicon TTS with MLX / PyTorch, quantization and voice cloning. |
| Style-BERT-VITS2-GeneLab-Blackwell | Windows-native Blackwell support, with CPU/GPU fallback. |
| Area | Projects | Focus |
|---|---|---|
| Video editing | JoyAI-Video-Edit | Instruction-based editing with a persistent worker. |
| Fast video | SparkDiffusion · LongLive-Plug | Sparse attention, few-step sampling and adapter/scheduler verification. |
| Video attention | MonarchRT | Monarch-matrix attention vs. a same-weight dense baseline. |
| Causal video | CausalForcing · NVIDIA-CMD | Autoregressive generation, image-to-video and long rollouts. |
| Talking heads | LeapTalk · TBDub | Speech-driven portraits and video redubbing. |
| Image generation | LoopedDiT | Pixel-space diffusion and controlled loop-depth comparisons. |
| Image restoration | PixRestore · MiRipple | One-step restoration and verified local artifact repair. |
| Music | AceMusic | ACE-Step song generation, editing and LoRA support. |
| Windows setup | Win-Blackwell | ComfyUI installation and tested workflows for RTX 50-series GPUs. |
| Project | What it provides |
|---|---|
| gassan | React modal components built on native <dialog>, with accessible activation gates and explicit close reasons. |
| lensing | Dependency-free JavaScript/CSS glass refraction effects using signed distance fields and Snell's law. |
| MermGraph | A Tauri Mermaid editor with live preview and PNG/SVG export. |
| VMagic | Tauri video conversion with FFmpeg, RIFE frame interpolation and Real-ESRGAN upscaling. |
| aitxt | A Go CLI for text processing and developer workflows through OpenAI, Anthropic and Gemini APIs. |
| RustConv / dtx | A Rust CLI for converting, querying, validating and merging structured data. |
| imgai · ghstat | Go CLIs for batch image/EXIF processing and GitHub statistics. |
Languages & scripting
AI & model runtimes
Web & desktop
Build, delivery & platforms
More libraries, APIs and development tools
Model tooling
Scientific computing & media
UI, graphics & services
CLI, APIs & native code
Testing & packaging
Operating systems
Built across my ML projects, React libraries, desktop apps, web tooling and Go CLIs. Docker Compose is also part of my everyday workflow. Main hardware: RTX 5090 (Blackwell) and Apple Silicon. Badges without a brand logo use a generic code icon.
Reproducible checks for answer-order sensitivity in open decision models:
- Laya: score-position bias report, regression checks, non-English inputs and choice questions.
- kev / lev: kev measurements, lev measurements and presentation checks / order averaging.
- Mi-Ripple: OpenCV runtime handling, found while validating the ComfyUI integration.
DiM-2 explores Mamba-2 / Structured State Space Duality for image and video diffusion. These preprints present architectural proposals and experimental directions; implementation results are reported separately in the OSS repositories above.
Related DiM-2 / SSD preprints
| Preprint | Direction |
|---|---|
| SSD-CM | Few-step consistency distillation |
| SSD-Control | Depth, pose and edge conditioning |
| SSD-Portrait | Audio-driven portraits and lip synchronization |
| SSD-SR | Image and video super-resolution |
| SSD-Flow | Flow matching with Mamba-2 |
| SSD-Edit | Instruction-guided image and video editing |
I write about implementation and experiments in Japanese and English.
Zenn · Qiita · note · Medium · dev.to · X
If these tools save you time, a coffee helps me maintain them.




