仓库
hacksider 的仓库
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
ACE-Step: A Step Towards Music Generation Foundation Model
The most powerful local music generation model that outperforms most commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
AI-Driven Research Assistant: An advanced multi-agent system for automating complex research processes. Leveraging LangChain, OpenAI GPT, and LangGraph, this tool streamlines hypothesis generation, data analysis, visualization, and report writing. Perfect for researchers and data scientists seeking to enhance their workflow and productivity.
A full stack app for interruptible, low-latency and near-human quality AI phone calls built from stitching LLMs, speech understanding tools, text-to-speech models, and Twilio’s phone API.
Detect and fix skew in images containing text
Notebooks and recipes for creating custom entity recognizer for Amazon comprehend.
AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Build like a team of hundreds_
Pytorch implementation of Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
Open-source unified multimodal model
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Pytorch-Named-Entity-Recognition-with-BERT