Today
Trending in AI today
Today's trending AI news, launches and research from five sources people actually watch — Hacker News, Hugging Face, Product Hunt, arXiv and GitHub Trending — in one place, refreshed through the day.
- Items right now
- 237
Hacker News
58 items1. Void: Open-source Cursor alternative
1 day ago
⭐️
873
2. ALICE detects the conversion of lead into gold at the LHC
22 hours ago
⭐️
575
3. LegoGPT: Generating Physically Stable and Buildable Lego
1 day ago
⭐️
559
4. NSF faces shake-up as officials abolish its 37 divisions
1 day ago
⭐️
536
5. Business books are entertainment, not strategic tools
16 hours ago
⭐️
401
6. Rust’s dependencies are starting to worry me
1 day ago
⭐️
345
7. Sofie: open-source web based system for automating live TV news production
23 hours ago
⭐️
339
8. Vision Now Available in Llama.cpp
9 hours ago
⭐️
338
9. Show HN: Hyvector – A fast and modern SVG editor
1 day ago
⭐️
303
10. 21 GB/s CSV Parsing Using SIMD on AMD 9950X
23 hours ago
⭐️
296
11. Starlink User Terminal Teardown
1 day ago
⭐️
291
12. Itter.sh – Micro-Blogging via Terminal
23 hours ago
⭐️
258
13. Progress toward fusion energy gain as measured against the Lawson criteria
1 day ago
⭐️
227
14. CryptPad: An Alternative to the Google Suite
1 day ago
⭐️
213
15. Show HN: Aberdeen – An elegant approach to reactive UIs
1 day ago
⭐️
212
16. Data manipulations alleged in study that paved way for Microsoft's quantum chip
1 day ago
⭐️
210
17. Audiobookshelf: Self-hosted audiobook and podcast server
1 day ago
⭐️
193
18. All BART trains were stopped due to ‘computer networking problem’
22 hours ago
⭐️
192
19. WebGL Water (2010)
13 hours ago
⭐️
174
20. What’s new in Swift 6.2
16 hours ago
⭐️
167
21. Internet Roadtrip: Vote to steer
11 hours ago
⭐️
164
22. Brandon's Semiconductor Simulator
12 hours ago
⭐️
161
23. Gmail to SQLite
8 hours ago
⭐️
159
24. Fui: C library for interacting with the framebuffer in a TTY context
1 day ago
⭐️
158
25. A Formal Analysis of Apple's iMessage PQ3 Protocol [pdf]
1 day ago
⭐️
152
26. How to start a school with your friends
1 day ago
⭐️
149
27. Launch HN: Nao Labs (YC X25) – Cursor for Data
20 hours ago
⭐️
146
28. A simple 16x16 dot animation from simple math rules
10 hours ago
⭐️
141
29. Fleurs du Mal
14 hours ago
⭐️
130
30. Past, present, and future of Sorbet type syntax
21 hours ago
⭐️
127
31. Implementing a Struct of Arrays
1 day ago
⭐️
118
32. Odin, a Pragmatic C Alternative with a Go Flavour
19 hours ago
⭐️
114
33. Private Japanese lunar lander enters orbit around moon ahead of a June touchdown
8 hours ago
⭐️
109
34. Why 536 was 'the worst year to be alive' (2018)
18 hours ago
⭐️
108
35. Slow software for a burning world
6 hours ago
⭐️
98
36. Show HN: Oliphaunt – A native Mastodon client for macOS
20 hours ago
⭐️
94
37. New Tool: lsds – List All Linux Block Devices and Settings in One Place
19 hours ago
⭐️
88
38. The Linux Kernel's PGP Web of Trust
1 day ago
⭐️
84
39. Reverse Engineering "DNA Sequences" in the Lost World: Jurassic Park Video Game
19 hours ago
⭐️
84
40. Show HN: A backend agnostic Ruby framework for building reactive desktop apps
21 hours ago
⭐️
82
41. Google Doc Templates for Startups
17 hours ago
⭐️
82
42. Zombieverter: Open source VCU for reusing salvage EV components
1 day ago
⭐️
75
43. Not a three-year-old chimney sweep (2022)
6 hours ago
⭐️
74
44. Charles Bukowski, William Burroughs, and the Computer (2009)
12 hours ago
⭐️
73
45. Europe launches program to lure scientists away from the US
6 hours ago
⭐️
72
46. PlainBudget – Minimalist Plain Text Budgeting
13 hours ago
⭐️
72
47. Math Machine – A notebook will show your kid how far they have travelled
18 hours ago
⭐️
71
48. "Night of the Living Dead" accidentally became public domain (2019)
1 day ago
⭐️
71
49. Inventing the Adventure Game (1984)
18 hours ago
⭐️
65
50. Show HN: Hyper – Standards-first React alternative
23 hours ago
⭐️
62
51. LTXVideo 13B AI video generation
1 hour ago
⭐️
61
52. Detect and crash Chromium bots
6 hours ago
⭐️
61
53. Some novelists are becoming video game writers – and vice-versa
15 hours ago
⭐️
57
54. Linear Programming for Fun and Profit
1 day ago
⭐️
56
55. Show HN: BlenderQ – A TUI for managing multiple Blender renders
21 hours ago
⭐️
56
56. Show HN: Hydra (YC W22) – Serverless Analytics on Postgres
20 hours ago
⭐️
53
57. The birth of AI poker? Letters from the 1984 WSOP
22 hours ago
⭐️
51
58. Stratolaunch Successfully Completes Reusable Hypersonic Flight and Recovery
12 hours ago
⭐️
51
Hugging Face
60 items1. ibm-granite/granite-4.0-tiny-preview
ibm-granite/granite-4.0-tiny-preview
Text Generation · 1.74k Download · 92 Like
2. lusxvr/nanoVLM-222M
lusxvr/nanoVLM-222M
Image-Text-to-Text · 496 Download · 53 Like
3. Qwen/Qwen3-0.6B
Qwen/Qwen3-0.6B
Text Generation · 314k Download · 220 Like
4. deepseek-ai/DeepSeek-R1
deepseek-ai/DeepSeek-R1
Text Generation · 1.27M Download
5. Qwen/Qwen2.5-Omni-3B
Qwen/Qwen2.5-Omni-3B
Any-to-Any · 16.4k Download · 188 Like
6. microsoft/Phi-4-reasoning
microsoft/Phi-4-reasoning
Text Generation · 6.89k Download · 161 Like
7. tencent/Hunyuan3D-2
tencent/Hunyuan3D-2
Image-to-3D · 419k Download · 1.4k Like
8. HiDream-ai/HiDream-E1-Full
HiDream-ai/HiDream-E1-Full
Any-to-Any · 3.84k Download · 143 Like
9. nvidia/parakeet-tdt-0.6b-v2
nvidia/parakeet-tdt-0.6b-v2
Automatic Speech Recognition · 72.7k Download · 659 Like
10. ACE-Step/ACE-Step-v1-3.5B
ACE-Step/ACE-Step-v1-3.5B
Text-to-Audio · 287 Download
11. nari-labs/Dia-1.6B
nari-labs/Dia-1.6B
Text-to-Speech · 148k Download
12. Lightricks/LTX-Video
Lightricks/LTX-Video
Text-to-Video · 214k Download
13. deepseek-ai/DeepSeek-Prover-V2-671B
deepseek-ai/DeepSeek-Prover-V2-671B
Text Generation · 7.38k Download
14. Qwen/Qwen3-235B-A22B
Qwen/Qwen3-235B-A22B
Text Generation · 82.4k Download
15. JetBrains/Mellum-4b-base
JetBrains/Mellum-4b-base
Text Generation · 2.81k Download · 304 Like
16. lodestones/Chroma
lodestones/Chroma
Text-to-Image · 392 Download
17. Qwen/Qwen3-30B-A3B
Qwen/Qwen3-30B-A3B
Text Generation · 150k Download
18. black-forest-labs/FLUX.1-dev
black-forest-labs/FLUX.1-dev
Text-to-Image · 2.66M Download
19. cognition-ai/Kevin-32B
cognition-ai/Kevin-32B
Updated
4 days ago · 87 Download
20. microsoft/Phi-4-reasoning-plus
microsoft/Phi-4-reasoning-plus
Text Generation · 10.8k Download · 236 Like
21. Tesslate/UIGEN-T2-7B-Q8_0-GGUF
Tesslate/UIGEN-T2-7B-Q8_0-GGUF
Text Generation · 2.59k Download · 113 Like
22. fdtn-ai/Foundation-Sec-8B
fdtn-ai/Foundation-Sec-8B
Text Generation · 21.5k Download · 143 Like
23. nvidia/parakeet-tdt-0.6b-v2
nvidia/parakeet-tdt-0.6b-v2
Automatic Speech Recognition · 72.7k Download · 659 Like
24. ACE-Step/ACE-Step-v1-3.5B
ACE-Step/ACE-Step-v1-3.5B
Text-to-Audio · 287 Download
25. nari-labs/Dia-1.6B
nari-labs/Dia-1.6B
Text-to-Speech · 148k Download
26. Lightricks/LTX-Video
Lightricks/LTX-Video
Text-to-Video · 214k Download
27. deepseek-ai/DeepSeek-Prover-V2-671B
deepseek-ai/DeepSeek-Prover-V2-671B
Text Generation · 7.38k Download
28. Qwen/Qwen3-235B-A22B
Qwen/Qwen3-235B-A22B
Text Generation · 82.4k Download
29. JetBrains/Mellum-4b-base
JetBrains/Mellum-4b-base
Text Generation · 2.81k Download · 304 Like
30. lodestones/Chroma
lodestones/Chroma
Text-to-Image · 392 Download
31. Qwen/Qwen3-30B-A3B
Qwen/Qwen3-30B-A3B
Text Generation · 150k Download
32. black-forest-labs/FLUX.1-dev
black-forest-labs/FLUX.1-dev
Text-to-Image · 2.66M Download
33. cognition-ai/Kevin-32B
cognition-ai/Kevin-32B
Updated
4 days ago · 87 Download
34. microsoft/Phi-4-reasoning-plus
microsoft/Phi-4-reasoning-plus
Text Generation · 10.8k Download · 236 Like
35. Tesslate/UIGEN-T2-7B-Q8_0-GGUF
Tesslate/UIGEN-T2-7B-Q8_0-GGUF
Text Generation · 2.59k Download · 113 Like
36. fdtn-ai/Foundation-Sec-8B
fdtn-ai/Foundation-Sec-8B
Text Generation · 21.5k Download · 143 Like
37. hexgrad/Kokoro-82M
hexgrad/Kokoro-82M
Text-to-Speech · 1.88M Download
38. tencent/HunyuanCustom
tencent/HunyuanCustom
Image-to-Video · 63 Download
39. Qwen/Qwen3-8B
Qwen/Qwen3-8B
Text Generation · 394k Download · 259 Like
40. ServiceNow-AI/Apriel-Nemotron-15b-Thinker
ServiceNow-AI/Apriel-Nemotron-15b-Thinker
Updated
4 days ago · 62 Download
41. XiaomiMiMo/MiMo-7B-RL
XiaomiMiMo/MiMo-7B-RL
Text Generation · 5.9k Download · 242 Like
42. Goekdeniz-Guelmez/Josiefied-Qwen3-8B-abliterated-v1
Goekdeniz-Guelmez/Josiefied-Qwen3-8B-abliterated-v1
Text Generation · 427 Download · 63 Like
43. sanaka87/ICEdit-MoE-LoRA
sanaka87/ICEdit-MoE-LoRA
Image-to-Image · 3.04k Download · 86 Like
44. ds4sd/SmolDocling-256M-preview
ds4sd/SmolDocling-256M-preview
Image-Text-to-Text · 84.8k Download · 1.34k Like
45. hexgrad/Kokoro-82M
hexgrad/Kokoro-82M
Text-to-Speech · 1.88M Download
46. tencent/HunyuanCustom
tencent/HunyuanCustom
Image-to-Video · 63 Download
47. Qwen/Qwen3-8B
Qwen/Qwen3-8B
Text Generation · 394k Download · 259 Like
48. ServiceNow-AI/Apriel-Nemotron-15b-Thinker
ServiceNow-AI/Apriel-Nemotron-15b-Thinker
Updated
4 days ago · 62 Download
49. XiaomiMiMo/MiMo-7B-RL
XiaomiMiMo/MiMo-7B-RL
Text Generation · 5.9k Download · 242 Like
50. Goekdeniz-Guelmez/Josiefied-Qwen3-8B-abliterated-v1
Goekdeniz-Guelmez/Josiefied-Qwen3-8B-abliterated-v1
Text Generation · 427 Download · 63 Like
51. sanaka87/ICEdit-MoE-LoRA
sanaka87/ICEdit-MoE-LoRA
Image-to-Image · 3.04k Download · 86 Like
52. ds4sd/SmolDocling-256M-preview
ds4sd/SmolDocling-256M-preview
Image-Text-to-Text · 84.8k Download · 1.34k Like
53. ibm-granite/granite-4.0-tiny-preview
ibm-granite/granite-4.0-tiny-preview
Text Generation · 1.74k Download · 92 Like
54. lusxvr/nanoVLM-222M
lusxvr/nanoVLM-222M
Image-Text-to-Text · 496 Download · 53 Like
55. Qwen/Qwen3-0.6B
Qwen/Qwen3-0.6B
Text Generation · 314k Download · 220 Like
56. deepseek-ai/DeepSeek-R1
deepseek-ai/DeepSeek-R1
Text Generation · 1.27M Download
57. Qwen/Qwen2.5-Omni-3B
Qwen/Qwen2.5-Omni-3B
Any-to-Any · 16.4k Download · 188 Like
58. microsoft/Phi-4-reasoning
microsoft/Phi-4-reasoning
Text Generation · 6.89k Download · 161 Like
59. tencent/Hunyuan3D-2
tencent/Hunyuan3D-2
Image-to-3D · 419k Download · 1.4k Like
60. HiDream-ai/HiDream-E1-Full
HiDream-ai/HiDream-E1-Full
Any-to-Any · 3.84k Download · 143 Like
Product Hunt
9 items1. SeeMuseums
The first agentive art museum tour guide
🔼
195
2. Rubber Duck Armada
A fun clicker-game that's all about rubber ducks
🔼
159
3. Aladin
Supercharge your browser with AI
🔼
150
4. Photo AI 4
Studio-level focus, detail, and clarity anywhere you shoot
🔼
137
5. QU3
Quantum-safe MCP servers for private, verifiable inference
🔼
129
6. AI Voice Cloning
Clone any voice in 3 seconds – hyper-realistic and free
🔼
126
7. Aurascope
Scan the world, collect aura points, compete with friends
🔼
126
8. Voila
Open-source AI for real-time, expressive voice role-play
🔼
114
9. Autograph Fair Offer
Estimate fair compensation for any role
🔼
109
arXiv
100 items1.
Generating Physically Stable and Buildable LEGO Designs from Text
Abstract:
We introduce LegoGPT, the first approach for generating physically stable LEGO brick models from text prompts. To achieve this, we construct a…
▽ More
| Submitted 8 May, 2025
2.
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
Abstract:
We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in adapting existing…
▽ More
| Submitted 8 May, 2025
3.
ComPO: Preference Alignment via Comparison Oracles
Abstract:
Direct alignment methods are increasingly used for aligning large…
▽ More
| Submitted 8 May, 2025
4.
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
Abstract:
Vision-Language…
▽ More
| Submitted 8 May, 2025
5.
UKElectionNarratives: A Dataset of Misleading Narratives Surrounding Recent UK General Elections
Abstract:
…the first dataset of human-annotated misleading narratives which circulated during the UK General Elections in 2019 and 2024. We also benchmark Pre-trained and Large Language Models (focusing on GPT-4o), studying their effectiveness in detecting election-related misleading narrat…
▽ More
| Submitted 8 May, 2025
6.
SITE: towards Spatial Intelligence Thorough Evaluation
Abstract:
…to robotics. We introduce SITE, a benchmark dataset towards SI Thorough Evaluation in a standardized format of multi-choice visual question-answering, designed to assess large vision-…
▽ More
| Submitted 8 May, 2025
7.
Conversational Process Model Redesign
Abstract:
With the recent success of large…
▽ More
| Submitted 8 May, 2025
8.
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
Abstract:
The emergence of instruction-tuned large…
▽ More
| Submitted 8 May, 2025
9.
MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
10.
GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality
Abstract:
Large Language Model (LLM)-based copilots have shown great potential in Extended Reality (XR) applications. However, the user faces challenges when describing the 3D environments to the copilots due to the complexity of conveying spatial-temporal information through text or speec…
▽ More
| Submitted 8 May, 2025
11.
EcoAgent: An Efficient Edge-Cloud Collaborative Multi-Agent Framework for Mobile Automation
Abstract:
Cloud-based mobile agents powered by (multimodal) large language models ((M)LLMs) offer strong reasoning abilities but suffer from high latency and cost. While fine-tuned (M)SLMs enable edge deployment, they often lose general capabilities and struggle with complex tasks. To addr…
▽ More
| Submitted 8 May, 2025
12.
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
Abstract:
Data quality has become a key factor in enhancing model performance with the rapid development of large language models (LLMs). Model-driven data filtering has increasingly become a primary approach f…
▽ More
| Submitted 8 May, 2025
13.
TransProQA: an LLM-based literary Translation evaluation metric with Professional Question Answering
Abstract:
The impact of Large…
▽ More
| Submitted 8 May, 2025
14.
Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
15.
Crosslingual Reasoning through Test-Time Scaling
Abstract:
Reasoning capabilities of large…
▽ More
| Submitted 8 May, 2025
16.
Frame In, Frame Out: Do LLMs Generate More Biased News Headlines than Humans?
Abstract:
Framing in media critically shapes public perception by selectively emphasizing some details while downplaying others. With the rise of large…
▽ More
| Submitted 8 May, 2025
17.
Let's Ask GNN: Empowering Large Language Model for Graph In-Context Learning
Abstract:
Textual Attributed Graphs (TAGs) are crucial for modeling complex real-world systems, yet leveraging large language models (LLMs) for TAGs presents unique challenges due to the gap between sequential text processing and graph-structured dat…
▽ More
| Submitted 8 May, 2025
18.
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
19.
DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning
Abstract:
…vehicles into a unified framework. DSDrive leverages a compact LLM that employs a distillation method to preserve the enhanced reasoning capabilities of a larger-sized vision language…
▽ More
| Submitted 8 May, 2025
20.
Towards the Worst-case Robustness of Large Language Models
Abstract:
Recent studies have revealed the vulnerability of large language models to adversarial attacks, where adversaries craft specific input sequences to induce harmful, violent, private, or incorrect outputs. In this work, we study their worst-case robustness, i.e., whether an adversa…
▽ More
| Submitted 8 May, 2025
21.
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
Abstract:
Large-scale vision-…
▽ More
| Submitted 8 May, 2025
22.
FLAM: Frame-Wise Language-Audio Modeling
Abstract:
Recent multi-modal audio-language…
▽ More
| Submitted 8 May, 2025
23.
ICon: In-Context Contribution for Automatic Data Selection
Abstract:
Data selection for instruction tuning is essential for improving the performance of Large…
▽ More
| Submitted 8 May, 2025
24.
Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects
Abstract:
The rapid adoption of Vision Language Models (VLMs), pre-trained on large image-text and video-text datasets, calls for protecting and informing users about when to trust these systems. This survey reviews studies on trust dynamics in user-VLM interactions, through a multi-discip…
▽ More
| Submitted 8 May, 2025
25.
Scalable Chain of Thoughts via Elastic Reasoning
Abstract:
Large reasoning…
▽ More
| Submitted 8 May, 2025
26.
Toward Reasonable Parrots: Why Large Language Models Should Argue with Us by Design
Abstract:
…paper, we advocate for the development of conversational technology that is inherently designed to support and facilitate argumentative processes. We argue that, at present, large language models (LLMs) are inadequate for this purpose, and we propose an ideal technology design ai…
▽ More
| Submitted 8 May, 2025
27.
HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow
Abstract:
Recent advances in leveraging the agentic paradigm of large language models (LLMs) utilization have significantly enhanced Text-to-SQL capabilities, enabling users without specialized database expertise to query data intuitively. However, deploying these agentic LLM-based Text-to…
▽ More
| Submitted 8 May, 2025
28.
Software Development Life Cycle Perspective: A Survey of Benchmarks for CodeLLMs and Agents
Abstract:
Code large…
▽ More
| Submitted 8 May, 2025
29.
A Two-Sample Test of Text Generation Similarity
Abstract:
…across two groups of documents. The proposed test aims to assess text similarity by comparing the entropy of the documents. Entropy is estimated using neural network-based language…
▽ More
| Submitted 8 May, 2025
30.
Generating Symbolic World Models via Test-time Scaling of Large Language Models
Abstract:
Solving complex planning problems requires Large…
▽ More
| Submitted 8 May, 2025
31.
PADriver: Towards Personalized Autonomous Driving
Abstract:
In this paper, we propose PADriver, a novel closed-loop framework for personalized autonomous driving (PAD). Built upon Multi-modal Large Language Model (MLLM), PADriver takes streaming frames and personalized textual prompts as inputs. It autoaggressively performs scene understa…
▽ More
| Submitted 8 May, 2025
32.
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
33.
Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning
Abstract:
Few-shot tabular learning, in which machine learning models are trained with a limited amount of labeled data, provides a cost-effective approach to addressing real-world challenges. The advent of…
▽ More
| Submitted 8 May, 2025
34.
ChemRxivQuest: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv Preprints
Abstract:
…chemistry literature poses significant challenges for researchers seeking to efficiently access domain-specific knowledge. To support advancements in chemistry-focused natural language processing (NLP), we present ChemRxivQuest, a curated dataset of 970 high-quality question-answer (QA) pairs derived from 155 ChemRxiv preprints across 17 subfields of chemist…
▽ More
| Submitted 8 May, 2025
35.
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
Abstract:
The rapid advancement of Chinese large…
▽ More
| Submitted 8 May, 2025
36.
SmallPlan: Leverage Small Language Models for Sequential Path Planning with Simulation-Powered, LLM-Guided Distillation
Abstract:
Efficient path planning in robotics, particularly within large-scale, dynamic environments, remains a significant hurdle. While…
▽ More
| Submitted 8 May, 2025
37.
Stealthy LLM-Driven Data Poisoning Attacks Against Embedding-Based Retrieval-Augmented Recommender Systems
Abstract:
…constraints, and we examine their effectiveness in both promotion (long-tail items) and demotion (short-head items) scenarios. Our experiments on MovieLens, using two large language model (LLM) retrieval modules, show that even subtle attacks shift final rankings and item exposur…
▽ More
| Submitted 8 May, 2025
38.
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
Abstract:
Text watermarking aims to subtly embed statistical signals into text by controlling the Large…
▽ More
| Submitted 8 May, 2025
39.
Biomed-DPT: Dual Modality Prompt Tuning for Biomedical Vision-Language Models
Abstract:
Prompt learning is one of the most effective paradigms for adapting pre-trained vision-language…
▽ More
| Submitted 8 May, 2025
40.
MARK: Memory Augmented Refinement of Knowledge
Abstract:
Large Language Models (LLMs) assist in specialized tasks but struggle to align with evolving domain knowledge without costly fine-tuning. Domain knowledge consists of: Knowledge: Immutable facts (e.g., 'A stone is solid') and generally accepted principles (e.g., ethical s…
▽ More
| Submitted 8 May, 2025
41.
Re-evaluating Open-ended Evaluation of Large Language Models
Abstract:
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as…
▽ More
| Submitted 8 May, 2025
42.
Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks
Abstract:
The application of large…
▽ More
| Submitted 8 May, 2025
43.
Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models
Abstract:
Vision-Language…
▽ More
| Submitted 8 May, 2025
44.
FedTDP: A Privacy-Preserving and Unified Framework for Trajectory Data Preparation via Federated Learning
Abstract:
…(i) they do not address data privacy concerns, particularly in federated settings where trajectory data sharing is prohibited, and (ii) they typically design task-specific models that lack generalizability across diverse TDP scenarios. To overcome these challenges, we propose FedTDP, a privacy-preserving and unified framework that leverages the capabilities…
▽ More
| Submitted 8 May, 2025
45.
CacheFL: Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
Abstract:
Large pre-trained Vision-…
▽ More
| Submitted 8 May, 2025
46.
Faster, Cheaper, Better: Multi-Objective Hyperparameter Optimization for LLM and RAG Systems
Abstract:
While Retrieval Augmented Generation (RAG) has emerged as a popular technique for improving Large…
▽ More
| Submitted 8 May, 2025
47.
Large Language Models Understanding: an Inherent Ambiguity Barrier
Abstract:
A lively ongoing debate is taking place, since the extraordinary emergence of Large Language Models (LLMs) with regards to their capability to understand the world and capture the meaning of the dialogues in which they are involved. Arguments and counter-arguments have been propo…
▽ More
| Submitted 8 May, 2025
48.
Text2Cypher: Data Pruning using Hard Example Selection
Abstract:
Database query languages such as SQL for relational databases and Cypher for graph databases have been widely adopted. Recent advancements in…
▽ More
| Submitted 8 May, 2025
49.
Enhancing Text2Cypher with Schema Filtering
Abstract:
Knowledge graphs represent complex data using nodes, relationships, and properties. Cypher, a powerful query language for graph databases, enables efficient…
▽ More
| Submitted 8 May, 2025
50.
How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape Game
Abstract:
The rapid advancing of Multimodal Large…
▽ More
| Submitted 8 May, 2025
51.
Generating Physically Stable and Buildable LEGO Designs from Text
Abstract:
We introduce LegoGPT, the first approach for generating physically stable LEGO brick models from text prompts. To achieve this, we construct a…
▽ More
| Submitted 8 May, 2025
52.
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
Abstract:
We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in adapting existing…
▽ More
| Submitted 8 May, 2025
53.
ComPO: Preference Alignment via Comparison Oracles
Abstract:
Direct alignment methods are increasingly used for aligning large…
▽ More
| Submitted 8 May, 2025
54.
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
Abstract:
Vision-Language…
▽ More
| Submitted 8 May, 2025
55.
UKElectionNarratives: A Dataset of Misleading Narratives Surrounding Recent UK General Elections
Abstract:
…the first dataset of human-annotated misleading narratives which circulated during the UK General Elections in 2019 and 2024. We also benchmark Pre-trained and Large Language Models (focusing on GPT-4o), studying their effectiveness in detecting election-related misleading narrat…
▽ More
| Submitted 8 May, 2025
56.
SITE: towards Spatial Intelligence Thorough Evaluation
Abstract:
…to robotics. We introduce SITE, a benchmark dataset towards SI Thorough Evaluation in a standardized format of multi-choice visual question-answering, designed to assess large vision-…
▽ More
| Submitted 8 May, 2025
57.
Conversational Process Model Redesign
Abstract:
With the recent success of large…
▽ More
| Submitted 8 May, 2025
58.
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
Abstract:
The emergence of instruction-tuned large…
▽ More
| Submitted 8 May, 2025
59.
MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
60.
GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality
Abstract:
Large Language Model (LLM)-based copilots have shown great potential in Extended Reality (XR) applications. However, the user faces challenges when describing the 3D environments to the copilots due to the complexity of conveying spatial-temporal information through text or speec…
▽ More
| Submitted 8 May, 2025
61.
EcoAgent: An Efficient Edge-Cloud Collaborative Multi-Agent Framework for Mobile Automation
Abstract:
Cloud-based mobile agents powered by (multimodal) large language models ((M)LLMs) offer strong reasoning abilities but suffer from high latency and cost. While fine-tuned (M)SLMs enable edge deployment, they often lose general capabilities and struggle with complex tasks. To addr…
▽ More
| Submitted 8 May, 2025
62.
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
Abstract:
Data quality has become a key factor in enhancing model performance with the rapid development of large language models (LLMs). Model-driven data filtering has increasingly become a primary approach f…
▽ More
| Submitted 8 May, 2025
63.
TransProQA: an LLM-based literary Translation evaluation metric with Professional Question Answering
Abstract:
The impact of Large…
▽ More
| Submitted 8 May, 2025
64.
Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
65.
Crosslingual Reasoning through Test-Time Scaling
Abstract:
Reasoning capabilities of large…
▽ More
| Submitted 8 May, 2025
66.
Frame In, Frame Out: Do LLMs Generate More Biased News Headlines than Humans?
Abstract:
Framing in media critically shapes public perception by selectively emphasizing some details while downplaying others. With the rise of large…
▽ More
| Submitted 8 May, 2025
67.
Let's Ask GNN: Empowering Large Language Model for Graph In-Context Learning
Abstract:
Textual Attributed Graphs (TAGs) are crucial for modeling complex real-world systems, yet leveraging large language models (LLMs) for TAGs presents unique challenges due to the gap between sequential text processing and graph-structured dat…
▽ More
| Submitted 8 May, 2025
68.
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
69.
DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning
Abstract:
…vehicles into a unified framework. DSDrive leverages a compact LLM that employs a distillation method to preserve the enhanced reasoning capabilities of a larger-sized vision language…
▽ More
| Submitted 8 May, 2025
70.
Towards the Worst-case Robustness of Large Language Models
Abstract:
Recent studies have revealed the vulnerability of large language models to adversarial attacks, where adversaries craft specific input sequences to induce harmful, violent, private, or incorrect outputs. In this work, we study their worst-case robustness, i.e., whether an adversa…
▽ More
| Submitted 8 May, 2025
71.
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
Abstract:
Large-scale vision-…
▽ More
| Submitted 8 May, 2025
72.
FLAM: Frame-Wise Language-Audio Modeling
Abstract:
Recent multi-modal audio-language…
▽ More
| Submitted 8 May, 2025
73.
ICon: In-Context Contribution for Automatic Data Selection
Abstract:
Data selection for instruction tuning is essential for improving the performance of Large…
▽ More
| Submitted 8 May, 2025
74.
Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects
Abstract:
The rapid adoption of Vision Language Models (VLMs), pre-trained on large image-text and video-text datasets, calls for protecting and informing users about when to trust these systems. This survey reviews studies on trust dynamics in user-VLM interactions, through a multi-discip…
▽ More
| Submitted 8 May, 2025
75.
Scalable Chain of Thoughts via Elastic Reasoning
Abstract:
Large reasoning…
▽ More
| Submitted 8 May, 2025
76.
Toward Reasonable Parrots: Why Large Language Models Should Argue with Us by Design
Abstract:
…paper, we advocate for the development of conversational technology that is inherently designed to support and facilitate argumentative processes. We argue that, at present, large language models (LLMs) are inadequate for this purpose, and we propose an ideal technology design ai…
▽ More
| Submitted 8 May, 2025
77.
HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow
Abstract:
Recent advances in leveraging the agentic paradigm of large language models (LLMs) utilization have significantly enhanced Text-to-SQL capabilities, enabling users without specialized database expertise to query data intuitively. However, deploying these agentic LLM-based Text-to…
▽ More
| Submitted 8 May, 2025
78.
Software Development Life Cycle Perspective: A Survey of Benchmarks for CodeLLMs and Agents
Abstract:
Code large…
▽ More
| Submitted 8 May, 2025
79.
A Two-Sample Test of Text Generation Similarity
Abstract:
…across two groups of documents. The proposed test aims to assess text similarity by comparing the entropy of the documents. Entropy is estimated using neural network-based language…
▽ More
| Submitted 8 May, 2025
80.
Generating Symbolic World Models via Test-time Scaling of Large Language Models
Abstract:
Solving complex planning problems requires Large…
▽ More
| Submitted 8 May, 2025
81.
PADriver: Towards Personalized Autonomous Driving
Abstract:
In this paper, we propose PADriver, a novel closed-loop framework for personalized autonomous driving (PAD). Built upon Multi-modal Large Language Model (MLLM), PADriver takes streaming frames and personalized textual prompts as inputs. It autoaggressively performs scene understa…
▽ More
| Submitted 8 May, 2025
82.
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
Abstract:
Large…
▽ More
| Submitted 8 May, 2025
83.
Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning
Abstract:
Few-shot tabular learning, in which machine learning models are trained with a limited amount of labeled data, provides a cost-effective approach to addressing real-world challenges. The advent of…
▽ More
| Submitted 8 May, 2025
84.
ChemRxivQuest: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv Preprints
Abstract:
…chemistry literature poses significant challenges for researchers seeking to efficiently access domain-specific knowledge. To support advancements in chemistry-focused natural language processing (NLP), we present ChemRxivQuest, a curated dataset of 970 high-quality question-answer (QA) pairs derived from 155 ChemRxiv preprints across 17 subfields of chemist…
▽ More
| Submitted 8 May, 2025
85.
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
Abstract:
The rapid advancement of Chinese large…
▽ More
| Submitted 8 May, 2025
86.
SmallPlan: Leverage Small Language Models for Sequential Path Planning with Simulation-Powered, LLM-Guided Distillation
Abstract:
Efficient path planning in robotics, particularly within large-scale, dynamic environments, remains a significant hurdle. While…
▽ More
| Submitted 8 May, 2025
87.
Stealthy LLM-Driven Data Poisoning Attacks Against Embedding-Based Retrieval-Augmented Recommender Systems
Abstract:
…constraints, and we examine their effectiveness in both promotion (long-tail items) and demotion (short-head items) scenarios. Our experiments on MovieLens, using two large language model (LLM) retrieval modules, show that even subtle attacks shift final rankings and item exposur…
▽ More
| Submitted 8 May, 2025
88.
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
Abstract:
Text watermarking aims to subtly embed statistical signals into text by controlling the Large…
▽ More
| Submitted 8 May, 2025
89.
Biomed-DPT: Dual Modality Prompt Tuning for Biomedical Vision-Language Models
Abstract:
Prompt learning is one of the most effective paradigms for adapting pre-trained vision-language…
▽ More
| Submitted 8 May, 2025
90.
MARK: Memory Augmented Refinement of Knowledge
Abstract:
Large Language Models (LLMs) assist in specialized tasks but struggle to align with evolving domain knowledge without costly fine-tuning. Domain knowledge consists of: Knowledge: Immutable facts (e.g., 'A stone is solid') and generally accepted principles (e.g., ethical s…
▽ More
| Submitted 8 May, 2025
91.
Re-evaluating Open-ended Evaluation of Large Language Models
Abstract:
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as…
▽ More
| Submitted 8 May, 2025
92.
Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks
Abstract:
The application of large…
▽ More
| Submitted 8 May, 2025
93.
Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models
Abstract:
Vision-Language…
▽ More
| Submitted 8 May, 2025
94.
FedTDP: A Privacy-Preserving and Unified Framework for Trajectory Data Preparation via Federated Learning
Abstract:
…(i) they do not address data privacy concerns, particularly in federated settings where trajectory data sharing is prohibited, and (ii) they typically design task-specific models that lack generalizability across diverse TDP scenarios. To overcome these challenges, we propose FedTDP, a privacy-preserving and unified framework that leverages the capabilities…
▽ More
| Submitted 8 May, 2025
95.
CacheFL: Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
Abstract:
Large pre-trained Vision-…
▽ More
| Submitted 8 May, 2025
96.
Faster, Cheaper, Better: Multi-Objective Hyperparameter Optimization for LLM and RAG Systems
Abstract:
While Retrieval Augmented Generation (RAG) has emerged as a popular technique for improving Large…
▽ More
| Submitted 8 May, 2025
97.
Large Language Models Understanding: an Inherent Ambiguity Barrier
Abstract:
A lively ongoing debate is taking place, since the extraordinary emergence of Large Language Models (LLMs) with regards to their capability to understand the world and capture the meaning of the dialogues in which they are involved. Arguments and counter-arguments have been propo…
▽ More
| Submitted 8 May, 2025
98.
Text2Cypher: Data Pruning using Hard Example Selection
Abstract:
Database query languages such as SQL for relational databases and Cypher for graph databases have been widely adopted. Recent advancements in…
▽ More
| Submitted 8 May, 2025
99.
Enhancing Text2Cypher with Schema Filtering
Abstract:
Knowledge graphs represent complex data using nodes, relationships, and properties. Cypher, a powerful query language for graph databases, enables efficient…
▽ More
| Submitted 8 May, 2025
100.
How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape Game
Abstract:
The rapid advancing of Multimodal Large…
▽ More
| Submitted 8 May, 2025
GitHub Trending
10 items1. harry0703 /
MoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM. | like: 28,430 fork: 4,175
⭐️today
435
2. Blaizzy /
mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon. | like: 1,386 fork: 95
⭐️today
174
3. voideditor /
void
| like: 16,011 fork: 931
⭐️today
1190
4. zed-industries /
zed
Code at the speed of thought – Zed is a high-performance, multiplayer code editor from the creators of Atom and Tree-sitter. | like: 58,937 fork: 4,122
⭐️today
256
5. Peterande /
D-FINE
D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement [ICLR 2025 Spotlight] | like: 2,106 fork: 177
⭐️today
23
6. shane-mason /
FieldStation42
Broadcast TV simulator | like: 398 fork: 19
⭐️today
41
7. wolfpld /
tracy
Frame profiler | like: 11,466 fork: 786
⭐️today
35
8. Lightricks /
LTX-Video
Official repository for LTX-Video | like: 4,429 fork: 363
⭐️today
263
9. longbridge /
gpui-component
UI components for building fantastic desktop application by using GPUI. | like: 1,960 fork: 95
⭐️today
503
10. panaversity /
learn-agentic-ai
Learn Agentic AI using Dapr Agentic Cloud Ascent (DACA) Design Pattern and Agent-Native Cloud Technologies: OpenAI Agents SDK, Memory, MCP, A2A, Knowledge Graphs, Dapr, Rancher Desktop, and Kubernetes. | like: 1,442 fork: 437
⭐️today
13