The Fix That Let Us Run the Biggest Open Models Overnight

The Fix That Let Us Run the Biggest Open Models Overnight

May 18 2025
By: Rachel Wong

We hit a wall trying to run the latest open-source LLMs like Qwen3, Llama 4 and Deepseek R1, even on our multi-GPU AWS instances.

At Basis Set, our engineering team is always chasing faster ways to experiment with the newest models. As models become larger, hardware limitations are one of the biggest hurdles, specifically GPU memory. For large models (72B parameters+), the memory demands are simply too immense.

As a small team, deploying these models is often infeasible due to resource intensity. Frustrated by memory constraints limiting our model performance, we were eager to find a smarter, more scalable solution.

The Use Case: Reddit Sentiment Analyzer

We built a tool internally to track sentiment and proactively surface reactions across community conversations within Reddit threads. Initially, we used OpenAI’s GPT-4.1, but it came with two problems:

We started looking for alternatives.

Try our tool

Enter Parasail + Open Source Models

We turned to Parasail for model hosting. It gave us a plug-and-play way to run the latest open-source LLMs like the new Qwen3 model they released last week and the Llama4 model released last month with the click of a button. We were eager to integrate these models into our AI systems and Parasail was the first provider to launch both of these new models.

Our first test was to see if we could replace OpenAI’s GPT 4.1 with Qwen3, because of recent developer chatter, to see what impact it would have on output quality on our Reddit sentiment analyzer. We built this tool to track public sentiment and community reactions across specific Reddit threads.

Using Parasail we were able to implement Qwen3 to test against GPT 4.1 and other popular models. We found that Qwen3 achieved comparable results to GPT 4.1 at a fraction of the cost.

Pricing Comparison

Open AI GPT-4.1

Qwen3 via Parasail Serverless

Input Output
$2.00 / 1M tokens $8.00 / 1M tokens
$0.10 / 1M tokens $0.50 / 1M tokens

Why This Matters for Builders

As the market for large AI models matures, we expect continued dramatic swings in both performance and costs of new models. Switching models and running experiments is a bottleneck. What worked for us was decoupling infra from model experimentation. Tools like Parasail make that easy. Now we can swap models in near real-time while evaluating quality and cost.

What Are You Using?

If you’re building with LLMs and hitting the same GPU or infra walls, we’d love to hear from you. We’re constantly testing new tools and workflows and always looking to exchange tips with fellow builders. Shoot us a note at bsvtech@basisset.ventures

Check out our tool outputs:

GPT 4.1

Cross-Subreddit Trends This Week

These themes appeared across multiple AI subreddits this week

Theme Sentiment Subreddits
Concerns Over AI's Impact on Employment Negative /r/ArtificialIntelligence, /r/programming
AI Hallucination and Reliability Mixed /r/singularity, /r/OpenA
Advancements in AI and Open Source Models Positive /r/OpenAI, /r/LocalLLaMA
Deepfakes and Trust Issues Negative /r/singularity
Monetization and Accessibility of Open Source Projects Mixed /r/selfhosted, /r/programming
AI in Personal and Creative Applications Positive /r/ArtificialIntelligence, /r/selfhosted

Key Takeaways

The discussion centers around Grok (an LLM by xAI) generating political outputs that some users interpret as biased or reflective of its training, especially in response to loaded or politically charged prompts. Many participants debate whether LLMs can truly 'know' their own biases or training details, with some arguing that Grok and similar models simply reflect patterns observed in training data and human discourse, while others suggest that recent research shows LLMs can develop some self-awareness of their behaviors. The thread also goes off on tangents about AI bias, the realities and limits of capitalism and communism, and the nature of political discourse online.

Positive Insights

Concerns & Criticisms

Qwen3

Cross-Subreddit Trends This Week

These themes appeared across multiple AI subreddits this week

Theme Sentiment Subreddits
AI's Impact on Jobs Negative /r/programming, /r/OpenAI
AI Agents and LLMs Mixed /r/singularity, /r/ArtificialIntelligence, /r/LocalLLaMA, /r/aiagents
Community and Open Source Contribution Positive /r/selfhosted, /r/programming
Deepfakes and Misinformation Negative /r/singularity
Open Source AI Projects Positive /r/LocalLLaMA, /r/selfhosted
Monopolies and Tech Control Negative /r/programming

Key Takeaways

The discussion revolves around LLMs like Grok's potential hallucinations, their lack of self-awareness regarding training data, and debates about political bias in AI outputs. Key themes include the limitations of LLMs in understanding their own training, the influence of training data on responses, and skepticism about claims of 'liberal bias' in AI or reality. Users also explore how system prompts, user input, and societal biases shape AI behavior.

Positive Insights

Concerns & Criticisms