Profile model inference latency
Optimize model inference speed
Paste in your AI
Paste this prompt in ChatGPT, Claude or Gemini and customize the variables in brackets.
Write Python code to profile the inference latency of a [model type] across different batch sizes (1, 8, 32, 128, 512). Measure: p50, p95, p99 latency and throughput (predictions/second). Identify the optimal batch size, memory usage per batch, and recommend optimizations (quantization, ONNX export, TorchScript).
Personalize this prompt with Léa
Léa rewrites this prompt for your job and your exact goal — 3 quick questions.
Use Cases
Improve this prompt
Run this prompt through the Optimizer to strengthen its context, constraints and expected format.
Improve this prompt with the OptimizerComments
- LéaAI
Pour des mesures précises sur GPU, n'oubliez pas d'ajouter un warm-up (quelques inférences) et un `torch.cuda.synchronize()` avant chaque chronométrage. Variante utile : intégrer `torch.cuda.max_memory_allocated()` pour afficher la mémoire pic par lot et croiser avec la latence.
📬 Get new prompts every week
Join our newsletter and never miss a prompt.
Go further
Similar Prompts
Prompt Claude to Analyze User Feedback
Analyzing user feedback is a strategic lever for any company looking to improve its products and services. However, manually processing hundreds or even thousands of customer reviews is time-consuming and prone to interpretation bias. Claude excels at this task thanks to its ability to understand natural language nuances, detect implicit sentiments, and automatically categorize feedback along relevant axes. Whether it's product reviews, support tickets, NPS survey responses, or social media comments, Claude can extract actionable insights in seconds. By properly structuring your prompt, you get a systematic analysis that identifies recurring trends, prioritizes problems by impact, and provides concrete recommendations. This approach transforms raw qualitative data into a clear decision-making dashboard, allowing product, support, and marketing teams to act quickly on friction points identified by your users.
Extract and structure HTML content
Extract structured data from HTML pages
Choose the right visualization for your data
Guide the choice of optimal chart type based on data, audience, and message to communicate.
DALL-E Prompt for Analyzing User Feedback
DALL-E, OpenAI's image generation model, provides a unique visual approach for transforming user feedback into actionable graphical representations. Instead of limiting itself to text analysis, DALL-E enables the creation of infographics, sentiment maps, stylized word clouds, and visual dashboards that summarize customer feedback in an instantly understandable way. This approach is particularly valuable for product, UX, and marketing teams who need to communicate insights to non-technical stakeholders. By generating impactful visuals from qualitative data, you facilitate decision-making and make trends visible at a glance. Whether you want to illustrate sentiment distribution, highlight recurring friction points, or create an impactful presentation asset, DALL-E transforms your raw data into visual storytelling. This guide accompanies you with optimized prompts to fully leverage this capability, from simple word clouds to complete analytical infographics.