Skip to content

What is LLARS?

LLARS (LLM Assisted Research System) is an open-source web platform for researchers. In LLARS, research teams jointly rate, compare and label texts, e-mail threads and counselling conversations. Human judgements and LLM raters (LLM evaluators) run in one workflow: design prompts, generate outputs, have them evaluated and analyse how well the raters agree. LLARS is developed at the Technische Hochschule Nürnberg and is free to use for research and teaching.

Who is LLARS for?

  • Researchers who set up evaluation studies with LLM outputs or communication data, e.g. "Which answer is better?", "Was this text written by a human or an AI?" or "Which category fits this passage?".
  • Raters (evaluators) who are invited to studies and rate items there.
  • Teams building chatbots on their own documents or websites.

What can you do with LLARS?

  • Create scenarios (evaluation studies): import data, define the rating task, invite a team – see Scenario Wizard and Scenario Manager.
  • Rate: work through items in the matching rating interface – see Evaluation.
  • Use LLM evaluators: LLMs rate the same items as humans, for a direct comparison.
  • Analyse: progress, inter-rater reliability (e.g. Krippendorff's alpha) and CSV or JSON exports in the Scenario Manager.
  • Prompt Engineering: design, version and test prompts together in real time – see Prompt Engineering.
  • Batch Generation: generate outputs across many prompts, models and inputs and evaluate them directly – see Batch Generation.
  • Build chatbots with RAG: create chatbots from websites or documents and chat with them – see Chatbot Wizard.
  • Use your own LLM access: store and share your own API keys – see LLM Provider.

Which evaluation types does LLARS offer?

LLARS has eight evaluation types. The evaluation type is chosen when creating a scenario in the Scenario Wizard; the wizard suggests a type based on the uploaded data. The technical name (function_type) is given in brackets.

  • Ranking (ranking): sort items or place them in ordered categories (buckets such as good, medium, bad).
  • Rating (rating): multi-dimensional rating on Likert scales, e.g. coherence, fluency, relevance and consistency.
  • Mail rating (mail_rating): rate the quality of entire counselling e-mail threads.
  • Pairwise comparison (comparison): two answers side by side – A is better, B is better or tie.
  • Authenticity (authenticity): was the text written by a human or by an AI?
  • Labeling (labeling): assign a category to a text or conversation, binary or multi-class.
  • Communication comparison (communication_comparison): A/B comparison of counselling answers – which answer would I send?
  • Conversation labeling (conversation_labeling): label passages (spans) within a conversation one after another – see Conversation Labeling.

How do I create a scenario?

You create a new scenario (an evaluation study) in the Scenario Manager via the "+ New Scenario" button. This opens the Scenario Wizard with five steps: upload data, choose the evaluation type, configuration (dimensions, scales, buckets), invite the team (humans and LLM evaluators) and summary. The roles researcher and admin may create scenarios. The detailed guide is Scenario Wizard.

Scenario Wizard or Chatbot Wizard?

  • The Scenario Wizard creates an evaluation study. It opens in the Scenario Manager via "+ New Scenario".
  • The Chatbot Wizard builds a chatbot with its own knowledge base (RAG): crawl a website or use documents, create embeddings, generate name and system prompt. It lives in the Chatbot Admin area and does not create scenarios.

Which roles exist in LLARS?

  • researcher: create and analyse scenarios, Prompt Engineering, Batch Generation.
  • evaluator: rate in scenarios they have been invited to.
  • chatbot_manager: manage chatbots and RAG collections, Prompt Engineering.
  • admin: everything, including user and permission management in the Admin Dashboard.

Which tiles appear on the home page depends on the role.

Where do I find what in the interface?

The home page shows tiles for the individual areas:

  • Scenario Manager: own scenarios, invitations, progress and analysis; "+ New Scenario" starts the Scenario Wizard.
  • Evaluation: all scenarios you rate in; from there into the rating interface.
  • Prompt Engineering: collaborative prompt editing with versions and tests against LLMs.
  • Batch Generation: bulk generation of outputs for later evaluation.
  • Chatbot: chat with shared chatbots, including this assistant.
  • Chatbot Admin and RAG Admin: create chatbots (Chatbot Wizard) and manage knowledge collections.
  • User Settings: profile, preferences, personal API keys and own LLM providers.
  • Admin Dashboard: users, roles and system settings (admins only).

Further reading