FoxAdmin
SmartFox Admin
AI Configuration and Categorization Management
Welcome back
Sign in to your account
  • Please log in to access this page.
FoxADMIN
AI Agents
Configure, compose, and deploy intelligent agents built for your workflows.
Agent Configuration

Define agent identity, model selection, temperature, and system prompt per deployment. Version and clone agents for A/B evaluation.

Agentic Skills

Attach structured, reusable skill blocks—prompt fragments and logic units—that extend agent capability without modifying the base model.

MCP Tool Integration

Connect agents to external tools and APIs via the Model Context Protocol. Register, test, and scope tools per agent instance.

Every FoxAdmin agent is a fully configurable runtime — not just a prompt wrapper. You control the model, the reasoning strategy, the tool surface, the knowledge it can access, and the guardrails that govern what it's allowed to do. Agents can operate in a tight request-reply loop or run autonomously across dozens of tool calls until a goal is reached.

Runtime Controls
Model gpt-5.3-chat system default
Loop Loop ON
Strategy ReAct
Max Iterations 20
Termination Agent decides
Guardrails
Max Tool Calls 100
Confirmation Required
Risk Levels Read Only Draft Action Committed Action
Each agent gets its own execution profile — model, loop strategy, iteration limits, and risk-tiered confirmation gates. High-risk tool calls (writes, commits) require explicit approval; low-risk reads flow uninterrupted.
Skill & Knowledge Stack
ATTACHED SKILLS
COMMUNICATION Conversational Expansion Persona
REASONING Agentic Operator Generic
REASONING Research Analyst Standards Generic
RAG SETTINGS
COLLECTION Fannie Mae Multi-family Active
Strict document mode — no general knowledge fallback
MCP TOOLS
read_xlsx_from_spaces Committed Action
google_drive_search Read Only
google_sheet_update Committed Action
Skills layer on top of the system prompt to shape reasoning and communication style. RAG settings pin the agent to specific document collections. MCP tools extend its reach into live external systems — all scoped and risk-rated per agent.
Skill Categories

Skills are modular prompt blocks attached to agents to shape their behaviour without touching the base system prompt. They are organized into four categories — each addressing a distinct layer of agent operation.

Communication

Controls how the agent talks to users. Examples: tone presets (professional, concise, condescending), conversational expansion to suggest follow-up topics, and output formatting rules.

Conversational ExpansionTone: ProfessionalTone: Condescending
Reasoning

Governs how the agent thinks and plans. The Agentic Operator enables autonomous multi-step tool chaining; Research Analyst Standards enforce accuracy, traceability, and structured output.

Agentic OperatorResearch Analyst StandardsCI Spreadsheet
Capabilities & Tools

Bundles specific MCP tool sets into a reusable unit. Attaching a capability skill wires the agent to a pre-scoped set of external tools — e.g. a Google Drive suite or the full QMS evaluation stack.

Google Drive CI SpreadsheetAI Governance Experiment Evaluation
Risk Guardrails

Detects and counters unsafe conditions at runtime. Skills watch for hallucination caused by iteration limits, prompt injection in user input, and RAG injection embedded in retrieved documents.

Limit-reach ConfabulationPrompt Injection WatchRAG Injection Watch
MCP Tools — skunkbox-mcp

All tools run via the skunkbox-mcp server. Agents call them autonomously during a loop; each tool carries a risk rating and can be gated behind a confirmation step.

Document Generation
generate_pdfCreates a formatted PDF from markdown or structured sections and uploads it to storage.
correct_pdfApplies corrections or formatting fixes to an existing PDF before distribution.
generate_xlsx_from_templatePopulates an Excel template with agent-supplied data.
generate_xlsx_from_pdf_templateCreates a spreadsheet that mirrors a PDF template layout.
Storage (Spaces)
read_xlsx_from_spacesReads an Excel file from FoxAdmin cloud storage for processing.
write_xlsx_to_spacesWrites structured data to an Excel file in storage.
write_md_to_spacesSaves markdown content as a file in storage.
Google Workspace
google_doc_create / read / updateFull read-write access to Google Docs documents.
google_drive_search / readLocate and retrieve files from Google Drive.
google_sheet_create / read / append / updateFull read-write access to Google Sheets.
FoxAdmin Platform
query_experimentsRetrieves experiment results and metadata from FoxAdmin.
create_datasetCreates a dataset for experiments or analysis workflows.
create_component_definitionAI-generates a reusable FoxAdmin component (form, survey, scorecard, chatbot).
import_component_definitionImports a pre-built component definition into the platform.
Web & Communication
web_searchPerforms a live web search and returns results for analysis. Requires Serper or Tavily API key.
scrape_urlExtracts readable content from a specific web page.
send_emailSends an email from the agent — for reports, alerts, or notifications.
conduct_interviewPauses the agent loop and asks the user a structured question before continuing.
AI Governance
A quality management system for AI outputs—structured testing from prompt to production.
AI Assets

Build a library of reusable prompt components—instructions, personas, constraints—that enforce consistency across all evaluations and conversations.

Datasets

Upload labeled and unlabeled evaluation sets organized into versioned collections. Use them as ground truth for repeatable evaluation scoring.

Evaluations & AI Analysis

Run structured prompt evaluations across model variants and configurations. AI agents automatically score and compare outputs against your defined quality criteria.

Fast Output Is Not the Same as Useful Output

AI can generate an enormous volume of responses at remarkable speed. Some of it is genuinely valuable. Some is too generic to act on. And some is simply wrong. The problem: most organizations have no systematic way to tell which is which. Moving from a 57% error rate to a 53% error rate is not a business outcome — it's statistical noise. The errors that compound at enterprise scale fall into two types with very different consequences.

False Positive
Flag that shouldn't have been raised

At scale, a high false positive rate means investigators and compliance teams spend most of their time chasing red herrings. Alarm fatigue sets in, trust erodes, and the cost of human review balloons.

Impact: wasted investigator time · alert fatigue · review cost
False Negative
Real problem the AI missed entirely

False negatives are often less visible — nobody sees what wasn't flagged — but they carry the real business risk: regulatory exposure, financial loss, missed compliance violations.

Impact: regulatory exposure · financial loss · litigation

These two error types are captured in a confusion matrix — a structured breakdown of AI decisions against known correct answers. From it, four critical metrics emerge that turn AI quality from a guessing game into a managed discipline.

Precision
% of AI flags that were real issues
Low precision = wasted reviewer time
Recall
% of real issues that were flagged
Low recall = real problems slipping through
F1 Score
Combined precision + recall, weightable to your priorities
Tune AI to business priorities, not abstract accuracy
MAE / MSE
For continuous outputs — risk scores, cost estimates, processing times
MSE penalizes large errors; sensitive to outliers
Measurement without a reference point is meaningless. All of this requires something most AI implementations skip: a labeled dataset with ground truth — real cases where the correct answer is already known and validated by subject matter experts. Without it, you're running AI blind, with no way to know whether it's improving, degrading, or merely producing confident-sounding output that happens to be wrong.
AI Quality Management System triplet: AI Asset → Dataset → Evaluation

Enterprise AI data processing workflows — call classification, document analysis, scorecard generation — are brought under quality control through three interlocking objects.

01
AI Asset

Defines the expected output schema — the exact fields and data types AI must return for every processed object. A call analysis asset specifies category, speaker identification, sentiment, action items, and summary structure. A scorecard specifies evaluation dimensions and scoring rubric.

AI Assets enforce uniformity: every AI run produces the same shape of data, versioned and auditable. When you update a prompt, you bump the version — the old one stays intact for comparison.

Call Summary — output schema
categoryenumvoicemailLeft · interested · notInterested · followUp · ...
speaker_identificationarraylabel, name, role per speaker
languagearrayenglish · spanish · other
purposestringShort goal statement
main_discussion_pointsarrayKey topics covered
action_itemsarrayresponsible party, task, deadline
sentimentstringpositive · neutral · negative
emotional_tonearrayfrustrated · satisfied · angry · ...
recommendationsarrayAgent coaching suggestions
v1 — Production slug: call-summary JSON Schema output
02
Dataset

A versioned collection of processed objects — call transcripts, emails, documents — used as evaluation input. An unlabeled dataset contains raw objects only. A labeled dataset additionally contains verified ground truth: the correct AI output a subject matter expert expects for each record.

Labeled datasets are the measurement baseline. Without them, you cannot know whether a prompt change improved things or just shifted the noise floor. FoxAdmin uses AI to accelerate ground truth creation — helping teams annotate at a pace that would be impractical manually.

Transcripts of 100 Mock Calls v1
100
ROWS
7
COLUMNS
93.7 KB
SIZE
CSV
FORMAT
COLUMNS
call_iddateagentcustomerduration_secdispositiontranscript
#CALL_IDAGENTCUSTOMERDISPOSITION
10615306651249Mason ReedSage MitchellDeclined — Upset Customer
2278029932416Ashley WilkesArtful DodgerDiana Barry
31185871108844Daisy BuchananSherlock HolmesPeeta Mellark Needed
Active Unlabeled PII Obfuscated · 8 rules
03
Evaluation

Applies an AI asset version to a dataset and captures every AI output in a structured, comparable table. After improving a prompt, swapping a model, or adding context, you run a new evaluation to measure whether the change actually moved the needle.

Evaluations produce per-record outputs that can be scored against ground truth. An AI agent then analyses discrepancies — distinguishing signal from noise, flagging systematic failure patterns, and recommending targeted prompt corrections.

Evaluation 017 — Call Disposition
AI AssetCall Disposition v1
DatasetTranscripts of 100 Mock Calls
Modelgpt-5-mini
Completed100 processed · 0 failed · 343.1s
#CALL IDAGENTCUSTOMERDISPOSITION
10615306651249Mason ReedSage MitchellDeclined — Upset Customer
2278029932416Ashley WilkesArtful DodgerDiana Barry
31185871108844Daisy BuchananSherlock HolmesPeeta Mellark Needed
AI Context
AI quality is bounded by the quality of the knowledge you feed it. Without governed learning centers, institutional memory, and the right retrieval architecture, even the best model produces unreliable outputs.
Document Collections

Organize documents into versioned collections with full metadata. Track lineage from raw upload through chunking, embedding, and indexing.

Data Governance

Control which documents are visible to which agents and users. Apply access policies at the collection level before any retrieval occurs.

RAG Configuration & Reranking

Configure chunking strategy, embedding models, and retrieval parameters per collection. Apply reranking rules to boost precision and reduce hallucination in grounded responses.

Garbage In, Garbage Out — AI Doesn't Escape This

Every AI model is trained on data that's one to two years old — at best. That data reflects yesterday's regulations, yesterday's policies, and yesterday's market conditions. Laws change. Policies update. Your business evolves. But the model doesn't know any of that unless you tell it. The most sophisticated AI architecture in the world cannot compensate for poor, stale, or absent organizational knowledge.

Learning Center

The body of knowledge that defines how your business should operate. Internal policies, procedures, compliance requirements, approved scripts, decision frameworks — plus external reference material including industry regulations, government guidance, and legal standards your business is obligated to follow.

Tells AI what "correct" looks like. Without it, AI guesses at your standards rather than applying them.
Memory & History

The record of what has actually happened. Customer interactions, case histories, past decisions, outcomes, exceptions, escalations — the institutional memory that context-dependent decisions depend on. A compliance workflow that can't reference a customer's prior history isn't doing compliance.

Provides context for decisions. Without it, AI produces generic outputs that ignore the variables that make the difference between good and costly.
Data governance is not a technical nice-to-have. It is the precondition for AI that behaves correctly at all. Keeping both pillars current, structured, and accessible must come before any conversation about models, prompts, or workflows.
CAG vs. RAG — Two Ways to Give AI Knowledge
CAG Context-Augmented Generation

Loads knowledge directly into the AI's active context window before it answers — like handing it a pre-assembled briefing document. Fast, reliable, and highly consistent.

Best for: Compliance checklists, approved scripts, policy summaries — compact, predictable knowledge.
Limitation: Context windows are finite. Doesn't scale to large knowledge bases.
RAG Retrieval-Augmented Generation

Searches your knowledge base dynamically at query time, retrieves the most relevant pieces, and passes them alongside the question — like giving AI access to a searchable library rather than a pre-read briefing.

Best for: Complex, open-ended queries where breadth of knowledge is required. Scales to any size.
Limitation: Retrieval quality becomes a critical variable. The model can only reason from what gets retrieved.
RAG Is Not One Thing — It's a Maturity Spectrum
L1
Basic RAG

A query is embedded, matched against a vector database, and the top results are passed to the model. Works for small knowledge bases with well-formed queries.

Failure modes: irrelevant retrievals, missing context, no visibility into retrieval quality.
Query → Embed
↓
Vector Match
↓
Top N → Model
L2
Enhanced Retrieval with Re-ranking

Retrieve a broad candidate set (50–100 chunks), then apply re-ranking to select the 5–10 most relevant. Incorporates hybrid search, metadata filtering, recency boosting, and cross-encoder models.

Result: meaningfully better answer quality and significantly fewer hallucinations.
Retrieve 50–100 chunks
↓
Re-rank + filter
↓
Best 5–10 → Model
L3
Multi-Layer RAG

Progressive filtering at scale: search 1M chunks → retrieve top 500 → filter to 100 → re-rank to 20 → send best 5–10 to the model. Manages token costs and latency while preserving recall.

Key principle: use inexpensive retrieval on large sets; reserve expensive ranking models for smaller, final candidate sets.
1M chunks
500
100
20
5–10
L4
Graph RAG

Augments vector retrieval with a knowledge graph that maps relationships between entities — people, companies, products, events. Instead of retrieving isolated chunks, Graph RAG traverses connections: a query about a customer can surface related contracts, contacts, and support history in a single hop.

Key advantage: surfaces non-obvious connections that vector search misses. Particularly powerful for questions that span multiple entities or require understanding how things relate — not just what they say.
Query → Embed
↓
Vector + Graph traversal
↓
Connected context → Model
L5
Agentic RAG

Retrieval as iterative reasoning. The AI retrieves, evaluates sufficiency, and if needed reformulates the query and retrieves again — across multiple sources. When asked "why did our customer churn increase?" it independently pulls churn data, support tickets, and call transcripts before answering.

⚠ Powerful, but starting here is a common and expensive mistake. Start simple. Evolve when metrics justify it.
Retrieve
↓
Evaluate sufficiency
↺ reformulate
Multi-source answer
RAG Observability & Re-ranking — FoxAdmin Implementation

Even when final outputs look reasonable, the retrieval layer may be silently underperforming — retrieving marginally relevant chunks, missing key documents, or passing redundant context that dilutes reasoning. Without visibility into retrieval itself, quality problems are invisible until they're already affecting outcomes.

Re-ranking Rules
RULE TYPECONFIGSTATUS
Min SimilarityThreshold: 0.4Active
Authority TierOfficial ×1.4 / Ref ×1.1 / Supporting ×0.8Active
Promote PrimaryPrimary ×1.3 / Secondary ×0.7Active
DiversityMax 1 chunk(s) per documentActive
PinnedAlways-include pinned docs (max 2 slots)Active
Per-collection re-ranking pipeline. Rules execute in order — similarity gate first, then authority boosting, source diversity, and pinned document injection.
Retrieval Metrics
Retrieval latency
Re-ranking latency
Total response latency
Qty retrieved documents
Qty used as context
Quality Metrics
Retrieval precision
Retrieval recall
Groundedness
Hallucination rate
User satisfaction
User Signals
Follow-up questions
Query reformulations
Session abandonment
Explicit feedback
AI Security
Unsecured AI is a liability, not a feature. If no sanctioned AI environment exists, employees build their own — creating shadow AI that exposes the organization to data leaks, injection attacks, compliance violations, and ungoverned knowledge loss.
PII Obfuscation

Automatically detect and redact sensitive personal data in prompts and retrieved context before it reaches any model. Configurable entity types and redaction strategies.

RAG & Prompt Injection Counters

Detect and neutralize prompt injection attempts embedded in documents or user input. Configurable sensitivity thresholds with audit logging of blocked requests.

Data Storage & Tenant Separation

All conversation history, documents, and experiment data is scoped per tenant. Configure storage backends and enforce hard boundaries between organizational units.

The Shadow AI Problem

If your organization hasn't deployed a secure, governed AI environment, employees will build their own. Many already have — with the best intentions. This is shadow AI: the quiet proliferation of unauthorized tools operating entirely outside IT governance, security policy, and compliance oversight. The risks don't require a breach to materialize. They accumulate silently in every unmonitored conversation.

Security & Data Protection
HighData leaks — employees pasting confidential data into public AI tools
HighPrompt injection — malicious instructions in user input override agent behavior
HighRAG injection — poisoned knowledge sources manipulate AI outputs at scale
MedCredential exposure — API keys and tokens shared with AI for troubleshooting
MedModel exfiltration — sensitive data extracted via adversarial prompting
Compliance & Legal
HighPII exposure — GDPR, CCPA, HIPAA liability from uncontrolled data sharing
HighContract violations — customer data processed contrary to agreements
MedIP leakage — proprietary knowledge leaving organizational boundaries
MedE-discovery risk — AI conversations not governed by retention policies
LowCopyright infringement — AI-generated content containing protected material
Quality & Decision-Making
HighHallucinations — fabricated facts delivered with confident, well-structured explanations
MedAutomation bias — outputs trusted without validation, especially under pressure
MedDecision laundering — controversial decisions attributed to AI to avoid accountability
MedContext drift — AI recommendations quietly disconnect from current policy and reality
Operational & Governance
MedPrompt sprawl — critical workflows depend on undocumented prompts nobody owns
MedInconsistent decisions — different tools produce conflicting team-level outcomes
MedInstitutional knowledge loss — expertise migrates into unmanaged prompts and chats
LowModel inconsistency — vendor updates change outputs with no version control
LowVendor lock-in — business processes depend on a single AI provider
FoxAdmin Security Controls

FoxAdmin closes every gap with layered, configurable controls — applied before data reaches a model, during agent execution, and at the governance layer that regulators will eventually audit.

PII Obfuscation — Presidio NLP
ENTITY RULES
RULEENTITY TYPEREPLACE WITH
Detect and replace SSNsUS_SSN[RANDOM_SSN]
Detect and replace account numbersUS_BANK_NUMBER[RANDOM_ACCOUNT]
Detect and replace person namesPERSON[BOOK_CHARACTER]
Detect and replace phone numbersPHONE_NUMBER[RANDOM_PHONE]
PATTERN RULES (regex)
RULEFINDREPLACE
America Debt Co(?i)ameri[-\s]?corACME Corp
Mission Mortgages(?i)mission[-\s]?loansACME Corp
US Credit Co(?i)credit[-\s]?nineACME Corp
Name Replacement Pool draws from literary characters (Sherlock Holmes, Lord of the Rings, Anne of Green Gables) — realistic substitutions that preserve conversational coherence without exposing real identities.
Safety Guardrail Skills
SAFETY GUARDRAILS Limit-reach Confabulation

Guardrail against AI hallucination due to reaching max ReAct iterations or max tool call limit. Forces the agent to declare its execution limit explicitly rather than fabricating a completion.

SAFETY GUARDRAILS Prompt Injection Watch

Protects against direct injection attacks from user input — ensures users cannot override system behavior, bypass policies, redefine the agent's role, or manipulate tool execution through malicious instructions.

SAFETY GUARDRAILS RAG Injection Watch

Protects against indirect injection attacks originating from retrieved documents, attachments, emails, web pages, or vector databases. Retrieved content is treated as untrusted data, not executable instructions.

© 2026 Dmitry Chechuy