소설 사용 사례를 해결하기 위해 사용자 정의 가능한 큰 언어 모델을 개발하는 초보자 가이드
** 저자:** 레이첼 마르티네스 ** 역할:** 개발자 ** 태그:** AI 에이전트
소개
기술 블로그에 오신 것을 환영합니다. 복잡한 개념을 모든 사람들에게 접근할 수 있도록 합니다. 오늘날 우리는 흥미로운 프로젝트에 뛰어들었습니다. 다양한 작업에 적응할 수 있는 큰 언어 모델 (LLM) 에이전트를 개발합니다.
우리는 각 작업에 가장 적합한 LLM를 선택하도록 강력한 검색, 증강 및 생성 (RAG) 인프라를 만들 것입니다. H2O와 HuggingFace 같은 회사들은 이미 비슷한 시스템을 출시하고 있습니다. 그리고 우리는 AI 기술을 그들의 애플리케이션에 통합하고자 하는 스타트업들로부터 관심을 받았습니다.
우리의 특정 사용 사례는 단일 입력 (하나의 샷) 또는 여러 개의 펌프를 묶음으로써 LLM를 촉구할 수 있는 능력을 요구합니다. 이것은 다양한 시나리오에 맞게 사용자 정의 가능하고 사용자 친화적인 LLM 모듈을 만드는 것이 필요합니다.
이 여행을 시작하면서 각 단계를 간단하고 매력적으로 정리해 주시기 바랍니다.
현재 사용 사례
현재 사용 사례에서 우리는 2 단계의 과정을 통해 사용자 서사진의 내용을 향상시킵니다.
- ** 입력**: 사용자는 자기소개서의 프로페셔널 서미과 경험 섹션들을 시스템에 입력하고 LLM에 보다 정교한 텍스트 버전을 생성하도록 요청합니다.
- ** 리뷰**: 사용자는 LLM에서 생성된 콘텐츠를 검토하고, 출력을 정비하고 개인화하기 위해 추가적인 충고를 또는 지침을 제공하여, 그들의 독특한 기술과 경험에 맞게 조화를 이루도록 한다.
이러한 접근 방식을 통해 사용자는 LLM의 힘을 활용하여 잠재적인 고용주들에게 눈에 띄는 설득력 있는 레지뷰를 작성할 수 있습니다.
목표
우리는 우리의 사용 사례를 지원하고 적어도 부분적으로 리서미 생성 프로세스를 자동화하는 완전히 사용자 정의 가능한 LLM 모듈을 구축하는 것을 목표로합니다.
시행
단계 1: LLM를 선택
필요에 맞는 LLM를 선택할 때 두 가지 주요 범위를 고려해야 합니다. 오픈소스 LLM와 독자적인 LLM. 우리는 두 그룹에서 인기 있는 선택들을 논의할 것이며, 오픈소스 LLM를 선택하는 것의 이점에 초점을 맞추어 논의할 것입니다.
오픈소스 LLM
** 예제**: 미스트랄 AI, Phi 3, LLaMa 3.
** 장점**:
- 공동체 지원: 개발자 및 사용자들의 크고 활동적인 커뮤니티는 지침을 제공하고 자원을 공유하고 개선에 협력합니다.
- ** 사용자 정의**: 소스 코드 접근은 특정 사용 사례에 맞게 수정 및 적응을 허용합니다.
- ** 투명성**: 엄격한 윤리 또는 개인정보 보호 지침을 가진 응용 프로그램에서 모델의 내부 기능을 이해하는 것이 중요합니다.
독점적 LLM
** 예제**: GPT-4.
** 장점**: 인상적 인 성능과 독특한 특징.
** 제한 사항**:
- 변편성이 부족: 모델의 사용에 제한되어 있습니다.
- 비용: 사용에 따른 수수료는 시간이 지남에 따라 증가할 수 있습니다.
- 오파크 자연: 그 근본적인 메커니즘을 이해하기 어렵기 때문에 투명성과 설명성을 우선시하는 사람들에게는 우려가 될 수 있습니다.
결론적으로, 오픈소스 및 독자적인 LLM 모두 장단점을 가지고 있지만, Mistral AI, Phi 3 및 LLaMa 3와 같은 오픈소스 LLM는 광범위한 사용자에게 매우 유익한 접근성, 사용자 정의 및 투명성을 제공합니다.
2단계: 비용 과 특징 을 살펴
오픈소스 LLM는 일반적으로 비용 효율적이기 때문에 자신의 서버에 호스팅 될 수 있습니다 또는 Azure 또는 AWS와 같은 제공자와 함께 지불-이-보-가 모델. 다음은 다양한 LLM를 개최하는 비용과 특징의 요약입니다.
| LLM | 세부 사항 | 컨텍스트 길이는 | 비용 |
|---|---|---|---|
| 믹스트랄 8x7B | MoE LLM 미스트랄 A | 훈련: 8192개의 토큰 | $0.7/1M 토큰 (치아) |
| 믹스트랄 8x22B | 미크스트랄의 후계자 | 훈련: 32768 토큰 | 2달러 1백만 달러의 토큰 (가용) |
| LLaMa 3 70B | Met의 비MoE 모델 | 훈련: 2048 토큰 | $0.02/1K 토큰 (Azu) |
| Phi 3 128K | 멀티모달 SLM | 훈련 및 총: 3 | $0.002/1K 토큰 (Az) |
| GPT-4 터보 | O의 소유자 LLM | 16K 기본, 더 높습니다 | AWS에 사용할 수 없습니다 |
단계 3: OpenRouter에 준하는 공개 API를 사용하여 테스트
가장 좋은 LLM를 선택하고 테스트하는 것은 다양한 모델에 접근하고 평가하는 효율적인 방법을 필요로 한다. 오픈 소스 API 표준인 OpenRouter는 다양한 AI 모델과 서비스 사이의 원활한 전환을 가능하게 한다. 우리는 파이어워크 AI를 우리의 공급자로 선택했습니다.
- ** 모델 다양성**: Mistral AI, Phi 3 및 LLaMa 3와 같은 인기있는 오픈소스 LLM를 포함하여 다양한 AI 모델을 제공합니다.
- ** 통합의 용이성**: OpenRouter 표준을 준수하여 기존 인프라에 통합을 단순화합니다.
- ** 확장성**: 높은 요청량을 처리하고 원활하게 확장하도록 설계되었습니다.
다음은 동일한 하드웨어를 사용하여 모델 응답과 지연성 결과를 호출하는 샘플 함수입니다:
import json
def get_sanitized_json_response(user_input, guardrail):
final = chain.invoke({"text": f"Sanitize {user_input}. Guardrail: {guardrail}"})
final = chain.invoke({"text": final.content})
return json.loads(final.content.replace("<json>", "").replace("</json>", "").replace("\n",""))
| 모델 | 지연성 | 언급 |
|---|---|---|
| 믹스트랄 8x7B | 2.1s | 가장 가벼운 모델, PERF |
| 믹스트랄 8x22B | 5.3s | 후속 모델, sim |
| LLaMa 3 70B | 2.6s | 괜찮은 반응이지만 |
| Phi 3 128K | 1.5 분 | 멀티모델 모델 |
매우 간단한 보호 릴과 입력으로 위와 같은 테스트를 수행했습니다.
guardrail = """Format details as bullets and use the XYZ method to format the relevant detail.
For example:
- I did X boosting Y percent resulting in Z.
"""
우리는 다음과 같은 해제 된 JSON 응답을 얻었습니다:
get_sanitized_json_response("""I served as Junior as a Software Engineer at TechCorp in Sydney in January 2011 and worked there until December 2018. During my tenure,
저는 AI에 기반한 금융 분석 도구를 개발한 팀을 이끌었습니다.
분산 환경과 회사의 독점적인 데이터 파이프라인의 초기 버전을 만들었습니다. 제 역할에서 저는
솔루션 설계 및 여러 API를 구현하는 것. 지난 해 블록체인 솔루션에 집중한 팀을 관리했습니다.
{'경험': {'명칭': '소프트웨어 엔지니어',
"회사" - "TechCorp",
"지구"는 "시드니"
"start_date": "2011-01",
"End_date": "2018-12"
"하이라이트": ['- 클라이언트 데이터 처리 속도를 40% 향상시키는 주요 프로젝트의 확장 가능한 데이터베이스 시스템을 설계.',
'- AI 기반 솔루션에 집중한 팀을 이끌었다.',
"- 분산 환경에서 100M+ 기록을 처리하는 AI 기반의 금융 분석 도구를 개발함.',
'- 창업 회사의 독자적인 데이터 파이프라인.',
'- 마지막 해에 블록체인 솔루션에 대한 관리팀.'
In conclusion, since the response quality was fairly comparable among all these models, it came down to latency and cost. The experiment was run 3 times achieving very similar results. So we chose Mistral AI’s Mixtral 8x7B due to its favorable balance of latency and cost.
Step 4: Building an Orchestration Layer
Managing and coordinating multiple large-scale language models or a single model across various applications is crucial for optimal performance and adaptability. Instead of building from scratch using vanilla PyTorch, employing an orchestration framework is more practical and efficient. Here are three options:
Langchain
Advantages: Vast array of integrations, supports Python and JavaScript, extensive community support.
Drawbacks: Introduces an additional dependency, potential learning curve.
LlamaIndex
Advantages: Backed by Meta, offers interoperability with Langchain.
Drawbacks: Less developer-friendly syntax.
Haystack
Advantages: User-friendly syntax, supported by prominent tech corporations.
Drawbacks: Limited integrations as it’s still maturing.
Using an orchestration framework simplifies development, improves scalability, and provides access to a supportive community, making it a more favorable choice for our project.
Design and Data Flow
The following image shows the design diagram:
The data flow within our AI-driven system is an essential aspect of ensuring a seamless and efficient user experience. While the design diagram provides a visual representation of the process, a detailed explanation can help clarify the steps involved:
- User Input: The journey begins when the user enters their details, such as the “about” and “experience” sections of their resume, into the system.
- Prompt Generation: The user’s input is then processed and transformed into one or more prompts, which will be used to query the Large Language Model (LLM) for a more polished and impressive version of the text.
- LLM Integration via Langchain: The generated prompts are passed to Langchain, our chosen orchestrator framework, which facilitates the interaction with the LLM client. Langchain leverages a Pydantic schema in the form of a class to parse and validate the data using kor, ensuring that it adheres to the required format.
- JSONResponse Generation: Once the data is validated, it is converted into a final JSONResponse, a structured and easy-to-work-with data format that simplifies further processing and manipulation.
- Content Extraction and Indexing: The content part of the JSONResponse is extracted using kor and indexed by computing embeddings, which are mathematical representations of the text. These embeddings are then stored in a vector database, such as Qdrant, for efficient retrieval and comparison during subsequent queries.
- Caching for Context Module: The entire JSONResponse is cached for the context module to function correctly. This ensures that the system can maintain the context of the user’s input and the LLM’s output during follow-up queries.
- Follow-up Queries: When the user provides additional input or requests further refinement of the generated text, the cached JSONResponse is modified according to the new context and chained to the updated prompt. This modified prompt then goes through the entire process again, from LLM integration to content extraction and indexing, to provide the user with a tailored and accurate response.
By understanding the data flow within our system, we can appreciate the importance of each step and the role of the orchestrator framework in streamlining the process, ultimately resulting in an efficient and user-friendly AI-driven solution.
LCEL를 이용한 단속 구축 및 체인
이것은 PromptTemplate 샘플의 예제입니다.
from langchain.core.output_parsers import PydanticOutputParser
prompt = PromptTemplate(
template="Please extract the following information from the given text and format it as a JSON object according to the schema:\n\n{format_instructions}\n{query}\n",
input_variables=["query"],
partial_variables={"format_instructions": output_parser.get_format_instructions()},
)
chain = prompt | chat_model | output_parser
전형적인 템플릿은 두 가지 주요 주장 (이 블로그의 범위에 해당하지 않는 더 많은 주장이 있습니다) 으로 구성됩니다. 첫 번째 정수는 실행 시 시킬 템플릿에 주입되는 문자열 목록을 받아들이고 있습니다. 이 변수는 일반적으로 Jinja2Template 또는 f-줄의 변수와 비슷합니다. chat_model 여기서 우리가 사용하고 있는 LLM가 prompt에 연쇄되어 있습니다. BASH와 유사한 `` (파이프) 운영자는 연쇄를 나타냅니다. 이것은 LCEL (LangChain 표현 언어) 스펙에서 제공되는 편리한 단편입니다. 이것은 하나의 명령어에서 다른 명령어에 대한 응답을 연쇄할 수 있게 해줍니다. 또한 함수 호출을 지원합니다. 즉, 일반적으로 출력을 정화하기 위해 람바 표현이나 일부 함수를 연쇄할 수 있습니다. 이것은 단순히 큰 함수 대신 체인을 작성함으로써 코드 복잡성을 줄입니다.
단계 5: VectorDBs을 사용
벡터 데이터베이스 (VectorDBs) 는 LLM에서 생성되는 텍스트 임베디션과 같은 고차원 벡터를 저장하고 관리하고 효율적으로 검색합니다. 다음은 인기 있는 옵션의 비교입니다.
Qdrant
** 장점**: 높은 성능, 쉬운 통합, 적극적인 커뮤니티.
** 드래그스**: 중요하지 않은 것
밀브우스
** 장점**: 여러 거리 측정, 데이터 시각화 도구.
** 회귀**: 더 급한 학습 곡선
크로마
** 장점**: 사용하기 쉬운 Python API, 모듈형 디자인.
** 드래그스**: 비교적 새로운, 덜 광범위한 커뮤니티 지원.
PGVector
** 장점**: PostgreSQLs 기능, SQL 기반의 질의를 활용합니다.
** 드래브스**: 고차원 벡터 검색에 덜 성능.
FAISS
** 장점**: 고성능 검색, GPU 지원
** 드래그스**: 다른 데이터베이스와 통합을 요구합니다.
Qdrant의 확장성, 통합 용이성, 그리고 적극적인 개발은 AI 기반 솔루션에 적합한 선택으로 만듭니다.
IR 솔루션으로 VectorDBs를 활용하는
VectorDBs를 이용한 정보 검색 (IR) 솔루션을 만드는 것은 다음과 같습니다.
- ** 텍스트 사전 처리**: 원본 데이터들을 깨끗하게 처리하고 사전 처리한다.
- ** 텍스트 임베디션 세트**: LLM을 사용하여 사전 처리 된 텍스트를 고차원 벡터로 변환합니다.
- VectorDB에 저장된 임베디션: 효율적인 검색을 위한 인덱스 임베디션
- ** 질의 처리**: 사용자 질의를 임베디드로 미리 처리하고 변환합니다.
- ** 유사성 검색**: VectorDB에서 유사성 검색을 수행하여 관련 임베디션을 찾습니다.
- ** 순위 및 표시 결과**: 검색된 임베디션을 순위화하고 다시 텍스트로 변환하여 표시합니다.
이 과정을 따라 VectorDBs를 활용하면 효율적이고 효과적인 IR 시스템을 만들 수 있습니다.